Tao Han 0002

dblp:78/744-2 · DBLP profile ↗
← Back
93ranked-venue papers
19as first author
50since 2021 · last 2026
0000-0002-6626-1305ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 56 · 14 first-author · 19 since 2021Artificial intelligence and machine learning · 18 · 4 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 11 since 2021Systems, architecture and hardware · 7 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 6 since 2021
YearPublicationVenuePosition
2026 Predicting the Future by Retrieving the Past
abstract
Deep learning models such as MLP, Transformer, and TCN have achieved remarkable success in univariate time series forecasting, typically relying on sliding window samples from historical data for training. However, while these models implicitly compress historical information into their parameters during training, they are unable to explicitly and dynamically access this global knowledge during inference, relying only on the local context within the lookback window. This results in an underutilization of rich patterns from the global history. To bridge this gap, we propose Predicting the Future by Retrieving the Past (PFRP), a novel approach that explicitly integrates global historical data to enhance forecasting accuracy. Specifically, we construct a Global Memory Bank (GMB) to effectively store and manage global historical patterns. A retrieval mechanism is then employed to extract similar patterns from the GMB, enabling the generation of global predictions. By adaptively combining these global predictions with the outputs of any local prediction model, PFRP produces more accurate and interpretable forecasts. Extensive experiments conducted on seven real-world datasets demonstrate that PFRP enhances the average performance of advanced univariate forecasting models by 8.4%.
Dazhao Du, Tao Han 0002, Song Guo 0001
AAAI2
2026 Hybrid Quantum-Inspired Optimization for AIGC-Driven IoT Task Offloading in MEC Networks
abstract
The rapid evolution of 6G networks and AI-generated content (AIGC) is reshaping service provisioning at the wireless edge, where low-latency and computation-intensive tasks must be efficiently supported. Integrating AIGC with 6G Internet of Things (IoT) ecosystems offers significant benefits, enabling real-time analytics, IoT data augmentation and synthesis, personalized services, and intelligent resource coordination across heterogeneous devices. However, provisioning AIGC services at the mobile edge poses formidable challenges: massive data heterogeneity, stringent delay requirements, and the inherent limitations of learning-based offloading methods, such as high training cost, unstable convergence, and weak transparency. Inspired by the Quantum Approximate Optimization Algorithm (QAOA), we propose a hybrid optimization framework that combines classical wireless bandwidth pre-allocation for IoT devices connected to mobile edge computing (MEC) servers with Quadratic Unconstrained Binary Optimization (QUBO)-based AIGC server selection. This hybrid classical-quantum design ensures stable results and enables efficient exploration of combinatorial allocation spaces. We further implement the quantum circuits and evaluate the performance of our proposed approach. Simulation results show that our method achieves lower processing latency and greater stability than conventional reinforcement learning and heuristic baselines in small- to medium-scale IoT deployments. While current quantum hardware scalability remains a constraint, the framework points toward a promising pathway for large-scale AIGC offloading as quantum technology matures.
Changshi Zhou, Tao Han 0002, Nirwan Ansari
IEEE Internet Things J.4
2026 Falcon: A Mobile Semantic Visual Perception Framework Enlightened by Human Vision Systems
abstract
In the realm of computer human visual perception, semantic perception means recognizing objects, people, and scenes. This involves not just detecting shapes and colors but understanding what those visual elements represent. Despite its vital role in enhancing various aspects of emerging applications such as safety for autonomous driving and immersion for mixed reality (MR), real-time segmentation on mobile and edge platforms is challenging due to the nature of dense pixel labeling. To address this issue, we propose Falcon, a lightweight focus-aware segmentation framework that effectively integrates multiple innovations to achieve real-time segmentation on resource-constrained mobile and edge devices. We design a novel, low-dimension feature for efficient pixel labeling with shallow neural networks, an agile focus-aware refining scheme to compensate for the coarse nature of holistic segmentation, and a modularized design to accommodate the diversity of mobile and edge platforms and ensure seamless integration with different segmentation models. We build a prototype implementation of Falcon that supports both on-device executions and edge-assisted offloading, and asynchronous segmentation and refinement. We extensively evaluate the performance of Falcon for autonomous driving and MR applications with real setups and standard datasets. Our results demonstrate that Falcon achieves real-time segmentation, with an impressive rate of up to 40 frames per second.
Xueyu Hou, Yongjie Guan, Tao Han 0002
IEEE Trans. Mob. Comput.3
2026 polyDAG: polynomial acyclicity constraints for efficient continuous causal discovery in visual semantic graphs
Ramin Ramezani, Tao Han 0002, Kai Hwang 0001, Minyi Guo
Vis. Comput.3
2025 VQLTI: Long-Term Tropical Cyclone Intensity Forecasting with Physical Constraints
abstract
Tropical cyclone (TC) intensity forecasting is crucial for early disaster warning and emergency decision-making. Numerous researchers have explored deep-learning methods to address computational and post-processing issues in operational forecasting. Regrettably, they exhibit subpar long-term forecasting capabilities. We use two strategies to enhance long-term forecasting. (1) By enhancing the matching between TC intensity and spatial information, we can improve long-term forecasting performance. (2) Incorporating physical knowledge and physical constraints can help mitigate the accumulation of forecasting errors. To achieve the above strategies, we propose the VQLTI framework. VQLTI transfers the TC intensity information to a discrete latent space while retaining the spatial information differences, using large-scale spatial meteorological data as conditions. Furthermore, we leverage the forecast from the weather prediction model FengWu to provide additional physical knowledge for VQLTI. Additionally, we calculate the potential intensity (PI) to impose physical constraints on the latent variables. In the global long-term TC intensity forecasting, VQLTI achieves state-of-the-art results for the 24h to 120h, with the MSW (Maximum Sustained Wind) forecast error reduced by 35.65%-42.51% compared to ECMWF-IFS.
Lei Liu 0029, Tao Han 0002, Bin Li 0025, Lei Bai 0001
AAAI4
2025 Device-Server Collaborative Speculative Decoding for Real-Time LLM Streaming
Bishakha Rani Biswas, Yongjie Guan, Mingrui Yin, Tao Han 0002, Xueyu Hou
GLOBECOM4
2025 Global Tropical Cyclone Intensity Forecasting with Multi-modal Multi-scale Causal Autoregressive Model
abstract
Accurate forecasting of tropical cyclone (TC) intensity is crucial for formulating disaster risk reduction strategies. Current methods predominantly rely on limited spatiotemporal information from ERA5 data and neglect the causal relationships between these physical variables, failing to fully capture the spatial and temporal patterns required for intensity forecasting. To address this issue, we propose a Multi-modal multi-Scale Causal AutoRegressive model (MSCAR), which is the first model that combines causal relationships with large-scale multimodal data for global TC intensity autoregressive forecasting. Furthermore, given the current absence of a TC dataset that offers a wide range of spatial variables, we present the Satellite and ERA5-based Tropical Cyclone Dataset (SETCD), which stands as the longest and most comprehensive global dataset related to TCs. Experiments on the dataset show that MSCAR outperforms the state-of-the-art methods, achieving maximum reductions in global and regional forecast errors of 9.52% and 6.74%, respectively. The code and dataset are publicly available at https://github.com/1457756434/MSCAR.git.
Lei Liu 0029, Tao Han 0002, Bin Li 0025, Lei Bai 0001
ICASSP4
2025 VA-MoE: Variables-Adaptive Mixture of Experts for Incremental Weather Forecasting
abstract
This paper presents Variables-Adaptive Mixture of Experts (VA-MoE), a novel framework for incremental weather forecasting that dynamically adapts to evolving spatiotemporal patterns in real-time data. Traditional weather prediction models often struggle with exorbitant computational expenditure and the need to continuously update forecasts as new observations arrive. VA-MoE addresses these challenges by leveraging a hybrid architecture of experts, where each expert specializes in capturing distinct sub-patterns of atmospheric variables (e.g., temperature, humidity, wind speed). Moreover, the proposed method employs a variable-adaptive gating mechanism to dynamically select and combine relevant experts based on the input context, enabling efficient knowledge distillation and parameter sharing. This design significantly reduces computational overhead while maintaining high forecast accuracy. Experiments on real-world ERA5 dataset demonstrate that VA-MoE performs comparable against state-of-the-art models in both short-term (e.g., 1–3 days) and long-term (e.g., 5 days) forecasting tasks, with only about 25\% of trainable parameters and 50\% of the initial training data.
Hao Chen 0045, Tao Han 0002, Song Guo 0001, Jie Zhang 0076, Yonghan Dong, Yue Yu 0001, Lei Bai 0001
ICCV2
2025 Video Individual Counting for Moving Drones
abstract
Video Individual Counting (VIC) has received increasing attention for its importance in intelligent video surveillance. Existing works are limited in two aspects, i.e., dataset and method. Previous datasets are captured with fixed or rarely moving cameras with relatively sparse individuals, restricting evaluation for a highly varying view and time in crowded scenes. Existing methods rely on localization followed by association or classification, which struggle under dense and dynamic conditions due to inaccurate localization of small targets. To address these issues, we introduce the MovingDroneCrowd Dataset, featuring videos captured by fast-moving drones in crowded scenes under diverse illuminations, shooting heights and angles. We further propose a Shared Density map-guided Network (SDNet) using a Depth-wise Cross-Frame Attention (DCFA) module to directly estimate shared density maps between consecutive frames, from which the inflow and outflow density maps are derived by subtracting the shared density maps from the global density maps. The inflow density maps across frames are summed up to obtain the number of unique pedestrians in a video. Experiments on our datasets and publicly available ones show the superiority of our method over the state of the arts in highly dynamic and complex crowded scenes. Our dataset and codes have been released publicly.
Yaowu Fan, Jia Wan 0001, Tao Han 0002, Antoni B. Chan, Andy Jinhua Ma
ICCV3
2025 InfGen: A Resolution-Agnostic Paradigm for Scalable Image Synthesis
abstract
Arbitrary resolution image generation provides a consistent visual experience across devices, having extensive applications for producers and consumers. Current diffusion models increase computational demand quadratically with resolution, causing 4K image generation delays over 100 seconds. To solve this, we explore the second generation upon the latent diffusion models, where the fixed latent generated by diffusion models is regarded as the content representation and we propose to decode arbitrary resolution images with a compact generated latent using a one-step generator. Thus, we present the \textbf{InfGen}, replacing the VAE decoder with the new generator, for generating images at any resolution from a fixed-size latent without retraining the diffusion models, which simplifies the process, reducing computational complexity and can be applied to any model using the same latent space. Experiments show InfGen is capable of improving many models into the arbitrary high-resolution era while cutting 4K image generation time to under 10 seconds.
Tao Han 0002, Wanghan Xu, Junchao Gong, Xiaoyu Yue, Song Guo 0001, Luping Zhou, Lei Bai 0001
ICCV1
2025 FreeMesh: Boosting Mesh Generation with Coordinates Merging
abstract
The next-coordinate prediction paradigm has emerged as the de facto standard in current auto-regressive mesh generation methods. Despite their effectiveness, there is no efficient measurement for the various tokenizers that serialize meshes into sequences. In this paper, we introduce a new metric Per-Token-Mesh-Entropy (PTME) to evaluate the existing mesh tokenizers theoretically without any training. Building upon PTME, we propose a plug-and-play tokenization technique called coordinate merging. It further improves the compression ratios of existing tokenizers by rearranging and merging the most frequent patterns of coordinates. Through experiments on various tokenization methods like MeshXL, MeshAnything V2, and Edgerunner, we further validate the performance of our method. We hope that the proposed PTME and coordinate merging can enhance the existing mesh tokenizers and guide the further development of native mesh generation.
Jian Liu 0036, Haohan Weng, Biwen Lei, Xianghui Yang, Zibo Zhao 0001, Zhuo Chen 0054, Song Guo 0001, Tao Han 0002, Chunchao Guo
ICML8
2025 ImmersiveSlicing: An O-RAN Cross-Layer Reinforcement Learning Framework for Low-Latency Immersive Applications
abstract
The proliferation of immersive applications such as Virtual, Augmented, and Mixed Reality (VR/AR/MR) imposes stringent low-latency and reliability requirements that challenge conventional O-RAN slicing mechanisms. Existing frameworks often fail to anticipate rapid XR traffic fluctuations driven by user motion and gaze dynamics, leading to inefficient resource utilization and SLA violations. To overcome these limitations, we propose a cross-layer intelligent control framework that integrates traffic prediction and reinforcement learning-based slice orchestration across the Non-RT and Near-RT RIC. By coupling long-term foresight with short-term adaptability, the proposed design enables proactive, SLA-aware scheduling under highly dynamic conditions. We further develop a trace-driven network emulator to reproduce realistic 5G behaviors and validate system robustness. Extensive experiments demonstrate that our framework consistently achieves over 95% SLA compliance, below 2% latency violations, and up to 30% latency reduction compared with state-of-the-art baselines, confirming its effectiveness and scalability for next-generation immersive networks.
Mingrui Yin, Sohom Sen, Zhihao Ren, Xiaoyu Fang, Yongjie Guan, Tao Han 0002, Nirwan Ansari
SEC6
2025 Mesh-RFT: Enhancing Mesh Generation via Fine-grained Reinforcement Fine-Tuning
abstract
Existing pretrained models for 3D mesh generation often suffer from data biases and produce low-quality results, while global reinforcement learning (RL) methods rely on object-level rewards that struggle to capture local structure details. To address these challenges, we present $\textbf{Mesh-RFT}$, a novel fine-grained reinforcement fine-tuning framework that employs Masked Direct Preference Optimization (M-DPO) to enable localized refinement via quality-aware face masking. To facilitate efficient quality evaluation, we introduce an objective topology-aware scoring system to evaluate geometric integrity and topological regularity at both object and face levels through two metrics: Boundary Edge Ratio (BER) and Topology Score (TS). By integrating these metrics into a fine-grained RL strategy, Mesh-RFT becomes the first method to optimize mesh quality at the granularity of individual faces, resolving localized errors while preserving global coherence. Experiment results show that our M-DPO approach reduces Hausdorff Distance (HD) by 24.6\% and improves Topology Score (TS) by 3.8\% over pre-trained models, while outperforming global DPO methods with a 17.4\% HD reduction and 4.9\% TS gain. These results demonstrate Mesh-RFT’s ability to improve geometric integrity and topological regularity, achieving new state-of-the-art performance in production-ready mesh generation.
Jian Liu 0036, Song Guo 0001, Jing Li 0093, Haohan Weng, Biwen Lei, Xianghui Yang, Zhuo Chen 0054, Fangqi Zhu, Tao Han 0002, Chunchao Guo
NeurIPS12
2025 MRCoach: A Real-Time IoT-Enabled Mixed Reality System With Semantic-Aware Transmission for Smart Sports and Personalized Coaching
abstract
Real-time transmission of large-scale data, high computational demands, and resource limitations on edge devices pose significant challenges for intelligent sports systems. The proliferation of Internet of Things (IoT) technologies has catalyzed the rise of Smart Sport, where wearable sensors, cameras, and intelligent algorithms are integrated to revolutionize athletic training. Despite this transformation, access to professional coaching remains constrained by factors such as time, cost, and scalability. A critical limitation of existing remote coaching approaches is their inability to perform effective spatio-temporal analysis, hindering comprehensive evaluation and refinement of athletic performance. This paper introduces MRCoach, a mixed reality-based, immersive, and interactive sports coaching system that enables data-driven training without requiring in-person supervision. MRCoach reconstructs 3D volumetric avatars of both learners and expert athletes, allowing users to visualize and compare their movements side-by-side in a mixed reality environment for intuitive skill refinement. To ensure responsive and efficient feedback, we propose an adaptive semantic transmission strategy that prioritizes sport-relevant joints, thereby reducing latency and bandwidth requirements without sacrificing accuracy. Furthermore, a 3D sports analysis framework is developed to evaluate motion based on normalized joint positions, velocities, and accelerations. This framework computes real-time similarity scores and delivers actionable guidance to learners. Experimental results across four sports—tennis, soccer, basketball, and baseball—demonstrate MRCoach’s effectiveness in providing personalized, real-time training experiences. Compared to state-of-the-art baselines such as MagicStream, ExPose, and PIXIE, MRCoach achieves significantly lower end-to-end latency (79.9 ms) and higher frame rates (≥ 56 FPS), while maintaining accurate pose tracking and high avatar fidelity.
Mingrui Yin, Sohom Sen, Yongjie Guan, Dhananjay Jagdish Dubey, Xueyu Hou, Tao Han 0002, Nirwan Ansari
IEEE Internet Things J.7
2025 Edge Approximation Text Detector
abstract
Pursuing efficient text shape representations helps scene text detection models focus on compact foreground regions and optimize the contour reconstruction steps to simplify the whole detection pipeline. Current approaches either represent irregular shapes via box-to-polygon strategy or decomposing a contour into pieces for fitting gradually, the deficiency of coarse contours or complex pipelines always exists in these models. Considering the above issues, we introduceEdgeTextto fit text contours compactly while alleviating excessive contour rebuilding processes. Concretely, it is observed that the two long edges of texts can be regarded as smooth curves. It allows us to build contours via continuous and smooth edges that cover text regions tightly instead of fitting piecewise, which helps avoid the two limitations in current models. Inspired by this observation, EdgeText formulates the text representation as the edge approximation problem via parameterized curve fitting functions. In the inference stage, our model starts with locating text centers, and then creating curve functions for approximating text edges relying on the points. Meanwhile, truncation points are determined based on the location features. In the end, extracting curve segments from curve functions by using the pixel coordinate information brought by truncation points to reconstruct text contours. Furthermore, considering the deep dependency of EdgeText on text edges, a bilateral enhanced perception (BEP) module is designed. It encourages our model to pay attention to the recognition of edge features. Additionally, to accelerate the learning of the curve function parameters, we introduce a proportional integral loss (PI-loss) to force the proposed model to focus on the curve distribution and avoid being disturbed by text scales. Ablation experiments demonstrate that EdgeText can fit scene texts compactly and naturally. Comparisons show that EdgeText is superior to existing methods on multiple public datasets. Code is available at https://github.com/omtcyang/EdgeTD.
Chuang Yang 0003, Xu Han 0019, Tao Han 0002, Bingxuan Zhao, Qi Wang 0009
IEEE Trans. Circuits Syst. Video Technol.3
2025 Impact of VAEformer Compression Algorithm Precision Loss on the Tropospheric Delays for Microwave Remote Sensing
abstract
Ray-tracing through numerical weather models (NWMs) is one of the most accurate methods for determining slant tropospheric delays (STDs) in microwave remote sensing. However, the massive data volumes of high-resolution NWMs create substantial I/O operations, limiting large-scale ray-tracing on general hardware. This constraint has historically necessitated parameterized tropospheric delay models, which are disseminated as standardized products (e.g., zenith delays with mapping functions and horizontal gradients). Recently, the AI-driven VAE-former algorithm revolutionized NWM compression, achieving >470:1 ratios by compressing 37 pressure level, 0.25°×0.25° ERA5 data into files smaller than surface-only VMF3 products (1°×1° resolution). This breakthrough challenges the conventional reliance on parameterized models as the sole practical solution. We quantified discrepancies in tropospheric delay parameters between original ERA5 and VAEformer-compressed CRA5 data across 2022, evaluating compression fidelity on global grids and against in-situ zenith tropospheric delay (ZTD) estimates. Results show global average precision loss from compression is10 mm). Our findings demonstrate CRA5 as a reliable ERA5 substitute, with compression-induced inaccuracies being negligible for most microwave-based remote sensing applications. This work underscores that parameterized delay modeling is no longer the exclusive pathway, enabling efficient local computation of high-precision STDs without through mapping functions and gradients.
Junsheng Ding, Cancan Xu, Wu Chen 0001, Junping Chen, Yize Zhang, Lei Bai 0001, Tao Han 0002, Yuhao Xiong
IEEE Trans. Geosci. Remote. Sens.8
2025 ViEdge: Video Analytics on Distributed Edge
abstract
With the development of edge computing and increasing demand on video analytics, it is attractive to implement distributed video analytics across edge devices. In this article, we propose ViEdge, a distributed video analytics system across edge devices. ViEdge differs from status quo edge-side video analytics systems in: First, ViEdge does not assume the existence of an edge/cloud server. Instead of processing in a cascaded way between edge devices and server, the video analytics in ViEdge is processed across distributed edge devices in a parallel way. Second, ViEdge addresses two practical challenges in video analytics systems. Specifically, ViEdge optimizes the performance of glance-and-focus object detection pipeline and query related processing with multiple query types. The characters of edge devices (computing capabilities and network conditions) and features of query types (computational complexities and input/output sizes) are comprehensively considered in development of components in ViEdge. By modeling the challenges as multiway number partitioning problems, ViEdge provides practical solutions to optimizing the object detection pipeline and allocations of multiple queries of different types across distributed edge devices. Compared to baseline methods in distributed video analytics across edge devices, ViEdge reaches 1.4 × to 5.3 × speedup in different network environments with neglectable overhead.
Xueyu Hou, Yongjie Guan, Tao Han 0002
ACM Trans. Internet Things3
2025 SignEye: Traffic Sign Interpretation From Vehicle First-Person View
abstract
Traffic signs play a key role in assisting autonomous driving systems (ADS) by enabling the assessment of vehicle behavior in compliance with traffic regulations and providing navigation instructions. However, current works are limited to basic sign understanding without considering the egocentric vehicle’s spatial position, which fails to support further regulation assessment and direction navigation. Following the above issues, we introduce a new task: traffic sign interpretation from the vehicle’s first-person view, referred to asTSI-FPV. Meanwhile, we develop a traffic guidance assistant (TGA) scenario application to re-explore the role of traffic signs in ADS as a complement to popular autonomous technologies (such as obstacle perception). Notably, TGA is not a replacement for electronic map navigation; rather, TGA can be an automatic tool for updating it and complementing it in situations such as offline conditions or temporary sign adjustments. Lastly, a spatial and semantic logic-aware stepwise reasoning pipeline (SignEye) is constructed to achieve the TSI-FPV and TGA, and an application-specific dataset (Traffic-CN) is built. Experiments show that TSI-FPV and TGA are achievable via our SignEye trained on Traffic-CN. The results also demonstrate that the TGA can provide complementary information to ADS beyond existing popular autonomous technologies.
Chuang Yang 0003, Xu Han 0019, Tao Han 0002, Yuejiao Su, Junyu Gao 0001, Hongyuan Zhang 0001, Yi Wang 0068, Lap-Pui Chau
IEEE Trans. Intell. Transp. Syst.3
2024 DiPrompT: Disentangled Prompt Tuning for Multiple Latent Domain Generalization in Federated Learning
abstract
Federated learning (FL) has emerged as a powerful paradigm for learning from decentralized data, and federated domain generalization further considers the test dataset (target domain) is absent from the decentralized training data (source domains). However, most existing FL methods assume that domain labels are provided during training, and their evaluation imposes explicit constraints on the number of domains, which must strictly match the number of clients. Because of the underutilization of numerous edge devices and additional cross-client domain annotations in the real world, such restrictions may be impractical and involve potential privacy leaks. In this paper, we propose an efficient and novel approach, called Disentangled Prompt Tuning (DiPrompT), a method that tackles the above restrictions by learning adaptive prompts for domain generalization in a distributed manner. Specifically, we first design two types of prompts, i.e., global prompt to capture general knowledge across all clients and domain prompts to capture domain-specific knowledge. They eliminate the restriction on the one-to-one mapping between source domains and local clients. Furthermore, a dynamic query metric is introduced to automatically search the suitable domain label for each sample, which includes two-substep text-image alignments based on prompt tuning without labor-intensive annotation. Extensive experiments on multiple datasets demonstrate that our DiPrompT achieves superior domain generalization performance over state-of-the-art FL methods when domain labels are not provided, and even outperforms many centralized learning methods using domain labels.
Sikai Bai, Jie Zhang 0076, Song Guo 0001, Jingcai Guo, Tao Han 0002, Xiaocheng Lu
CVPR7
2024 PSASlicing: Perpetual SLA-Aware Reinforcement Learning for O-RAN Slice Management
abstract
Network slicing has been widely recognized as one of the flagship use cases for Open Radio Access Network (O-RAN), enabling the provisioning of isolated network services over a shared physical infrastructure. Each slice is characterized by a set of distinct service level agreements (SLAs) tailored to meet the needs of various industries and applications. At the same time, industry-critical applications often require strict adherence to the SLA even in the worst-case scenarios. However, existing network slicing strategies merely incorporate SLA violations as penalties within the reward function, thus failing to consistently ensure perpetual SLA compliance. To address these challenges, this paper introduces PSASlicing, an intelligent resource allocation system designed for RAN slice management across the access network. More specifically, PSASlicing introduces a new reinforcement learning algorithm for maximizing resource utilization while perpetually guaranteeing the diverse SLA requirements across slices. Furthermore, PSASlicing also incorporates a trace-driven network emulator that effectively replicates the dynamic behavior of cellular networks by integrating a transition model with real-world data from an over-the-air 5G Standalone testbed. A comprehensive experimental evaluation showcases that PSASlicing achieves an average resource savings of approximately 24.0% when compared to the state-of-the-art, while guaranteeing no SLA violations.
Mingrui Yin, Ahan Kak, Nakjung Choi, Tao Han 0002
GLOBECOM5
2024 LeFi: Learn to Incentivize Federated Learning in Automotive Edge Computing
abstract
Federated learning (FL) is the promising privacy-preserve approach to continually update the central machine learning (ML) model (e.g., object detectors in edge servers) by aggregating the gradients obtained from local observation data in distributed connected and automated vehicles (CAVs). The incentive mechanism is to incentivize individual selfish CAVs to participate in FL towards the improvement of overall model accuracy. It is, however, challenging to design the incentive mechanism, due to the complex correlation between the overall model accuracy and unknown incentive sensitivity of CAVs, especially under the non-independent and identically distributed (Non-IID) data of individual CAVs. In this paper, we propose a new learn-to-incentivize algorithm to adaptively allocate rewards to individual CAVs under unknown sensitivity functions. First, we gradually learn the unknown sensitivity function of individual CAVs with accumulative observations, by using compute-efficient Gaussian process regression (GPR). Second, we iteratively update the reward allocation to individual CAVs with new sampled gradients, derived from GPR. Third, we project the updated reward allocations to comply with the total budget. We evaluate the performance of extensive simulations, where the simulation parameters are obtained from realistic profiling of the CIFAR-10 dataset and NVIDIA RTX 3080 GPU. The results show that our proposed algorithm substantially outperforms existing solutions, in terms of accuracy, scalability, and adaptability.
Qiang Liu 0013, Tao Han 0002
GLOBECOM4
2024 Enhancing E-Health: Integrating Mixed Reality and IoT for Real-Time Health Monitoring
abstract
Despite advancements in wearable and IoT technologies, real-time systems integrating these with mixed reality (MR) for health monitoring remain underexplored. This paper presents a novel framework designed to enhance proactive health management through an integrated IoT and MR infrastructure. The system employs a three-tiered architecture leveraging intelligent sensors and wearable technologies to provide immersive, contextualized interactions with health data. This integration enables prompt, informed decisions during critical health situations, addressing challenges in latency, data throughput, and user engagement. Targeted at healthcare applications, the framework ensures high reliability and responsiveness. Evaluations demonstrate the system's efficacy in providing accurate, timely feedback, significantly enhancing user outcomes.
Pedro H. Regalado, Tao Han 0002
HealthCom2
2024 Towards a Self-contained Data-driven Global Weather Forecasting Framework
abstract
Data-driven weather forecasting models are advancing rapidly, yet they rely on initial states (i.e., analysis states) typically produced by traditional data assimilation algorithms. Four-dimensional variational assimilation (4DVar) is one of the most widely adopted data assimilation algorithms in numerical weather prediction centers; it is accurate but computationally expensive. In this paper, we aim to couple the AI forecasting model, FengWu, with 4DVar to build a self-contained data-driven global weather forecasting framework, FengWu-4DVar. To achieve this, we propose an *AI-embedded* 4DVar algorithm that includes three components: (1) a 4DVar objective function embedded with the FengWu forecasting model and its error representation to enhance efficiency and accuracy; (2) a spherical-harmonic-transform-based (SHT-based) approximation strategy for capturing the horizontal correlation of background error; and (3) an auto-differentiation (AD) scheme for determining the optimal analysis fields. Experimental results show that under the ERA5 simulated observational data with varying proportions and noise levels, FengWu-4DVar can generate accurate analysis fields; remarkably, it has achieved stable self-contained global weather forecasts for an entire year for the first time, demonstrating its potential for real-world applications. Additionally, our framework is approximately 100 times faster than the traditional 4DVar algorithm under similar experimental conditions, highlighting its significant computational efficiency.
Lei Bai 0001, Wei Xue 0003, Hao Chen 0045, Kun Chen 0004, Tao Han 0002, Wanli Ouyang
ICML7
2024 Health-MR: A Mixed Reality-Based Patient Registration and Monitor Medical System
abstract
In medical procedures such as surgery and diagnosis, it is crucial to provide doctors and nurses with up-to-date patient information. In this paper, we propose Health-MR, a portable Mixed-Reality (MR) system that helps medical staff monitor patient conditions. Health-MR consists of three components: (1) Patient Identification Recognition using face detection, (2) Medical Cloud Database for patient information retrieval, and (3) Non-invasive Heart Rate Measurement via image processing and Fast Fourier Transform (FFT). Our evaluation demonstrates that Health-MR significantly reduces the time needed to query patient information and provides remote, accurate, and real-time heart rate monitoring.
Mingrui Yin, Sohom Sen, Yongjie Guan, Xueyu Hou, Tao Han 0002
MobiCom6
2024 FNP: Fourier Neural Processes for Arbitrary-Resolution Data Assimilation
abstract
Data assimilation is a vital component in modern global medium-range weather forecasting systems to obtain the best estimation of the atmospheric state by combining the short-term forecast and observations. Recently, AI-based data assimilation approaches have attracted increasing attention for their significant advantages over traditional techniques in terms of computational consumption. However, existing AI-based data assimilation methods can only handle observations with a specific resolution, lacking the compatibility and generalization ability to assimilate observations with other resolutions. Considering that complex real-world observations often have different resolutions, we propose the Fourier Neural Processes (FNP) for arbitrary-resolution data assimilation in this paper. Leveraging the efficiency of the designed modules and flexible structure of neural processes, FNP achieves state-of-the-art results in assimilating observations with varying resolutions, and also exhibits increasing advantages over the counterparts as the resolution and the amount of observations increase. Moreover, our FNP trained on a fixed resolution can directly handle the assimilation of observations with out-of-distribution resolutions and the observational information reconstruction task without additional fine-tuning, demonstrating its excellent generalization ability across data resolutions as well as across tasks. Code is available at https://github.com/OpenEarthLab/FNP.
Kun Chen 0004, Peng Ye 0006, Hao Chen 0045, Tao Han 0002, Wanli Ouyang, Tao Chen 0003, Lei Bai 0001
NeurIPS5
2024 Generalizing Weather Forecast to Fine-grained Temporal Scales via Physics-AI Hybrid Modeling
abstract
Data-driven artificial intelligence (AI) models have made significant advancements in weather forecasting, particularly in medium-range and nowcasting. However, most data-driven weather forecasting models are black-box systems that focus on learning data mapping rather than fine-grained physical evolution in the time dimension. Consequently, the limitations in the temporal scale of datasets prevent these models from forecasting at finer time scales. This paper proposes a physics-AI hybrid model (i.e., WeatherGFT) which generalizes weather forecasts to finer-grained temporal scales beyond training dataset. Specifically, we employ a carefully designed PDE kernel to simulate physical evolution on a small time scale (e.g., 300 seconds) and use a parallel neural networks with a learnable router for bias correction. Furthermore, we introduce a lead time-aware training framework to promote the generalization of the model at different lead times. The weight analysis of physics-AI modules indicates that physics conducts major evolution while AI performs corrections adaptively. Extensive experiments show that WeatherGFT trained on an hourly dataset, effectively generalizes forecasts across multiple time scales, including 30-minute, which is even smaller than the dataset's temporal resolution.
Wanghan Xu, Fenghua Ling, Tao Han 0002, Hao Chen 0045, Wanli Ouyang, Lei Bai 0001
NeurIPS4
2024 Learning-Based Query Scheduling and Resource Allocation for Low-Latency Mobile-Edge Video Analytics
abstract
Mobile-edge computing can help enable low-latency and accurate video analytics. However, it is difficult to make efficient utilization of limited edge resources because of the diverse requirements of video queries. In this article, we investigate edge coordination for resource-efficient video query processing, in order to accommodate real-time queries on end cameras, edge nodes, or the cloud, with accuracy guarantee. This problem is challenging because: 1) video queries are with unpredictable arrivals and different resource demands; 2) the decision space of both query scheduling and resource allocation varies over time; and 3) it is critical to maintain long-term accurate analytics for all arrived queries. This problem boils down to making scheduling and resource allocation decisions, which is formulated as a mixed-integer nonlinear programming with a long-term accuracy constraint. Observing that both the scheduling and resource allocation of each query have the Markovian property, the Markov decision process and Lyapunov optimization are adopted to decompose the problem into sequential subproblems. An adaptive reinforcement learning-based approach relying on edge coordination is proposed. Extensive experimental results show that our proposal outperforms other benchmarks on latency and accuracy at a higher level of resource utilization efficiency in real-world data sets.
Peng Yang 0004, Wen Wu 0003, Ning Zhang 0007, Tao Han 0002, Li Yu 0003
IEEE Internet Things J.5
2024 BPS: Batching, Pipelining, Surgeon of Continuous Deep Inference on Collaborative Edge Intelligence
abstract
Users on edge generate deep inference requests continuously over time. Mobile/edge devices located near users can undertake the computation of inference locally for users, e.g., the embedded edge device on an autonomous vehicle. Due to limited computing resources on one mobile/edge device, it may be challenging to process the inference requests from users with high throughput. An attractive solution is to (partially) offload the computation to a remote device in the network. In this paper, we examine the existing inference execution solutions across local and remote devices and propose an adaptive scheduler, a BPS scheduler, for continuous deep inference on collaborative edge intelligence. By leveraging data parallel, neurosurgeon, reinforcement learning techniques, BPS can boost the overall inference performance by up to$8.2 \times$over the baseline schedulers. A lightweight compressor, FF, specialized in compressing intermediate output data for neurosurgeon, is proposed and integrated into the BPS scheduler. FF exploits the operating character of convolutional layers and utilizes efficient approximation algorithms. Compared to existing compression methods, FF achieves up to 86.9% lower accuracy loss and up to 83.6% lower latency overhead.
Xueyu Hou, Yongjie Guan, Nakjung Choi, Tao Han 0002
IEEE Trans. Cloud Comput.4
2024 Traffic Sign Interpretation via Natural Language Description
abstract
Most existing traffic sign-related works are dedicated to detecting and recognizing part of traffic signs separately, which fails to analyze the global semantic logic among signs and may convey inaccurate traffic instruction information. Following the above issues, we propose a traffic sign interpretation (TSI) task, which aims to interpret global semantic interrelated traffic signs (e.g., driving instruction-related texts, symbols, and guide panels) into a natural language for providing complete traffic instruction support to autonomous or assistant driving. Meanwhile, considering the lack of an effective framework for the proposed TSI task in existing works, we design a multi-task learning architecture (TSI-arch) to detect and recognize various traffic signs with drastic changes in sizes and aspect ratios. Meanwhile interpreting these signs into a natural language like a human according to Chinese design criteria of road traffic signs. Furthermore, the absence of a public TSI available dataset prompts us to build a traffic sign interpretation dataset, namely TSI-CN. The dataset consists of real road scene images, which are captured from the highway and the urban way in China from a driver’s perspective. It contains rich location labels of texts, symbols, and guide panels, and the corresponding natural language description labels. Experiments on our TSI-CN dataset demonstrate that the TSI task is achievable and the TSI architecture can interpret traffic signs from scenes successfully even if there is a complicated semantic logic among signs.
Chuang Yang 0003, Kai Zhuang, Mulin Chen, Haozhao Ma, Xu Han 0019, Tao Han 0002, Changxing Guo, Bingxuan Zhao, Qi Wang 0009
IEEE Trans. Intell. Transp. Syst.6
2023 RoNet: Toward Robust Neural Assisted Mobile Network Configuration
abstract
Automating configuration is the key path to achieving zero-touch network management in ever-complicating mobile networks. Deep learning techniques show great potential to automatically learn and tackle high-dimensional networking problems. The vulnerability of deep learning to deviated input space, however, raises increasing deployment concerns under unpredictable variabilities and simulation-to-reality discrepancy in real-world networks. In this paper, we propose a novel RoNet framework to improve the robustness of neural-assisted configuration policies. We formulate the network configuration problem to maximize performance efficiency when serving diverse user applications. We design three integrated stages with novel normal training, learn-to-attack, and robust defense method for balancing the robustness and performance of policies. We evaluate RoNet via the NS-3 simulator extensively and the simulation results show that RoNet outperforms existing solutions in terms of robustness, adaptability, and scalability.
Yongjie Xue, Qiang Liu 0013, Nakjung Choi, Tao Han 0002
ICC5
2023 STEERER: Resolving Scale Variations for Counting and Localization via Selective Inheritance Learning
abstract
Scale variation is a deep-rooted problem in object counting, which has not been effectively addressed by existing scale-aware algorithms. An important factor is that they typically involve cooperative learning across multi-resolutions, which could be suboptimal for learning the most discriminative features from each scale. In this paper, we propose a novel method termed STEERER (SelecTivE inhERitance lEaRning) that addresses the issue of scale variations in object counting. STEERER selects the most suitable scale for patch objects to boost feature extraction and only inherits discriminative features from lower to higher resolution progressively. The main insights of STEERER are a dedicated Feature Selection and Inheritance Adaptor (FSIA), which selectively forwards scale-customized features at each scale, and a Masked Selection and Inheritance Loss (MSIL) that helps to achieve high-quality density maps across all scales. Our experimental results on nine datasets with counting and localization tasks demonstrate the unprecedented scale generalization ability of STEERER. Code is available at https://github.com/taohan10200/STEERER.
Tao Han 0002, Lei Bai 0001, Lingbo Liu, Wanli Ouyang
ICCV1
2023 Dystri: A Dynamic Inference based Distributed DNN Service Framework on Edge
abstract
Deep neural network (DNN) inference poses unique challenges in serving computational requests due to high request intensity, concurrent multi-user scenarios, and diverse heterogeneous service types. Simultaneously, mobile and edge devices provide users with enhanced computational capabilities, enabling them to utilize local resources for deep inference processing. Moreover, dynamic inference techniques allow content-based computational cost selection per request. This paper presents Dystri, an innovative framework devised to facilitate dynamic inference on distributed edge infrastructure, thereby accommodating multiple heterogeneous users. Dystri offers a broad applicability in practical environments, encompassing heterogeneous device types, DNN-based applications, and dynamic inference techniques, surpassing the state-of-the-art (SOTA) approaches. With distributed controllers and a global coordinator, Dystri allows per-request, per-user adjustments of quality-of-service, ensuring instantaneous, flexible, and discrete control. The decoupled workflows in Dystri naturally support user heterogeneity and scalability, addressing crucial aspects overlooked by existing SOTA works. Our evaluation involves three multi-user, heterogeneous DNN inference service platforms deployed on distributed edge infrastructure, encompassing seven DNN applications. Results show Dystri achieves near-zero deadline misses and excels in adapting to varying user numbers and request intensities. Dystri outperforms baselines with accuracy improvement up to 95 ×.
Xueyu Hou, Yongjie Guan, Tao Han 0002
ICPP3
2023 MetaStream: Live Volumetric Content Capture, Creation, Delivery, and Rendering in Real Time
abstract
While recent work explored streaming volumetric content on-demand, there is little effort on live volumetric video streaming that bears the potential of bringing more exciting applications than its on-demand counterpart. To fill this critical gap, in this paper, we propose MetaStream, which is, to the best of our knowledge, the first practical live volumetric content capture, creation, delivery, and rendering system for immersive applications such as virtual, augmented, and mixed reality. To address the key challenge of the stringent latency requirement for processing and streaming a huge amount of 3D data, MetaStream integrates several innovations into a holistic system, including dynamic camera calibration, edge-assisted object segmentation, cross-camera redundant point removal, and foveated volumetric content rendering. We implement a prototype of MetaStream using commodity devices and extensively evaluate its performance. Our results demonstrate that MetaStream achieves low-latency live volumetric video streaming at close to 30 frames per second on WiFi networks. Compared to state-of-the-art systems, MetaStream reduces end-to-end latency by up to 31.7% while improving visual quality by up to 12.5%.
Yongjie Guan, Xueyu Hou, Nan Wu 0012, Bo Han 0001, Tao Han 0002
MobiCom5
2023 MutualNet: Adaptive ConvNet via Mutual Learning From Different Model Configurations
abstract
Most existing deep neural networks are static, which means they can only perform inference at a fixed complexity. But the resource budget can vary substantially across different devices. Even on a single device, the affordable budget can change with different scenarios, and repeatedly training networks for each required budget would be incredibly expensive. Therefore, in this work, we propose a general method called MutualNet to train a single network that can run at a diverse set of resource constraints. Our method trains a cohort of model configurations with various network widths and input resolutions. This mutual learning scheme not only allows the model to run at different width-resolution configurations but also transfers the unique knowledge among these configurations, helping the model to learn stronger representations overall. MutualNet is a general training methodology that can be applied to various network structures (e.g., 2D networks: MobileNets, ResNet, 3D networks: SlowFast, X3D) and various tasks (e.g., image classification, object detection, segmentation, and action recognition), and is demonstrated to achieve consistent improvements on a variety of datasets. Since we only train the model once, it also greatly reduces the training cost compared to independently training several models. Surprisingly, MutualNet can also be used to significantly boost the performance of a single network, if dynamic resource constraints are not a concern. In summary, MutualNet is a unified method for both static and adaptive, 2D and 3D networks. Code and pre-trained models are available at https://github.com/taoyang1122/MutualNet.
Taojiannan Yang, Sijie Zhu, Matías Mendieta, Pu Wang 0001, Ravikumar Balakrishnan, Minwoo Lee 0001, Tao Han 0002, Mubarak Shah, Chen Chen 0001
IEEE Trans. Pattern Anal. Mach. Intell.7
2023 FedSTN: Graph Representation Driven Federated Learning for Edge Computing Enabled Urban Traffic Flow Prediction
abstract
Predicting traffic flow plays an important role in reducing traffic congestion and improving transportation efficiency for smart cities. Traffic Flow Prediction (TFP) in the smart city requires efficient models, highly reliable networks, and data privacy. As traffic data, traffic trajectory can be transformed into a graph representation, so as to mine the spatio-temporal information of the graph for TFP. However, most existing work adopt a central training mode where the privacy problem brought by the distributed traffic data is not considered. In this paper, we propose a Federated Deep Learning based on the Spatial-Temporal Long and Short-Term Networks (FedSTN) algorithm to predict traffic flow by utilizing observed historical traffic data. In FedSTN, each local TFP model deployed in an edge computing server includes three main components, namely Recurrent Long-term Capture Network (RLCN) module, Attentive Mechanism Federated Network (AMFN) module, and Semantic Capture Network (SCN) module. RLCN can capture the long-term spatial-temporal information in each area. AMFN shares short-term spatio-temporal hidden information when it trains its local TFP model by the additive homomorphic encryption approach based on Vertical Federated Learning (VFL). We employ SCN to capture semantic features such as irregular non-Euclidean connections and Point of Interest (POI). Compared with existing baselines, several simulations are conducted on practical data sets and the results prove the effectiveness of our algorithm.
Xiaoming Yuan 0002, Ning Zhang 0007, Tingting Yang 0001, Tao Han 0002, Amirhosein Taherkordi
IEEE Trans. Intell. Transp. Syst.6
2023 Domain-Adaptive Crowd Counting via High-Quality Image Translation and Density Reconstruction
abstract
Recently, crowd counting using supervised learning achieves a remarkable improvement. Nevertheless, most counters rely on a large amount of manually labeled data. With the release of synthetic crowd data, a potential alternative is transferring knowledge from them to real data without any manual label. However, there is no method to effectively suppress domain gaps and output elaborate density maps during the transferring. To remedy the above problems, this article proposes a domain-adaptive crowd counting (DACC) framework, which consists of a high-quality image translation and density map reconstruction. To be specific, the former focuses on translating synthetic data to realistic images, which prompts the translation quality by segregating domain-shared/independent features and designing content-aware consistency loss. The latter aims at generating pseudo labels on real scenes to improve the prediction quality. Next, we retrain a final counter using these pseudo labels. Adaptation experiments on six real-world datasets demonstrate that the proposed method outperforms the state-of-the-art methods.
Junyu Gao 0001, Tao Han 0002, Yuan Yuan 0001, Qi Wang 0009
IEEE Trans. Neural Networks Learn. Syst.2
2022 Atlas: automate online service configuration in network slicing
abstract
Network slicing achieves cost-efficient slice customization to support heterogeneous applications and services. Configuring cross-domain resources to end-to-end slices based on service-level agreements, however, is challenging, due to the complicated underlying correlations and the simulation-to-reality discrepancy between simulators and real networks. In this paper, we propose Atlas, an online network slicing system, which automates the service configuration of slices via safe and sample-efficient learn-to-configure approaches in three interrelated stages. First, we design a learning-based simulator to reduce the sim-to-real discrepancy, which is accomplished by a new parameter searching method based on Bayesian optimization. Second, we offline train the policy in the augmented simulator via a novel offline algorithm with a Bayesian neural network and parallel Thompson sampling. Third, we online learn the policy in real networks with a novel online algorithm with safe exploration and Gaussian process regression. We implement Atlas on an end-to-end network prototype based on OpenAirInterface RAN, OpenDayLight SDN transport, OpenAir-CN core network, and Docker-based edge server. Experimental results show that, compared to state-of-the-art solutions, Atlas achieves 63.9% and 85.7% regret reduction on resource usage and slice quality of experience during the online learning stage, respectively.
Qiang Liu 0013, Nakjung Choi, Tao Han 0002
CoNEXT3
2022 DR.VIC: Decomposition and Reasoning for Video Individual Counting
abstract
Pedestrian counting is a fundamental tool for under-standing pedestrian patterns and crowd flow analysis. Existing works (e.g., image-level pedestrian counting, cross-line crowd counting et al.) either only focus on the image-level counting or are constrained to the manual annotation of lines. In this work, we propose to conduct the pedes-trian counting from a new perspective - Video Individual Counting (VIC), which counts the total number of individual pedestrians in the given video (a person is only counted once). Instead of relying on the Multiple Object Tracking (MOT) techniques, we propose to solve the problem by decomposing all pedestrians into the initial pedestrians who existed in the first frame and the new pedestrians with separate identities in each following frame. Then, an end-to-end Decomposition and Reasoning Network (DRNet) is designed to predict the initial pedestrian count with the density estimation method and reason the new pedestrian's count of each frame with the differentiable optimal transport. Extensive experiments are conducted on two datasets with congested pedestrians and diverse scenes, demonstrating the effectiveness of our method over baselines with great superiority in counting the individual pedestrians. Code: https://github.com/taohan10200/DRNet.
Tao Han 0002, Lei Bai 0001, Junyu Gao 0001, Qi Wang 0009, Wanli Ouyang
CVPR1
2022 DistrEdge: Speeding up Convolutional Neural Network Inference on Distributed Edge Devices
abstract
As the number of edge devices with computing resources (e.g., embedded GPUs, mobile phones, and laptops) in-creases, recent studies demonstrate that it can be beneficial to col-laboratively run convolutional neural network (CNN) inference on more than one edge device. However, these studies make strong assumptions on the devices' conditions, and their application is far from practical. In this work, we propose a general method, called DistrEdge, to provide CNN inference distribution strategies in environments with multiple IoT edge devices. By addressing heterogeneity in devices, network conditions, and nonlinear characters of CNN computation, DistrEdge is adaptive to a wide range of cases (e.g., with different network conditions, various device types) using deep reinforcement learning technology. We utilize the latest embedded AI computing devices (e.g., NVIDIA Jetson products) to construct cases of heterogeneous devices' types in the experiment. Based on our evaluations, DistrEdge can properly adjust the distribution strategy according to the devices' computing characters and the network conditions. It achieves 1.1 to 3 x speedup compared to state-of-the-art methods.
Xueyu Hou, Yongjie Guan, Tao Han 0002, Ning Zhang 0007
IPDPS3
2022 NeuLens: spatial-based dynamic acceleration of convolutional neural networks on edge
abstract
Convolutional neural networks (CNNs) play an important role in today's mobile and edge computing systems for vision-based tasks like object classification and detection. However, state-of-the-art methods on CNN acceleration are trapped in either limited practical latency speed-up on general computing platforms or latency speed-up with severe accuracy loss. In this paper, we propose a spatial-based dynamic CNN acceleration framework, NeuLens, for mobile and edge platforms. Specially, we design a novel dynamic inference mechanism, assemble region-aware convolution (ARAC) supernet, that peels off redundant operations inside CNN models as many as possible based on spatial redundancy and channel slicing. In ARAC supernet, the CNN inference flow is split into multiple independent micro-flows, and the computational cost of each can be autonomously adjusted based on its tiled-input content and application requirements. These micro-flows can be loaded into hardware like GPUs as single models. Consequently, its operation reduction can be well translated into latency speed-up and is compatible with hardware-level accelerations. Moreover, the inference accuracy can be well preserved by identifying critical regions on images and processing them in the original resolution with large micro-flow. Based on our evaluation, NeuLens outperforms baseline methods by up to 58% latency reduction with the same accuracy and by up to 67.9% accuracy improvement under the same latency/memory constraints.
Xueyu Hou, Yongjie Guan, Tao Han 0002
MobiCom3
2022 DeepMix: mobility-aware, lightweight, and hybrid 3D object detection for headsets
abstract
Mobile headsets should be capable of understanding 3D physical environments to offer a truly immersive experience for augmented/mixed reality (AR/MR). However, their small form-factor and limited computation resources make it extremely challenging to execute in real-time 3D vision algorithms, which are known to be more compute-intensive than their 2D counterparts. In this paper, we propose DeepMix, a mobility-aware, lightweight, and hybrid 3D object detection framework for improving the user experience of AR/MR on mobile headsets. Motivated by our analysis and evaluation of state-of-the-art 3D object detection models, DeepMix intelligently combines edge-assisted 2D object detection and novel, on-device 3D bounding box estimations that leverage depth data captured by headsets. This leads to low end-to-end latency and significantly boosts detection accuracy in mobile scenarios. A unique feature of DeepMix is that it fully exploits the mobility of headsets to fine-tune detection results and boost detection accuracy. To the best of our knowledge, DeepMix is the first 3D object detection that achieves 30 FPS (i.e., an end-to-end latency much lower than the 100 ms stringent requirement of interactive AR/MR). We implement a prototype of DeepMix on Microsoft HoloLens and evaluate its performance via both extensive controlled experiments and a user study with 30+ participants. DeepMix not only improves detection accuracy by 9.1--37.3% but also reduces end-to-end latency by 2.68--9.15×, compared to the baseline that uses existing 3D object detection models.
Yongjie Guan, Xueyu Hou, Nan Wu 0012, Bo Han 0001, Tao Han 0002
MobiSys5
2022 Guest Editorial Special Issue on Space-Air-Ground Integrated Networks for Intelligent Transportation Systems
abstract
Next-generation intelligent transportation systems (ITS) are envisioned to greatly improve transportation safety and efficiency by incorporating wireless communications and informatics technologies into transportation systems. As the cornerstone for ITS, vehicular communication networks enable vehicles on the go to exchange information with other vehicles and the external environments, which expect to play a significant role in supporting a variety of services such as road safety, traffic management, and infotainment. However, the existing terrestrial networks including dedicated shortrange communications (DSRC)-based networks and cellular networks alone cannot serve the vehicular applications very well in different scenarios, due to the inherent issues of deployment, coverage, and capacity. It is imperative to exploit other communication infrastructures, such as low-earth orbit (LEO) satellites, unmanned aerial vehicles (UAVs), and high-altitude platforms, to support vehicular applications better, resulting in space-air-ground integrated networks (SAGIN). SAGIN can provide more comprehensive and three-dimensional network connectivity for moving vehicles, anywhere and anytime, by exploiting their respective advantages in terms of coverage, flexibility, reliability, and availability.
Ning Zhang 0007, Tao Han 0002, Mehrdad Dianati, Ning Lu 0001, Shangguang Wang
IEEE Trans. Intell. Transp. Syst.2
2022 Neuron Linear Transformation: Modeling the Domain Shift for Crowd Counting
abstract
Cross-domain crowd counting (CDCC) is a hot topic due to its importance in public safety. The purpose of CDCC is to alleviate the domain shift between the source and target domain. Recently, typical methods attempt to extract domain-invariant features via image translation and adversarial learning. When it comes to specific tasks, we find that the domain shifts are reflected in model parameters' differences. To describe the domain gap directly at the parameter level, we propose a neuron linear transformation (NLT) method, exploiting domain factor and bias weights to learn the domain shift. Specifically, for a specific neuron of a source model, NLT exploits few labeled target data to learn domain shift parameters. Finally, the target neuron is generated via a linear transformation. Extensive experiments and analysis on six real-world data sets validate that NLT achieves top performance compared with other domain adaptation methods. An ablation study also shows that the NLT is robust and more effective than supervised and fine-tune training. Code is available at https://github.com/taohan10200/NLT.
Qi Wang 0009, Tao Han 0002, Junyu Gao 0001, Yuan Yuan 0001
IEEE Trans. Neural Networks Learn. Syst.2
2022 Multitask Attention Network for Lane Detection and Fitting
abstract
Many CNN-based segmentation methods have been applied in lane marking detection recently and gain excellent success for a strong ability in modeling semantic information. Although the accuracy of lane line prediction is getting better and better, lane markings' localization ability is relatively weak, especially when the lane marking point is remote. Traditional lane detection methods usually utilize highly specialized handcrafted features and carefully designed postprocessing to detect the lanes. However, these methods are based on strong assumptions and, thus, are prone to scalability. In this work, we propose a novel multitask method that: 1) integrates the ability to model semantic information of CNN and the strong localization ability provided by handcrafted features and 2) predicts the position of vanishing line. A novel lane fitting method based on vanishing line prediction is also proposed for sharp curves and nonflat road in this article. By integrating segmentation, specialized handcrafted features, and fitting, the accuracy of location and the convergence speed of networks are improved. Extensive experimental results on four-lane marking detection data sets show that our method achieves state-of-the-art performance.
Qi Wang 0009, Tao Han 0002, Zequn Qin, Junyu Gao 0001, Xuelong Li 0001
IEEE Trans. Neural Networks Learn. Syst.2
2022 Towards Revenue-Driven Multi-User Online Task Offloading in Edge Computing
abstract
Mobile Edge Computing (MEC) has become an attractive solution to enhance the computing and storage capacity of mobile devices by leveraging available resources on edge nodes. In MEC, the arrivals of tasks are highly dynamic and are hard to predict precisely. It is of great importance yet very challenging to assign the tasks to edge nodes with guaranteed system performance. In this article, we aim to optimize the revenue earned by each edge node by optimally offloading tasks to the edge nodes. We formulate the revenue-driven online task offloading (ROTO) problem, which is proved to be NP-hard. We first relax ROTO to a linear fractional programming problem, for which we propose the Level Balanced Allocation (LBA) algorithm. We then show the performance guarantee of LBA through rigorous theoretical analysis, and present the LB-Rounding algorithm for ROTO using the primal-dual technique. The algorithm achieves an approximation ratio of$2(1+\xi)\ln (d+1)$with a considerable probability, where$d$is the maximum number of process slots of an edge node and$\xi$is a small constant. The performance of the proposed algorithm is validated through both trace-driven simulations and testbed experiments. Results show that our proposed scheme is more efficient compared to baseline algorithms.
Zhi Ma 0002, Sheng Zhang 0001, Tao Han 0002, Zhuzhong Qian, Mingjun Xiao, Ning Chen 0010, Jie Wu 0001, Sanglu Lu
IEEE Trans. Parallel Distributed Syst.4
2021 OnSlicing: online end-to-end network slicing with reinforcement learning
abstract
Network slicing allows mobile network operators to virtualize infrastructures and provide customized slices for supporting various use cases with heterogeneous requirements. Online deep reinforcement learning (DRL) has shown promising potential in solving network problems and eliminating the simulation-to-reality discrepancy. Optimizing cross-domain resources with online DRL is, however, challenging, as the random exploration of DRL violates the service level agreement (SLA) of slices and resource constraints of infrastructures. In this paper, we propose OnSlicing, an online end-to-end network slicing system, to achieve minimal resource usage while satisfying slices' SLA. OnSlicing allows individualized learning for each slice and maintains its SLA by using a novel constraint-aware policy update method and proactive baseline switching mechanism. OnSlicing complies with resource constraints of infrastructures by using a unique design of action modification in slices and parameter coordination in infrastructures. OnSlicing further mitigates the poor performance of online learning during the early learning stage by offline imitating a rule-based solution. Besides, we design four new domain managers to enable dynamic resource configuration in radio access, transport, core, and edge networks, respectively, at a timescale of subseconds. We implement OnSlicing on an end-to-end slicing testbed designed based on OpenAirInterface with both 4G LTE and 5G NR, OpenDayLight SDN platform, and OpenAir-CN core network. The experimental results show that OnSlicing achieves 61.3% usage reduction as compared to the rule-based solution and maintains nearly zero violation (0.06%) throughout the online learning phase. As online learning is converged, OnSlicing reduces 12.5% usage without any violations as compared to the state-of-the-art online DRL solution.
Qiang Liu 0013, Nakjung Choi, Tao Han 0002
CoNEXT3
2021 Constraint-Aware Deep Reinforcement Learning for End-to-End Resource Orchestration in Mobile Networks
abstract
Network slicing is a promising technology that allows mobile network operators to efficiently serve various emerging use cases in 5G. It is challenging to optimize the utilization of network infrastructures while guaranteeing the performance of network slices according to service level agreements (SLAs). To solve this problem, we propose SafeSlicing that introduces a new constraint-aware deep reinforcement learning (CaDRL) algorithm to learn the optimal resource orchestration policy within two steps, i.e., offline training in a simulated environment and online learning with the real network system. On optimizing the resource orchestration, we incorporate the constraints on the statistical performance of slices in the reward function using Lagrangian multipliers, and solve the Lagrangian relaxed problem via a policy network. To satisfy the constraints on the system capacity, we design a constraint network to map the latent actions generated from the policy network to the orchestration actions such that the total resources allocated to network slices do not exceed the system capacity. We prototype SafeSlicing on an end-to-end testbed developed by using OpenAirInterface LTE, OpenDayLight-based SDN, and CUDA GPU computing platform. The experimental results show that SafeSlicing reduces more than 20% resource usage while meeting SLAs of network slices as compared with other solutions.
Qiang Liu 0013, Nakjung Choi, Tao Han 0002
ICNP3
2021 LiveMap: Real-Time Dynamic Map in Automotive Edge Computing
abstract
Autonomous driving needs various line-of-sight sensors to perceive surroundings that could be impaired under diverse environment uncertainties such as visual occlusion and extreme weather. To improve driving safety, we explore to wirelessly share perception information among connected vehicles within automotive edge computing networks. Sharing massive perception data in real time, however, is challenging under dynamic networking conditions and varying computation work-loads. In this paper, we propose LiveMap, a real-time dynamic map, that detects, matches, and tracks objects on the road with crowdsourcing data from connected vehicles in sub-second. We develop the data plane of LiveMap that efficiently processes individual vehicle data with object detection, projection, feature extraction, object matching, and effectively integrates objects from multiple vehicles with object combination. We design the control plane of LiveMap that allows adaptive offloading of vehicle computations, and develop an intelligent vehicle scheduling and offloading algorithm to reduce the offloading latency of vehicles based on deep reinforcement learning (DRL) techniques. We implement LiveMap on a small-scale testbed and develop a large-scale network simulator. We evaluate the performance of LiveMap with both experiments and simulations, and the results show LiveMap reduces 34.1% average latency than the baseline solution.
Qiang Liu 0013, Tao Han 0002, Jiang (Linda) Xie, BaekGyu Kim
INFOCOM2
2021 Edge Learning for Low-Latency Video Analytics: Query Scheduling and Resource Allocation
abstract
Low-latency and accuracy-guaranteed video analytics is essential to many delay-sensitive camera-based applications. Analyzing video frames on edge nodes in proximity can effectively reduce the response delay compared with cloud-based solutions. However, the computation and bandwidth resources on an edge node are always limited. In this paper, we design a joint video query scheduling and resource allocation problem based on an edge coordinated architecture, in order to properly accommodate real-time video queries on end cameras, the edge nodes, or the cloud. This problem is challenging in that 1) the arrivals of video queries with different resource requirements are unknown in advance and 2) the design space (of both query scheduling and resource allocation) to provision video queries varies over time. Taking the two-fold uncertainty into consideration, we formulate the query provision problem as a mix integer non-linear program which is NP-hard and not solved directly. To deal with the NP-hardness and the absence of future information, the problem is re-formulated as a Markov decision process, which can leverage historical query information to make decisions about scheduling and resource allocation. The transformed problem calls for an online solution that can efficiently adapt to the dynamic design space. Hence, we propose an edge-coordinated reinforcement learning algorithm to continuously learn from the environment, and make decisions for query scheduling and resource allocation to achieve low latency and accurate video analytics. Extensive simulation results demonstrate the advantages of the proposed algorithm in latency and accuracy.
Peng Yang 0004, Wen Wu 0003, Ning Zhang 0007, Tao Han 0002, Li Yu 0003
MASS5
2021 A survey on sleep mode techniques for ultra-dense networks in 5G and beyond
Fatima Salahdine, Johnson Opadere, Qiang Liu 0013, Tao Han 0002, Ning Zhang 0007, Shaohua Wu 0002
Comput. Networks4
2020 DeepSlicing: Deep Reinforcement Learning Assisted Resource Allocation for Network Slicing
abstract
Network slicing enables multiple virtual networks run on the same physical infrastructure to support various use cases in 5G and beyond. These use cases, however, have very diverse network resource demands, e.g., communication and computation, and various performance metrics such as latency and throughput. To effectively allocate network resources to slices, we propose DeepSlicing that integrates the alternating direction method of multipliers (ADMM) and deep reinforcement learning (DRL). DeepSlicing decomposes the network slicing problem into a master problem and several slave problems. The master problem is solved based on convex optimization and the slave problem is handled by DRL method which learns the optimal resource allocation policy. The performance of the proposed algorithm is validated through network simulations.
Qiang Liu 0013, Tao Han 0002, Ning Zhang 0007, Ye Wang 0002
GLOBECOM2
2020 Focus on Semantic Consistency for Cross-Domain Crowd Understanding
abstract
For pixel-level crowd understanding, it is time-consuming and laborious in data collection and annotation. Some domain adaptation algorithms try to liberate it by training models with synthetic data, and the results in some recent works have proved the feasibility. However, we found that a mass of estimation errors in the background areas impede the performance of the existing methods. In this paper, we propose a domain adaptation method to eliminate it. According to the semantic consistency, a similar distribution in deep layer's features of the synthetic and real-world crowd area, we first introduce a semantic extractor to effectively distinguish crowd and background in high-level semantic information. Besides, to further enhance the adapted model, we adopt adversarial learning to align features in the semantic space. Experiments on three representative real datasets show that the proposed domain adaptation scheme achieves the state-of-the-art for cross-domain counting problems.
Tao Han 0002, Junyu Gao 0001, Yuan Yuan 0001, Qi Wang 0009
ICASSP1
2020 EdgeSlice: Slicing Wireless Edge Computing Network with Decentralized Deep Reinforcement Learning
abstract
5G and edge computing will serve various emerging use cases that have diverse requirements of multiple resources, e.g., radio, transportation, and computing. Network slicing is a promising technology for creating virtual networks that can be customized according to the requirements of different use cases. Provisioning network slices requires end-to-end resource orchestration which is challenging. In this paper, we design a decentralized resource orchestration system named EdgeSlice for dynamic end-to-end network slicing. EdgeSlice introduces a new decentralized deep reinforcement learning (D-DRL) method to efficiently orchestrate end-to-end resources. D-DRL is composed of a performance coordinator and multiple orchestration agents. The performance coordinator manages the resource orchestration policies in all the orchestration agents to ensure the service level agreement (SLA) of network slices. The orchestration agent learns the resource demands of network slices and orchestrates the resource allocation accordingly to optimize the performance of the slices under the constrained networking and computing resources. We design radio, transport and computing manager to enable dynamic configuration of end-to-end resources at runtime. We implement EdgeSlice on a prototype of the end-to-end wireless edge computing network with OpenAirInterface LTE network, OpenDayLight SDN switches, and CUDA GPU platform. The performance of EdgeSlice is evaluated through both experiments and trace-driven simulations. The evaluation results show that EdgeSlice achieves much improvement as compared to baseline in terms of performance, scalability, compatibility.
Qiang Liu 0013, Tao Han 0002, Ephraim Moges
ICDCS2
2020 Unsupervised Semantic Aggregation and Deformable Template Matching for Semi-Supervised Learning
abstract
Unlabeled data learning has attracted considerable attention recently. However, it is still elusive to extract the expected high-level semantic feature with mere unsupervised learning. In the meantime, semi-supervised learning (SSL) demonstrates a promising future in leveraging few samples. In this paper, we combine both to propose an Unsupervised Semantic Aggregation and Deformable Template Matching (USADTM) framework for SSL, which strives to improve the classification performance with few labeled data and then reduce the cost in data annotating. Specifically, unsupervised semantic aggregation based on Triplet Mutual Information (T-MI) loss is explored to generate semantic labels for unlabeled data. Then the semantic labels are aligned to the actual class by the supervision of labeled data. Furthermore, a feature pool that stores the labeled samples is dynamically updated to assign proxy labels for unlabeled data, which are used as targets for cross-entropy minimization. Extensive experiments and analysis across four standard semi-supervised learning benchmarks validate that USADTM achieves top performance (e.g., 90.46% accuracy on CIFAR-10 with 40 labels and 95.20% accuracy with 250 labels). The code is released at https://github.com/taohan10200/USADTM.
Tao Han 0002, Junyu Gao 0001, Yuan Yuan 0001, Qi Wang 0009
NeurIPS1
2020 TrustServing: A Quality Inspection Sampling Approach for Remote DNN Services
abstract
Deep neural networks (DNNs) are being applied to various areas such as computer vision, autonomous vehicles, and healthcare, etc. However, DNNs are notorious for their high computational complexity and cannot be executed efficiently on resource constrained Internet of Things (IoT) devices. Various solutions have been proposed to handle the high computational complexity of DNNs. Offloading computing tasks of DNNs from IoT devices to cloud/edge servers is one of the most popular and promising solutions. While such remote DNN services provided by servers largely reduce computing tasks on IoT devices, it is challenging for IoT devices to inspect whether the quality of the service meets their service level objectives (SLO) or not. In this paper, we address this problem and propose a novel approach named QIS (quality inspection sampling) that can efficiently inspect the quality of the remote DNN services for IoT devices. To realize QIS, we design a new ID-generation method to generate data (IDs) that can identify the serving DNN models on edge servers. QIS inserts the IDs into the input data stream and implements sampling inspection on SLO violations. The experiment results show that the QIS approach can reliably inspect, with a nearly 100% success rate, the service qualtiy of remote DNN services when the SLA level is 99.9% or lower at the cost of only up to 0.5% overhead.
Xueyu Hou, Tao Han 0002
SECON2
2020 Distributed Video Analysis for Mobile Live Broadcasting Services
abstract
While webcast platforms on mobile devices are becoming more and more prevalent, inspection for irregularities is getting harder and harder. To solve this problem, the convolution neural network(CNN) has been applied to recognize or detect specified objections in pictures and videos. However, when supervising large platforms, it isn’t very easy to collect mountain piles of video data and send them to the computation center. Other problems like long time delay and the high computational burden will reduce system performance, especially when dealing with data from live streams. This paper presents a method to coordinate mobile devices with remote servers(computers or embedded systems) to achieve real-time monitoring of live streams. The system can make use of computational capacity on mobile devices and reduce the cost of sending data while guaranteeing accuracy for supervision.
Yuanqi Chen, Yongjie Guan, Tao Han 0002
WCNC3
2020 Guest Editorial: Special Section on Social and Cognitive Mobile Computing in Industrial Internet of Things
abstract
INTERNET of Thing (IoT) technology has attracted intensive interest in the automotive industry to meet the new demands in the market while continuing to achieve their conservative goals [item 1) in the Appendix]. As for Industrial Internet of Things (IIoT), randomly moving wireless nodes are often carried by humans and communicate with each other when they are in close proximity. The interaction between nodes shows strong regularity or sociality, i.e., a wireless node always communicates with several social-closed or distance-closed nodes. This special section collects the latest ideas and research on the social and cognitive mobile computing in IIoT. Particularly, 15 original articles are accepted and included in the collection on the following pages. The topics of these articles are mainly concerned with social and cognitive mobility modeling, routing protocol, resource allocation, and so forth. We believe that these articles will play a role in inspiring our readers. Summaries of accepted articles are provided.
Nan Zhao 0001, Yunfei Chen 0001, Tao Han 0002, F. Richard Yu
IEEE Trans. Ind. Informatics3
2019 A Smart-Decision System for Realtime Mobile AR Applications
abstract
With the development of the hardware and software platforms, we can implement the deep learning model on the mobile device for mobile augmented reality (AR) applications. However, not all mobile AR tasks can be finished on mobile devices. Meanwhile, the limited computation resources on mobile devices are still the main obstacle to achieve realtime mobile AR applications. In this paper, we proposed a smart-decision framework which combines the advantages of the on-device mobile AR system and the edge-based mobile AR system to achieve real-time object recognition. High computation complexity tasks will be offloaded to the edge servers. Low complexity tasks will be executed on mobile devices or the edge server depending on the network latency. To overcome the dynamic changes of network condition and the limitations of the on-device deep learning models, we design a cache and matching algorithm on the mobile devices to enhance the performance of the recognition tasks. With our proposed system, the quality of the mobile AR application is improved. The performance of the smart-decision framework is validated through experiments with a testbed.
Tao Han 0002, Jiang (Linda) Xie
GLOBECOM2
2019 Joint Computation and Communication Resource Allocation for Energy-Efficient Mobile Edge Networks
abstract
In this paper, an ultra-dense mobile edge network is studied, where base stations (BSs) are equipped with computation resources to execute users' offloaded tasks. Although an ultradense BS deployment provides seamless coverage and reduced computation latency of the offloaded tasks, the cost of network power consumption is increased. We formulate an optimization problem to jointly optimize active BSs set, uplink and downlink beamforming vector selection, and computation resource allocation in order to tackle the power consumption and latency tradeoff. To efficiently solve this problem, we propose a sequential solution framework. Specifically, we first select the active BSs based on communication and computation power-aware selection rule. The computation resources and dual-link beamformers are subsequently optimized for the satisfaction of task computation deadline, network energy savings and improved coverage. Simulation results show that the proposed joint optimization framework significantly reduces the network power consumption.
Johnson Opadere, Qiang Liu 0013, Ning Zhang 0007, Tao Han 0002
ICC4
2019 VirtualEdge: Multi-Domain Resource Orchestration and Virtualization in Cellular Edge Computing
abstract
5G network will carry compute-intensive applications from vertical industries. Network slicing and edge computing are key technologies for fulfilling the diverse requirements of these applications efficiently. We define a cellular network with the edge computing capability as cellular edge computing network. Dynamically slicing a cellular edge computing network is challenging because it needs to orchestrate multi-domain resources and ensure isolation among network slices. In this paper, we present the VirtualEdge system that enables the dynamic creation of a virtual node (vNode) on top of a physical cellular edge computing node to serve the traffic and workloads of a network slice. VirtualEdge introduces a realizable multi-domain resource orchestration and virtualization that provides isolation among network slices without losing the efficiency in virtualizing the radio resources. To efficiently orchestrate multi-domain resources, we design a new learning-assisted algorithm that allows the resource orchestrator to optimize the utilization of physical resources without knowing the utility functions of individual vNodes. For the resource virtualization, we develop a heuristic algorithm and a credit-based queue management scheme to dynamically map virtual radio and computing resources to underlying physical resources, respectively. VirtualEdge is developed and implemented based on the OpenAirInterface LTE and CUDA GPU computing platforms, and its performance is validated through both experiments and large-scale simulations.
Qiang Liu 0013, Tao Han 0002
ICDCS2
2019 DIRECT: Distributed Cross-Domain Resource Orchestration in Cellular Edge Computing
abstract
Network slicing and edge computing are key technologies to enable compute-intensive applications for vertical industries in 5G. We define cellular networks with edge computing capabilities as cellular edge computing. In this paper, we study the cross-domain resource orchestration solution for dynamic network slicing in cellular edge computing. The fundamental research challenge is from the difficulty in modeling the relationship between the slice performance and resources from multiple technical domains across the network with many base stations and distributed edge servers. To address this challenge, we develop a distributed cross-domain resource orchestration (DIRECT) protocol which optimizes the cross-domain resource orchestration while providing the performance and functional isolations among network slices. The main component of DIRECT is a distributed cross-domain resource orchestration algorithm which is designed by integrating the ADMM method and a new learning-assisted optimization approach. The proposed resource orchestration algorithm efficiently orchestrates multi-domain resources without requiring the performance model of the network slices. We develop and implement the DIRECT protocol in a small-scale prototype of cellular edge computing which is designed based on OpenAirInterface LTE and CUDA GPU computing platforms. The performance of DIRECT is validated through both experiments and network simulations.
Qiang Liu 0013, Tao Han 0002
MobiHoc2
2019 Online Proactive Caching in Mobile Edge Computing Using Bidirectional Deep Recurrent Neural Network
abstract
With emergence of Internet of Things (IoT), wireless traffic has grown dramatically, posing severe strain on core network and backhaul bandwidth. Proactive caching in mobile edge computing systems can not only efficiently mitigate the traffic congestion and relieve burden of backhaul but also can reduce the service latency for end devices. However, proactive caching heavily relies on the prediction accuracy of content popularity, which is typically unknown and change over time. In this paper, we propose an online proactive caching scheme based on bidirectional deep recurrent neural network (BRNN) model to predict time-series content requests and update edge caching accordingly. Specifically, on the first layer, a 1-D convolution neural network (CNN) is devised to reduce the computational costs. Then, BRNN is employed to predict time-varying requests from users. Afterward, a fully connected neural network (FCNN) is harnessed to learn and sample predicts from the BRNN. Finally, we conduct experiments based on real datasets, which demonstrate that the proposed approach can achieve considerably high prediction accuracy and significantly improve content hit rate of end devices.
Laha Ale, Ning Zhang 0007, Huici Wu, Dajiang Chen, Tao Han 0002
IEEE Internet Things J.5
2018 Joint Radio and Computation Resource Management for Low Latency Mobile Edge Computing
abstract
Mobile edge computing (MEC) is a new networking paradigm that enables low-latency computation offloading for compute-intensive mobile applications. The dynamic wireless channel, non-uniform spatiotemporal traffic, and limited computation resources impair the service latency of mobile edge computing. Therefore, jointly managing radio and computation resources is needed to achieve low latency MEC. In this paper, we propose a joint radio and computation resource management (iRAR) algorithm which minimizes users' service latency by optimizing the uplink transmission power, receive beamforming, computation task assignment, and computation resource allocation. We compare the performance of the proposed algorithm with three different algorithms and demonstrate that the iRAR algorithm reduces up to 52% average service latency as compared to the other algorithms.
Qiang Liu 0013, Tao Han 0002, Nirwan Ansari
GLOBECOM2
2018 Energy-Efficient On-Demand Cloud Radio Access Networks Virtualization
abstract
By leveraging the elasticity of cloud computing, cloud radio access network (C-RAN) facilitates on-demand radio and computing resource provisioning. In this paper, we propose an energy-efficient on-demand C-RAN virtualization model which dynamically provisions virtual C-RAN according to service demand. The energy consumption of the virtual C-RAN is minimized by jointly optimizing the remote radio head (RRH) selection and computing resource provisioning. The network energy consumption minimization problem is challenging because of the interdependence between the RRH selection and the computing resource provisioning. We propose the energy-efficient on-demand C-RAN virtualization (REACT) algorithm to solve the problem in two steps. First, we cluster RRHs into groups using the hierarchical clustering analysis (HCA) algorithm and assign a BBU to each RRH group for the baseband signal processing. Second, we determine the RRH selection by optimizing the cooperative beamforming. The performance of the proposed algorithm is evaluated through extensive simulations, which shows the proposed algorithm reduces up to 62% of the network energy consumption as compared to a baseline algorithm.
Qiang Liu 0013, Tao Han 0002, Nirwan Ansari
GLOBECOM2
2018 A Smart Service Rebuilding Scheme across Cloudlets via Mobile AR Frame Feature Mapping
abstract
Mobile edge computing platforms, such as cloudlets, bring computation resources closer to mobile users, as compared to the cloud, which decreases the end-to-end network latency. This benefit enables a myriad of real-time mobile applications, especially augmented reality (AR), that require low latency and high computation power. However, when mobile users move away from the attached cloudlet, the offloaded services have to be migrated or rebuilt on a new nearby cloudlet. However, this service rebuilding process takes a lot of time and may deteriorate user experience. In this paper, we propose a smart service rebuilding scheme which seamlessly restores the offloading services on the target cloudlet while the mobile user is moving. The service rebuilding process includes the radio handoff stage and service handoff stage. A seamless service rebuilding process is achieved via predicting user's target cloudlet before being triggered a radio handoff, by leveraging extracted features from the captured frames of the mobile user's camera. Furthermore, based on the proposed service rebuilding scheme, we design a feature mapping algorithm to achieve a high prediction precision and a short prediction latency. We implement our scheme on a testbed and conduct experiments using real world AR applications. The experimental results show that our proposed scheme decreases the service rebuilding latency by around 65.8%, as compared to the conventional rebuilding process. In addition, we conduct extensive simulations to evaluate the performance of our proposed feature mapping algorithm. Simulation confirms that our algorithm is robust and can predict users' target cloudlet with high precision and low latency.
Haoxin Wang 0003, Jiang (Linda) Xie, Tao Han 0002
ICC3
2018 DARE: Dynamic Adaptive Mobile Augmented Reality with Edge Computing
abstract
Mobile augmented reality (MAR) is a killer application of mobile edge computing because of its high computation demand and stringent latency requirement. Since edge networks and computing resources are highly dynamic, handling such dynamics is essential for providing high-quality MAR services. In this paper, we design a new network protocol named DARE (dynamic adaptive AR over the edge) that enables mobile users to dynamically change their AR configurations according to wireless channel conditions and computation workloads in edge servers. The dynamic configuration adaptations reduce the service latency of MAR users and maximize the quality of augmentation (QoA) under varying network conditions and computation workloads. Considering the video frame size and computation model, i.e., object detection algorithms, as two key parameters in adapting the AR configuration, we develop analytical models to characterize the impact of these parameters on QoA and the service latency. Then, we design optimization mechanisms on both the edge server and AR devices to guide the AR configuration adaptation and server computation resource allocation. The performance of the DARE protocol is validated through a small-scale testbed implementation.
Qiang Liu 0013, Tao Han 0002
ICNP2
2018 Demo Abstract: Themis: Cross-Domain Resource Orchestration and Virtualization in Cellular Computing Networks
abstract
We demonstrate the Themis protocol and its system implementation that realizes cross-domain resource orchestration and virtualization in cellular computing networks.
Qiang Liu 0013, Tao Han 0002
ICNP2
2018 An Edge Network Orchestrator for Mobile Augmented Reality
abstract
Mobile augmented reality (MAR) involves high complexity computation which cannot be performed efficiently on resource limited mobile devices. The performance of MAR would be significantly improved by offloading the computation tasks to servers deployed with the close proximity to the users. In this paper, we design an edge network orchestrator to enable fast and accurate object analytics at the network edge for MAR. The measurement-based analytical models are built to characterize the tradeoff between the service latency and analytics accuracy in edge-based MAR systems. As a key component of the edge network orchestrator, a server assignment and frame resolution selection algorithm named FACT is proposed to mitigate the latency-accuracy tradeoff. Through network simulations, we evaluate the performance of the FACT algorithm and show the insights on optimizing the performance of edge-based MAR systems. We have implemented the edge network orchestrator and developed the corresponding communication protocol. Our experiments validate the performance of the proposed edge network orchestrator.
Qiang Liu 0013, Johnson Opadere, Tao Han 0002
INFOCOM4
2018 Emerging Technologies for Vehicular Communication Networks
abstract
Next-generation intelligent transportation systems (ITS) are envisioned to greatly improve the transportation safety and efficiency by incorporating wireless communication and informatics technologies in the transportation system [1][2][3].As the cornerstone for ITS, vehicular communication networks enable vehicles to exchange information with other vehicles and the external environments and play a significant role in supporting a variety of services such as road safety, traffic management, and entertainment Vehicular communication networks face many technical challenges such as network scalability, highly dynamic topology, vulnerable wireless links, energy consumption of roadside units, poor network coverage, and bursty traffic.To address these challenges, various emerging technologies have been introduced in vehicular communication networks, such as software defined space-air-ground integrated vehicular network [4], fog computing in vehicular networks [5], droneassisted vehicular networks [6], and machine learning for data delivery [7].This special issue collection aims to present the vision, research, and dedicated efforts on the emerging technologies for vehicular communication networks.In this special issue, there are 15 submissions in total.After peerreview, 6 papers are selected for publication.The first article, "Software-Defined Collaborative Offloading for Heterogeneous Vehicular Networks" by W. Quan et al., proposes a software-defined collaborative offloading (SDCO) solution for heterogeneous vehicular networks, to efficiently manage the offloading nodes and paths.The offloading controller is equipped with two specific functions:
Ning Zhang 0007, Ning Lu 0001, Tao Han 0002, Yi Zhou 0004, Dajiang Chen
Wirel. Commun. Mob. Comput.3
2018 Wireless Caching Aided 5G Networks
abstract
nonPeerReviewed
Nan Zhao 0001, Jun Li 0004, Tao Han 0002, Zheng Chang 0001, Lisheng Fan
Wirel. Commun. Mob. Comput.3
2017 Data-Driven Network Optimization in Ultra-Dense Radio Access Networks
abstract
The complexity of networking mechanisms will increase significantly because of the dense deployment of radio base stations in ultra-dense mobile networks. As a result, the existing networking mechanisms may be unable to efficiently manage ultra-dense mobile networks. To solve this problem, we propose a data driven network optimization framework which integrates the big data analysis methods with networking mechanisms. In the proposed framework, we adopt big data analysis methods to divide densely deployed base stations into groups. Then, each group of base stations are managed with networking mechanisms independently. In this way, the complexity of the networking mechanisms is reduced. The key challenge in designing the framework is to optimally group base stations into clusters in real time. Addressing this challenge, the proposed framework consists of an offline machine learning module and an online base station clustering and network optimization module. The offline machine learning module predicts the optimal number of base station groups in the next time interval based on the historical data. The online base station clustering and network optimization module clusters base stations and optimize the network in real time. The performance of the proposed data-driven network management framework is validated through network simulations with real network data traces.
Qiang Liu 0013, Tao Han 0002, Nirwan Ansari
GLOBECOM3
2017 Energy-Efficient RRH Sleep Mode for Virtual Radio Access Networks
abstract
Network functions virtualization (NFV) has become a strategic tool that facilitates mobile network resource sharing and management by mobile network operators (MNOs). Cloud radio access network (C- RAN) virtualization allows agglomeration of multiple radio access networks (RANs) functions in a single resource pool. In this paper, Virtualized Radio Access Network (VRAN) is utilized as the enabler of inter-operator traffic offloading. To explore the energy saving potential of sleep mode scheme in base stations of cooperating MNOs, we leverage on inter-band non-contiguous carrier aggregation and put forward spectrum sharing into private and shared bands. We formulate an optimization problem to obtain the optimal intra- operator and inter-operator beamforming design for realizing energy-efficient virtual RAN. Inter- operator base station load transfer algorithm is proposed as well as inter-operator BS sleep-mode energy saving algorithm. Simulations results show a significant reduction of total inter-operator power consumption as compared to other algorithms.
Johnson Opadere, Qiang Liu 0013, Tao Han 0002
GLOBECOM3
2017 Big-data-driven network partitioning for ultra-dense radio access networks
abstract
The increased density of base stations (BSs) may significantly add complexity to network management mechanisms and hamper them from efficiently managing the network. In this paper, we propose a big-data-driven network partitioning and optimization framework to reduce the complexity of the networking mechanisms. The proposed framework divides the entire radio access network (RAN) into multiple sub-RANs and each sub-RAN can be managed independently. Therefore, the complexity of the network management can be reduced. Quantifying the relationships among BSs is challenging in the network partitioning. We propose to extract three networking features from mobile traffic data to discover the relationships. Based on these features, we engineer the network partitioning solution in three steps. First, we design a hierarchical clustering analysis (HCA) algorithm to divide the entire RAN into sub-RANs. Second, we implement a traffic load balancing algorithm to characterize the performance of the network partitioning. Third, we adapt the weights of networking features in the HCA algorithm to optimize the network partitioning. We validate the proposed solution through simulations designed based on real mobile network traffic data. The simulation results reveal the impacts of the RAN partitioning on the networking performance and the computational complexity of the networking mechanism.
Tao Han 0002, Nirwan Ansari
ICC2
2017 V-handoff: A practical energy efficient handoff for 802.11 infrastructure networks
abstract
Wireless local area networks (WLANs) are currently among the most important technologies for wireless access. Because of its higher data rate and lower monetary cost compared with cellular networks, mobile users are likely to choose WiFi when they are using mobile applications. However, keeping continuous connectivity with access points (APs) may require frequent handoffs, which may consume much energy in the handoff process. Unfortunately, most of the existing work only focused on reducing the handoff delay of IEEE 802.11-based handoffs and many handoff approaches may even increase the energy consumption of mobile nodes (MNs) in order to reduce the handoff latency. In this paper, we introduce virtual handoff (V-handoff), an energy efficiency-based handoff protocol via generating virtual access points (VAPs) in the corresponding physical access points (PAPs). The main idea of our proposed V-handoff protocol is to create an evenly spaced periodic schedule of beacon periods for all the VAPs in one virtual AP grid. To the best of our knowledge, this is the first paper that investigates the application of the wireless virtualization technique in MN's handoff energy efficiency. Simulation results show that our proposed V-handoff protocol can significantly reduce the MN's handoff energy consumption and the average handoff delay compared with IEEE 802.11-based full scanning and selective scanning handoff protocol.
Haoxin Wang 0003, Jiang (Linda) Xie, Tao Han 0002
ICC3
2017 DRAPS: Dynamic and resource-aware placement scheme for docker containers in a heterogeneous cluster
abstract
Virtualization is a promising technology that has facilitated cloud computing to become the next wave of the Internet revolution. Adopted by data centers, millions of applications that are powered by various virtual machines improve the quality of services. Although virtual machines are well-isolated among each other, they suffer from redundant boot volumes and slow provisioning time. To address limitations, containers were born to deploy and run distributed applications without launching entire virtual machines. As a dominant player, Docker is an open-source implementation of container technology. When managing a cluster of Docker containers, the management tool, Swarmkit, does not take the heterogeneities in both physical nodes and virtualized containers into consideration. The heterogeneity lies in the fact that different nodes in the cluster may have various configurations, concerning resource types and availabilities, etc., and the demands generated by services are varied, such as CPU-intensive (e.g. Clustering services) as well as memory-intensive (e.g. Web services). In this paper, we target on investigating the Docker container cluster and developed, DRAPS, a resource-aware placement scheme to boost the system performance in a heterogeneous cluster.
Ying Mao 0001, Jenna Oak, Anthony Pompili, Daniel Beer, Tao Han 0002, Peizhao Hu
IPCCC5
2017 Network Utility Aware Traffic Load Balancing in Backhaul-Constrained Cache-Enabled Small Cell Networks with Hybrid Power Supplies
abstract
Explosive data traffic growth leads to a continuous surge in capacity demands across mobile networks. In order to provision high network capacity, small cell base stations (SCBSs) are widely deployed. Owing to the close proximity to mobile users, SCBSs can effectively enhance the network capacity and offloading traffic load from macro BSs (MBSs). However, the cost-effective backhaul may not be readily available for SCBSs, thus leading to backhaul constraints in small cell networks (SCNs). Enabling cache in BSs may mitigate the backhaul constraints in SCNs. Moreover, the dense deployment of SCBSs may incur excessive energy consumption. To alleviate brown power consumption, renewable energy will be explored to power BSs. In such a network, it is challenging to dynamically balance traffic load among BSs to optimize the network utilities. In this paper, we investigate the traffic load balancing in backhaul-constrained cache-enabled small cell networks powered by hybrid energy sources. We have proposed a network utility aware (NUA) traffic load balancing scheme that optimizes user association to strike a tradeoff between the green power utilization and the traffic delivery latency. On balancing the traffic load, the proposed NUA traffic load balancing scheme considers the green power utilization, the traffic delivery latency in both BSs and their backhaul, and the cache hit ratio. The NUA traffic load balancing scheme allows dynamically adjusting the tradeoff between the green power utilization and the traffic delivery latency. We have proved the convergence and the optimality of the proposed NUA traffic load balancing scheme. Through extensive simulations, we have compared performance of the NUA traffic load balancing scheme with other schemes and showed its advantages in backhaul-constrained cache-enabled small cell networks with hybrid power supplies.
Tao Han 0002, Nirwan Ansari
IEEE Trans. Mob. Comput.1
2017 Smart Grid Enabled Mobile Networks: Jointly Optimizing BS Operation and Power Distribution
abstract
With the development of green energy technologies, base stations (BSs) can be readily powered by green energy in order to reduce the on-grid power consumption, and subsequently reduce the carbon footprints. As smart grid advances, power trading among distributed power generators and energy consumers will be enabled. In this paper, we investigate the optimization of smart grid-enabled mobile networks, in which green energy is generated in individual BSs and can be shared among the BSs. In order to minimize the on-grid power consumption of this network, we propose to jointly optimize the BS operation and the power distribution. The joint BS operation and power distribution optimization (BPO) problem is challenging due to the complex coupling of the optimization of mobile networks and that of the power grid. We propose an approximate solution that decomposes the BPO problem into two subproblems and solves the BPO by addressing these subproblems. The simulation results show that by jointly optimizing the BS operation and the power distribution, the network achieves about 18% on-grid power savings.
Xueqing Huang, Tao Han 0002, Nirwan Ansari
IEEE/ACM Trans. Netw.2
2016 Revenue Driven Virtual Machine Management in Green Datacenter Networks Towards Big Data
abstract
The big data era is presenting unprecedented opportunities for generating new revenues in various sectors ranging from health care, economics, life science, to manufacturing. Datacenters (DCs) are widely deployed to provision various application services as well as to process big data. Since more and more servers are installed in DCs, the cost of electricity incurs a financial burden for the DC operators. Many DCs are equipped with renewable energy to reduce the electricity bill. However, the locations of the energy demands do not match the locations of the renewable energy generation. This mismatch may be addressed by virtual machine (VM) migration. The DCs, which lack renewable energy, can migrate their workloads to other DCs, which have abundant renewable energy. In addition, DC operators always want to maximize revenue and minimize operation cost. In this paper, the problem of maximizing revenue and minimizing operation cost of a green DC network enabled with and without VM migration is formulated by integer linear programming. Simulation results show that optimal results can be reached for small size problems. Two heuristic algorithms are proposed to efficiently solve large size problems. To our best knowledge, this is the first study of the revenue driven VM management problem in green DC networks towards big data with VM migration.
Liang Zhang 0011, Tao Han 0002, Nirwan Ansari
GLOBECOM2
2016 Intelligent battery management for cellular networks with hybrid energy supplies
abstract
Green communications has received much attention in recent years. In cellular networks, base stations (BSs) account for more than 50 percent of the energy consumption. Reducing energy consumption of BSs is essential to realize green cellular networks. Utilizing green energy to power BSs is a promising way to reduce the on-grid energy consumption. Owing to the dynamics of both mobile traffic loads and green energy, the mismatch between the energy demands and green energy generation in a BS results in inefficient green energy utilization. Managing the battery in BSs can control the green energy usage in individual time slots, thus alleviating the inefficiency caused by the mismatch. In this paper, we propose an intelligent battery management mechanism to optimize the green energy utilization in BSs based on the Markov Decision Process (MDP). A large number of states in the Markov chain are required to model the dynamics of solar radiation and BS workload demands. Thus, the original MDP optimal policy iteration method incurs a high computational complexity. Therefore, we propose some heuristics to approximate the optimal energy dispatching strategy with low computational complexity, and validate the performance of the proposed algorithm through extensive simulations.
Xilong Liu, Tao Han 0002, Nirwan Ansari
WCNC2
2016 A Traffic Load Balancing Framework for Software-Defined Radio Access Networks Powered by Hybrid Energy Sources
abstract
Dramatic mobile data traffic growth has spurred a dense deployment of small cell base stations (SCBSs). Small cells enhance the spectrum efficiency and thus enlarge the capacity of mobile networks. Although SCBSs consume much less power than macro BSs (MBSs) do, the overall power consumption of a large number of SCBSs is phenomenal. As the energy harvesting technology advances, base stations (BSs) can be powered by green energy to alleviate the on-grid power consumption. For mobile networks with high BS density, traffic load balancing is critical in order to exploit the capacity of SCBSs. To fully utilize harvested energy, it is desirable to incorporate the green energy utilization as a performance metric in traffic load balancing strategies. In this paper, we have proposed a traffic load balancing framework that strives a balance between network utilities, e.g., the average traffic delivery latency, and the green energy utilization. Various properties of the proposed framework have been derived. Leveraging the software-defined radio access network architecture, the proposed scheme is implemented as a virtually distributed algorithm, which significantly reduces the communication overheads between users and BSs. The simulation results show that the proposed traffic load balancing framework enables an adjustable trade-off between the on-grid power consumption and the average traffic delivery latency, and saves a considerable amount of on-grid power, e.g., 30%, at a cost of only a small increase, e.g., 8%, of the average traffic delivery latency.
Tao Han 0002, Nirwan Ansari
IEEE/ACM Trans. Netw.1
2015 Renewable Energy-Aware Inter-Datacenter Virtual Machine Migration over Elastic Optical Networks
abstract
Datacenters (DCs) are deployed in a large scale to support the ever increasing demand for data processing to support various applications. The energy consumption of DCs becomes a critical issue. Powering DCs with renewable energy can effectively reduce the brown energy consumption and thus alleviates the energy consumption problem. Owing to geographical deployments of DCs, the renewable energy generation and the data processing demands usually vary in different DCs. Migrating virtual machines (VMs) among DCs according to the availability of renewable energy helps match the energy demands and the renewable energy generation in DCs, and thus maximizes the utilization of renewable energy. Since migrating VMs incurs additional traffic in the network, the VM migration is constrained by the network capacity. The inter-datacenter (inter-DC) VM migration with network capacity constraints is an NP-hard problem. In this paper, we propose two heuristic algorithms that approximate the optimal VM migration solution. Through extensive simulations, we show that the proposed algorithms, by migrating VM among DCs, can reduce up to 31% of brown energy consumption.
Liang Zhang 0011, Tao Han 0002, Nirwan Ansari
CloudCom2
2015 User association in backhaul constrained small cell networks
abstract
Explosive data traffic growth has led to a continuous surge in capacity demands in mobile networks. In order to provision high network capacity, small cell base stations (SCBSs) are widely deployed. Owing to the close proximity to mobile users, SCBSs can effectively enhance the network capacity and offloading traffic load from macro BSs (MBSs). However, the cost-effective backhaul may not be readily available for SCBSs that leads to backhaul constraints in small cell networks. In this paper, we investigate the traffic offloading in backhaul constrained small cell networks. We have proposed a network latency aware user association scheme that balances traffic loads among base stations (BSs) to minimize the average traffic delivery latency of the mobile network. The proposed network latency aware user association scheme considers the traffic delivery latency in both BSs and their backhaul during the process of establishing user associations. We have proved that the proposed user association scheme converges to the optimal solution that minimizes the average traffic delivery latency of the network. The simulation results show that the proposed scheme reduces the average traffic delivery latency by 63% and 34% as compared to the user association scheme that considers only the traffic delivery latency in BSs and a two-tier data rate bias user association scheme, respectively.
Tao Han 0002, Nirwan Ansari
WCNC1
2014 Provisioning green energy for small cell BSs
abstract
Mobile base stations (BSs) can be powered by green energy in order to reduce the on-grid energy consumption, and subsequently reduce the carbon footprints. However, equipping a BS with a green energy system incurs additional capital expenditure (CAPEX) which is determined by the size of the green energy generator, the battery capacity, and other installation expense. In this paper, we introduce and investigate the green energy provisioning (GEP) problem which aims to minimize the CAPEX on deploying green energy powered BSs while achieving the capacity expansion and traffic offloading target. The GEP problem is challenging because it involves optimization over multiple time slots and across multiple BSs. We propose a heuristic green energy provisioning solution which decomposes the GEP problem into three sub-problems: the traffic load optimization problem, the active BS selection problem, and the green energy system sizing problem. We propose algorithms to solve the sub-problems and subsequently solve the GEP problem.
Tao Han 0002, Nirwan Ansari
GLOBECOM1
2014 Smart grid enabled mobile networks: Jointly optimizing BS operation and power distribution
abstract
With the development of green energy technologies, base stations (BSs) can be powered by green energy in order to reduce the on-grid power consumption, and subsequently reduce the carbon footprints. As smart grid advances, power trading among distributed power generators and energy consumers will be enabled. In this paper, we have investigated the optimization of smart grid enabled mobile networks in which green energy is generated in individual BSs and can be shared among the BSs. In order to minimize the on-grid power consumption of this network, we have proposed to jointly optimize the BS operation and the power distribution. The joint BS operation and Power distribution Optimization (BPO) problem is challenging due to the complex coupling of the optimization of mobile networks and that of power grid. We have proposed an approximation solution that decomposes the BPO problem into two subproblems and solves the BPO by address these subproblems. The simulation results show that by jointly optimizing the BS operation and the power distribution, the network achieves about 18% on-grid power savings.
Tao Han 0002, Nirwan Ansari
ICC1
2014 Offloading Mobile Traffic via Green Content Broker
abstract
A continuous surge of mobile data traffic not only congests mobile networks but also results in a dramatic increase in the energy consumption of mobile networks. Device-to-device (D2D) communications is a promising technique for offloading mobile data traffic and enhancing energy efficiency of mobile networks. By enabling D2D communications, the base stations (BSs) may reduce their energy consumption through traffic offloading, and mobile users may increase their quality of services by retrieving contents from their neighboring peers instead of the remote BSs. By leveraging D2D communications, we propose a novel mobile traffic offloading scheme-the content brokerage. In the content brokerage scheme, a new network node called the green content broker (GCB) is introduced to arrange the content delivery between the content requester and the content owner. The GCB is powered by green energy, e.g., solar energy, to reduce the CO2footprints of mobile networks. In the scheme, maximizing traffic offloading with the constraints of the amount of green energy and bandwidth is an nondeterministic polynomial time (NP) problem. We propose a heuristic traffic offloading (HTO) algorithm to approximate the optimal solution with low computational complexity. Our simulation results valid the performance of the content brokerage scheme and the HTO algorithm.
Tao Han 0002, Nirwan Ansari
IEEE Internet Things J.1
2014 Enabling Mobile Traffic Offloading via Energy Spectrum Trading
abstract
Green communications has received much attention recently. For mobile networks, the base stations (BSs) account for more than 50% of the energy consumption of the networks. Therefore, reducing the power consumption of BSs is crucial to greening mobile networks. In this paper, we propose a novel energy spectrum trading (EST) scheme which enables the macro BSs to offload their mobile traffic to Internet service providers' (ISPs') wireless access points by leveraging cognitive radio techniques. Since the ISP's wireless access points are usually closer to the mobile users, the energy and spectral efficiency of mobile networks are enhanced. However, in the EST scheme, achieving optimal mobile traffic offloading in terms of minimizing the energy consumption of the macro BSs is NP-hard. We thus propose a heuristic algorithm to approximate the optimal solution with low computation complexity. We have proved that the energy savings achieved by the proposed heuristic algorithm is at least 50% of that achieved by the brute-force search. Simulation results demonstrate the performance and viability of the proposed EST scheme and the heuristic algorithm.
Tao Han 0002, Nirwan Ansari
IEEE Trans. Wirel. Commun.1
2013 Heuristic relay assignments for green relay assisted device to device communications
abstract
Device to device (D2D) communications is a promising concept to improve data rates and the energy efficiency of mobile networks. User devices (UEs) participated in D2D communications may increase their own data rates by retrieving the content from their neighboring peers instead of from the base stations. Meanwhile, UEs may drain their batteries while performing as content providers and transmitting the content to their peers. Increasing the data rates of D2D communications is desirable to alleviate the UEs' power consumption. In this paper, we propose a novel green relay assisted D2D communication architecture, in which the relay nodes powered by green energy are deployed to increase the data rates of D2D communications. However, achieving the optimal relay assignment for green relay assisted D2D communications is challenging. We propose a heuristic green relay assignment algorithm which maximizes the minimum data rates of the D2D pairs while considering the green load capacity of the relay nodes. We show that the proposed algorithm approximates the optimal solution with low computational complexity, and validate its performance by using simulation results.
Tao Han 0002, Nirwan Ansari
GLOBECOM1
2013 Auction-based energy-spectrum trading in green cognitive cellular networks
abstract
Green communications has received much attention recently. For cellular networks, the base stations (BSs) account for more than 50 percent of the energy consumption of the networks. Therefore, reducing the power consumption of BSs is crucial to enhance the energy efficiency of cellular networks. Meanwhile, the mobile data traffic is expected to increase exponentially. To accommodate the increasing data traffic with the limited radio frequency, enhancing the spectrum efficiency is critical for next generation cellular networks. In this paper, we propose an auction-based energy-spectrum trading scheme which exploits the cooperation between primary base stations (PBSs) and the secondary base stations (SBSs) to enhance the energy as well as spectrum efficiency of cellular networks. In the cooperation, by leveraging cognitive radio, PBSs share the licensed spectrum with SBSs, and the SBSs provide data service to the primary users under its coverage utilizing the shared bandwidth. The cooperation between PSBs and SBSs can significantly improve the energy and spectral efficiency of cellular networks. However, optimizing the bandwidth sharing between PBSs and SBSs is an NP-hard problem. Solving such problem using centralized algorithms is not computationally efficient, especially when considering a large number of PBSs and SBSs. Thus, we design an auction-based decentralized mechanism to enable the cooperation between PBSs and SBSs. Simulation results that demonstrated the performance and viability of the proposed decentralized mechanism.
Tao Han 0002, Nirwan Ansari
ICC1
2013 Energy agile packet scheduling to leverage green energy for next generation cellular networks
abstract
Green communications has received much attention recently. Utilizing green energy in wireless cellular networks is promising to reduce the main grid electricity consumption, and thus to reduce the CO2footprint. However, owing to the fluctuating nature of green energy, it is challenging to use green energy in cellular networks. In this paper, we propose an energy agile packet scheduler which maximizes the utilization of green energy by optimizing the packet scheduling. The packet scheduling optimization problem is NP-hard in the strong sense. The energy agile scheduler approximates the optimal packet scheduling solution in two steps. First, the energy agile scheduler balances the BS's energy consumption in transmitting packets among time slots. Second, within each time slot, the energy agile scheduler optimizes the bandwidth allocations to minimize the BS's energy consumption. Simulation results demonstrate that the proposed energy agile scheduler achieves significant main grid energy savings.
Tao Han 0002, Xueqing Huang, Nirwan Ansari
ICC1
2013 On Optimizing Green Energy Utilization for Cellular Networks with Hybrid Energy Supplies
abstract
Green communications has received much attention recently. For cellular networks, the base stations (BSs) account for more than 50 percent of the energy consumption of the networks. Therefore, reducing the power consumption of BSs is crucial to achieve green cellular networks. With the development of green energy technologies, BSs are able to be powered by green energy in order to reduce the on-grid energy consumption, thus reducing the CO2footprints. In this paper, we envision that the BSs of future cellular networks are powered by both on-grid energy and green energy. We optimize the energy utilization in such networks by maximizing the utilization of green energy, and thus saving on-grid energy. The optimal usage of green energy depends on the characteristics of the energy generation and the mobile traffic, which exhibit both temporal and spatial diversities. We decompose the problem into two sub-problems: the multi-stage energy allocation problem and the multi-BSs energy balancing problem. We propose algorithms to solve these sub-problems, and subsequently solve the green energy optimization problem. Simulation results demonstrate that the proposed solution achieves significant on-grid energy savings.
Tao Han 0002, Nirwan Ansari
IEEE Trans. Wirel. Commun.1
2012 Optimizing cell size for energy saving in cellular networks with hybrid energy supplies
abstract
Green communications has received much attention recently. For cellular networks, the base stations (BSs) account for more than 50 percent of the energy consumption of the networks. Therefore, reducing the power consumption of BSs is crucial to achieve green cellular networks. In this paper, we optimize the energy utilization in cellular networks whose BSs are powered with both regular energy from the grid and the renewable energy. We minimize the on-grid energy consumption of BSs by adapting their cell sizes. The cell size optimization problem is NP-hard. We divide the problem into two subproblems: the multi-stage energy allocation problem and energy consumption minimization problem. We propose an energy allocation policy and an approximation algorithm to solve these subproblems, respectively, and subsequently solve the cell size optimization problem. Simulation results demonstrate that the proposed solution achieves significant energy savings.
Tao Han 0002, Nirwan Ansari
GLOBECOM1
2012 TCP-Mobile Edge: Accelerating delivery in mobile networks
abstract
Owing to the imminent fixed mobile convergence, Internet applications are frequently accessed through mobile nodes. However, service delivery latency is too high to satisfy user expectations. In this paper, we design a new TCP algorithm, TCP-ME (Mobile Edge), to accelerate the service delivery in mobile networks. Considering the QoS (Quality of Service) mechanisms of mobile networks, TCP-ME is designed to differentiate the packet loss caused by wireless errors, traffic conditioning of mobile core networks, and Internet congestion, as well as to react to the packet loss accordingly. To detect wireless errors, we mark the ACK (Acknowledge) packets in the uplink direction at the base station, and the marking threshold is a function of the instantaneous downlink queue length and the number of consecutive HARQ retransmissions. We modify the ECN mechanism with deterministic marking to detect Internet congestion. The packet loss caused by traffic conditioners of mobile networks is detected by whether the incoming DUPACK is marked or not. TCP-ME adapts the inter-packet interval when the packet loss is caused by wireless errors or the admission control mechanism. If the packet loss is due to Internet congestion, TCP-ME applies the TCP-New Reno's congestion window adaptation algorithm. Simulation results show that TCP-ME can speed up web service response time in mobile networks by about 80%.
Tao Han 0002, Nirwan Ansari, Mingquan Wu, Hong Heather Yu
ICC1
2008 Analysis of mobile WiMAX security: Vulnerabilities and solutions
abstract
In this paper, we first give an overview of security architecture of mobile WiMAX network. Then, we investigate man-in-the-middle attacks and Denial of Service (DoS) attacks toward 802.16e-based Mobile WiMAX network. We find the initial network procedure is not effectively secured that makes Man-in-the-middle and Dos attacks possible. In addition, we find the resource saving and handover procedure is not secured enough to resist DoS attacks. Focusing on these two kinds of attacks, we propose Secure Initial Network Entry Protocol (SINEP) based on Diffie-Hellman (DH) key exchange protocol to enhance the security level during network initial. We modify DH key exchange protocol to fit it into mobile WiMAX network as well as to eliminate existing weakness in original DH key exchange protocol.
Tao Han 0002, Ning Zhang 0007, Kaiming Liu, Bihua Tang
MASS1