VLDB 2026 Research / reviewers in the wild / expert
Yonggang Wen 0001
dblp:33/885
· DBLP profile ↗
244ranked-venue papers
11as first author
69since 2021 · last 2026
0000-0002-2751-5114ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 87 · 7 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 72 · 3 first-author · 19 since 2021Artificial intelligence and machine learning · 34 · 17 since 2021Systems, architecture and hardware · 33 · 15 since 2021Databases, data management, data science and information retrieval · 12 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 7 · 4 since 2021Security and privacy · 5 · 2 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Physics-Informed Multi-Task Learning for Battery State of Health Prediction with Uncertainty QuantificationabstractExisting battery State of Health (SOH) prediction approaches often struggle to provide both accurate predictions and reliable uncertainty estimates. This paper presents a novel Multi-Task Learning (MTL) framework that jointly tackles SOH prediction and provides a proxy metric for uncertainty through a unified architecture. The framework combines a Physics-Informed Neural Network (PINN) for SOH prediction with a deep autoencoding Gaussian mixture model for uncertainty modeling. Particularly, the energy score from the Gaussian mixture model serves as a proxy metric for uncertainty, where a higher score indicates potential prediction unreliability. Moreover, to enhance task-specific learning, we employ a multi-head attention mechanism that adaptively captures distinct feature relationships. Our experiments show improvements in prediction performance compared to the state-of-the-art baseline. A comprehensive evaluation on six XJTU battery benchmark datasets demonstrates that our framework achieves a prediction accuracy of 99.50% (MAPE: 0.0050) while providing reliable uncertainty quantification through the proxy metric. Tianwen Zhu, Ruihang Wang, Jimin Jia, Yong Luo 0002, Yonggang Wen 0001 |
AAAI | 7 |
| 2026 | SPPO: Making Million-Token LLM Training Practical on Modest GPU ClustersabstractIn recent years, Large Language Models (LLMs) have exhibited remarkable capabilities, driving advancements in real-world applications. However, training LLMs on increasingly long input sequences imposes significant challenges due to high GPU memory and computational demands. Existing solutions face two key limitations: (1) memory reduction techniques, such as activation recomputation and CPU offloading, compromise training efficiency; (2) distributed parallelism strategies require excessive GPU resources, limiting the scalability of input sequence length. Qiaoling chen, Shenggui Li, Wei Gao 0064, Peng Sun 0006, Yonggang Wen 0001, Tianwei Zhang 0004 |
ICS | 5 |
| 2026 | Efficient Fuzzy Private Set Intersection from Secret-Shared OPRF
Xinpeng Yang, Meng Hao 0001, Chenkai Weng, Robert H. Deng, Yonggang Wen 0001, Tianwei Zhang 0004 |
SP | 5 |
| 2026 | Knowledge-Aware Modeling With Frequency-Adaptive Learning for Battery Health PrognosticsabstractBattery health prognostics are critical for ensuring safety, efficiency and sustainability in modern energy systems. However, it has been challenging to achieve accurate and robust prognostics due to complex battery degradation behaviors with nonlinearity, noises, capacity regeneration, etc. Existing data-driven models capture temporal degradation features but often lack knowledge guidance, which leads to unreliable long-term health prognostics. To overcome these limitations, we propose KARMA, a knowledge-aware model with frequency-adaptive learning for battery capacity estimation and remaining useful life prediction. The model first performs signal decomposition to derive battery signals in different frequency bands. A dual-stream deep learning architecture is developed, where one stream captures long-term low-frequency degradation trends and the other models high-frequency short-term dynamics. KARMA regulates the prognostics with knowledge, where battery degradation is modeled as a double exponential function based on empirical studies. Our dual-stream model is used to optimize the parameters of the knowledge with particle filters to ensure physically consistent and reliable prognostics and uncertainty quantification. Experimental study demonstrates KARMA’s superior performance, achieving average error reductions of 50.6% and 33.3% over state-of-the-art algorithms for battery health prediction on two mainstream datasets, respectively. These results highlight KARMA’s robustness, generalizability and potential for safer and reliable battery management across diverse applications. Vijay Babu Pamshetti, Wei Zhang 0082, Sumei Sun, Jie Zhang 0002, Yonggang Wen 0001, Qingyu Yan |
IEEE Internet Things J. | 5 |
| 2026 | Advancing Generative Artificial Intelligence and Large Language Models for Demand Side Management With Internet of Electric VehiclesabstractThe energy optimization and demand side management (DSM) of Internet of Things (IoT)-enabled microgrids are being transformed by generative artificial intelligence, such as large language models (LLMs). This paper explores an integration of LLMs into energy management, and emphasizes their roles in automating the optimization of DSM strategies with Internet of Electric Vehicles (IoEV) as a representative example of the Internet of Vehicles (IoV). We investigate challenges and solutions associated with DSM and explore new opportunities presented by leveraging LLMs. Then, we propose an innovative solution that enhances LLMs with retrieval-augmented generation for automatic problem formulation, code generation, and customizing optimization. The results demonstrate the effectiveness of our proposed solution in charging scheduling and optimization for electric vehicles, and highlight our solution’s significant advancements in energy efficiency and user adaptability. This work shows LLMs’ potential in energy optimization of the IoT-enabled microgrids and promotes intelligent DSM solutions. Hanwen Zhang 0004, Ruichen Zhang 0001, Wei Zhang 0082, Dusit Niyato, Yonggang Wen 0001, Chunyan Miao |
IEEE Internet Things J. | 5 |
| 2026 | ChatDC: Geometric-aware Data Center Digital Twin Generation via Large Language ModelsabstractA modern data center is supported by an Internet of Assets (IoA), a specialized Internet of Things (IoT), where the sim-ready assets encompass physical properties and connected data streams for passive data collection and proactive simulation. A digital twin (DT) is a virtual replica of the IoA system with a proper level of abstraction, integrating asset-level dynamics models to simulate system-wide behavior, prototype designs, and conduct what-if analysis. However, creating high-fidelity DTs is hindered by the manual effort needed to encode complex geometric layouts and domain-specific design constraints. While large language models (LLMs) offer automation potential, existing methods struggle to generate geometrically plausible and functionally valid DT scenes due to limited domain integration. To bridge the gaps, we propose ChatDC, a conversational system that leverages LLMs to automate data center digital twin generation through a S egment- G enerate- O ptimize (SGO) workflow. ChatDC integrates domain knowledge via a dedicated code library named DCBuild and employs SGO to decompose user prompts, generate initial structures, and optimize layouts in compliance with data center design constraints. Evaluation shows that ChatDC outperforms other baselines with 98% success rate on scratch generation tasks and reduces the average makespan by 10×. Ablation study reveals that the SGO design increases the generation success rate by 65% at most. Furthermore, computational fluid dynamics simulations validate the physical plausibility, confirming the readiness of generated DTs for real-world analysis. Minghao Li 0005, Ruihang Wang, Xin Zhou 0003, Zhaomeng Zhu, Yonggang Wen 0001, Rui Tan 0001, Huiwen Zheng, Stuart Kennedy |
ACM Trans. Internet Things | 5 |
| 2025 | ADEM-VL: Adaptive and Embedded Fusion for Efficient Vision-Language Tuning
Zhiwei Hao 0001, Jianyuan Guo, Li Shen 0008, Yong Luo 0002, Han Hu 0003, Yonggang Wen 0001 |
Int. J. Comput. Vis. | 6 |
| 2025 | Aligning Text-to-Image Diffusion Models With Constrained Reinforcement LearningabstractReward finetuning has emerged as a powerful technique for aligning diffusion models with specific downstream objectives or user preferences. However, current approaches suffer from a persistent challenge of reward overoptimization, where models exploit imperfect reward feedback at the expense of overall performance. In this work, we identify three key contributors to overoptimization: (1) a granularity mismatch between the multi-step diffusion process and sparse rewards; (2) a loss of plasticity that limits the model's ability to adapt and generalize; and (3) an overly narrow focus on a single reward objective that neglects complementary performance criteria. Accordingly, we introduce Constrained Diffusion Policy Optimization (CDPO), a novel reinforcement learning framework that addresses reward overoptimization from multiple angles. Firstly, CDPO tackles the granularity mismatch through a temporal policy optimization strategy that delivers step-specific rewards throughout the entire diffusion trajectory, thereby reducing the risk of overfitting to sparse final-step rewards. Then we incorporate a neuron reset strategy that selectively resets overactive neurons in the model, preventing overoptimization induced by plasticity loss. Finally, to avoid overfitting to a narrow reward objective, we integrate constrained reinforcement learning with auxiliary reward objectives serving as explicit constraints, ensuring a balanced optimization across diverse performance metrics. Ziyi Zhang 0001, Sen Zhang 0006, Li Shen 0008, Yibing Zhan, Yong Luo 0002, Han Hu 0003, Bo Du 0001, Yonggang Wen 0001, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2025 | CoFormer: Collaborating With Heterogeneous Edge Devices for Scalable Transformer InferenceabstractThe impressive performance of transformer models has sparked the deployment of intelligent applications on resource-constrained edge devices. However, ensuring high-quality service for real-time edge systems is a significant challenge due to the considerable computational demands and resource requirements of these models. Existing strategies typically either offload transformer computations to other devices or directly deploy compressed models on individual edge devices. These strategies, however, result in either considerable communication overhead or suboptimal trade-offs between accuracy and efficiency. To tackle these challenges, we propose a collaborative inference system for general transformer models, termed CoFormer. The central idea behind CoFormer is to exploit the divisibility and integrability of transformer. An off-the-shelf large transformer can be decomposed into multiple smaller models for distributed inference, and their intermediate results are aggregated to generate the final output. We formulate an optimization problem to minimize both inference latency and accuracy degradation under heterogeneous hardware constraints. DeBo algorithm is proposed to first solve the optimization problem to derive the decomposition policy, and then progressively calibrate decomposed models to restore performance. We demonstrate the capability to support a wide range of transformer models on heterogeneous edge devices, achieving up to 3.1× inference speedup with large transformer models. Notably, CoFormer enables the efficient inference of GPT2-XL with 1.6 billion parameters on edge devices, reducing memory requirements by 76.3%. CoFormer can also reduce energy consumption by approximately 40% while maintaining satisfactory inference performance. Guanyu Xu, Zhiwei Hao 0001, Li Shen 0008, Yong Luo 0002, Fuhui Sun, Han Hu 0003, Yonggang Wen 0001 |
IEEE Trans. Computers | 8 |
| 2025 | Federated Learning With Only Positive Labels by Exploring Label CorrelationsabstractFederated learning (FL) aims to collaboratively learn a model by using the data from multiple users under privacy constraints. In this article, we study the multilabel classification (MLC) problem under the FL setting, where trivial solution and extremely poor performance may be obtained, especially when only positive data with respect to a single class label is provided for each client. This issue can be addressed by adding a specially designed regularizer on the server side. Although effective sometimes, the label correlations are simply ignored and thus suboptimal performance may be obtained. Besides, it is expensive and unsafe to exchange user's private embeddings between server and clients frequently, especially when training model in the contrastive way. To remedy these drawbacks, we propose a novel and generic method termed federated averaging (FedAvg) by exploring label correlations (FedALCs). Specifically, FedALC estimates the label correlations in the class embedding learning for different label pairs and utilizes it to improve the model training. To further improve the safety and also reduce the communication overhead, we propose a variant to learn fixed class embedding for each client, so that the server and clients only need to exchange class embeddings once. Extensive experiments on multiple popular datasets demonstrate that our FedALC can significantly outperform the existing counterparts. Xuming An 0001, Dui Wang, Li Shen 0008, Yong Luo 0002, Han Hu 0003, Bo Du 0001, Yonggang Wen 0001, Dacheng Tao |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2025 | IceFrog: A Layer-Elastic Scheduling System for Deep Learning Training in GPU ClustersabstractThe high resource demand of deep learning training (DLT) workloads necessitates the design of efficient schedulers. While most existing schedulers expedite DLT workloads by considering GPU sharing and elastic training, they neglectlayer elasticity, which dynamically freezes certain layers of a network. This technique has been shown to significantly speed up individual workloads. In this paper, we explore how to incorporatelayer elasticityinto DLT scheduler designs to achieve higher cluster-wide efficiency. A key factor that hinders the application of layer elasticity in GPU clusters is the potential loss in model accuracy, making users reluctant to enable layer elasticity for their workloads. It is necessary to have an efficient layer-elastic system, which can well balance training accuracy and speed for layer elasticity. We introduceIceFrog, the first scheduling system that utilizes layer elasticity to improve the efficiency of DLT workloads in GPU clusters. It achieves this goal with superior algorithmic designs and intelligent resource management. In particular, (1) we model the frozen penalty and layer-aware throughput to measure the effective progress metric of layer-elastic workloads. (2) We design a novel scheduler to further improve the efficiency of layer elasticity. We implement and deployIceFrogin a physical cluster of 48 GPUs. Extensive evaluations and large-scale simulations show thatIceFrogreduces average job completion times by 36-48% relative to state-of-the-art DL schedulers. Wei Gao 0064, Zhuoyuan Ouyang, Peng Sun 0006, Tianwei Zhang 0004, Yonggang Wen 0001 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2025 | Adaptive Capacity Provisioning for Carbon-Aware Data Centers: A Digital Twin-Based ApproachabstractThis paper considers the carbon-aware data center (DC) capacity provisioning problem under uncertain green energy availability and computing demand. To address it, accurate carbon emissions estimation and robust capacity provisioning are necessary. Existing studies mainly consider the carbon footprint of the computing system and merely consider that of the physical facilities, which also contribute significant carbon emissions. Furthermore, their capacity provisioning is neither uncertainty-aware nor adaptive to the dynamic computing demand. To bridge these gaps, we propose an adaptive capacity provisioning framework based on the physics-informed digital twin. We design the digital twin to holistically capture a DC's operational carbon footprint, including both the computing system and the physical facilities. The digital twin is differentiable and established with a collection of physics-informed learnable models that are learned with online operational data. We further address the challenge of capacity provisioning under uncertainties by designing a shrinking horizon model predictive control. The designed capacity planner updates its estimation of future computing demand based on the observable computing system states. At each capacity provisioning round, we solve the capacity provisioning problem using a gradient-based optimization technique with the gradient provided by the digital twin. We extensively evaluate our approach usingrealoperational data from a large-scale production data center. First, our digital twin accurately predicts holistic DC energy usage with a relative absolute error of less than 5%, which is accurate according to the industrial rule of thumb. Second, we show that our solution is comparable to the oracle solution with perfect knowledge about all uncertainties, outperforming the state-of-the-art Predict-then-Plan approach significantly in terms of SLO violation reduction. Furthermore, our approach reduces carbon footprint by 27% compared with the over-provisioning scheme currently adopted by the industry. Ruihang Wang, Xin Zhou 0003, Rui Tan 0001, Yonggang Wen 0001, Yuejun Yan |
IEEE Trans. Sustain. Comput. | 5 |
| 2024 | Sylvie: 3D-Adaptive and Universal System for Large-Scale Graph Neural Network TrainingabstractDistributed full-graph training of Graph Neural Networks (GNNs) has been widely adopted to learn large-scale graphs. While recent system advancements can improve the training throughput of GNNs, their practical adoption is limited by the potential accuracy decline. This concern is particularly prominent in deeper and more intricate GNN architectures, where noticeable performance degradation becomes apparent. Moreover, existing works fail to comprehensively consider diverse opportunities for acceleration. Motivated by these deficiencies, we propose Sylvie,a full-graph training system that not only improves the training throughput substantially but also maintains the model quality for universal GNNs. By harnessing the inherent information embedded in the graph data and model structure, Sylvie intelligently optimizes GNN training across three key dimensions: data, time, and execution. It identifies performance-relevant features of the input graph offline as subsequent optimization guidance. Subsequently, Sylvie devises an online convergence-maintenance strategy that adaptively integrates and aligns GNN-specific quantization and inter-epoch asynchronous training with the real-time training characteristics. Extensive experiments demonstrate that Sylvie surpasses existing GNN training systems by up to 17.2× speedup for both shallow and deep GNNs, without compromising the model accuracy. Meng Zhang 0045, Qinghao Hu 0004, Cheng Wan 0005, Haozhao Wang, Peng Sun 0006, Yonggang Wen 0001, Tianwei Zhang 0001 |
ICDE | 6 |
| 2024 | Confronting Reward Overoptimization for Diffusion Models: A Perspective of Inductive and Primacy BiasesabstractBridging the gap between diffusion models and human preferences is crucial for their integration into practical generative workflows. While optimizing downstream reward models has emerged as a promising alignment strategy, concerns arise regarding the risk of excessive optimization with learned reward models, which potentially compromises ground-truth performance. In this work, we confront the reward overoptimization problem in diffusion model alignment through the lenses of both inductive and primacy biases. We first identify a mismatch between current methods and the temporal inductive bias inherent in the multi-step denoising process of diffusion models, as a potential source of reward overoptimization. Then, we surprisingly discover that dormant neurons in our critic model act as a regularization against reward overoptimization while active neurons reflect primacy bias. Motivated by these observations, we propose Temporal Diffusion Policy Optimization with critic active neuron Reset (TDPO-R), a policy gradient algorithm that exploits the temporal inductive bias of diffusion models and mitigates the primacy bias stemming from active neurons. Empirical results demonstrate the superior efficacy of our methods in mitigating reward overoptimization. Code is avaliable at https://github.com/ZiyiZhang27/tdpo. Ziyi Zhang 0001, Sen Zhang 0006, Yibing Zhan, Yong Luo 0002, Yonggang Wen 0001, Dacheng Tao |
ICML | 5 |
| 2024 | AutoSched: An Adaptive Self-configured Framework for Scheduling Deep Learning Training WorkloadsabstractModern Deep Learning Training (DLT) schedulers in GPU datacenters are designed to be very sophisticated with many configurations. These configurations need to be adjusted delicately as they can significantly affect the scheduling performance. Existing schedulers require the datacenter operator to tune the configurations only once before they are deployed, based on the historical workload traces. Unfortunately, workloads in a datacenter would experience dynamic changes and deviate a lot from the historical ones over time, making the pre-determined configurations less effective. Wei Gao 0064, Shangwei Guo, Peng Sun 0006, Yonggang Wen 0001, Tianwei Zhang 0004 |
ICS | 6 |
| 2024 | Ymir: A Scheduler for Foundation Model Fine-tuning Workloads in DatacentersabstractThe breakthrough of foundation models makes foundation model fine-tuning (FMF) workloads prevalent in modern GPU datacenters. However, existing schedulers tailored for model training do not consider the unique characteristics of FMs, making them inefficient in handling FMF workloads. To bridge the gap, we propose Ymir, a scheduler to improve the efficiency of FMF workloads in GPU datacenters. Ymir leverages the shared FM backbone architecture to expedite FMF workloads from two aspects: (1) Ymir investigates the task transferability among different FMF workloads and automatically merges FMF workloads with the same FM into one to improve the cluster-wide efficiency via transfer learning. (2) Ymir reuses the fine-tuning runtime of FMF workloads to reduce the significant context switch overhead. We conduct 32-GPU physical experiments and 240-GPU trace-driven simulations to validate the effectiveness of Ymir. Ymir can reduce the average job completion time by up to 4.3 × compared with existing state-of-the-art schedulers. It also promotes scheduling fairness by fully exploiting the task transferability. More supplementary materials can be found on our project website https://sites.google.com/view/ymir-project. Wei Gao 0064, Weiming Zhuang, Minghao Li 0005, Peng Sun 0006, Yonggang Wen 0001, Tianwei Zhang 0004 |
ICS | 5 |
| 2024 | Joint Input and Output Coordination for Class-Incremental Learning
Shuai Wang 0011, Yibing Zhan, Yong Luo 0002, Han Hu 0003, Wei Yu 0004, Yonggang Wen 0001, Dacheng Tao |
IJCAI | 6 |
| 2024 | Lins: Reducing Communication Overhead of ZeRO for Efficient LLM TrainingabstractTraining large language models (LLMs) encounters challenges in GPU memory consumption due to the high memory requirements of model states. The widely used Zero Redundancy Optimizer (ZeRO) addresses this issue through strategic sharding but introduces communication challenges at scale. To tackle this problem, we propose Lins, a system designed to optimize ZeRO for scalable LLM training. Lins incorporates three flexible sharding strategies: Full-Replica, Full-Sharding, and Partial-Sharding, and allows each component within the model states (Parameters, Gradients, Optimizer States) to independently choose a sharding strategy as well as the device mesh. We conduct a thorough analysis of communication costs, formulating an optimization problem to discover the optimal sharding strategy. Evaluations demonstrate up to 52% Model FLOPs Utilization (MFU) when training the LLaMA-based model on 1024 GPUs, resulting in a 1.56 times improvement in training throughput compared to newly proposed systems like MiCS and ZeRO++. Qiaoling Chen, Qinghao Hu 0004, Guoteng Wang, Yingtong Xiong, Yang Gao 0042, Hang Yan 0001, Yonggang Wen 0001, Tianwei Zhang 0004, Peng Sun 0006 |
IWQoS | 9 |
| 2024 | Characterization of Large Language Model Development in the Datacenter
Qinghao Hu 0004, Zhisheng Ye 0002, Zerui Wang, Guoteng Wang, Meng Zhang 0045, Qiaoling Chen, Peng Sun 0006, Dahua Lin, Xiaolin Wang 0001, Yingwei Luo, Yonggang Wen 0001, Tianwei Zhang 0004 |
NSDI | 11 |
| 2024 | TorchGT: A Holistic System for Large-Scale Graph Transformer TrainingabstractGraph Transformer is a new architecture that surpasses GNNs in graph learning. While there emerge inspiring algorithm advancements, their practical adoption is still limited, particularly on real-world graphs involving up to millions of nodes. We observe existing graph transformers fail on large-scale graphs mainly due to heavy computation, limited scalability and inferior model quality. Motivated by these observations, we propose TORCHGT, the first efficient, scalable, and accurate graph transformer training system. TORCHGT optimizes training at three different levels. At algorithm level, by harnessing the graph sparsity, TORCHGT introduces a Dual-interleaved Attention which is computation-efficient and accuracy-maintained. At runtime level, TORCHGT scales training across workers with a communicationlight Cluster-aware Graph Parallelism. At kernel level, an Elastic Computation Reformation further optimizes the computation by reducing memory access latency in a dynamic way. Extensive experiments demonstrate that TORCHGT boosts training by up to 62.7× and supports graph sequence lengths of up to 1M. Meng Zhang 0045, Jie Sun 0017, Qinghao Hu 0004, Peng Sun 0006, Zeke Wang, Yonggang Wen 0001, Tianwei Zhang 0004 |
SC | 6 |
| 2024 | FedDSE: Distribution-aware Sub-model Extraction for Federated Learning over Resource-constrained DevicesabstractSub-model extraction based federated learning has emerged as a popular strategy for training models on resource-constrained devices. However, existing methods treat all clients equally and extract sub-models using predetermined rules, which disregard the statistical heterogeneity across clients and may lead to fierce competition among them. Specifically, this paper identifies that when making predictions, different clients tend to activate different neurons of the entire model related to their respective distributions. If highly activated neurons from some clients with one distribution are incorporated into the sub-model allocated to other clients with different distributions, they will be forced to fit the new distributions, which can hinder their activation over the previous clients and result in a performance reduction. Motivated by this finding, we propose a novel method called FedDSE, which can reduce the conflicts among clients by extracting sub-models based on the data distribution of each client. The core idea of FedDSE is to empower each client to adaptively extract neurons from the entire model based on their activation over the local dataset. We theoretically show that FedDSE can achieve an improved classification score and convergence over general neural networks with the ReLU activation function. Experimental results on various datasets and models show that FedDSE outperforms all state-of-the-art baselines. Haozhao Wang, Yabo Jia, Meng Zhang 0045, Qinghao Hu 0004, Hao Ren 0001, Peng Sun 0006, Yonggang Wen 0001, Tianwei Zhang 0004 |
WWW | 7 |
| 2024 | UniSched: A Unified Scheduler for Deep Learning Training Jobs With Different User DemandsabstractThe growth of deep learning training (DLT) jobs in modern GPU clusters calls for efficient deep learning (DL) scheduler designs. Due to the extensive applications of DL technology, developers may have different demands for their DLT jobs. It is important for a GPU cluster to support all these demands and efficiently execute those DLT jobs. Unfortunately, existing DL schedulers mainly focus on part of those demands, and cannot provide comprehensive scheduling services.In this work, we present UniSched, a unified scheduler to optimize different types of scheduling objectives (e.g., guaranteeing the deadlines of SLO jobs, minimizing the latency of best-effort jobs). Meanwhile, UniSchedsupports different job stopping criteria (e.g., iteration-based, performance-based). UniSched includes two key components: Estimator for estimating the job duration, and Selector for selecting jobs and allocating resources. We perform large-scale simulations over the job traces from the production clusters. Compared to state-of-the-art schedulers, UniSchedcan significantly decrease the deadline miss rate of SLO jobs by up to 6.84×, and the latency of best-effort jobs by up to 4.02×, To demonstrate the practicality of UniSched, we implement and deploy a prototype on Kubernetes in a physical cluster consisting of 64 GPUs. Wei Gao 0064, Zhisheng Ye 0002, Peng Sun 0006, Tianwei Zhang 0004, Yonggang Wen 0001 |
IEEE Trans. Computers | 5 |
| 2024 | Green Data Center Cooling Control via Physics-guided Safe Reinforcement LearningabstractDeep reinforcement learning (DRL) has shown good performance in tackling Markov decision process (MDP) problems. As DRL optimizes a long-term reward, it is a promising approach to improving the energy efficiency of data-center cooling. However, enforcement of thermal safety constraints during DRL’s state exploration is a main challenge. The widely adopted reward-shaping approach adds negative reward when the exploratory action results in unsafety. Thus, it needs to experience sufficient unsafe states before it learns how to prevent unsafety. In this article, we propose a safety-aware DRL framework for data-center cooling control. It applies offline imitation learning and online post-hoc rectification to holistically prevent thermal unsafety during online DRL. In particular, the post-hoc rectification searches for the minimum modification to the DRL-recommended action such that the rectified action will not result in unsafety. The rectification is designed based on a thermal state transition model that is fitted using historical safe operation traces and able to extrapolate the transitions to unsafe states explored by DRL. Extensive evaluation for chilled water and direct expansion-cooled data centers in two climate conditions show that our approach saves 18% to 26.6% of total data-center power compared with conventional control and reduces safety violations by 94.5% to 99% compared with reward shaping. We also extend the proposed framework to address data centers with non-uniform temperature distributions for detailed safety considerations. The evaluation shows that our approach saves 14% power usage compared with the PID control while addressing safety compliance during the training. Ruihang Wang, Xin Zhou 0003, Yonggang Wen 0001, Rui Tan 0001 |
ACM Trans. Cyber Phys. Syst. | 4 |
| 2024 | Exploring the Practicality of Differentially Private Federated Learning: A Local Iteration Tuning ApproachabstractAlthough Federated Learning (FL) prevents the exposure of original data samples when collaboratively training machine learning models among decentralized clients, it has been revealed that vanilla FL is still susceptible to adversarial attacks if model parameters are leaked to malicious attackers. To enhance the protection level of FL, Differential Private Federated Learning (DPFL) has been proposed in recent years. DPFL injects zero-mean noises randomly generated by differential private (DP) mechanisms on local model parameters before they are disclosed. Nevertheless, DP noises can significantly deteriorate model utility jeopardizing the practicality of DPFL. In this paper, we are among the first to explore how to improve the model utility of DPFL by tuning the number of local iterations (LIs) on DPFL clients. Our work shows that such a local iteration tuning approach can well mitigate the adverse influence of DP noises on the final model utility. Formally, we derive the sensitivity (a measure of the maximum change of the output given two adjacent inputs) with respect to the number of LIs conducted on DPFL clients for the Laplace mechanism, and the aggregated variances of Laplace noises at the server side. We further conduct convergence rate analysis to quantify the influence of the Laplace noises on the final model accuracy and determine how to optimally set the number of LIs. Finally, to verify our theoretical findings, we perform extensive experiments using three real-world datasets, namely, Lending Club, MNIST and Fashion-MNIST. The results not only corroborate our analysis, but also demonstrate that our approach significantly improves the practicality of DPFL. Yipeng Zhou, Jiahao Liu 0001, Di Wu 0001, Shui Yu 0001, Yonggang Wen 0001 |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2024 | Not All Instances Contribute Equally: Instance-Adaptive Class Representation Learning for Few-Shot Visual RecognitionabstractFew-shot visual recognition refers to recognize novel visual concepts from a few labeled instances. Many few-shot visual recognition methods adopt the metric-based meta-learning paradigm by comparing the query representation with class representations to predict the category of query instance. However, the current metric-based methods generally treat all instances equally and consequently often obtain biased class representation, considering not all instances are equally significant when summarizing the instance-level representations for the class-level representation. For example, some instances may contain unrepresentative information, such as too much background and information of unrelated concepts, which skew the results. To address the above issues, we propose a novel metric-based meta-learning framework termed instance-adaptive class representation learning network (ICRL-Net) for few-shot visual recognition. Specifically, we develop an adaptive instance revaluing network (AIRN) with the capability to address the biased representation issue when generating the class representation, by learning and assigning adaptive weights for different instances according to their relative significance in the support set of corresponding class. In addition, we design an improved bilinear instance representation and incorporate two novel structural losses, i.e., intraclass instance clustering loss and interclass representation distinguishing loss, to further regulate the instance revaluation process and refine the class representation. We conduct extensive experiments on four commonly adopted few-shot benchmarks: miniImageNet, tieredImageNet, CIFAR-FS, and FC100 datasets. The experimental results compared with the state-of-the-art approaches demonstrate the superiority of our ICRL-Net. Mengya Han, Yibing Zhan, Yong Luo 0002, Bo Du 0001, Han Hu 0003, Yonggang Wen 0001, Dacheng Tao |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | Web3 Metaverse: State-of-the-Art and VisionabstractThe metaverse, as a rapidly evolving socio-technical phenomenon, exhibits significant potential across diverse domains by leveraging Web3 (a.k.a. Web 3.0) technologies such as blockchain, smart contracts, and non-fungible tokens (NFTs). This survey aims to provide a comprehensive overview of the Web3 metaverse from a human-centered perspective. We (i) systematically review the development of the metaverse over the past 30 years, highlighting the balanced contributions from its core components: Web3, immersive convergence, and crowd intelligence communities, (ii) define the metaverse that integrates the Web3 community as the Web3 metaverse and propose an analysis framework from the community, society, and human layers to describe the features, missions, and relationships for each community and their overlapping sections, (iii) survey the state-of-the-art of the Web3 metaverse from a human-centered perspective, namely, the identity, field, and behavior aspects, and (iv) provide supplementary technical reviews. To the best of our knowledge, this work represents the first systematic, interdisciplinary survey on the Web3 metaverse. Specifically, we commence by discussing the potential for establishing decentralized identities (DID) utilizing mechanisms such as profile picture (PFP) NFTs, domain name NFTs, and soulbound tokens (SBTs). Subsequently, we examine land, utility, and equipment NFTs within the Web3 metaverse, highlighting interoperable and full on-chain solutions for existing centralization challenges. Lastly, we spotlight current research and practices about individual, intra-group, and inter-group behaviors within the Web3 metaverse, such as Creative Commons Zero license (CC0) NFTs, decentralized education, decentralized science (DeSci), and decentralized autonomous organizations (DAO). Furthermore, we share our insights into several promising directions, encompassing three key socio-technical facets of Web3 metaverse development. Hongzhou Chen, Haihan Duan, Maha Abdallah, Yufeng Zhu, Yonggang Wen 0001, Abdulmotaleb El Saddik, Wei Cai 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2024 | Personalized Federated Mutual Learning for Unsupervised Camera-Aware Person Re-IdentificationabstractPerson re-identification (ReID) is essential for enhancing security and tracking in multi-camera surveillance systems. To achieve effective ReID performance across diverse datasets, the Federated Unsupervised Person Re-identification via Camera-aware Clustering (FedUCA) approach has made strides in utilizing distributed datasets while ensuring data privacy. Nevertheless, its uniform model may not adequately cater to the specific characteristics of each participant’s data, given the diversity in camera perspectives and client-specific data variances, thus obtaining degraded results. To address this issue, we propose an advanced framework, Personalized Federated Dual-Model Learning for Camera-Aware Person Re-Identification (PerFedDual), which introduces knowledge-sharing techniques inherent to mutual learning for FedUCA with the camera-centric clustering process. PerFedDual supports a dual-model training approach that creates a cooperative learning space that improves both the global model and client-specific models by exchanging knowledge both ways. The methodology adopted is precisely adjusted to the distinctive data environment of each client, ensuring the protection of privacy while simultaneously enhancing the accuracy and flexibility of ReID models in diverse camera configurations. The empirical evaluation reveals that PerFedDual outperforms FedUCA and alternative federated learning strategies, highlighting the benefits of our technique that leverages collective intelligence to enhance unsupervised ReID. Jiabei Liu, Weiming Zhuang, Yuanyuan Liu 0004, Yonggang Wen 0001, Jun Huang 0007, Wei Lin 0016 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2024 | Data Center Sustainability: Revisits and OutlooksabstractAs energy-intensive entities, data centers are associated with significant environmental impacts, making their sustainability a subject of growing interest in recent years. In this article, we revisit data center sustainability and propose a forward-looking vision for improving data center sustainability. We argue that data center sustainability encompasses more than just energy efficiency and must be evaluated and optimized through a multi-faceted approach. To this end, we first present an overview of the sustainability metrics from five aspects. After that, we demonstrate the sustainability status of the latest data centers utilizing publicly available data center sustainability ratings. Furthermore, we examine the evolution of data center sustainability standards in Singapore to highlight several trending features. Based on the analysis, we identify several key elements of sustainable data centers. We then propose the Cognitive Digital Twin (CDT) architecture, which incorporates a digital twin engine for system-wide simulation and a decision engine for optimal control to improve data center sustainability. A case study is performed to optimize the chiller plant efficiency of a production data center in Singapore. The results demonstrate that the CDT can improve chiller plant energy efficiency by 5%, indicating around 140 metric tons of annual carbon emission savings. Xin Zhou 0003, Zhaomeng Zhu, Tracy Liu, Jeffery Neng, Yonggang Wen 0001 |
IEEE Trans. Sustain. Comput. | 7 |
| 2023 | FedABC: Targeting Fair Competition in Personalized Federated LearningabstractFederated learning aims to collaboratively train models without accessing their client's local private data. The data may be Non-IID for different clients and thus resulting in poor performance. Recently, personalized federated learning (PFL) has achieved great success in handling Non-IID data by enforcing regularization in local optimization or improving the model aggregation scheme on the server. However, most of the PFL approaches do not take into account the unfair competition issue caused by the imbalanced data distribution and lack of positive samples for some classes in each client. To address this issue, we propose a novel and generic PFL framework termed Federated Averaging via Binary Classification, dubbed FedABC. In particular, we adopt the ``one-vs-all'' training strategy in each client to alleviate the unfair competition between classes by constructing a personalized binary classification problem for each class. This may aggravate the class imbalance challenge and thus a novel personalized binary classification loss that incorporates both the under-sampling and hard sample mining strategies is designed. Extensive experiments are conducted on two popular datasets under different settings, and the results demonstrate that our FedABC can significantly outperform the existing counterparts. Dui Wang, Li Shen 0008, Yong Luo 0002, Han Hu 0003, Kehua Su, Yonggang Wen 0001, Dacheng Tao |
AAAI | 6 |
| 2023 | Discriminative Reasoning with Sparse Event Representation for Document-level Event-Event Relation ExtractionabstractDocument-level Event-Event Relation Extraction (DERE) aims to extract relations between events in a document.It challenges conventional sentence-level task (SERE) with difficult long-text understanding.In this paper, we propose a novel DERE model (SENDIR) for better document-level reasoning.Different from existing works that build an event graph via linguistic tools, SENDIR does not require any prior knowledge.The basic idea is to discriminate event pairs in the same sentence or span multiple sentences by assuming their different information density: 1) low density in the document suggests sparse attention to skip irrelevant information.Our module 1 designs various types of attention for event representation learning to capture long-distance dependence.2) High density in a sentence makes SERE relatively easy.Module 2 uses different weights to highlight the roles and contributions of intra-and intersentential reasoning, which introduces supportive event pairs for joint modeling.Extensive experiments demonstrate great improvements in SENDIR and the effectiveness of various sparse attention for document-level representations.Codes will be released later. Changsen Yuan, Heyan Huang, Yixin Cao 0002, Yonggang Wen 0001 |
ACL (1) | 4 |
| 2023 | Lucid: A Non-intrusive, Scalable and Interpretable Scheduler for Deep Learning Training JobsabstractWhile recent deep learning workload schedulers exhibit excellent performance, it is arduous to deploy them in practice due to some substantial defects, including inflexible intrusive manner, exorbitant integration and maintenance cost, limited scalability, as well as opaque decision processes. Motivated by these issues, we design and implement Lucid, a non-intrusive deep learning workload scheduler based on interpretable models. It consists of three innovative modules. First, a two-dimensional optimized profiler is introduced for efficient job metric collection and timely debugging job feedback. Second, Lucid utilizes an indolent packing strategy to circumvent interference. Third, Lucid orchestrates resources based on estimated job priority values and sharing scores to achieve efficient scheduling. Additionally, Lucid promotes model performance maintenance and system transparent adjustment via a well-designed system optimizer. Our evaluation shows that Lucid reduces the average job completion time by up to 1.3× compared with state-of-the-art preemptive scheduler Tiresias. Furthermore, it provides explicit system interpretations and excellent scalability for practical deployment. Qinghao Hu 0004, Meng Zhang 0045, Peng Sun 0006, Yonggang Wen 0001, Tianwei Zhang 0004 |
ASPLOS (2) | 4 |
| 2023 | MAS: Towards Resource-Efficient Federated Multiple-Task LearningabstractFederated learning (FL) is an emerging distributed machine learning method that empowers in-situ model training on decentralized edge devices. However, multiple simultaneous FL tasks could overload resource-constrained devices. In this work, we propose the first FL system to effectively coordinate and train multiple simultaneous FL tasks. We first formalize the problem of training simultaneous FL tasks. Then, we present our new approach, MAS (Merge and Split), to optimize the performance of training multiple simultaneous FL tasks. MAS starts by merging FL tasks into an all-in-one FL task with a multi-task architecture. After training for a few rounds, MAS splits the all-in-one FL task into two or more FL tasks by using the affinities among tasks measured during the all-in-one training. It then continues training each split of FL tasks based on model parameters from the all-in-one training. Extensive experiments demonstrate that MAS outperforms other methods while reducing training time by 2× and reducing energy consumption by 40%. We hope this work will inspire the community to further study and optimize training simultaneous FL tasks. Weiming Zhuang, Yonggang Wen 0001, Lingjuan Lyu, Shuai Zhang 0004 |
ICCV | 2 |
| 2023 | Rethinking the Localization in Weakly Supervised Object LocalizationabstractWeakly supervised object localization (WSOL) is one of the most popular and challenging tasks in computer vision. This task is to localize the objects in the images given only the image-level supervision. Recently, dividing WSOL into two parts (class-agnostic object localization and object classification) has become the state-of-the-art pipeline for this task. However, existing solutions under this pipeline usually suffer from the following drawbacks: 1) they are not flexible since they can only localize one object for each image due to the adopted single-class regression (SCR) for localization; 2) the generated pseudo bounding boxes may be noisy, but the negative impact of such noise is not well addressed. To remedy these drawbacks, we first propose to replace SCR with a binary-class detector (BCD) for localizing multiple objects, where the detector is trained by discriminating the foreground and background. Then we design a weighted entropy (WE) loss using the unlabeled data to reduce the negative impact of noisy bounding boxes. Extensive experiments on the popular CUB-200-2011 and ImageNet-1K datasets demonstrate the effectiveness of our method. Rui Xu 0031, Yong Luo 0002, Han Hu 0003, Bo Du 0001, Jialie Shen 0001, Yonggang Wen 0001 |
ACM Multimedia | 6 |
| 2023 | Hydro: Surrogate-Based Hyperparameter Tuning Service in Datacenters
Qinghao Hu 0004, Zhisheng Ye 0002, Meng Zhang 0045, Qiaoling Chen, Peng Sun 0006, Yonggang Wen 0001, Tianwei Zhang 0004 |
OSDI | 6 |
| 2023 | Automatic Transformation Search Against Deep Leakage From GradientsabstractCollaborative learning has gained great popularity due to its benefit of data privacy protection: participants can jointly train a Deep Learning model without sharing their training sets. However, recent works discovered that an adversary can fully recover the sensitive training samples from the shared gradients. Such reconstruction attacks pose severe threats to collaborative learning. Hence, effective mitigation solutions are urgently desired. In this paper, we systematically analyze existing reconstruction attacks and propose to leverage data augmentation to defeat these attacks: by preprocessing sensitive images with carefully-selected transformation policies, it becomes infeasible for the adversary to extract training samples from the corresponding gradients. We first design two new metrics to quantify the impacts of transformations on data privacy and model usability. With the two metrics, we design a novel search method to automatically discover qualified policies from a given data augmentation library. Our defense method can be further combined with existing collaborative training systems without modifying the training protocols. We conduct comprehensive experiments on various system settings. Evaluation results demonstrate that the policies discovered by our method can defeat state-of-the-art reconstruction attacks in collaborative learning, with high efficiency and negligible impact on the model performance. Wei Gao 0064, Shangwei Guo, Tianwei Zhang 0004, Tao Xiang 0001, Han Qiu 0001, Yonggang Wen 0001, Yang Liu 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2023 | Learning Relation Prototype From Unlabeled Texts for Long-Tail Relation ExtractionabstractRelation Extraction (RE) is a vital step to complete Knowledge Graph (KG) by extracting entity relations from texts. However, it usually suffers from the long-tail issue. This paper proposes a novel approach to learn relation prototypes from unlabeled texts, to facilitate long-tail RE by transferring knowledge from relation types with sufficient training data. We learn relation prototypes as an implicit factor between entities, which reflects meanings of relations and their proximities. We construct a co-occurrence graph from texts, and capture both first-order and second-order entity proximities for embedding learning. By optimize the distance from entity pairs to corresponding prototypes, our method can be easily adapted to almost arbitrary RE frameworks. Thus, the learning of infrequent or even unseen relation types will benefit from semantically proximate relations through pairs of entities and large-scale textual information. Extensive experiments on two publicly available datasets present promising improvements (4.1% F1 on average). Ablation studies on long-tail relations, main components, and different RE models demonstrate the effectiveness of the learned relation prototypes. Finally, we analyze several example cases to give intuitive impressions as qualitative analysis. Our codes and data can be found in https://github.com/CrisJk/PA-TRP. Yixin Cao 0002, Jun Kuang, Ming Gao 0001, Aoying Zhou, Yonggang Wen 0001, Tat-Seng Chua |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2023 | A Review on Generative Adversarial Networks: Algorithms, Theory, and ApplicationsabstractGenerative adversarial networks (GANs) have recently become a hot research topic; however, they have been studied since 2014, and a large number of algorithms have been proposed. Nevertheless, few comprehensive studies explain the connections among different GAN variants and how they have evolved. In this paper, we attempt to provide a review of the various GAN methods from the perspectives of algorithms, theory, and applications. First, the motivations, mathematical representations, and structures of most GAN algorithms are introduced in detail, and we compare their commonalities and differences. Second, theoretical issues related to GANs are investigated. Finally, typical applications of GANs in image processing and computer vision, natural language processing, music, speech and audio, the medical field, and data science are discussed. Jie Gui, Zhenan Sun, Yonggang Wen 0001, Dacheng Tao, Jieping Ye |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2023 | Editorial
Yonggang Wen 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | Two-Stream Prototype Learning Network for Few-Shot Face Recognition Under OcclusionsabstractFew-shot face recognition under occlusion (FSFRO) aims to recognize novel subjects given only a few, probably occluded face images, and it is challenging and common in real-world scenarios. Unknown occlusions may deteriorate the class prototypes, while an occluded image in the support set may be critical for recognition if the query image is occluded. This motivates us to propose a novel Two-stream Prototype Learning Network (TSPLN) for FSFR under occlusions by simultaneously considering the quality of support images and their relevance to the query i mage. Specifically, we design a two-stream architecture, which mainly consists of a support-centered stream and query-centered stream, to learn the optimal class prototypes. The former stream is to reduce the negative impact of occluded images on the prototype. This is achieved by exploring the similarities between different images in the support set. In the query-centered stream, we exploit the relevance between the query and support set based on feature alignment (FA). We conduct extensive experiments on two popular datasets: CASIA-WebFace and RMFRD. The experimental results show that our proposed method achieves the state-of-the-art performance for occluded face recognition in the few-shot setting. Mengya Han, Yong Luo 0002, Han Hu 0003, Yonggang Wen 0001 |
IEEE Trans. Multim. | 5 |
| 2023 | Distributed Energy Trading and Scheduling Among Microgrids via Multiagent Reinforcement LearningabstractRenewable energy technologies empower microgrids to generate electricity to supply themselves and trade with others. Under this paradigm, microgrids have become autonomous entities that must intelligently determine their policies for energy trading and scheduling. Many factors influence a microgrid's decision-making, such as the complex microgrid infrastructure, the uncertain energy yield and demand, and the competition among the energy market players. These factors are usually hard to precisely model, and deriving the optimal policy for a microgrid is challenging. We propose a multiagent reinforcement learning (MARL) approach with an attention mechanism to learn the optimal policies for the microgrids without complex system modeling. We model each microgrid as an autonomous agent, which learns how to schedule energy resources and trade with others by collaborating with other agents. We adopt attention mechanism to enable intelligently selecting contextual information for the training of each agent. After training, an agent can make control decisions using only its local information, which can well preserve the microgrids' privacy and reduce the communication overhead among microgrids to facilitate distributed control. We implement a simulation environment and evaluate the performances of our proposed method using real-world datasets. The experimental results show that our method can significantly reduce the cost of the microgrids compared with the baseline methods. Guanyu Gao, Yonggang Wen 0001, Dacheng Tao |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Optimizing Performance of Federated Person Re-identification: Benchmarking and AnalysisabstractIncreasingly stringent data privacy regulations limit the development of person re-identification (ReID) because person ReID training requires centralizing an enormous amount of data that contains sensitive personal information. To address this problem, we introduce federated person re-identification ( FedReID )—implementing federated learning, an emerging distributed training method, to person ReID. FedReID preserves data privacy by aggregating model updates, instead of raw data, from clients to a central server. Furthermore, we optimize the performance of FedReID under statistical heterogeneity via benchmark analysis. We first construct a benchmark with an enhanced algorithm, two architectures, and nine person ReID datasets with large variances to simulate the real-world statistical heterogeneity. The benchmark results present insights and bottlenecks of FedReID under statistical heterogeneity, including challenges in convergence and poor performance on datasets with large volumes. Based on these insights, we propose three optimization approaches: (1) we adopt knowledge distillation to facilitate the convergence of FedReID by better transferring knowledge from clients to the server, (2) we introduce client clustering to improve the performance of large datasets by aggregating clients with similar data distributions, and (3) we propose cosine distance weight to elevate performance by dynamically updating the weights for aggregation depending on how well models are trained in clients. Extensive experiments demonstrate that these approaches achieve satisfying convergence with much better performance on all datasets. We believe that FedReID will shed light on implementing and optimizing federated learning on more computer vision applications. Weiming Zhuang, Xin Gan, Yonggang Wen 0001, Shuai Zhang 0004 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2023 | Optimizing Energy Efficiency for Data Center via Parameterized Deep Reinforcement LearningabstractThe rapid advancements in cloud computing, Big Data and their related applications have led to a skyrocketing increase in data center energy consumption year by year. The prior approaches for improving data center energy efficiency mostly suffer from high system dynamics or the complexity of data centers. In this paper, we propose an optimization framework based on deep reinforcement learning, named DeepEE, to jointly optimize energy consumption from the perspectives of task scheduling and cooling control. In DeepEE, a PArameterized action space based Deep Q-Network (PADQN) algorithm is proposed to tackle the hybrid action space problem. Then, a dynamic time factor mechanism for adjusting cooling control interval is introduced into PADQN (PADQN-D) to achieve more accurate and efficient coordination of IT and cooling subsystems. Finally, in order to train and evaluate the proposed algorithms safely and quickly, a simulation platform is built to model the dynamics of IT and cooling subsystems. Extensive real-trace based experiments illustrate that: 1) the proposed PADQN algorithm can save up to 15% and 10% energy consumption compared with the baseline siloed and joint optimization approaches respectively; 2) the proposed PADQN-D algorithm with dynamic cooling control interval can better adapt to the change of IT workload; 3) our proposed algorithms achieve more stable performance gain in terms of power consumption by adopting the parameterized action space. Yongyi Ran, Han Hu 0003, Yonggang Wen 0001, Xin Zhou 0003 |
IEEE Trans. Serv. Comput. | 3 |
| 2023 | Optimizing Data Center Energy Efficiency via Event-Driven Deep Reinforcement LearningabstractTo reduce the skyrocketing energy consumption of data centers, the prevailing approaches adopt the time-driven manner to control IT and cooling subsystems. These methods suffer from highly dynamic system states, complex action spaces and the risk of instability caused by frequent and unnecessary control operations. To tackle these problems, we propose a novel event-driven control paradigm and an optimization algorithm, under the deep reinforcement learning (DRL) framework. The principle is to make decisions based on certain critical events (e.g., overheating), rather than fixed periodic control. Specifically, we design an event-driven optimization framework to trigger control operations. Then, we present several models to describe IT and cooling subsystems, and mathematically define events to capture four types of prior factors that impact system performance. Furthermore, we develop an event-driven DRL (E-DRL) optimization algorithm to dispatch jobs and regulate cooling facilities for energy efficiency. Using two different types of real workload traces, we conduct extensive experiments to demonstrate that: 1) E-DRL reduces the number of regulating decisions by 70%$\sim$95% while achieving a comparable or even better energy efficiency in comparison with the state-of-the-art algorithm; and 2) E-DRL can adapt the control frequency to the changing operational conditions and diverse workloads. Yongyi Ran, Xin Zhou 0003, Han Hu 0003, Yonggang Wen 0001 |
IEEE Trans. Serv. Comput. | 4 |
| 2022 | Titan: a scheduler for foundation model fine-tuning workloadsabstractThe recent breakthrough of foundation model (FM) research raises a new trend to acquire efficient DL models by fine-tuning FMs with low-resource datasets. Current GPU clusters are mainly established to develop DL models by training from scratch. How to tailor a GPU cluster scheduler for FM fine-tuning workloads is still not explored. Wei Gao 0064, Peng Sun 0006, Yonggang Wen 0001, Tianwei Zhang 0004 |
SoCC | 3 |
| 2022 | Divergence-aware Federated Self-Supervised Learning
Weiming Zhuang, Yonggang Wen 0001, Shuai Zhang 0004 |
ICLR | 2 |
| 2022 | Federated Unsupervised Domain Adaptation for Face RecognitionabstractGiven labeled data in a source domain, unsupervised domain adaptation has been widely adopted to generalize models for unlabeled data in a target domain, whose data distributions are different. However, existing works are inapplicable to face recognition under privacy constraints because they re-quire sharing of sensitive face images between domains. To address this problem, we propose federated unsupervised do-main adaptation for face recognition, FedFR. FedFR jointly optimizes clustering-based domain adaptation and federated learning to elevate performance on the target domain. Specif-ically, for unlabeled data in the target domain, we enhance a clustering algorithm with distance constrain to improve the quality of predicted pseudo labels. Besides, we propose a new domain constraint loss (DCL) to regularize source do-main training in federated learning. Extensive experiments on a newly constructed benchmark demonstrate that FedFR outperforms the baseline and classic methods on the target domain by 3% to 14% on different evaluation metrics. Weiming Zhuang, Xin Gan, Xuesen Zhang, Yonggang Wen 0001, Shuai Zhang 0004, Shuai Yi |
ICME | 4 |
| 2022 | Optimizing Federated Unsupervised Person Re-identification via Camera-aware ClusteringabstractPerson re-identification (ReID) is a critical computer vision problem which identifies individuals from non-overlapping cameras. Many recent works on person ReID achieve remarkable performance by extracting features from large amounts of data using deep neural networks. However, the growing awareness of privacy concerns limits the development of person ReID. Prior studies employ federated person ReID to learn from decentralized edges without sharing raw data, but they overlook the variation of identities in different camera views. Concerning this issue, we propose a federated unsupervised person ReID (FedUCA) that leverages camera information to improve learning from decentralized unlabeled data. Specifically, FedUCA jointly learns person ReID models by transmitting training updates instead of raw data. We generate pseudo-labels for unlabeled local datasets on edges by clustering them into multiple groups according to different cameras. We then introduce contrastive learning with an intra-camera loss and an inter-camera loss to enhance the discrimination ability. In extensive experiments on eight person ReID datasets, our proposed approach significantly outperforms the state-of-the-art federated learning based method. It improves performance by 6% to 32% on these datasets, and notably by over 25 % on large datasets. We hope this paper will shed light on optimizing federated learning across a broader range of multimedia applications. Jiabei Liu, Weiming Zhuang, Yonggang Wen 0001, Jun Huang 0007, Wei Lin 0016 |
MMSP | 3 |
| 2022 | Nonlinear Multi-Model ReuseabstractThe goal of model reuse is to build a model in a new target domain by reusing some pre-trained source models. It can significantly reduce the training costs and the data required for training, and hence has various potential applications. Most of the existing model reuse approaches only reuse the output features or labels of the source model, and more information contained in the model are ignored. Besides, only a single model can be utilized in these approaches. A recently proposed multi-model reuse method is able to remedy these drawbacks by utilizing the hidden layer representations of multiple source models to help improve the representations in the target model, but it assumes that there are linear connections between the source and target models. This assumption is too restrictive and may be not valid in real-world applications. In this paper, we relax this assumption by introducing the manifold regularization scheme to exploit arbitrary nonlinear relationships between the source and target models. Effectiveness of our method is demonstrated empirically by the extensive experiments in the popular person re-identification task for smart city application. Yong Luo 0002, Ling-Yu Duan, Tongliang Liu, Yihang Lou, Yonggang Wen 0001 |
MMSP | 6 |
| 2022 | Optimizing Data Centre Energy Efficiency via Event Driven Deep Reinforcement Learningabstract[J1C2 Presentation Abstract at IEEE SERVICES 2022 for IEEE Transactions on Services Computing DOI 10.1109/TSC.2022.3157145] Yongyi Ran, Xin Zhou 0003, Han Hu 0003, Yonggang Wen 0001 |
SERVICES | 4 |
| 2022 | Primo: Practical Learning-Augmented Systems with Interpretable Models
Qinghao Hu 0004, Harsha Nori, Peng Sun 0006, Yonggang Wen 0001, Tianwei Zhang 0004 |
USENIX ATC | 4 |
| 2022 | FogChain: A Blockchain-Based Peer-to-Peer Solar Power Trading System Powered by Fog AIabstractMicrogrids, gaining traction from rising distributed generation for carbon reduction, demand novel solutions to regulate on- and off-grid operations, as well as both energy and monetary transfers between the microgrid and the central grid and among different microgrid participants. This research aims to develop and validate an intelligent microgrid management system to secure the competitiveness of Singapore’s energy market, by leveraging the inherent synergy between two emerging technologies, i.e., blockchain for Peer-to-Peer (P2P) solar power trading and fog computing for grid infrastructure management. For this vision, we have developed FogChain, an integrative, cost-effective, and scalable microgrid operating system (MGOS), consisting of three technical service layers: 1) a novel microgrid information infrastructure based on the fog-computing paradigm (i.e., intelligence on edge); 2) a blockchain-based microgrid service layer, providing smart contract and decentralized control capabilities for grid application development; and 3) a microgrid application layer (i.e., P2P energy trading) over the blockchain-based grid service. This MGOS would fundamentally transform how solar power is traded among participating electricity prosumers, leading to potentially new operational and business models. We have implemented the FogChain system and conducted extensive experiments to verify its performance advantages. Our results demonstrate that FogChain can efficiently process energy auction among 1000 participants with 1.1 s delay on average, reduce transmission cost up to 20% under the loss-aware trading mechanism, and reduce the solar yield prediction error to 0.11. Our system prototype suggests that FogChain provides a promising solution for efficient decentralized energy trading and intelligent distributed control for microgrids. Guanyu Gao, Chengru Song, T. G. Thusitha Asela Bandara, Meng Shen 0002, Fan Yang 0172, Wolf Posdorfer, Dacheng Tao, Yonggang Wen 0001 |
IEEE Internet Things J. | 8 |
| 2022 | EasyFL: A Low-Code Federated Learning Platform for DummiesabstractAcademia and industry have developed several platforms to support the popular privacy-preserving distributed learning method—federated learning (FL). However, these platforms are complex to use and require a deep understanding of FL, which imposes high barriers to entry for beginners, limits the productivity of researchers, and compromises deployment efficiency. In this article, we propose the first low-code FL platform,EasyFL, to enable users with various levels of expertise to experiment and prototype FL applications with little coding. We achieve this goal while ensuring great flexibility and extensibility for customization by unifying simple API design, modular design, and granular training flow abstraction. With only a few lines of code (LOC), EasyFL empowers them with many out-of-the-box functionalities to accelerate experimentation and deployment. These practical functionalities are heterogeneity simulation, comprehensive tracking, distributed training optimization, and seamless deployment. They are proposed based on challenges identified in the proposed FL life cycle. Compared with other platforms, EasyFL not only requires just three LOC (at least$10\times $lesser) to build a vanilla FL application but also incurs lower training overhead. Besides, our evaluations demonstrate that EasyFL expedites distributed training by$1.5\times $. It also improves the efficiency of deployment. We believe that EasyFL will increase the productivity of researchers and democratize FL to wider audiences. Weiming Zhuang, Xin Gan, Yonggang Wen 0001, Shuai Zhang 0004 |
IEEE Internet Things J. | 3 |
| 2022 | GradientFlow: Optimizing Network Performance for Large-Scale Distributed DNN TrainingabstractIt is important to scale out deep neural network (DNN) training for reducing model training time. The high communication overhead is one of the major performance bottlenecks for distributed DNN training across multiple GPUs. Our investigations have shown that popular open-source DNN systems could only achieve 2.5 speedup ratio on 64 GPUs connected by 56 Gbps network. To address this problem, we propose a communication backend named GradientFlow for distributed DNN training, and employ a set of network optimization techniques. First, we integrate ring-based allreduce, mixed-precision training, and computation/communication overlap into GradientFlow. Second, we propose lazy allreduce to improve network throughput by fusing multiple communication operations into a single one, and design coarse-grained sparse communication to reduce network traffic by only transmitting important gradient chunks. When training AlexNet and ResNet-50 on the ImageNet dataset using 512 GPUs, our approach could achieve 410.2 and 434.1 speedup ratio, respectively. Peng Sun 0006, Yonggang Wen 0001, Ruobing Han, Wansen Feng, Shengen Yan |
IEEE Trans. Big Data | 2 |
| 2021 | Are Missing Links Predictable? An Inferential Benchmark for Knowledge Graph CompletionabstractYixin Cao, Xiang Ji, Xin Lv, Juanzi Li, Yonggang Wen, Hanwang Zhang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yixin Cao 0002, Xiang Ji 0005, Juan-Zi Li, Yonggang Wen 0001, Hanwang Zhang |
ACL/IJCNLP (1) | 5 |
| 2021 | Chronus: A Novel Deadline-aware Scheduler for Deep Learning Training JobsabstractModern GPU clusters support Deep Learning training (DLT) jobs in a distributed manner. Job scheduling is the key to improve the training performance, resource utilization and fairness across users. Different training jobs may require various objectives and demands in terms of completion time. How to efficiently satisfy all these requirements is not extensively studied. Wei Gao 0064, Zhisheng Ye 0002, Peng Sun 0006, Yonggang Wen 0001, Tianwei Zhang 0004 |
SoCC | 4 |
| 2021 | Privacy-Preserving Collaborative Learning With Automatic Transformation SearchabstractCollaborative learning has gained great popularity due to its benefit of data privacy protection: participants can jointly train a Deep Learning model without sharing their training sets. However, recent works discovered that an adversary can fully recover the sensitive training samples from the shared gradients. Such reconstruction attacks pose severe threats to collaborative learning. Hence, effective mitigation solutions are urgently desired.In this paper, we propose to leverage data augmentation to defeat reconstruction attacks: by preprocessing sensitive images with carefully-selected transformation policies, it becomes infeasible for the adversary to extract any useful information from the corresponding gradients. We design a novel search method to automatically discover qualified policies. We adopt two new metrics to quantify the impacts of transformations on data privacy and model usability, which can significantly accelerate the search speed. Comprehensive evaluations demonstrate that the policies discovered by our method can defeat existing reconstruction attacks in collaborative learning, with high efficiency and negligible impact on the model performance. Wei Gao 0064, Shangwei Guo, Tianwei Zhang 0004, Han Qiu 0001, Yonggang Wen 0001, Yang Liu 0003 |
CVPR | 5 |
| 2021 | Towards Cost-Optimal Energy Procurement for Cooling as a Service: A Data-Driven ApproachabstractCoolingas a Service (CaaS) is an emerging business that provides air conditioning services for buildings. With the rapid development of the business and the continuous increase of energy load, CaaS providers need cost-effective energy procurement to meet the service requirements. In this paper, we propose a data-driven approach for energy procurement for CaaS providers. First, we focus on two dominant variables of cooling energy cost, including outdoor temperature and electricity price. We predict their trends in the next day and accordingly, we estimate the energy usage and purchase the energy in the day-ahead energy market, one day before the actual usage. During the real-time operation, we can obtain the actual temperature and price, and we use this information to adjust the quality of service without violating the service standards. The adjustment serves as the demand response to the real-time energy market and can be cost beneficial. We conducted experimental studies to verify the performance of the proposed solution. The results show that our solution provides high-quality cooling services with minimal energy expenditure and helps improve the stability of the power grid. Wei Zhang 0082, Yonggang Wen 0001, Fang Liu 0009 |
GLOBECOM | 2 |
| 2021 | Collaborative Unsupervised Visual Representation Learning from Decentralized DataabstractUnsupervised representation learning has achieved outstanding performances using centralized data available on the Internet. However, the increasing awareness of privacy protection limits sharing of decentralized unlabeled image data that grows explosively in multiple parties (e.g., mobile phones and cameras). As such, a natural problem is how to leverage these data to learn visual representations for downstream tasks while preserving data privacy. To address this problem, we propose a novel federated unsupervised learning framework, FedU. In this framework, each party trains models from unlabeled data independently using contrastive learning with an online network and a target network. Then, a central server aggregates trained models and updates clients’ models with the aggregated model. It preserves data privacy as each party only has access to its raw data. Decentralized data among multiple parties are normally non-independent and identically distributed (non-IID), leading to performance degradation. To tackle this challenge, we propose two simple but effective methods: 1) We design the communication protocol to upload only the encoders of online networks for server aggregation and update them with the aggregated encoder; 2) We introduce a new module to dynamically decide how to update predictors based on the divergence caused by non-IID. The predictor is the other component of the online network. Extensive experiments and ablations demonstrate the effectiveness and significance of FedU. It outperforms training with only one party by over 5% and other methods by over 14% in linear and semi-supervised evaluation on non-IID data. Weiming Zhuang, Xin Gan, Yonggang Wen 0001, Shuai Zhang 0004, Shuai Yi |
ICCV | 3 |
| 2021 | Data-Free Ensemble Knowledge Distillation for Privacy-conscious Multimedia Model CompressionabstractRecent advances in deep learning bring impressive performance for multimedia applications. Hence, compressing and deploying these applications on resource-limited edge devices via model compression becomes attractive. Knowledge distillation (KD) is one of the most popular model compression techniques. However, most well-behaved KD approaches require the original dataset, which is usually unavailable due to privacy issues, while existing data-free KD methods perform much worse than data-required counterparts. In this paper, we analyze previous data-free KD methods from the data perspective and point out that using a single pre-trained model limits the performance of these approaches. We then propose a Data-Free Ensemble knowledge Distillation (DFED) framework, which contains a student network, a generator network, and multiple pre-trained teacher networks. During training, the student mimics behaviors of the ensemble of teachers using samples synthesized by a generator, which aims to enlarge the prediction discrepancy between the student and teachers. A moment matching loss term assists the generator training by minimizing the distance between activations of synthesized samples and real samples. We evaluate DFED on three popular image classification datasets. Results demonstrate that our method achieves significant performance improvements compared with previous works. We also design an ablation study to verify the effectiveness of each component of the proposed framework. Zhiwei Hao 0001, Yong Luo 0002, Han Hu 0003, Jianping An, Yonggang Wen 0001 |
ACM Multimedia | 5 |
| 2021 | Missing Data Imputation for Solar Yield Prediction using Temporal Multi-Modal Variational Auto-EncoderabstractThe accurate and robust prediction of short-term solar power generation is significant for the management of modern smart grids, where solar power has become a major energy source due to its green and economical nature. However, the solar yield prediction can be difficult to conduct in the real world where hardware and network issues can make the sensors unreachable. Such data missing problem is so prevalent that it degrades the performance of deployed prediction models and even fails the model execution. In this paper, we propose a novel temporal multi-modal variational auto-encoder (TMMVAE) model, to enhance the robustness of short-term solar power yield prediction with missing data. It can impute the missing values in time-series sensor data, and reconstruct them by consolidating multi-modality data, which then facilitates more accurate solar power yield prediction. TMMVAE can be deployed efficiently with an end-to-end framework. The framework is verified at our real-world testbed on campus. The results of extensive experiments show that our proposed framework can significantly improve the imputation accuracy when the inference data is severely corrupted, and can hence dramatically improve the robustness of short-term solar energy yield forecasting. Meng Shen 0002, Huaizheng Zhang, Yixin Cao 0002, Fan Yang 0172, Yonggang Wen 0001 |
ACM Multimedia | 5 |
| 2021 | Exploring Sequence Feature Alignment for Domain Adaptive Detection TransformersabstractDetection transformers have recently shown promising object detection results and attracted increasing attention. However, how to develop effective domain adaptation techniques to improve its cross-domain performance remains unexplored and unclear. In this paper, we delve into this topic and empirically find that direct feature distribution alignment on the CNN backbone only brings limited improvements, as it does not guarantee domain-invariant sequence features in the transformer for prediction. To address this issue, we propose a novel Sequence Feature Alignment (SFA) method that is specially designed for the adaptation of detection transformers. Technically, SFA consists of a domain query-based feature alignment (DQFA) module and a token-wise feature alignment (TDA) module. In DQFA, a novel domain query is used to aggregate and align global context from the token sequence of both domains. DQFA reduces the domain discrepancy in global feature representations and object relations when deploying in the transformer encoder and decoder, respectively. Meanwhile, TDA aligns token features in the sequence from both domains, which reduces the domain gaps in local and instance-level feature representations in the transformer encoder and decoder, respectively. Besides, a novel bipartite matching consistency loss is proposed to enhance the feature discriminability for robust object detection. Experiments on three challenging benchmarks show that SFA outperforms state-of-the-art domain adaptive object detection methods. Code has been made available at: https://github.com/encounter1997/SFA. Wen Wang 0009, Yang Cao 0010, Jing Zhang 0037, Fengxiang He, Zhengjun Zha, Yonggang Wen 0001, Dacheng Tao |
ACM Multimedia | 6 |
| 2021 | Joint Optimization in Edge-Cloud Continuum for Federated Unsupervised Person Re-identificationabstractPerson re-identification (ReID) aims to re-identify a person from non-overlapping camera views. Since person ReID data contains sensitive personal information, researchers have adopted federated learning, an emerging distributed training method, to mitigate the privacy leakage risks. However, existing studies rely on data labels that are laborious and time-consuming to obtain. We present FedUReID, a federated unsupervised person ReID system to learn person ReID models without any labels while preserving privacy. FedUReID enables in-situ model training on edges with unlabeled data. A cloud server aggregates models from edges instead of centralizing raw data to preserve data privacy. Moreover, to tackle the problem that edges vary in data volumes and distributions, we personalize training in edges with joint optimization of cloud and edge. Specifically, we propose personalized epoch to reassign computation throughout training, personalized clustering to iteratively predict suitable labels for unlabeled data, and personalized update to adapt the server aggregated model to each edge. Extensive experiments on eight person ReID datasets demonstrate that FedUReID not only achieves higher accuracy but also reduces computation cost by 29%. Our FedUReID system with the joint optimization will shed light on implementing federated learning to more multimedia tasks without data labels. Weiming Zhuang, Yonggang Wen 0001, Shuai Zhang 0004 |
ACM Multimedia | 2 |
| 2021 | Characterization and prediction of deep learning workloads in large-scale GPU datacentersabstractModern GPU datacenters are critical for delivering Deep Learning (DL) models and services in both the research community and industry. When operating a datacenter, optimization of resource scheduling and management can bring significant financial benefits. Achieving this goal requires a deep understanding of the job features and user behaviors. We present a comprehensive study about the characteristics of DL jobs and resource management. First, we perform a large-scale analysis of real-world job traces from SenseTime. We uncover some interesting conclusions from the perspectives of clusters, jobs and users, which can facilitate the cluster system designs. Second, we introduce a general-purpose framework, which manages resources based on historical data. As case studies, we design (1) a Quasi-Shortest-Service-First scheduling service, which can minimize the cluster-wide average job completion time by up to 6.5×; (2) a Cluster Energy Saving service, which improves overall cluster utilization by up to 13%. Qinghao Hu 0004, Peng Sun 0006, Shengen Yan, Yonggang Wen 0001, Tianwei Zhang 0004 |
SC | 4 |
| 2021 | Toward Intelligent Multizone Thermal Control With Multiagent Deep Reinforcement LearningabstractEnergy usage and thermal comfort are the pillars of smart buildings. Many research works have been proposed to save energy while maintaining a comfortable thermal condition. However, most of them either make the oversimplified assumption on thermal comfort with unsatisfied comfort performance or deal with the single-zone thermal control only with limited practical impact. A few preliminary pieces of research on multizone control are available, but they fail to keep pace with the latest advancements in the deep-learning-based control techniques. In this article, we investigate the multizone thermal control with optimized energy usage and canonical thermal comfort modeling. We adopt the emerging multiagent deep reinforcement learning techniques and propose to model each zone as an agent. A multiagent framework is established to support the information exchange among the agents and enable intelligent thermal control in the heterogeneous zones. Accordingly, we mathematically formulate a problem to optimize both energy and comfort. A multizone thermal control algorithm (MOCA) is proposed to solve the problem by deriving optimal control policies. We validate the performance of MOCA through simulation in professional TRNSYS, configured based on our real-world laboratory. The results are promising with up to 15.4% energy saving as well as satisfied thermal comfort in different zones. Jie Li 0043, Wei Zhang 0082, Guanyu Gao, Yonggang Wen 0001, Guangyu Jin, George I. Christopoulos |
IEEE Internet Things J. | 4 |
| 2021 | Demystifying Thermal Comfort in Smart Buildings: An Interpretable Machine Learning ApproachabstractThermal comfort is a key consideration in smart buildings and a number of comfort models are available nowadays to evaluate the comfort level of occupants. However, the models are often complex and hardly interpretable for the developers and operators. Indeed, the model interpretations are beneficial in multifold such as for system inspection and optimization. In this article, we propose an interpretable thermal comfort system to introduce interpretability to any black-box comfort models. First, we focus on the relationship between a model's input features and output comfort level. The feature impact on comfort is investigated and the impact patterns are shown to be diverse for different features. Second, we unveil the model mechanisms about the data processing inside the model by building the model surrogates based on the interpretable machine learning algorithms. The surrogates offer outstanding fidelity for simulating the actual model mechanisms and the interpretations based on the surrogates are intuitive and informative. Our interpretable comfort system can be integrated with the existing building management systems. Accordingly, we can ease building owner's concerns about adopting new black-box technologies and enable various smart building applications like smart energy management. Wei Zhang 0082, Yonggang Wen 0001, King-Jet Tseng, Guangyu Jin |
IEEE Internet Things J. | 2 |
| 2021 | Cost Optimal Data Center Servers: A Voltage Scaling ApproachabstractData centers have experienced dramatic growth in recent years in order to meet the ever-increasing demand for computing. As a result, minimizing the electrical cost to operate data centers has become a crucial issue. In this paper, we observe that electricity prices change over time, and that we can take advantage of periods with low prices by scaling up processor speeds to perform more work, while scaling down speeds during high price periods to reduce cost. We apply this observation to several settings. First, we consider an offline setting which assumes future electricity prices are given, and propose an efficient algorithm for optimally scaling a processor's speed in order to minimize the total electrical cost for completing a task by a deadline. We then consider a more realistic stochastic setting in which future prices are not known, but vary according to a Markov model. We present another efficient algorithm for minimizing the expected cost to meet a deadline. We performed a number of experiments using real electricity price traces to test the performance of our algorithms. We show that our stochastic algorithm is light-weight and relies only on easily obtainable price data, but that it achieves excellent performance, with only a 1 percent cost difference on average from the optimal offline algorithm. In addition, the stochastic algorithm significantly reduced costs compared to several candidate algorithms. Wei Zhang 0082, Yonggang Wen 0001, Loi Lei Lai, Fang Liu 0009, Rui Fan 0004 |
IEEE Trans. Cloud Comput. | 2 |
| 2021 | Joint Deep Multi-View Learning for Image ClusteringabstractIn this paper, a novelDeepMulti-viewJointClustering (DMJC) framework is proposed, where multiple deep embedded features, multi-view fusion mechanism, and clustering assignments can be learned simultaneously. Through the joint learning strategy, the clustering-friendly multi-view features and useful multi-view complementary information can be exploited effectively to improve the clustering performance. Under the proposed joint learning framework, we design two ingenious variants of deep multi-view joint clustering models, whose multi-view fusion is implemented by two kinds of simple yet effective schemes. The first model, called DMJC-S, performs multi-view fusion in an implicit way via a novel multi-view soft assignment distribution. The second model, termed DMJC-T, defines a novel multi-view auxiliary target distribution to conduct the multi-view fusion explicitly. Both DMJC-S and DMJC-T are optimized under a KL divergence objective. Experiments on eight challenging image datasets demonstrate the superiority of both DMJC-S and DMJC-T over single/multi-view baselines and the state-of-the-art multi-view clustering methods, which proves the effectiveness of the proposed DMJC framework. To the best of our knowledge, this is the first work to model the multi-view clustering in a deep joint framework, which will provide a meaningful thinking in unsupervised multi-view learning. Yuan Xie 0006, Bingqian Lin, Yanyun Qu, Cuihua Li, Wensheng Zhang 0002, Lizhuang Ma, Yonggang Wen 0001, Dacheng Tao |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2021 | Intelligent Trainer for Dyna-Style Model-Based Deep Reinforcement LearningabstractModel-based reinforcement learning (MBRL) has been proposed as a promising alternative solution to tackle the high sampling cost challenge in the canonical RL, by leveraging a system dynamics model to generate synthetic data for policy training purpose. The MBRL framework, nevertheless, is inherently limited by the convoluted process of jointly optimizing control policy, learning system dynamics, and sampling data from two sources controlled by complicated hyperparameters. As such, the training process involves overwhelmingly manual tuning and is prohibitively costly. In this research, we propose a "reinforcement on reinforcement" (RoR) architecture to decompose the convoluted tasks into two decoupled layers of RL. The inner layer is the canonical MBRL training process which is formulated as a Markov decision process, called training process environment (TPE). The outer layer serves as an RL agent, called intelligent trainer, to learn an optimal hyperparameter configuration for the inner TPE. This decomposition approach provides much-needed flexibility to implement different trainer designs, referred to "train the trainer." In our research, we propose and optimize two alternative trainer designs: 1) an unihead trainer and 2) a multihead trainer. Our proposed RoR framework is evaluated for five tasks in the OpenAI gym. Compared with three other baseline methods, our proposed intelligent trainer methods have a competitive performance in autotuning capability, with up to 56% expected sampling cost saving without knowing the best parameter configurations in advance. The proposed trainer framework can be easily extended to tasks that require costly hyperparameter tuning. Linsen Dong, Xin Zhou 0003, Yonggang Wen 0001, Kyle Guan |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2021 | Deep Reinforcement Learning for Tropical Air Free-cooled Data Center ControlabstractAir free-cooled data centers (DCs) have not existed in the tropical zone due to the unique challenges of year-round high ambient temperature and relative humidity (RH). The increasing availability of servers that can tolerate higher temperatures and RH due to the regulatory bodies’ prompts to raise DC temperature setpoints sheds light upon the feasibility of air free-cooled DCs in the tropics. However, due to the complex psychrometric dynamics, operating the air free-cooled DC in the tropics generally requires adaptive control of supply air condition to maintain the computing performance and reliability of the servers. This article studies the problem of controlling the supply air temperature and RH in a free-cooled tropical DC below certain thresholds. To achieve the goal, we formulate the control problem as Markov decision processes and apply deep reinforcement learning (DRL) to learn the control policy that minimizes the cooling energy while satisfying the requirements on the supply air temperature and RH. We also develop a constrained DRL solution for performance improvements. Extensive evaluation based on real data traces collected from an air free-cooled testbed and comparisons among the unconstrained and constrained DRL approaches as well as two other baseline approaches show the superior performance of our proposed solutions. Duc Van Le, Rui Tan 0001, Yew-Wah Wong, Yonggang Wen 0001 |
ACM Trans. Sens. Networks | 6 |
| 2020 | Deep Heterogeneous Multi-Task Metric Learning for Visual Recognition and RetrievalabstractHow to estimate the distance between data instances is a fundamental problem in many artificial intelligence algorithms, and critical in diverse multimedia applications. A major challenge in the estimation is how to find an appropriate distance function when labeled data are insufficient for a certain task. Multi-task metric learning (MTML) is able to alleviate such data deficiency issue by learning distance metrics for multiple tasks together and sharing information between the different tasks. Recently, heterogeneous MTML (HMTML) has attracted much attention since it can handle multiple tasks with varied data representations. A major drawback of the current HMTML approaches is that only linear transformations are learned to connect different domains. This is suboptimal since the correlations between different domains may be very complex and highly nonlinear. To overcome this drawback, we propose a deep heterogeneous MTML (DHMTML) method, in which a nonlinear mapping is learned for each task by using a deep neural network. The correlations of different domains are exploited by sharing some parameters at the top layers of different networks. More importantly, the auto-encoder scheme and the adversarial learning mechanism are integrated and incorporated to help exploit the feature correlations in and between different tasks and the specific properties are preserved by learning additional task-specific layers together with the common layers. Experiments demonstrated that the proposed method outperforms single-task deep metric learning algorithms and other HMTML approaches consistently on several benchmark datasets. Shikang Gan, Yong Luo 0002, Yonggang Wen 0001, Tongliang Liu, Han Hu 0003 |
ACM Multimedia | 3 |
| 2020 | Semi-supervised Online Multi-Task Metric Learning for Visual Recognition and RetrievalabstractDistance metric learning (DML) is critial in many multimedia application tasks. However, it is hard to learn a satisfactory distance metric given only a few labeled samples for each task. In this paper, we proposed a novel semi-supervised online multi-task DML method termed SOMTML, which enables the models describing different tasks to help each other during the metric learning procedure and thus improving their respective performance. Besides, unlabeled data are leveraged to further help alleviate the data deficiency issue in different tasks by designing a novel regularization term, which also allows some prior information to be incorporated. More importantly, a quite efficient algorithm is developed to update the metrics of all tasks adaptively. The proposed SOMTML is experimentally validated in two popular visual analytic-based applications: handwriting digits recognition and face retrieval. We compared the proposed method with competitive single-task and multi-task metric learning approaches. Extensive experimental results demonstrate the effectiveness and efficiency of the proposed SOMTML. Yangxi Li, Han Hu 0003, Jin Li 0014, Yong Luo 0002, Yonggang Wen 0001 |
ACM Multimedia | 5 |
| 2020 | Look, Read and Feel: Benchmarking Ads Understanding with Multimodal Multitask LearningabstractGiven the massive market of advertising and the sharply increasing online multimedia content (such as videos), it is now fashionable to promote advertisements (ads) together with the multimedia content. However, manually finding relevant ads to match the provided content is labor-intensive, and hence some automatic advertising techniques are developed. Since ads are usually hard to understand only according to its visual appearance due to the contained visual metaphor, some other modalities, such as the contained texts, should be exploited for understanding. To further improve user experience, it is necessary to understand both the ads' topic and sentiment. This motivates us to develop a novel deep multimodal multitask framework that integrates multiple modalities to achieve effective topic and sentiment prediction simultaneously for ads understanding. In particular, in our framework termed Deep$M^2$Ad, we first extract multimodal information from ads and learn high-level and comparable representations. The visual metaphor of the ad is decoded in an unsupervised manner. The obtained representations are then fed into the proposed hierarchical multimodal attention modules to learn task-specific representations for final prediction. A multitask loss function is also designed to jointly train both the topic and sentiment prediction models in an end-to-end manner, where bottom-layer parameters are shared to alleviate over-fitting. We conduct extensive experiments on a large-scale advertisement dataset and achieve state-of-the-art performance for both prediction tasks. The obtained results could be utilized as a benchmark for ads understanding. Huaizheng Zhang, Yong Luo 0002, Qiming Ai, Yonggang Wen 0001, Han Hu 0003 |
ACM Multimedia | 4 |
| 2020 | Hysia: Serving DNN-Based Video-to-Retail Applications in CloudabstractCombining video streaming and online retailing (V2R) has been a growing trend recently. In this paper, we provide practitioners and researchers in multimedia with a cloud-based platform named Hysia for easy development and deployment of V2R applications. The system consists of: 1) a back-end infrastructure providing optimized V2R related services including data engine, model repository, model serving and content matching; and 2) an application layer which enables rapid V2R application prototyping. Hysia addresses industry and academic needs in large-scale multimedia by: 1) seamlessly integrating state-of-the-art libraries including NVIDIA video SDK, Facebook faiss, and gRPC; 2) efficiently utilizing GPU computation; and 3) allowing developers to bind new models easily to meet the rapidly changing deep learning (DL) techniques. On top of that, we implement an orchestrator for further optimizing DL model serving performance. Hysia has been released as an open source project on GitHub, and attracted considerable attention. We have published Hysia to DockerHub as an official image for seamless integration and deployment in current cloud environments. Huaizheng Zhang, Yuanming Li, Qiming Ai, Yong Luo 0002, Yonggang Wen 0001, Yichao Jin 0002, Ta Nguyen Binh Duong |
ACM Multimedia | 5 |
| 2020 | MLModelCI: An Automatic Cloud Platform for Efficient MLaaSabstractMLModelCI provides multimedia researchers and developers with a one-stop platform for efficient machine learning (ML) services. The system leverages DevOps techniques to optimize, test, and manage models. It also containerizes and deploys these optimized and validated models as cloud services (MLaaS). In its essence, MLModelCI serves as a housekeeper to help users publish models. The models are first automatically converted to optimized formats for production purpose and then profiled under different settings (e.g., batch size and hardware). The profiling information can be used as guidelines for balancing the trade-off between performance and cost of MLaaS. Finally, the system dockerizes the models for ease of deployment to cloud environments. A key feature of MLModelCI is the implementation of a controller, which allows elastic evaluation which only utilizes idle workers while maintaining online service quality. Our system bridges the gap between current ML training and serving systems and thus free developers from manual and tedious work often associated with service deployment. We release the platform as an open-source project on GitHub under Apache 2.0 license, with the aim that it will facilitate and streamline more large-scale ML applications and research projects. Huaizheng Zhang, Yuanming Li, Yizheng Huang 0001, Yonggang Wen 0001, Jianxiong Yin, Kyle Guan |
ACM Multimedia | 4 |
| 2020 | Performance Optimization of Federated Person Re-identification via Benchmark AnalysisabstractFederated learning is a privacy-preserving machine learning technique that learns a shared model across decentralized clients. It can alleviate privacy concerns of personal re-identification, an important computer vision task. In this work, we implement federated learning to person re-identification (FedReID) and optimize its performance affected by statistical heterogeneity in the real-world scenario. We first construct a new benchmark to investigate the performance of FedReID. This benchmark consists of (1) nine datasets with different volumes sourced from different domains to simulate the heterogeneous situation in reality, (2) two federated scenarios, and (3) an enhanced federated algorithm for FedReID. The benchmark analysis shows that the client-edge-cloud architecture, represented by the federated-by-dataset scenario, has better performance than client-server architecture in FedReID. It also reveals the bottlenecks of FedReID under the real-world scenario, including poor performance of large datasets caused by unbalanced weights in model aggregation and challenges in convergence. Then we propose two optimization methods: (1) To address the unbalanced weight problem, we propose a new method to dynamically change the weights according to the scale of model changes in clients in each training round; (2) To facilitate convergence, we adopt knowledge distillation to refine the server model with knowledge generated from client models on a public dataset. Experiment results demonstrate that our strategies can achieve much better convergence with superior performance on all datasets. We believe that our work will inspire the community to further explore the implementation of federated learning on more computer vision tasks in real-world scenarios. Weiming Zhuang, Yonggang Wen 0001, Xuesen Zhang, Xin Gan, Daiying Yin, Dongzhan Zhou, Shuai Zhang 0004, Shuai Yi |
ACM Multimedia | 2 |
| 2020 | Deep neural networks for emerging multimedia computing and applications
Shuqiang Jiang, Weiqing Min, Yonggang Wen 0001, Qingming Huang, Shuicheng Yan |
Neurocomputing | 3 |
| 2020 | DeepComfort: Energy-Efficient Thermal Comfort Control in Buildings Via Reinforcement LearningabstractHeating, ventilation, and air conditioning (HVAC) are extremely energy consuming, accounting for 40% of total building energy consumption. It is crucial to design some energy-efficient building thermal comfort control strategy which can reduce the energy consumption of the HVAC while maintaining the comfort of the occupants. However, implementing such a strategy is challenging, because the changes of the thermal states in a building environment are influenced by various factors. The relationships among these influencing factors are hard to model and are always different in different building environments. To address this challenge, we propose a deep-reinforcement-learning-based framework, DeepComfort, for thermal comfort control in buildings. We formulate the thermal comfort control as a cost-minimization problem by jointly considering the energy consumption of the HVAC and the occupants' thermal comfort. We first design a deep feedforward neural network (FNN)-based approach for predicting the occupants' thermal comfort and then propose a deep deterministic policy gradients (DDPGs)-based approach for learning the optimal thermal comfort control policy. We implement a building thermal comfort control simulation environment and evaluate the performance under various settings. The experimental results show that our approaches can improve the performance of thermal comfort prediction by 14.5% and reduce the energy consumption of HVAC by 4.31% while improving the occupants' thermal comfort by 13.6%. Guanyu Gao, Jie Li 0043, Yonggang Wen 0001 |
IEEE Internet Things J. | 3 |
| 2020 | Transforming Device Fingerprinting for Wireless Security via Online Multitask Metric LearningabstractDevice fingerprinting is a crucial part in the Internet of Things applications. Existing device-fingerprinting solutions either ignore the influence of the type of network traffic or separately learn a fingerprinting model for each traffic type. This often leads to suboptimal solutions, especially when training data are limited. Considering that the data distributions of different traffic types may be different but related, we propose a novel multitask learning method to learn the fingerprinting models for several traffic types simultaneously. Specifically, we first design a system for device fingerprinting using the popular k-nearest neighbor (KNN) approach. Then, a novel distance metric learning (DML) algorithm termed online multitask metric learning (OMTML) is developed to improve the distance estimation in our system. OMTML enables the models describing different traffic types to help each other during the metric learning procedure, and thus improving their respective accuracies. OMTML can also be updated adaptively, and the updating process is efficient. The experimental results show that the proposed KNN-based system outperforms the artificial neural network (ANN)-based counterpart significantly. Besides, the comparisons of our OMTML and other representative online and multitask DML approaches demonstrate both effectiveness and efficiency of the proposed metric learning method. Yong Luo 0002, Han Hu 0003, Yonggang Wen 0001, Dacheng Tao |
IEEE Internet Things J. | 3 |
| 2020 | Guest Editorial Special Issue on Advances in Artificial Intelligence and Machine Learning for Networkingabstracthttps://www.youtube.com/watch?v=SQmgSOi5oos Prosper Chemouil, Pan Hui 0001, Wolfgang Kellerer, Noura Limam, Rolf Stadler, Yonggang Wen 0001 |
IEEE J. Sel. Areas Commun. | 6 |
| 2020 | GraphMP: I/O-Efficient Big Graph Analytics on a Single Commodity MachineabstractRecent studies showed that single-machine graph processing systems can be as highly competitive as cluster-based approaches on large-scale problems. While several out-of-core graph processing systems and computation models have been proposed, the high disk I/O overhead could significantly reduce performance in many practical cases. In this paper, we propose GraphMP to tackle big graph analytics on a single machine. GraphMP achieves low disk I/O overhead with three techniques. First, we design a vertex-centric sliding window (VSW) computation model to avoid reading and writing vertices on disk. Second, we propose a selective scheduling method to skip loading and processing unnecessary edge shards on disk. Third, we use a compressed edge cache mechanism to fully utilize the available memory of a machine to reduce the amount of disk accesses for edges. Extensive evaluations have shown that GraphMP could outperform existing single-machine out-of-core systems such as GraphChi, X-Stream and GridGraph by up to 30, and can be as highly competitive as distributed graph engines like Pregel+, PowerGraph and Chaos. Peng Sun 0006, Yonggang Wen 0001, Ta Nguyen Binh Duong, Xiaokui Xiao |
IEEE Trans. Big Data | 2 |
| 2020 | Transforming Cooling Optimization for Green Data Center via Deep Reinforcement LearningabstractData center (DC) plays an important role to support services, such as e-commerce and cloud computing. The resulting energy consumption from this growing market has drawn significant attention, and noticeably almost half of the energy cost is used to cool the DC to a particular temperature. It is thus an critical operational challenge to curb the cooling energy cost without sacrificing the thermal safety of a DC. The existing solutions typically follow a two-step approach, in which the system is first modeled based on expert knowledge and, thus, the operational actions are determined with heuristics and/or best practices. These approaches are often hard to generalize and might result in suboptimal performances due to intrinsic model errors for large-scale systems. In this paper, we propose optimizing the DC cooling control via the emerging deep reinforcement learning (DRL) framework. Compared to the existing approaches, our solution lends itself an end-to-end cooling control algorithm (CCA) via an off-policy offline version of the deep deterministic policy gradient (DDPG) algorithm, in which an evaluation network is trained to predict the DC energy cost along with resulting cooling effects, and a policy network is trained to gauge optimized control settings. Moreover, we introduce a de-underestimation (DUE) validation mechanism for the critic network to reduce the potential underestimation of the risk caused by neural approximation. Our proposed algorithm is evaluated on an EnergyPlus simulation platform and on a real data trace collected from the National Super Computing Centre (NSCC) of Singapore. The resulting numerical results show that the proposed CCA can achieve up to 11% cooling cost reduction on the simulation platform compared with a manually configured baseline control algorithm. In the trace-based study of conservative nature, the proposed algorithm can achieve about 15% cooling energy savings on the NSCC data trace. Our pioneering approach can shed new light on the application of DRL to optimize and automate DC operations and management, potentially revolutionizing digital infrastructure management with intelligence. Yonggang Wen 0001, Dacheng Tao, Kyle Guan |
IEEE Trans. Cybern. | 2 |
| 2020 | DeepQoE: A Multimodal Learning Framework for Video Quality of Experience (QoE) PredictionabstractRecently, many models have been developed to predict video Quality of Experience (QoE), yet the applicability of these models still faces significant challenges. Firstly, many models rely on features that are unique to a specific dataset and thus lack the capability to generalize. Due to the intricate interactions among these features, a unified representation that is independent of datasets with different modalities is needed. Secondly, existing models often lack the configurability to perform both classification and regression tasks. Thirdly, the sample size of the available datasets to develop these models is often very small, and the impact of limited data on the performance of QoE models has not been adequately addressed. To address these issues, in this work we develop a novel and end-to-end framework termed as DeepQoE. The proposed framework first uses a combination of deep learning techniques, such as word embedding and 3D convolutional neural network (C3D), to extract generalized features. Next, these features are combined and fed into a neural network for representation learning. A learned representation will then serve as input for classification or regression tasks. We evaluate the performance of DeepQoE with three datasets. The results show that for small datasets (e.g., WHU-MVQoE2016 and Live-Netflix Video Database), the performance of state-of-the-art machine learning algorithms is greatly improved by using the QoE representation from DeepQoE (e.g., 35.71% to 44.82%); while for the large dataset (e.g., VideoSet), our DeepQoE framework achieves significant performance improvement in comparison to the best baseline method (90.94% vs. 82.84%). In addition to the much improved performance, DeepQoE has the flexibility to fit different datasets, to learn QoE representation, and to perform both classification and regression problems. We also develop a DeepQoE based adaptive bitrate streaming (ABR) system to verify that our framework can be easily applied to multimedia communication service. The software package of the DeepQoE framework has been released to facilitate the current research on QoE. Huaizheng Zhang, Linsen Dong, Guanyu Gao, Han Hu 0003, Yonggang Wen 0001, Kyle Guan |
IEEE Trans. Multim. | 5 |
| 2020 | Efficient Compute-Intensive Job Allocation in Data Centers via Deep Reinforcement LearningabstractReducing the energy consumption of the servers in a data center via proper job allocation is desirable. Existing advanced job allocation algorithms, based on constrained optimization formulations capturing servers' complex power consumption and thermal dynamics, often scale poorly with the data center size and optimization horizon. This article applies deep reinforcement learning to build an allocation algorithm for long-lasting and compute-intensive jobs that are increasingly seen among today's computation demands. Specifically, a deep Q-network is trained to allocate jobs, aiming to maximize a cumulative reward over long horizons. The training is performed offline using a computational model based on long short-term memory networks that capture the servers' power and thermal dynamics. This offline training approach avoids slow online convergence, low energy efficiency, and potential server overheating during the agent's extensive state-action space exploration if it directly interacts with the physical data center in the usually adopted online learning scheme. At run time, the trained Q-network is forward-propagated with little computation to allocate jobs. Evaluation based on eight months' physical state and job arrival records from a national supercomputing data center hosting 1,152 processors shows that our solution reduces computing power consumption by more than 10 percent and processor temperature by more than 4°C without sacrificing job processing throughput. Deliang Yi, Xin Zhou 0003, Yonggang Wen 0001, Rui Tan 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2020 | Coordinating Workload Scheduling of Geo-Distributed Data Centers and Electricity Generation of Smart GridabstractWith the rapidly increasing computing demand, data centers become more and more power-hungry, which incurs substantial electricity cost. Meanwhile, due to the time-dependent demand preference, power grid is suffering high load variations, which results in a large profit loss. In this paper, we consider a cost-efficient workload scheduling with a coordination between a cloud service provider operating multiple geo-distributed data centers and smart grids. The aim is to explore the flexibility of data center power demands to reduce the cost of the cloud service provider and smooth the load variations of smart grids simultaneously. We first present the penalty model of the computation workload scheduling at each data center, and introduce the cost model of smart grids, including power generation cost and the cost due to the power load variations. To jointly minimize the cost of smart grids and penalty of the cloud service provider resulted from workload scheduling, we formulate the objective function as a weighted sum of the cost and the penalty to study the tradeoffs, and obtain the optimal offline solution by the dual decomposition technique. In order to make the coordination implemented in an online fashion, we propose a Receding Horizon Control (RHC) based online algorithm to obtain the suboptimal workload management based on the predicted information, including the future amounts of interactive workload, batch workload, and power load, in the prediction horizon. The simulation results show that with the coordination between the cloud service provider and smart grids, the cost of smart grids can be significantly reduced, by up to 20 percent, and the load variations of smart grids can be well smoothed simultaneously. Han Hu 0003, Yonggang Wen 0001, Ling Qiu 0003, Dusit Niyato |
IEEE Trans. Serv. Comput. | 2 |
| 2020 | Electricity Cost Minimization for Interruptible Workload in Datacenter ServersabstractDatacenters have experienced dramatic growth in recent years, and the cost for powering them has become a significant problem. This paper proposes methods to minimize the energy cost for performing a task on a datacenter server before a deadline. We observe that energy prices fluctuate over time, and schedule the task to execute in periods of relatively low cost, despite not having knowledge of future costs during the execution. This problem is studied in several models, starting with an online setting where electricity prices can change arbitrarily. A$\sqrt{\varphi }$-competitive algorithm is proposed, where$\varphi$is the ratio between the maximum and minimum electricity prices, and this algorithm is also shown to be optimal by proving a matching lower bound. Next, we consider a stochastic setting in which prices vary in a Markovian fashion and propose an optimal algorithm based on dynamic programming. We then study the performance of our algorithms in practice using prices derived from real world data. The results show that the stochastic algorithm is very effective, and achieves cost that is within 3.4 percent of the optimum. Moreover, it performs well compared to several heuristics used in practice. Wei Zhang 0082, Yonggang Wen 0001, Loi Lei Lai, Fang Liu 0009, Rui Fan 0004 |
IEEE Trans. Serv. Comput. | 2 |
| 2019 | ResumeGAN: An Optimized Deep Representation Learning Framework for Talent-Job Fit via Adversarial LearningabstractNowadays, it is popular to utilize online recruitment services for talent recruitment and job recommendation. Given the vast amounts of online talent profiles and job-posts, it is labor-intensive and exhausted for recruiters to manually select only a few potential candidates for further consideration, and also nontrivial for talents to find the most matched job positions. Recently, some deep learning-based approaches are developed to automatically matching the talent resumes and job requirements, and have achieved encouraging performance. In this paper, we propose a novel framework that targets the same task, but integrate different types of information in a more sophisticated way and introduce adversarial learning to learn more expressive representation. In addition, we build a dataset for model evaluation and the effectiveness of our framework is demonstrated by extensive experiments. Yong Luo 0002, Huaizheng Zhang, Yonggang Wen 0001, Xinwen Zhang |
CIKM | 3 |
| 2019 | Content-Aware Personalised Rate Adaptation for Adaptive Streaming via Deep Video AnalysisabstractAdaptive bitrate (ABR) streaming is the de facto solution for achieving smooth viewing experiences under unstable network conditions. However, most of the existing rate adaptation approaches for ABR are content-agnostic, without considering the semantic information of the video content. Nevertheless, semantic information largely determines the informativeness and interestingness of the video content, and consequently affects the QoE for video streaming. One common case is that the user may expect higher quality for the parts of video content that are more interesting or informative so as to reduce overall subjective quality loss. This creates two main challenges for such a problem: First, how to determine which parts of the video content are more interesting? Second, how to allocate bitrate budgets for different parts of the video content with different significances? To address these challenges, we propose a Content-of-Interest (CoI) based rate adaptation scheme for ABR. We first design a deep learning approach for recognizing the interestingness of the video content, and then design a Deep Q-Network (DQN) approach for rate adaptation by incorporating video interestingness information. The experimental results show that our method can recognize video interestingness precisely, and the bitrate allocation for ABR can be aligned with the interestingness of video content while not compromising the performances on objective QoE metrics. Guanyu Gao, Linsen Dong, Huaizheng Zhang, Yonggang Wen 0001, Wenjun Zeng 0001 |
ICC | 4 |
| 2019 | Optimization for HTTP Adaptive Video Streaming in UAV-Enabled Relaying SystemabstractTo guarantee quality of experience (QoE) of video streaming for ground users with large obstacles that deteriorate the quality of links, unmanned aerial vehicle (UAV) is introduced in this paper as a relay to serve these users for HTTP adaptive streaming. By introducing the QoE utility model for users, we study the average QoE maximization problem in UAV-enabled relaying system by optimizing the UAV position along with the bandwidth and transmit power allocation for ground users, subject to the information causality constraint at the UAV relay. The optimization problem is formulated with a non-convex programming which is difficult to solve in general. By applying the successive convex approximation and block coordinate descent techniques, an efficient iterative algorithm is proposed to simultaneously update the UAV's position, transmit power and bandwidth allocation at each iteration, where the convergence is analyzed. Extensive simulations are conducted to show that the proposed solution can achieve significant gains for QoE in terms of the average streaming utility. Han Hu 0003, Cheng Zhan, Jianping An, Yonggang Wen 0001 |
ICC | 4 |
| 2019 | DeepEE: Joint Optimization of Job Scheduling and Cooling Control for Data Center Energy Efficiency Using Deep Reinforcement LearningabstractThe past decade witnessed the tremendous growth of power consumption in data centers due to the rapid development of cloud computing, big data analytics, and machine learning, etc. The prior approaches that optimize the power consumption of the information technology (IT) system and/or the cooling system always fail to capture the system dynamics or suffer from the complexity of system states and action spaces. In this paper, we propose a Deep Reinforcement Learning (DRL) based optimization framework, named DeepEE, to improve the energy efficiency for data centers by considering the IT and cooling systems concurrently. In DeepEE, we first propose a PArameterized action space based Deep Q-Network (PADQN) algorithm to solve the hybrid action space problem and jointly optimize the job scheduling for the IT system and the airflow rate adjustment for the cooling system. Then, a two-time-scale control mechanism is applied in PADQN to coordinate the IT and cooling systems more accurately and efficiently. In addition, to train and evaluate the proposed PADQN in a safe and quick way, we build a simulation platform to model the dynamics of IT workload and cooling systems simultaneously. Through extensive real-trace based simulations, we demonstrate that: 1) our algorithm can save up to 15% and 10% energy consumption in comparison with the baseline siloed and joint optimization approaches respectively; 2) our algorithm achieves more stable performance gain in terms of power consumption by adopting the parameterized action space; and 3) our algorithm leads to a better tradeoff between energy saving and service quality. Yongyi Ran, Han Hu 0003, Xin Zhou 0003, Yonggang Wen 0001 |
ICDCS | 4 |
| 2019 | Toward Efficient Compute-Intensive Job Allocation for Green Data Centers: A Deep Reinforcement Learning ApproachabstractReducing the energy consumption of the servers in a data center via proper job allocation is desirable. Existing advanced job allocation algorithms, based on constrained optimization formulations capturing servers' complex power consumption and thermal dynamics, often scale poorly with the data center size and optimization horizon. This paper applies deep reinforcement learning (DRL) to build an allocation algorithm for long-lasting and compute-intensive jobs that are increasingly seen among today's computation demands. Specifically, a deep Q-network is trained to allocate jobs, aiming to maximize a cumulative reward over long horizons. The training is performed offline using a computational model based on long short-term memory networks that capture the servers' power and thermal dynamics. This offline training approach avoids slow online convergence, low energy efficiency, and potential server overheating during the DRL's extensive state-action space exploration if it directly interacts with the physical data center in the usually adopted online learning scheme. At run time, the trained Q-network is forward-propagated with little computation to allocate jobs. Evaluation based on 8 months' physical state and job arrival records from a national supercomputing data center hosting 1,152 processors shows that our solution reduces computing power consumption by nearly 10% and processor temperature by more than 3°C without sacrificing job processing throughput. Deliang Yi, Xin Zhou 0003, Yonggang Wen 0001, Rui Tan 0001 |
ICDCS | 3 |
| 2019 | Incorporating Category Taxonomy in Deep Reinforcement Learning Based Image HashingabstractImage hashing is critical for large-scale image analytic-based applications, such as image retrieval. Although there have been dozens of hashing approaches, few of them take the hierarchical structure of the image categories into consideration. In this paper, we propose to incorporate the category taxonomy information in a deep reinforcement learning (DRL) model for image hashing. In particular, we learn an agent to predict the hashing codes sequentially under the DRL theme. Each coordinate of the hashing function can take the errors incurred by previous ones into consideration and hence more reliable hashing codes can be obtained than learning them independently. Besides, we design a novel level-specific reward function to gradually refine the hashing function according to the taxonomy information. Extensive experiments on two popular datasets demonstrate effectiveness of the proposed method. Qiang Fu 0006, Linsen Dong, Yong Luo 0002, Yonggang Wen 0001, Ying Li 0012, Ling-Yu Duan |
ICME | 5 |
| 2019 | QoE-Driven Mobile Streaming: A Location-Aware ApproachabstractIn this paper, we maximize the quality of experience (QoE) for mobile video streaming. QoE is modeled to capture user's video quality assessment as well as the freezing and bitrate variation during playback. Based on an observation that network bandwidth is correlated with location, we predict the future locations and accordingly the bandwidth along a trip. Then, the predicted information is utilized to dynamically adapt the video version and bitrate with maximized QoE. We show that our proposed solution well approximates the offline optimal performance by almost 98% on average. The proposed solution is also competitive compared to several popular streaming algorithms. Fang Liu 0009, Wei Zhang 0082, Yonggang Wen 0001 |
ICME | 3 |
| 2019 | Toward Intelligent Visual Sensing and Low-cost Analysis: A Collaborative Computing ApproachabstractIn the big data era, there has been an increasing consensus that the label information, computational resources and communication bandwidth are particularly precious. State-of-the-art research is revolutionizing the vision systems of the smart city, which converts the visual signals from sensory input into feature representations and conveys the compact feature for analysis by using the computational resources of both front and back ends. To deploy a robust model, large amounts of labeled data are usually required, and thereby heavy computational and communication resources are incurred in model training as well as inference. However, the computational resources in front-end devices are usually constrained, and heavy transmission burden is imposed when leveraging multiple models amongst different ends. In this work, we propose a novel collaborative computing approach for intelligent sensing and low-cost analysis, which reduces the requirement of labeled data and communication cost, and balances the computational load in model training and inference. By incorporating the adversarial learning mechanism into collaborative model training, knowledge of different domains can be better exploited. Moreover, the learned models are deployed for inference in a collaborative manner, in which part of model is placed in front-ends for extracting intermediate feature maps, and part of the model remains in back ends for inference with received feature maps. The effectiveness of the proposed approach has been validated in the context of an emerging digital retina system for smart city intelligent applications. Ling-Yu Duan, Yong Luo 0002, Shiqi Wang 0001, Yonggang Wen 0001, Wen Gao 0001 |
VCIP | 5 |
| 2019 | Thermal Comfort Modeling for Smart Buildings: A Fine-Grained Deep Learning ApproachabstractThe emerging Internet of Things (IoT) technology enables smart building management and operation to improve building energy efficiency and occupant thermal comfort. In this paper, we perform data analysis using the IoT generated building data to derive accurate thermal comfort model for smart building control. Deep neural network (DNN) is used to model the relationship between the controllable building operations and thermal comfort. As thermal comfort is determined by multiple comfort factors, a fine-grained architecture is proposed, where an exclusive model is trained for each factor and accordingly the corresponding thermal comfort can be evaluated. The experimental results show that the proposed fine-grained DNN outperforms its coarse-grained counterpart by 3.5× and is 1.7×, 2.5×, 2.4×, and 1.9× more accurate compared to four popular machine learning algorithms. Besides, DNN's performance promotes with deeper network topology and more neurons, and a simple topology with the same number of neurons per network hidden layer is sufficient to achieve high modeling accuracy. Finally, the derived thermal comfort model reveals a linear relationship between comfort and air conditioning setpoint. The linear property helps quickly and accurately search for the optimal controllable setpoint with the desired comfort. Wei Zhang 0082, Weizheng Hu, Yonggang Wen 0001 |
IEEE Internet Things J. | 3 |
| 2019 | Special Issue on Artificial Intelligence and Machine Learning for Networking and CommunicationsabstractResearch in large-scale networking systems has been shaped and will continue to be guided by specific characteristics of applications and the underlying platforms and infrastructures. On the one hand, applications are growing at an accelerated pace, which is fundamentally unpredictable in both breadth and depth. On the other hand, the underlying networking has been the focus of a huge transformation enabled by new models resulting from virtualization and cloud computing. This has led to a number of novel architectures supported by emerging technologies such as Software-Defined Networking (SDN), Network Function Virtualization (NFV), and more recently, edge cloud and fog networking, or network slicing[1],[2]. This evolution towards enhanced design opportunities along with increasing complexity in networking and its applications has fueled the need for improved network automation in agile infrastructures. At the same time, their complexity has dramatically increased. The networking dynamics have had the effect of making it even more important and challenging to design scalable network measurement and analysis techniques and associated tools. Critical applications such as resource allocation, network monitoring, security enforcement, or dynamic network management require real-time mechanisms for online analysis as well as efficient techniques for offline deep analysis of massive historical data. Prosper Chemouil, Pan Hui 0001, Wolfgang Kellerer, Yong Li 0008, Rolf Stadler, Dacheng Tao, Yonggang Wen 0001, Ying Zhang 0022 |
IEEE J. Sel. Areas Commun. | 7 |
| 2019 | Transferring Knowledge Fragments for Learning Distance Metric from a Heterogeneous DomainabstractThe goal of transfer learning is to improve the performance of target learning task by leveraging information (or transferring knowledge) from other related tasks. In this paper, we examine the problem of transfer distance metric learning (DML), which usually aims to mitigate the label information deficiency issue in the target DML. Most of the current Transfer DML (TDML) methods are not applicable to the scenario where data are drawn from heterogeneous domains. Some existing heterogeneous transfer learning (HTL) approaches can learn target distance metric by usually transforming the samples of source and target domain into a common subspace. However, these approaches lack flexibility in real-world applications, and the learned transformations are often restricted to be linear. This motivates us to develop a general flexible heterogeneous TDML (HTDML) framework. In particular, any (linear/nonlinear) DML algorithms can be employed to learn the source metric beforehand. Then the pre-learned source metric is represented as a set of knowledge fragments to help target metric learning. We show how generalization error in the target domain could be reduced using the proposed transfer strategy, and develop novel algorithm to learn either linear or nonlinear target metric. Extensive experiments on various applications demonstrate the effectiveness of the proposed method. Yong Luo 0002, Yonggang Wen 0001, Tongliang Liu, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | Dynamic Priority-Based Resource Provisioning for Video Transcoding With Heterogeneous QoSabstractVideo transcoding is widely adopted in online video services to transcode videos into multiple representations for dynamic adaptive bitrate streaming. This solution may consume significant resources and incur intolerable processing delays. Meanwhile, different videos have different quality-of-service (QoS) requirements for transcoding. Delay-sensitive videos must be transcoded within a strict deadline, whereas delay-tolerant videos are not required to be transcoded immediately. Some intelligent policies are required for provisioning the right amount of resources in the transcoding system to meet the heterogeneous QoS requirements, especially under dynamic workloads. To this end, we develop a robust dynamic priority-based resource provisioning scheme for video transcoding. We adopt the preemptive resume priority discipline to design a multiple-priority transcoding mechanism. The system performs the transcoding for delay-tolerant videos by utilizing idle resources for improving resource utilization while not affecting the transcoding for delay-sensitive videos. We adopt the model predictive control framework to design an online algorithm for dynamic resource provisioning to accommodate time-varying workloads by predicting future workloads. To seek performance robustness against prediction noise, we improve the performance of our online algorithm via robust design. The experimental results demonstrate that our proposed method can satisfy the heterogeneous QoS requirements while significantly reducing computing resource consumption. Guanyu Gao, Yonggang Wen 0001, Cédric Westphal |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Orchestrating Caching, Transcoding and Request Routing for Adaptive Video Streaming Over ICNabstractInformation-centric networking (ICN) has been touted as a revolutionary solution for the future of the Internet, which will be dominated by video traffic. This work investigates the challenge of distributing video content of adaptive bitrate (ABR) over ICN. In particular, we use the in-network caching capability of ICN routers to serve users; in addition, with the help of named function, we enable ICN routers to transcode videos to lower-bitrate versions to improve the cache hit ratio. Mathematically, we formulate this design challenge into a constrained optimization problem, which aims to maximize the cache hit ratio for service providers and minimize the service delay for endusers. We design a two-step iterative algorithm to find the optimum. First, given a content management scheme, we minimize the service delay via optimally configuring the routing scheme. Second, we maximize the cache hits for a given routing policy. Finally, we rigorously prove its convergence. Through extensive simulations, we verify the convergence and the performance gains over other algorithms. We also find that more resources should be allocated to ICN routers with a heavier request rate, and the routing scheme favors the shortest path to schedule more traffic. Han Hu 0003, Yichao Jin 0002, Yonggang Wen 0001, Cédric Westphal |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2018 | Speeding-Up Age Estimation in Intelligent Demographics System via Network OptimizationabstractAge estimation is a difficult task which requires the automatic detection and interpretation of facial features. Recently, Convolutional Neural Networks (CNNs) have made remarkable improvement on learning age patterns from benchmark datasets. However, for a face ``in the wild'' (from a video frame or Internet), the existing algorithms are not as accurate as for a frontal and neutral face. In addition, with the increasing number of in-the- wild aging data, the computation speed of existing deep learning platforms becomes another crucial issue. In this paper, we propose a high-efficient age estimation system with joint optimization of age estimation algorithm and deep learning system. Cooperated with the city surveillance network, this system can provide age group analysis for intelligent demographics. First, we build a three- tier fog computing architecture including an edge, a fog and a cloud layer, which directly processes age estimation from raw videos. Second, we optimize the age estimation algorithm based on CNNs with label distribution and K-L divergence distance embedded in the fog layer and evaluate the model on the latest wild aging dataset. Experimental results demonstrate that: 1. our system collects the demographics data dynamically at far-distance without contact, and makes the city population analysis automatically; and 2. the age model training has been speed-up without losing training progress or model quality. To our best knowledge, this is the first intelligent demographics system which has potential applications in improving the efficiency of smart cities and urban living. Zhenzhen Hu 0004, Peng Sun 0006, Yonggang Wen 0001 |
ICC | 3 |
| 2018 | ResumeNet: A Learning-Based Framework for Automatic Resume Quality AssessmentabstractRecruitment of appropriate people for certain positions is critical for any companies or organizations. Manually screening to select appropriate candidates from large amounts of resumes can be exhausted and time-consuming. However, there is no public tool that can be directly used for automatic resume quality assessment (RQA). This motivates us to develop a method for automatic RQA. Since there is also no public dataset for model training and evaluation, we build a dataset for RQA by collecting around 10K resumes, which are provided by a private resume management company. By investigating the dataset, we identify some factors or features that could be useful to discriminate good resumes from bad ones, e.g., the consistency between different parts of a resume. Then a neural-network model is designed to predict the quality of each resume, where some text processing techniques are incorporated. To deal with the label deficiency issue in the dataset, we propose several variants of the model by either utilizing the pair/triplet-based loss, or introducing some semi-supervised learning technique to make use of the abundant unlabeled data. Both the presented baseline model and its variants are general and easy to implement. Various popular criteria including the receiver operating characteristic (ROC) curve, F-measure and ranking-based average precision (AP) are adopted for model evaluation. We compare the different variants with our baseline model. Since there is no public algorithm for RQA, we further compare our results with those obtained from a website that can score a resume. Experimental results in terms of different criteria demonstrate effectiveness of the proposed method. We foresee that our approach would transform the way of future human resources management. Yong Luo 0002, Huaizheng Zhang, Yonggang Wen 0001, Xinwen Zhang |
ICDM | 4 |
| 2018 | Deepqoe: A Unified Framework for Learning to Predict Video QoEabstractMotivated by the prowess of deep learning (DL) based techniques in prediction, generalization, and representation learning, we develop a novel framework called DeepQoE to predict video quality of experience (QoE). The end-to-end framework first uses a combination of DL techniques (e.g., word embeddings) to extract generalized features. Next, these features are combined and fed into a neural network for representation learning. Such representations serve as inputs for classification or regression tasks. Evaluating the performance of DeepQoE with two datasets, we show that for the small dataset, the accuracy of all shallow learning algorithms is improved by using the representation derived from DeepQoE. For the large dataset, our DeepQoE framework achieves significant performance improvement in comparison to the best baseline method (90.94% vs. 82.84%). Moreover, DeepQoE, also released as an open source tool, provides video QoE research much-needed flexibility in fitting different datasets, extracting generalized features, and learning representations. Huaizheng Zhang, Han Hu 0003, Guanyu Gao, Yonggang Wen 0001, Kyle Guan |
ICME | 4 |
| 2018 | JALAD: Joint Accuracy-And Latency-Aware Deep Structure Decoupling for Edge-Cloud ExecutionabstractRecent years have witnessed a rapid growth of deep-network based services and applications. A practical and critical problem thus has emerged: how to effectively deploy the deep neural network models such that they can be executed efficiently. Conventional cloud-based approaches usually run the deep models in data center servers, causing large latency because a significant amount of data has to be transferred from the edge of network to the data center. In this paper, we propose JALAD, a joint accuracy- and latency-aware execution framework, which decouples a deep neural network so that a part of it will run at edge devices and the other part inside the conventional cloud, while only a minimum amount of data has to be transferred between them. Though the idea seems straightforward, we are facing challenges including i)how to find the best partition of a deep structure; ii)how to deploy the component at an edge device that only has limited computation power; and iii)how to minimize the overall execution latency. Our answers to these questions are a set of strategies in JALAD, including 1)A normalization based in-layer data compression strategy by jointly considering compression rate and model accuracy; 2)A latency-aware deep decoupling strategy to minimize the overall execution latency; and 3)An edge-cloud structure adaptation strategy that dynamically changes the decoupling for different network conditions. Experiments demonstrate that our solution can significantly reduce the execution latency: it speeds up the overall inference execution with a guaranteed model accuracy loss. Hongshan Li, Chenghao Hu, Jingyan Jiang, Zhi Wang 0001, Yonggang Wen 0001, Wenwu Zhu 0001 |
ICPADS | 5 |
| 2018 | Online Heterogeneous Transfer Metric LearningabstractDistance metric learning (DML) has been demonstrated to be successful and essential in diverse applications. Transfer metric learning (TML) can help DML in the target domain with limited label information by utilizing information from some related source domains. The heterogeneous TML (HTML), where the feature representations vary from the source to the target domain, is general and challenging. However, current HTML approaches are usually conducted in a batch manner and cannot handle sequential data. This motivates the proposed online HTML (OHTML) method. In particular, the distance metric in the source domain is pre-trained using some existing DML algorithms. To enable knowledge transfer, we assume there are large amounts of unlabeled corresponding data that have representations in both the source and target domains. By enforcing the distances (between these unlabeled samples) in the target domain to agree with those in the source domain under the manifold regularization theme, we learn an improved target metric. We formulate the problem in the online setting so that the optimization is efficient and the model can be adapted to new coming data. Experiments in diverse applications demonstrate both effectiveness and efficiency of the proposed method. Yong Luo 0002, Tongliang Liu, Yonggang Wen 0001, Dacheng Tao |
IJCAI | 3 |
| 2018 | Fast media caching for geo-distributed data centers
Wei Zhang 0082, Yonggang Wen 0001, Fang Liu 0009, Yiqiang Chen 0001, Rui Fan 0004 |
Comput. Commun. | 2 |
| 2018 | iTCM: Toward Learning-Based Thermal Comfort Modeling via Pervasive Sensing for Smart BuildingsabstractFor decades, ASHRAE Standard 55 has been using the Fanger's predicted mean vote (PMV) model to evaluate the indoor thermal comfort satisfaction. However, this canonical model has drawbacks in both data inadequacy and lack of inputs from test subjects. In this paper, we propose a learning-based solution for thermal comfort modeling via the emerging machine learning techniques and Internet of Things-based pervasive sensing technologies. First, we build an intelligent thermal comfort management (iTCM) system. It adopts the wireless sensor network to collect environmental data and utilizes the wearable device for vital sign monitoring. In addition, a cloud-based back-end system, with cost efficient deployment fees, is developed for data management and analysis. Second, we implement a black-box neural network (NN), namely the intelligent thermal comfort NN (ITCNN). To evaluate the performance of ITCNN, we compare it with the PMV model, three traditional white-box machine learning approaches and three classical black-box machine learning methods. Our preliminary results show that four black-box methods achieve better performance than the PMV model and the three white-box approaches. The ITCNN achieves the best performance and outperforms the PMV model by on average 13.1% and up to 17.8%. Third, with the iTCM system, we demonstrate a novel deep reinforcement learning-based application by encouraging human behavioral changes to form energy-saving habits for greener, smarter, and healthier building. Finally, we discuss the limitations of this paper and present the plan for our future research. Weizheng Hu, Yonggang Wen 0001, Kyle Guan, Guangyu Jin, King-Jet Tseng |
IEEE Internet Things J. | 2 |
| 2018 | MetaFlow: A Scalable Metadata Lookup Service for Distributed File Systems in Data CentersabstractIn large-scale distributed file systems, efficient metadata operations are critical since most file operations have to interact with metadata servers first. In existing distributed hash table (DHT) based metadata management systems, the lookup service could be a performance bottleneck due to its significant CPU overhead. Our investigations showed that the lookup service could reduce system throughput by up to 70 percent, and increase system latency by a factor of up to 8 compared to ideal scenarios. In this paper, we present MetaFlow, a scalable metadata lookup service utilizing software-defined networking (SDN) techniques to distribute lookup workload over network components. MetaFlow tackles the lookup bottleneck problem by leveraging B-tree, which is constructed over the physical topology, to manage flow tables for SDN-enabled switches. Therefore, metadata requests can be forwarded to appropriate servers using only switches. Extensive performance evaluations in both simulations and testbed showed that MetaFlow increases system throughput by a factor of up to 3.2, and reduce system latency by a factor of up to 5 compared to DHT-based systems. We also deployed MetaFlow in a distributed file system, and demonstrated significant performance improvement. Peng Sun 0006, Yonggang Wen 0001, Ta Nguyen Binh Duong, Haiyong Xie 0001 |
IEEE Trans. Big Data | 2 |
| 2018 | Energy-Efficient Task Execution for Application as a General Topology in Mobile Cloud ComputingabstractMobile cloud computing has been proposed as an effective solution to augment the capabilities of resource-poor mobile devices. In this paper, we investigate energy-efficient collaborative task execution to reduce the energy consumption on mobile devices. We model a mobile application as a general topology, consisting of a set of fine-grained tasks. Each task within the application can be either executed on the mobile device or on the cloud. We aim to find out the execution decision for each task to minimize the energy consumption on the mobile device while meeting a delay deadline. We formulate the collaborative task execution as a delay-constrained workflow scheduling problem. We leverage the partial critical path analysis for the workflow scheduling; for each path, we schedule the tasks using two algorithms based on different cases. For the special case without execution restriction, we adopt one-climb policy to obtain the solution. For the general case where there are some tasks that must be executed either on the mobile device or on the cloud, we adopt Lagrange Relaxation based Aggregated Cost (LARAC) algorithm to obtain the solution. We show by simulation that the collaborative task execution is more energy-efficient than local execution and remote execution. Yonggang Wen 0001 |
IEEE Trans. Cloud Comput. | 2 |
| 2018 | Toward Wi-Fi AP-Assisted Content Prefetching for an On-Demand TV Series: A Learning-Based ApproachabstractThe emergence of smart Wi-Fi access points (AP), which are equipped with huge storage space, opens a new research area on how to utilize these resources at the edge network to improve users' quality of experience (e.g., a short startup delay and smooth playback). One important research interest in this area is content prefetching which predicts and accurately fetches contents ahead of users' requests to shift the traffic away during peak periods. However, in practice, the different video watching patterns among users and the varying network connection status lead to the time-varying server load, which eventually makes the content prefetching problem challenging. To understand this challenge, this paper first performs a large-scale measurement study on users' AP connection and TV series watching patterns using real traces. Then, based on the obtained insights, we formulate the content prefetching problem as a Markov decision process. The objective is to strike a balance between the increased prefetching and storage cost incurred by incorrect prediction and the reduced content download delay because of successful prediction. A learning-based approach is proposed to solve this problem and another three algorithms are adopted as baselines. In particular, first we investigate the performance lower bound by using a random algorithm and the upper bound by using an ideal offline approach. Then, we present a heuristic algorithm as another baseline. Finally, we design a reinforcement learning algorithm that is more practical to work in the online manner. Through extensive trace-based experiments, we demonstrate the performance gain of our design. Remarkably, our learning-based algorithm achieves a better precision and hit ratio (e.g., 80%) with about 70% (resp. 50%) cost saving compared to the random (resp. heuristic) algorithm Wen Hu 0003, Yichao Jin 0002, Yonggang Wen 0001, Zhi Wang 0001, Lifeng Sun |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Budget-Efficient Viral Video Distribution Over Online Social Networks: Mining Topic-Aware Influential UsersabstractMarketing over online social networks (OSNs) has become an essential tool for spreading product information in a “word of mouth” way. In particular, campaigns normally adopt a pragmatic approach of seeding videos with a selected list of influential users, hoping to create a viral distribution to reach as many users as possible. In this paper, we propose a multitopic-aware influence maximization framework to identify a fixed number of influential users and assign video clips of specific topics to them, with an ultimate objective to maximize the number of message deliveries, defined as expected posting number (EPN). We first prove the submodularity of the EPN function, resulting in a general greedy algorithm with a performance bound of 1-1/e. We further develop two faster algorithms to accelerate the computing speed for large-scale social networks. The first algorithm leverages two estimation methods to compute the upper bound for marginal EPN without a loss of accuracy. The second algorithm generates an approximation solution based on the upper bound and lower bound estimation, with a performance bound of ε(1-1/e). We have implemented a prototype system based on a private data center at the Nanyang Technological University campus in Singapore to enable video clip extraction and sharing among social users. Furthermore, we conduct experiments on four real large-scale social networks (with different scales and structures) and the results show that the proposed methods are much faster than previous algorithms but with high accuracy. Han Hu 0003, Yonggang Wen 0001, Shanshan Feng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | Can We Speculate Running Application With Server Power Consumption Trace?abstractIn this paper, we propose to detect the running applications in a server by classifying the observed power consumption series for the purpose of data center energy consumption monitoring and analysis. Time series classification problem has been extensively studied with various distance measurements developed; also recently the deep learning-based sequence models have been proved to be promising. In this paper, we propose a novel distance measurement and build a time series classification algorithm hybridizing nearest neighbor and long short term memory (LSTM) neural network. More specifically, first we propose a new distance measurement termed as local time warping (LTW), which utilizes a user-specified index set for local warping, and is designed to be noncommutative and nondynamic programming. Second, we hybridize the 1-nearest neighbor (1NN)-LTW and LSTM together. In particular, we combine the prediction probability vector of 1NN-LTW and LSTM to determine the label of the test cases. Finally, using the power consumption data from a real data center, we show that the proposed LTW can improve the classification accuracy of dynamic time warping (DTW) from about 84% to 90%. Our experimental results prove that the proposed LTW is competitive on our data set compared with existed DTW variants and its noncommutative feature is indeed beneficial. We also test a linear version of LTW and find out that it can perform similar to state-of-the-art DTW-based method while it runs as fast as the linear runtime lower bound methods like LB_Keogh for our problem. With the hybrid algorithm, for the power series classification task we achieve an accuracy up to about 93%. Our research can inspire more studies on time series distance measurement and the hybrid of the deep learning models with other traditional models. Han Hu 0003, Yonggang Wen 0001, Jun Zhang 0003 |
IEEE Trans. Cybern. | 3 |
| 2018 | Cost-Sensitive Feature Selection by Optimizing F-MeasuresabstractFeature selection is beneficial for improving the performance of general machine learning tasks by extracting an informative subset from the high-dimensional features. Conventional feature selection methods usually ignore the class imbalance problem, thus the selected features will be biased towards the majority class. Considering that F-measure is a more reasonable performance measure than accuracy for imbalanced data, this paper presents an effective feature selection algorithm that explores the class imbalance issue by optimizing F-measures. Since F-measure optimization can be decomposed into a series of cost-sensitive classification problems, we investigate the cost-sensitive feature selection by generating and assigning different costs to each class with rigorous theory guidance. After solving a series of cost-sensitive feature selection problems, features corresponding to the best F-measure will be selected. In this way, the selected features will fully represent the properties of all classes. Experimental results on popular benchmarks and challenging real-world data sets demonstrate the significance of cost-sensitive feature selection for the imbalanced data setting and validate the effectiveness of the proposed method. Meng Liu 0003, Chang Xu 0002, Yong Luo 0002, Chao Xu 0006, Yonggang Wen 0001, Dacheng Tao |
IEEE Trans. Image Process. | 5 |
| 2018 | Toward Intelligent Product Retrieval for TV-to-Online (T2O) Application: A Transfer Metric Learning ApproachabstractIt is desired (especially for young people) to shop for the same or similar products shown in the multimedia contents (such as online TV programs). This indicates an urgent demand for improving the experience of TV-to-Online (T2O). In this paper, a transfer learning approach as well as a prototype system for effortless T2O experience is developed. In the system, a key component is high-precision product search, which is to fulfill exact matching between a query item and the database ones. The matching performance primarily relies on distance estimation, but the data characteristics cannot be well modeled and exploited by a simple Euclidean distance. This motivates us to introduce distance metric learning (DML) for improving the distance estimation. However, in traditional DML methods, the side information (such as the similar/dissimilar constraints or relevance/irrelevance judgements) in the target domain is leveraged. These methods may fail due to limited side information. Fortunately, this issue can be alleviated by utilizing transfer metric learning (TML) to exploit information from other related domains. In this paper, a novel manifold regularized heterogeneous multitask metric learning framework is proposed, in which each domain is treated equally. The proposed approach allows us to simultaneously exploit the information from other domains and the unlabeled information. Furthermore, the ranking-based loss is adopted to make our model more appropriate for search. Experiments on two challenging real-world datasets demonstrate the effectiveness of the proposed method. This TML approach is expected to impact the transformation of the emerging T2O trend in both TV and online video domains. Qiang Fu 0006, Yong Luo 0002, Yonggang Wen 0001, Dacheng Tao, Ying Li 0012, Ling-Yu Duan |
IEEE Trans. Multim. | 3 |
| 2018 | Optimizing Quality of Experience for Adaptive Bitrate Streaming via Viewer Interest InferenceabstractRate adaptation is widely adopted in video streaming to improve the quality of experience (QoE). However, most of the existing rate adaptation approaches neglect the underlying video semantic information. In fact, influenced by video semantics and viewer preferences, the viewer may have different degrees of interest on different parts of a video. The interesting parts of a video can draw more visual attention from the viewer and have higher visual importance. As such, delivering the parts of a video that are interesting to the viewer in a higher quality can improve the perceptual video quality, compared with the semantics-agnostic approaches that treat each part of a video equally. Thus, it is natural to wonder: how to allocate bitrate budgets temporally over a video session under time-varying bandwidth while considering viewer interest? As an exploratory study, we propose an interest-aware rate adaptation approach for improving QoE by inferring viewer interest based on video semantics. We adopt the deep learning method to recognize the scenes of video frames and leverage the term frequency-inverse document frequency method to analyze the degrees of an individual viewer's interest on different types of scenes. The bandwidth, buffer occupancy, and viewer interest are jointly considered under the model predictive control framework for selecting appropriate bitrates for maximizing QoE. The objective and subjective evaluations measured in a real environment show that our method can achieve a higher QoE compared with the semantics-agnostic approaches. Guanyu Gao, Huaizheng Zhang, Han Hu 0003, Yonggang Wen 0001, Jianfei Cai 0001, Chong Luo 0001, Wenjun Zeng 0001 |
IEEE Trans. Multim. | 4 |
| 2018 | Toward Rendering-Latency Reduction for Composable Web Services via Priority-Based Object CachingabstractWeb services serve as the cornerstone of the Internet for rendering webpages. The initial rendering latency of webpages, which depends on a subset of critical objects required by the webpage, is a key metric for web services. In this work, we propose to identify this set of critical objects systematically with the goal of caching them at a higher priority to reduce the initial rendering time. We first conduct a measurement study on a mainstream content delivery network provider, the results of which suggest that not all currently cached objects are critical and that only a small portion of the critical objects are cached. Thus, we model the critical-object aware caching scheme as a constrained optimization problem. Using the stochastic optimization framework, we decompose the problem into a set of one-shot optimization problems, which are proved to be NP-hard. We then develop two greedy algorithms with different computational complexity but the same performance bound. Finally, we integrate the resulting approximation algorithms into an online algorithm. Through trace-based simulations, we verify that our proposed algorithm can reduce service latency and network traffic by ensuring a higher cache hit ratio. Han Hu 0003, Yonggang Wen 0001 |
IEEE Trans. Multim. | 3 |
| 2018 | Heterogeneous Multitask Metric Learning Across Multiple DomainsabstractDistance metric learning plays a crucial role in diverse machine learning algorithms and applications. When the labeled information in a target domain is limited, transfer metric learning (TML) helps to learn the metric by leveraging the sufficient information from other related domains. Multitask metric learning (MTML), which can be regarded as a special case of TML, performs transfer across all related domains. Current TML tools usually assume that the same feature representation is exploited for different domains. However, in real-world applications, data may be drawn from heterogeneous domains. Heterogeneous transfer learning approaches can be adopted to remedy this drawback by deriving a metric from the learned transformation across different domains. However, they are often limited in that only two domains can be handled. To appropriately handle multiple domains, we develop a novel heterogeneous MTML (HMTML) framework. In HMTML, the metrics of all different domains are learned together. The transformations derived from the metrics are utilized to induce a common subspace, and the high-order covariance among the predictive structures of these domains is maximized in this subspace. There do exist a few heterogeneous transfer learning approaches that deal with multiple domains, but the high-order statistics (correlation information), which can only be exploited by simultaneously examining all domains, is ignored in these approaches. Compared with them, the proposed HMTML can effectively explore such high-order information, thus obtaining more reliable feature transformations and metrics. Effectiveness of our method is validated by the extensive and intensive experiments on text categorization, scene classification, and social image annotation. Yong Luo 0002, Yonggang Wen 0001, Dacheng Tao |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2017 | Cost-Sensitive Feature Selection via F-Measure Optimization ReductionabstractFeature selection aims to select a small subset from the high-dimensional features which can lead to better learning performance, lower computational complexity, and better model readability. The class imbalance problem has been neglected by traditional feature selection methods, therefore the selected features will be biased towards the majority classes. Because of the superiority of F-measure to accuracy for imbalanced data, we propose to use F-measure as the performance measure for feature selection algorithms. As a pseudo-linear function, the optimization of F-measure can be achieved by minimizing the total costs. In this paper, we present a novel cost-sensitive feature selection (CSFS) method which optimizes F-measure instead of accuracy to take class imbalance issue into account. The features will be selected according to optimal F-measure classifier after solving a series of cost-sensitive feature selection sub-problems. The features selected by our method will fully represent the characteristics of not only majority classes, but also minority classes. Extensive experimental results conducted on synthetic, multi-class and multi-label datasets validate the efficiency and significance of our feature selection method. Meng Liu 0003, Chang Xu 0002, Yong Luo 0002, Chao Xu 0006, Yonggang Wen 0001, Dacheng Tao |
AAAI | 5 |
| 2017 | GraphH: High Performance Big Graph Analytics in Small ClustersabstractIt is common for real-world applications to analyze big graphs using distributed graph processing systems. Popular in-memory systems require an enormous amount of resources to handle big graphs. While several out-of-core approaches have been proposed for processing big graphs on disk, the high disk I/O overhead could significantly reduce performance. In this paper, we propose GraphH to enable high-performance big graph analytics in small clusters. Specifically, we design a two-stage graph partition scheme to evenly divide the input graph into partitions, and propose a GAB (Gather-Apply-Broadcast) computation model to make each worker process a partition in memory at a time. We use an edge cache mechanism to reduce the disk I/O overhead, and design a hybrid strategy to improve the communication performance. GraphH can efficiently process big graphs in small clusters or even a single commodity server. Extensive evaluations have shown that GraphH could be up to 7.8x faster compared to popular in-memory systems, such as Pregel+ and PowerGraph when processing generic graphs, and more than 100x faster than recently proposed out-of-core systems, such as GraphD and Chaos when processing big graphs. Peng Sun 0006, Yonggang Wen 0001, Ta Nguyen Binh Duong, Xiaokui Xiao |
CLUSTER | 2 |
| 2017 | GraphMP: An Efficient Semi-External-Memory Big Graph Processing System on a Single MachineabstractRecent studies showed that single-machine graph processing systems can be as highly competitive as clusterbased approaches on large-scale problems. While several out-of-core graph processing systems and computation models have been proposed, the high disk I/O overhead could significantly reduce performance in many practical cases. In this paper, we propose GraphMP to tackle big graph analytics on a single machine. GraphMP achieves low disk I/O overhead with three techniques. First, we design a vertex-centric sliding window (VSW) computation model to avoid reading and writing vertices on disk. Second, we propose a selective scheduling method to skip loading and processing unnecessary edge shards on disk. Third, we use a compressed edge cache mechanism to fully utilize the available memory of a machine to reduce the amount of disk accesses for edges. Extensive evaluations have shown that GraphMP could outperform state-of-the-art systems such as GraphChi, X-Stream and GridGraph by 31.6x, 54.5x and 23.1x respectively, when running popular graph applications on a billion-vertex graph. Peng Sun 0006, Yonggang Wen 0001, Ta Nguyen Binh Duong, Xiaokui Xiao |
ICPADS | 2 |
| 2017 | Exploiting High-Order Information in Heterogeneous Multi-Task Feature LearningabstractMulti-task feature learning (MTFL) aims to improve the generalization performance of multiple related learning tasks by sharing features between them. It has been successfully applied to many pattern recognition and biometric prediction problems. Most of current MTFL methods assume that different tasks exploit the same feature representation, and thus are not applicable to the scenarios where data are drawn from heterogeneous domains. Existing heterogeneous transfer learning (including multi-task learning) approaches handle multiple heterogeneous domains by usually learning feature transformations across different domains, but they ignore the high-order statistics (correlation information) which can only be discovered by simultaneously exploring all domains. We therefore develop a tensor based heterogeneous MTFL (THMTFL) framework to exploit such high-order information. Specifically, feature transformations of all domains are learned together, and finally used to derive new representations. A connection between all domains is built by using the transformations to project the pre-learned predictive structures of different domains into a common subspace, and minimizing their divergence in the subspace. By exploring the high-order information, the proposed THMTFL can obtain more reliable feature transformations compared with existing heterogeneous transfer learning approaches. Extensive experiments on both text categorization and social image annotation demonstrate superiority of the proposed method. Yong Luo 0002, Dacheng Tao, Yonggang Wen 0001 |
IJCAI | 3 |
| 2017 | General Heterogeneous Transfer Distance Metric Learning via Knowledge Fragments TransferabstractTransfer learning aims to improve the performance of target learning task by leveraging information (or transferring knowledge) from other related tasks. Recently, transfer distance metric learning (TDML) has attracted lots of interests, but most of these methods assume that feature representations for the source and target learning tasks are the same. Hence, they are not suitable for the applications, in which the data are from heterogeneous domains (feature spaces, modalities and even semantics). Although some existing heterogeneous transfer learning (HTL) approaches is able to handle such domains, they lack flexibility in real-world applications, and the learned transformations are often restricted to be linear. We therefore develop a general and flexible heterogeneous TDML (HTDML) framework based on the knowledge fragment transfer strategy. In the proposed HTDML, any (linear or nonlinear) distance metric learning algorithms can be employed to learn the source metric beforehand. Then a set of knowledge fragments are extracted from the pre-learned source metric to help target metric learning. In addition, either linear or nonlinear distance metric can be learned for the target domain. Extensive experiments on both scene classification and object recognition demonstrate superiority of the proposed method. Yong Luo 0002, Yonggang Wen 0001, Tongliang Liu, Dacheng Tao |
IJCAI | 2 |
| 2017 | QDLCoding: QoS-differentiated low-cost video encoding scheme for online video serviceabstractAdaptive bitrate (ABR) streaming is the de facto solution in online video services to cope with heterogeneous devices and varying network connections. However, this solution is computation intensive, demanding a large number of servers for encoding videos. Moreover, due to the time-varying nature of video generation, intelligent strategies are required in order to determine the right amount of resources for encoding. The situation is further complicated by the fact that, the two types of co-existing video content, live content and Video-on-Demand (VoD) content, have different QoS requirements for encoding. These observations posit daunting challenges for meeting the heterogeneous QoS requirements with a minimum computing capacity. This paper proposes the QoS-differentiated low-cost video encoding (QDLCoding) scheme to address these challenges. We develop a framework for scheduling the encoding workloads of the two types of videos with statistical QoS guarantees. Each type of videos is specified with a QoS criterion and a QoS loss bound. The objective is to provision the minimum amount of resources while keeping the QoS loss probabilities within the prescribed bounds. We design an online algorithm that can determine the minimum required capacity by learning content arrival distributions. The experiment results demonstrate that our method can greatly reduce the required capacity for encoding online videos while controlling the likelihood of QoS loss precisely. Guanyu Gao, Yonggang Wen 0001, Han Hu 0003 |
INFOCOM | 2 |
| 2017 | MUSA: Wi-Fi AP-assisted video prefetching via Tensor LearningabstractDriven by the exponentially increasing amount of mobile video traffic, caching videos closer to the end users has become an appealing solution to reduce the traffic through the backbone network while improving users' perceived quality-of-experience (e.g., better video quality and reduced service delay). This research interest has been gaining lots of momentums due to the emergence of smart Access Points (APs), which are equipped with large storage space (several GBs). To address the “small population” problem involved in the prefetching at the edge, we propose to prefetch videos to APs ahead of users' requests via tensor learning: We first adopt the weighted tensor model to mine the hidden semantic pattern to characterize both users' preference for different types of videos and the dynamic video popularity over time; Then, based on the resulting low-dimension matrixes generated by the tensor factorization, we adopt an exponential smoothing model to capture the temporal pattern to predict users' propensity to unwatched videos; Finally, based on the predicted video popularity, we proactively replicate videos from the original CDN server to the APs at the edge. Through trace-driven simulations, we show that the proposed prefetching solution can outperform the baseline algorithms: compared with the SVD-based prefetching strategy, our design achieves a better hit ratio (e.g., surpassing about 10%) and accuracy (e.g., surpassing about 15%); compared with the history based strategy, our design also have about 40% (resp. 20%) improvement in terms of hit ratio (resp. accuracy). Wen Hu 0003, Zhi Wang 0001, Peng Wang 0012, Yonggang Wen 0001, Kaiyan Chu, Lifeng Sun |
IWQoS | 6 |
| 2017 | Towards Distributed Machine Learning in Shared Clusters: A Dynamically-Partitioned ApproachabstractMany cluster management systems (CMSs) have been proposed to share a single cluster with multiple distributed computing systems. However, none of the existing approaches can handle distributed machine learning (ML) workloads given the following criteria: high resource utilization, fair resource allocation and low sharing overhead. To solve this problem, we propose a new CMS named Dorm, incorporating a dynamically-partitioned cluster management mechanism and an utilization-fairness optimizer. Specifically, Dorm uses the container-based virtualization technique to partition a cluster, runs one application per partition, and can dynamically resize each partition at application runtime for resource efficiency and fairness. Each application directly launches its tasks on the assigned partition without petitioning for resources frequently, so Dorm imposes flat sharing overhead. Extensive performance evaluations showed that Dorm could simultaneously increase the resource utilization by a factor of up to 2.32, reduce the fairness loss by a factor of up to 1.52, and speed up popular distributed ML applications by a factor of up to 2.72, compared to existing approaches. Dorm's sharing overhead is less than 5% in most cases. Peng Sun 0006, Yonggang Wen 0001, Ta Nguyen Binh Duong, Shengen Yan |
SMARTCOMP | 2 |
| 2017 | Toward Joint Compression-Transmission Optimization for Green Wearable Devices: An Energy-Delay TradeoffabstractSmall-size and light-weight, as the modern design concept for the emerging wearable devices, has become a trend. However, such trend puts physical limitations to the battery, and the resulting short battery lifetime becomes the bottleneck for most wearable devices today. In this paper, we aim to optimize the energy usage through data compression and transmission rate control. We propose a novel joint compression-transmission approach, which not only minimizes the energy consumption of both compression and transmission, but also maintains the corresponding data distortion and transmission delay within a certain tolerant level. By adopting the Lyapunov framework, we develop an online algorithm to minimize the one-slot drift-plus-penalty function. We conduct numerical analysis and experimental study for our proposed approach. The results show that the size of queuing buffer has the significant impact on the energy cost. Next, we verify a fundamental tradeoff between the energy expenditure and transmission delay, and derive the theoretical performance bounds. After that, we show that the energy cost is also determined by the wireless channel gain and the data compression ratio. Finally, compared to a strategy without compression, our approach can save up to 92% of energy. Weizheng Hu, Wei Zhang 0082, Han Hu 0003, Yonggang Wen 0001, King-Jet Tseng |
IEEE Internet Things J. | 4 |
| 2017 | Guest Editorial Multimedia Communication in the Internet of ThingsabstractMultimedia communication in the Internet of Things (IoT) can potentially reach into a vast array of areas and touch people’s lives in profound and different ways. For example, real-time multimedia communication could be applied in the current U.S. 911 system to provide responders with detailed information about the nature and severity of an incident before they arrive on the scene, if the callers can transmit image and/or video of the incident site. City governments can also allow citizens to report traffic and road conditions by uploading real-time multimedia data via a specific smartphone app. Qing Yang 0003, Honggang Wang 0001, Mischa Dohler, Yonggang Wen 0001, Guoliang Xue |
IEEE Internet Things J. | 4 |
| 2017 | Public Cloud Storage-Assisted Mobile Social Video Sharing: A Supermodular Game ApproachabstractMobile social video sharing enables mobile users to create ultra-short video clips and instantly share them with social friends, which poses significant pressure to the content distribution infrastructure. In this paper, we propose a public cloud-assisted architecture to tackle this problem. In particular, by motivating mobile users to upload videos to the local public cloud to serve requests, and, therefore, having a permission to access friends' videos stored in the cloud, our method can alleviate the traffic burden to the social service providers, while reducing the service latency of mobile users. First, we present a general framework to model the information diffusion and utility function of each user on the proposed architecture, and formulate the problem as a decentralized social utility maximization game. Second, we show that this problem is a supermodular game and there exists at least one socially aware Nash equilibrium (SNE). We then develop two decentralized algorithms to solve this problem. The first algorithm can find an SNE with less computation complexity, and the second algorithm can find the Pareto-optimal SNE with better performance. Finally, through extensive experiments, we demonstrate that the overall system performance can be significantly improved by exploiting the selflessness among social friends. Han Hu 0003, Yonggang Wen 0001, Dusit Niyato |
IEEE J. Sel. Areas Commun. | 2 |
| 2017 | Spectrum Allocation and Bitrate Adjustment for Mobile Social Video Sharing: Potential Game With Online QoS Learning ApproachabstractWith the recent progress on mobile networking and devices, mobile social video sharing (MSVS) has emerged as one of the most important social media services. It enables mobile users to create ultra-short video clips and instantly share them with social friends. Due to the huge volume of videos and limited available bandwidth of wireless infrastructure, it is challenging to distribute these massive videos to mobile users with satisfactory quality of service (QoS). In this paper, we present a general framework to model the video diffusion among mobile users and user QoS of the MSVS service over the wireless infrastructure. Then, we utilize the hierarchical structure to decompose this problem into two subproblems, including a bitrate adjustment and spectrum allocation problems. For the bitrate adjustment problem, we propose a QoS estimation model based on the large deviation principle. By introducing a sliding window method to derive the online estimation, we develop an online bitrate adjustment strategy without relying on any prior knowledge of neither network environment nor video traffic. For the spectrum allocation problem, we prove that such a problem is a potential game. We devise a decentralized algorithm to find the Nash equilibrium, and analyze the convergence rate and the performance gap with the centralized optimization solution. Through extensive real trace driven simulations, we demonstrate that our proposed algorithm can guarantee smooth video playback with a higher PSNR. Han Hu 0003, Yonggang Wen 0001, Dusit Niyato |
IEEE J. Sel. Areas Commun. | 2 |
| 2017 | Editorial: Mobile Multimedia Communications
Zheng Yan 0002, Wei Wang 0015, Yonggang Wen 0001, Chonggang Wang, Honggang Wang 0001 |
Mob. Networks Appl. | 3 |
| 2017 | Energy consumption analysis of data stream processing: a benchmarking approachabstractSummary Energy efficiency of data analysis systems has become a very important issue in recent times because of the increasing costs of data center operations. Although distributed streaming workloads have increasingly been present in modern data centers, energy‐efficient scheduling of such applications remains as a significant challenge. In this paper, we conduct an energy consumption analysis of data stream processing systems in order to identify their energy consumption patterns. We follow stream system benchmarking approach to solve this issue. Specifically, we implement Linear Road benchmark on six stream processing environments (S4, Storm, ActiveMQ, Esper, Kafka, and Spark Streaming) and characterize these systems' performance on a real‐world data center. We study the energy consumption characteristics of each system with varying number of roads as well as with different types of component layouts. We also use a microbenchmark to capture raw energy consumption characteristics. We observed that S4, Esper, and Spark Streaming environments had highest average energy consumption efficiencies compared with the other systems. Using a neural networkbased technique with the power/performance information gathered from our experiments, we developed a model for the power consumption behavior of a streaming environment. We observed that energy‐efficient execution of streaming application cannot be specifically attributed to the system CPU usage. We observed that communication between compute nodes with moderate tuple sizes and scheduling plans with balanced system overhead produces better power consumption behaviors in the context of data stream processing systems. Copyright © 2016 John Wiley & Sons, Ltd. Miyuru Dayarathna, Yonggang Wen 0001, Rui Fan 0004 |
Softw. Pract. Exp. | 3 |
| 2017 | Delay-Optimized File Retrieval under LT-Based Cloud StorageabstractFountain-code based cloud storage system provides reliable online storage solution through placing unlabeled content blocks into multiple storage nodes. Luby Transform (LT) code is one of the popular fountain codes for storage systems due to its efficient recovery. However, to ensure high success decoding of fountain codes based storage, retrieval of additional fragments is required, and this requirement could introduce additional delay. In this paper, we show that multiple stage retrieval of fragments is effective to reduce the file-retrieval delay. We first develop a delay model for various multiple stage retrieval schemes applicable to our considered system. With the developed model, we study optimal retrieval schemes given requirements on success decodability. Our numerical results suggest a fundamental tradeoff between the file-retrieval delay and the target probability of successful file decoding, and that the file-retrieval delay can be significantly reduced by optimally scheduling packet requests in a multi-stage fashion. Haifeng Lu, Chuan Heng Foh, Yonggang Wen 0001, Jianfei Cai 0001 |
IEEE Trans. Cloud Comput. | 3 |
| 2017 | Guest Editorial Special Issue on Visual Computing in the Cloud: Mobile ComputingabstractRecent advances in mobile devices (e.g., smartphones and wearables) and wireless technologies are fueling a new wave of user demands for an improved user experience. Indeed, users are not only expecting ubiquitous network connections for traditional services (e.g., messaging and calling), but also demanding extensive access to a wealth of video contents and services. However, this growing demand is seriously hindered by the fact that the onboard resources with mobile devices are inherently limited and their growth rate falls behind that of their desktop counterparts. It follows that new solutions should be in order to resolve this fundamental tussle. Fortunately, the emerging cloud computing offers a natural solution to extend the desktop visual experience to mobile devices. It actually provides both computational and storage support for media-rich applications with both front-end and back-end functionalities. Yonggang Wen 0001, Jacob Chakareski, Pascal Frossard, Di Wu 0001, Wenjun Zeng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2017 | Facial Age Estimation With Age DifferenceabstractAge estimation based on the human face remains a significant problem in computer vision and pattern recognition. In order to estimate an accurate age or age group of a facial image, most of the existing algorithms require a huge face data set attached with age labels. This imposes a constraint on the utilization of the immensely unlabeled or weakly labeled training data, e.g., the huge amount of human photos in the social networks. These images may provide no age label, but it is easy to derive the age difference for an image pair of the same person. To improve the age estimation accuracy, we propose a novel learning scheme to take advantage of these weakly labeled data through the deep convolutional neural networks. For each image pair, Kullback-Leibler divergence is employed to embed the age difference information. The entropy loss and the cross entropy loss are adaptively applied on each image to make the distribution exhibit a single peak value. The combination of these losses is designed to drive the neural network to understand the age gradually from only the age difference information. We also contribute a data set, including more than 100 000 face images attached with their taken dates. Each image is both labeled with the timestamp and people identity. Experimental results on two aging face databases show the advantages of the proposed age difference learning system, and the state-of-the-art performance is gained. Zhenzhen Hu 0004, Yonggang Wen 0001, Meng Wang 0001, Richang Hong, Shuicheng Yan |
IEEE Trans. Image Process. | 2 |
| 2017 | Cost-Optimized Microblog Distribution over Geo-Distributed Data Centers: Insights from Cross-Media AnalysisabstractThe unprecedent growth of microblog services poses significant challenges on network traffic and service latency to the underlay infrastructure (i.e., geo-distributed data centers). Furthermore, the dynamic evolution in microblog status generates a huge workload on data consistence maintenance. In this article, motivated by insights of cross-media analysis-based propagation patterns, we propose a novel cache strategy for microblog service systems to reduce the inter-data center traffic and consistence maintenance cost, while achieving low service latency. Specifically, we first present a microblog classification method, which utilizes the external knowledge from correlated domains, to categorize microblogs. Then we conduct a large-scale measurement on a representative online social network system to study the category-based propagation diversity on region and time scales. These insights illustrate social common habits on creating and consuming microblogs and further motivate our architecture design. Finally, we formulate the content cache problem as a constrained optimization problem. By jointly using the Lyapunov optimization framework and simplex gradient method, we find the optimal online control strategy. Extensive trace-driven experiments further demonstrate that our algorithm reduces the system cost by 24.5% against traditional approaches with the same service latency. Han Hu 0003, Yonggang Wen 0001, Tat-Seng Chua, Xuelong Li 0001 |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2017 | Visual Classification of Furniture StylesabstractFurniture style describes the discriminative appearance characteristics of furniture. It plays an important role in real-world indoor decoration. In this article, we explore the furniture style features and study the problem of furniture style classification. Differing from traditional object classification, furniture style classification aims at classifying different furniture in terms of the “style” that describes its appearance (e.g., American style, Gothic style, Rococo style, etc.) rather than the “kind” that is more related to its functional structure (e.g., bed, desk, etc.). To pursue efficient furniture style features, we construct a novel dataset of furniture styles that contains 16 common style categories and implement three strategies with respect to two categories of classification, that is, handcrafted classification and learning-based classification. First, we follow the typical image classification pipeline to extract the handcrafted features and train the classifier by support vector machine. Then we use the convolutional neural network to extract learning-based features from training images. To obtain comprehensive furniture style features, we finally combine the handcrafted image classification pipeline and the learning-based network. We experimentally evaluate the performances of handcrafted features and learning-based features of each strategy, and the results show the superiority of learning-based features and also the comprehensiveness of handcrafted features. Zhenzhen Hu 0004, Yonggang Wen 0001, Luoqi Liu, Richang Hong, Meng Wang 0001, Shuicheng Yan |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2017 | Energy-Efficient Mobile Video Streaming: A Location-Aware ApproachabstractVideo streaming is one of the most widely used mobile applications today, and it also accounts for a large fraction of mobile battery usage. Much of the energy consumption is for wireless data transmission and is highly correlated to network bandwidth conditions. In periods of poor connectivity, up to 90% of mobile energy can be used for wireless data transfer. In this article, we study the problem of energy-efficient mobile video streaming. We make use of the observed correlation between bandwidth and user location , and also observe that a user’s location is predictable in many situations, such as when commuting to a known destination. Based on the user’s predicted locations and bandwidth conditions, we optimize wireless transmission times to achieve high quality video playback while minimizing energy use. We propose an optimal offline algorithm for this problem, which runs in O ( Tk ) time, where T is the duration of the video and k is the size of the video buffer. We also propose LAWS, a Location AWare Streaming algorithm. LAWS learns from historical location-aware bandwidth conditions and predicts future bandwidths along a planned route to make online wireless download decisions. We evaluate LAWS using real bandwidth traces, and show that LAWS closely approximates the performance of the optimal offline algorithm, achieving 90.6% of the optimal performance on average, and 97% in certain cases. LAWS also outperforms three popular strategies used in practice by, on average, 69%, 63%, and 38%, respectively. Lastly, we show that LAWS is able to deal with noisy data and can attain the stated performance after sampling bandwidth conditions only five times. Wei Zhang 0082, Rui Fan 0004, Yonggang Wen 0001, Fang Liu 0009 |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2017 | Resource Provisioning and Profit Maximization for Transcoding in Clouds: A Two-Timescale ApproachabstractTranscoding is widely adopted for content adaptation; however, it may incur excessive resource consumption and processing delays. Taking advantage of cloud infrastructure, cloud-based transcoding can elastically allocate resources under time-varying workloads and perform multiple transcodings in parallel to reduce delays. To provide transcoding as a cloud service, cloud transcoding systems require some intelligent mechanisms to provision resources and schedule tasks to satisfy user requirements while maximizing financial profit. To this end, we propose a two-timescale stochastic optimization framework for maximizing service profit while achieving performance requirements by jointly provisioning resources and scheduling tasks under a hierarchical control architecture. Our method analytically integrates service revenue, processing delay, and resource consumption in one optimization framework. We derive the offline exact solution and design some approximate online solutions for task scheduling and resource provisioning. We implement an open source cloud transcoding system, called Morph, and evaluate the performance of our method in a real environment. Empirical studies verify that our method can reduce resource consumption and achieve a higher profit compared with baseline schemes. Guanyu Gao, Han Hu 0003, Yonggang Wen 0001, Cédric Westphal |
IEEE Trans. Multim. | 3 |
| 2016 | Toward Effortless TV-to-Online (T2O) Experience: A Novel Metric Learning ApproachabstractShopping of the same or similar types of products as shown in the online TV programs has been highly desired by many people, especially the youth. To meet this eminent market need, we develop a prototype system to enable effortless TV-to-Online (T2O) experience. A key component of this system is the product search that maps specific items embedded in the video into a list of online merchants. The search performance mainly depends on the estimation of the distance or similarity between the queried item and all the curated items in the database. The simple Euclidean (EU) distance cannot capture the data characteristics, we therefore introduce distance metric learning (DML) to improve the distance estimation. Traditional DML methods only utilize the side information (e.g., similar/dissimilar constraints or relevance/irrelevance judgements) in the target domain, and may fail when the side information is scarce. Transfer metric learning (TML) can be adopted to leverage the side information from related domains. In this paper, we treat each domain equally and propose a novel Ranking-based Heterogeneous Multi-Task Metric Learning (RHMTML) framework, which adopts ranking-based loss, so the learned metric is particularly suitable for search. Extensive experiments demonstrate the effectiveness of our proposed method. We foresee that our approach would transform the emerging T2O trend in both TV and online video market. Yong Luo 0002, Yonggang Wen 0001, Dacheng Tao, Qiang Fu 0006 |
GLOBECOM | 2 |
| 2016 | Towards cost-efficient workload scheduling for a Tango between geo-distributed data center and power gridabstractNowadays, data centers consume substantial power, which takes up a considerable portion of local power supply (e.g., smart grid). In this paper, we leverage data center workload scheduling for the coordination between data centers and the smart grid, aiming to reduce the electricity cost of data centers and smooth the load variation of the smart grid simultaneously. We first build cost models of workload scheduling at data centers and the power generation and variation at the smart grid. We formulate the objective function as a weighted sum of the cost of the smart grid and the penalty caused by workload scheduling. Using the dual decomposition method, we then derive the optimal offline solution. To facilitate online implementation, we finally propose a Receding Horizon Control (RHC) based algorithm to obtain the suboptimal solution using limited predicted information. Extensive simulation results show that our proposed scheme can significantly reduce the cost of the smart grid, by up to 20%, while smoothing the load variation simultaneously. Han Hu 0003, Yonggang Wen 0001, Ling Qiu 0003 |
ICC | 2 |
| 2016 | Interference based virtual network embeddingabstractVirtual Network Embedding (VNE) is a key step towards network virtualization. In this paper, we first introduce a new link interference metric for each link to quantify the interference caused by its bandwidth scarcity to accept VN requests, and then an Interference-based VNE (I-VNE) algorithm is proposed. Benefited from the new metric, I-VNE can jointly consider the temporal and spatial topology information of networks, and tries to embed each virtual network request with low interference to avoid rejecting the future requests. Our simulations show that, I-VNE can significantly improve performance in terms of time-average revenue, acceptance ratio and average node utilization with more unbalanced average link utilization, compared with the three existing VNE algorithms with only the global resource information in the spatial dimension. Zheng Chen 0003, Ling Qiu 0003, Yonggang Wen 0001 |
ICC | 4 |
| 2016 | Tensor canonical correlation analysis for multi-view dimension reductionabstractCanonical correlation analysis (CCA) has proven an effective tool for two-view dimension reduction due to its profound theoretical foundation and success in practical applications. In respect of multi-view learning, however, it is limited by its capability of only handling data represented by two-view features, while in many real-world applications, the number of views is frequently many more. Although the ad hoc way of simultaneously exploring all possible pairs of features can numerically deal with multi-view data, it ignores the high order statistics (correlation information) which can only be discovered by simultaneously exploring all features. Therefore, in this work, we develop tensor CCA (TCCA) which straightforwardly yet naturally generalizes CCA to handle the data of an arbitrary number of views by analyzing the covariance tensor of the different views. TCCA aims to directly maximize the canonical correlation of multiple (more than two) views. Crucially, we prove that the main problem of multiview canonical correlation maximization is equivalent to finding the best rank-1 approximation of the data covariance tensor, which can be solved efficiently using the well-known alternating least squares (ALS) algorithm. As a consequence, the high order correlation information contained in the different views is explored and thus a more reliable common subspace shared by all features can be obtained. Yong Luo 0002, Dacheng Tao, Kotagiri Ramamohanarao, Chao Xu 0006, Yonggang Wen 0001 |
ICDE | 5 |
| 2016 | Timed Dataflow: Reducing Communication Overhead for Distributed Machine Learning SystemsabstractMany distributed machine learning (ML) systems exhibit high communication overhead when dealing with big data sets. Our investigations showed that popular distributed ML systems could spend about an order of magnitude more time on network communication than computation to train ML models containing millions of parameters. Such high communication overhead is mainly caused by two operations: pulling parameters and pushing gradients. In this paper, we propose an approach called Timed Dataflow (TDF) to deal with this problem via reducing network traffic using three techniques: a timed parameter storage system, a hybrid parameter filter and a hybrid gradient filter. In particular, the timed parameter storage technique and the hybrid parameter filter enable servers to discard unchanged parameters during the pull operation, and the hybrid gradient filter allows servers to drop gradients selectively during the push operation. Therefore, TDF could reduce the network traffic and communication time significantly. Extensive performance evaluations in a real testbed showed that TDF could reduce up to 77% and 79% of network traffic for the pull and push operations, respectively. As a result, TDF could speed up model training by a factor of up to 4 without sacrificing much accuracy for some popular ML models, compared to systems not using TDF. Peng Sun 0006, Yonggang Wen 0001, Ta Nguyen Binh Duong, Shengen Yan |
ICPADS | 2 |
| 2016 | Balanced Hashing and Efficient GPU Sparse General Matrix-Matrix MultiplicationabstractGeneral sparse matrix-matrix multiplication (SpGEMM) is a core component of many algorithms. A number of recent works have used high throughput graphics processing units (GPUs) to accelerate SpGEMM. However, exploiting the power of GPUs for SpGEMM requires addressing a number of challenges, including highly imbalanced workloads and large numbers of inefficient random global memory accesses. This paper presents a SpGEMM algorithm which uses several novel techniques to overcome these problems. We first propose two low cost methods to achieve perfect load balancing during the most expensive step in SpGEMM. Next, we show how to eliminate nearly all random global memory accesses using shared memory based hash tables. To optimize the performance of the hash tables, we propose a lightweight method to estimate the number of nonzeros in the output matrix. We compared our algorithm to the CUSP, CUSPARSE and the state-of-the-art BHSPARSE GPU SpGEMM algorithms, and show that it performs 5.6x, 2.4x and 1.5x better on average, and up to 11.8x, 9.5x and 2.5x better in the best case, respectively. Furthermore, we show that our algorithm performs especially well on highly imbalanced and unstructured matrices. Pham Nguyen Quang Anh, Rui Fan 0004, Yonggang Wen 0001 |
ICS | 3 |
| 2016 | On Combining Side Information and Unlabeled Data for Heterogeneous Multi-Task Metric Learning
Yong Luo 0002, Yonggang Wen 0001, Dacheng Tao |
IJCAI | 2 |
| 2016 | Morph: A Fast and Scalable Cloud Transcoding SystemabstractMorph is an open source cloud transcoding system. It can leverage the scalability of the cloud infrastructure to encode and transcode video contents in fast speed, and dynamically provision the resources in cloud to accommodate the workload. The system is composed of a master node that performs the video file segmentation, concentration, and task scheduling operations; and multiple worker nodes that perform the transcoding for video blocks. Morph can transcode the video blocks of a video file on multiple workers in parallel to achieve fast speed, and automatically manage the data transfers and communications between the master node and the worker nodes. The worker nodes can join into or leave the transcoding cluster at any time for dynamic resource provisioning. The system is very modular, and all of the algorithms can be easily modified or replaced. We release the source code of Morph under MIT License, hoping that it can be shared among various research communities. Guanyu Gao, Yonggang Wen 0001 |
ACM Multimedia | 2 |
| 2016 | Dynamic Resource Provisioning with QoS Guarantee for Video Transcoding in Online Video Sharing ServiceabstractVideo transcoding is widely adopted in online video sharing services to encode video content into multiple representations. This solution, however, could consume huge amount of computing resource and incur excessive processing delays. Moreover, content has heterogeneous QoS requirements for transcoding. Some content must be transcoded in real time, while some are deferrable for transcoding. It needs to determine the strategy for intelligently provisioning the right amount of resource under dynamic workload to meet the heterogeneous QoS requirements. To this end, this paper develops a robust dynamic resource provisioning scheme for transcoding with heterogeneous QoS criteria. We adopt the Preemptive Resume Priority discipline for scheduling, so that the transcoding-deferrable content can utilize idle resources for transcoding to maximize resource utilization while remain transparent to delay-sensitive content. We leverage Model Predictive Control to design the online algorithm for dynamic resource provisioning using predictions to accommodate time-varying workload. To seek robustness of system performance against prediction noises, we improve our online algorithm through Robust Design. The experiment results in a real environment demonstrate that our proposed framework can achieve the QoS requirements while reducing 50% of resource consumption on average. Guanyu Gao, Yonggang Wen 0001, Cédric Westphal |
ACM Multimedia | 2 |
| 2016 | Manifold regularized multi-view feature selection for social image annotation
Yangxi Li, Cuilan Du, Yang Liu 0003, Yonggang Wen 0001 |
Neurocomputing | 5 |
| 2016 | Multicolumn Bidirectional Long Short-Term Memory for Mobile Devices-Based Human Activity RecognitionabstractThe ever-growing popularity of mobile devices equipped with accelerometers has provided the opportunity to capture the semantic aspects of human activity and improve user experiences with behavior-based recommendations. These functions depend heavily on the accuracy of human activity recognition, and thus real applications that use mobile devices-based human activity recognition systems (MARSs) need to seamlessly incorporate the information carried by newly labeled training samples. Motivated by the success of the weightlessness feature, we propose a new two-directional feature for bidirectional long short-term memory (BLSTM) for incremental learning in human activity recognition. To further improve the performance, we also present a new ensemble classifier termed multicolumn BLSTM (MBLSTM), which effectively combines different acceleration signal features to further improve activity recognition accuracy. Experiments on the naturalistic mobile devices-based human activity dataset suggest that MBLSTM is superior to other state-of-the-art MARS methods. Dapeng Tao, Yonggang Wen 0001, Richang Hong |
IEEE Internet Things J. | 2 |
| 2016 | Big data meets multimedia analytics
Tat-Seng Chua, Xiangjian He, Weifeng Liu 0001, Massimo Piccardi, Yonggang Wen 0001, Dacheng Tao |
Signal Process. | 5 |
| 2016 | Introduction of New Associate EditorsabstractPresents a listing of the new Associate Editors for this issue of the publication. Nikolaos V. Boulgouris, David Bull 0001, Marco Cagnazzo, Andrea Cavallaro, Gene Cheung, Amit K. Roy-Chowdhury, Pedro Comesaña Alfaro, Sarp Ertürk, Markus Flierl, Gian Luca Foresti, Gang Hua 0001, Zhu Li 0001, Weisi Lin, Siwei Ma 0001, Pramod Kumar Meher, Debargha Mukherjee, Aleksandra Pizurica, Andrea Prati 0001, Paolo Remagnino, Arun Ross, Shin'ichi Satoh 0001, Andreas E. Savakis, Heiko Schwarz, Ling Shao 0001, Shervin Shirmohammadi, Giuseppe Valenzise, Meng Wang 0001, Zhou Wang 0001, Yonggang Wen 0001, Dong Xu 0001, Junsong Yuan 0001, Yuan Yuan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 30 |
| 2016 | Joint Content Replication and Request Routing for Social Video Distribution Over Cloud CDN: A Community Clustering MethodabstractThe increasing popularity of online social networks (OSNs) has been transforming the dissemination pattern of social video contents. We can utilize the social information propagation pattern to improve the efficiency of social video distribution. In this paper, motivated by the social community classification, we present a social video replication and user request dispatching mechanism in the cloud content delivery network architecture to reduce the system operational cost, while guaranteeing the averaged service latency. Specifically, we first present a community classification method that clusters social users with social relationships, close geolocations, and similar video watching interests into various communities. Then, we conduct a large-scale measurement on a real OSN system to study the diversities of social video propagation and the effectiveness of our communities on smoothing the diversity. Finally, we propose the community-based video replication and request dispatching strategy and formulate it as a constrained optimization problem. Based on a stochastic optimization framework, we derive an online solution and rigorously prove the optimality. We evaluate our algorithm on a real trace under realistic settings and demonstrate that our algorithm can reduce the monetary cost by 30% against traditional approaches with the same service latency. Han Hu 0003, Yonggang Wen 0001, Tat-Seng Chua, Wenwu Zhu 0001, Xuelong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2016 | Large Margin Multi-Modal Multi-Task Feature Extraction for Image ClassificationabstractThe features used in many image analysis-based applications are frequently of very high dimension. Feature extraction offers several advantages in high-dimensional cases, and many recent studies have used multi-task feature extraction approaches, which often outperform single-task feature extraction approaches. However, most of these methods are limited in that they only consider data represented by a single type of feature, even though features usually represent images from multiple modalities. We, therefore, propose a novel large margin multi-modal multi-task feature extraction (LM3FE) framework for handling multi-modal features for image classification. In particular, LM3FE simultaneously learns the feature extraction matrix for each modality and the modality combination coefficients. In this way, LM3FE not only handles correlated and noisy features, but also utilizes the complementarity of different modalities to further help reduce feature redundancy in each modality. The large margin principle employed also helps to extract strongly predictive features, so that they are more suitable for prediction (e.g., classification). An alternating algorithm is developed for problem optimization, and each subproblem can be efficiently solved. Experiments on two challenging real-world image data sets demonstrate the effectiveness and superiority of the proposed method. Yong Luo 0002, Yonggang Wen 0001, Dacheng Tao, Jie Gui, Chao Xu 0006 |
IEEE Trans. Image Process. | 2 |
| 2016 | Towards Information Diffusion in Mobile Social NetworksabstractThe emerging of mobile social networks opens opportunities for viral marketing. However, before fully utilizing mobile social networks as a platform for viral marketing, many challenges have to be addressed. In this paper, we address the problem of identifying a small number of individuals through whom the information can be diffused to the network as soon as possible, referred to as thediffusion minimizationproblem. Diffusion minimization under the probabilistic diffusion model can be formulated as an asymmetric$k$-center problem which is NP-hard, and the best known approximation algorithm for the asymmetric$k$-center problem has approximation ratio of$\log ^*n$and time complexity$O(n^5)$. Clearly, the performance and the time complexity of the approximation algorithm are not satisfiable in large-scale mobile social networks. To deal with this problem, we propose a community based algorithm and a distributed set-cover algorithm. The performance of the proposed algorithms is evaluated by extensive experiments on both synthetic networks and a real trace. The results show that the community based algorithm has the best performance in both synthetic networks and the real trace compared to existing algorithms, and the distributed set-cover algorithm outperforms the approximation algorithm in the real trace in terms of diffusion time. Zongqing Lu 0002, Yonggang Wen 0001, Weizhan Zhang, Guohong Cao |
IEEE Trans. Mob. Comput. | 2 |
| 2016 | Toward Cost-Efficient Content Placement in Media Cloud: Modeling and AnalysisabstractCloud-centric media network (CCMN) was previously proposed to provide cost-effective content distribution services for user-generated contents (UGCs) based on media cloud. CCMN service providers orchestrate cloud resources to deliver UGCs in a pay-per-use style, with an objective to minimize the operational monetary cost. The monetary cost depends on the actual usage of cloud resources (e.g., computing, storage, and bandwidth), which in turn, is affected by the content placement strategy. In this paper, we investigate this cost-optimal content placement problem. Specifically, it is formulated into a constrained optimization problem, in which the objective is to minimize the total monetary cost, with respect to the resource capacity. We tackle this problem via a two-step strategy. The first step focuses on the placement for a single content, which is mapped into a k-center problem. Using a graph-theoretic approach, we derive and verify a logarithmic model between the optimal mean hop distance from viewers to contents, and the optimal number of content replicas. The second step leverages this analytical result to solve the cost optimization problem, via a feasible direction method. The analysis is substantiated via numerical simulations, using a set of data traces from a top content website. This investigation suggests that the optimal number of content replica for each title follows a power-law distribution in respect to its popularity rank. Moreover, it reveals a fundamental tradeoff between the storage and bandwidth cost. Finally, compared to existing heuristics, our proposed algorithm is able to obtain the optimal placement strategy, with lower computational complexity. Yichao Jin 0002, Yonggang Wen 0001, Kyle Guan |
IEEE Trans. Multim. | 2 |
| 2016 | FUN Coding: Design and AnalysisabstractJoint FoUntain coding and Network coding (FUN) is proposed to boost information spreading over multi-hop lossy networks. The novelty of our FUN approach lies in combining the best features of fountain coding, intra-session network coding, and cross-next-hop network coding. This paper provides an in-depth study of FUN codes. First, we theoretically analyze the throughput of FUN codes. Second, we identify several practical issues that may undermine the actual performance, such as buffer overflow, and quantify the resulting throughput degradation. Finally, we propose a systematic design to overcome these issues. Simulation results in TDMA multi-hop networks show that our methods yield near-optimal throughput and are significantly better than fountain codes and existing network coding schemes. Huazi Zhang, Kairan Sun, Qiuyuan Huang, Yonggang Wen 0001, Dapeng Oliver Wu |
IEEE/ACM Trans. Netw. | 4 |
| 2016 | Rate-Adaptive Feedback With Bayesian Compressive Sensing in Multiuser MIMO Beamforming SystemsabstractMultiple-input multiple-output (MIMO) is a promising way to increase link capacity and energy efficiency in the next generation communication systems. However, the benefits of such an approach depend on proper channel state information (CSI) availability at the transmitter. The CSI is usually estimated at the receiver and fed back to the transmitter through a band-limited channel. Thus, an efficient feedback scheme is needed. In this paper, a comprehensive Bayesian compressive sensing (BCS) based feedback mechanism is proposed for time-varying spatially and temporally correlated vector autoregression (VAR) wireless channel, and the feedback rate distortion function is derived in closed form in statistics. The proposed BCS feedback scheme utilizes the sparse CSI features and prior knowledge to significantly compress the dimensionality of the feedback CSI. Furthermore, the relationship between the feedback rate and downlink capacity is derived in closed form in statistics to guide rate-adaptive feedback in MIMO system. We find out that the ergodic downlink capacity of a user is determined only by its own feedback rate in the proposed feedback scheme. Theoretical and simulation results all show that the proposed feedback scheme can realize efficient, rate-adaptive feedback based on downlink capacity requirement, and the proposed feedback performance is superior to other related works. Xin-Lin Huang, Jun Wu 0006, Yonggang Wen 0001, Fei Hu 0001, Yi Wang 0018, Tao Jiang 0002 |
IEEE Trans. Wirel. Commun. | 3 |
| 2015 | Low-Rank Multi-View Learning in Matrix Completion for Multi-Label Image ClassificationabstractMulti-label image classification is of significant interest due to its major role in real-world web image analysis applications such as large-scale image retrieval and browsing. Recently, matrix completion (MC) has been developed to deal with multi-label classification tasks. MC has distinct advantages, such as robustness to missing entries in the feature and label spaces and a natural ability to handle multi-label problems. However, current MC-based multi-label image classification methods only consider data represented by a single-view feature, therefore, do not precisely characterize images that contain several semantic concepts. An intuitive way to utilize multiple features taken from different views is to concatenate the different features into a long vector; however, this concatenation is prone to over-fitting and leads to high time complexity in MC-based image classification. Therefore, we present a novel multi-view learning model for MC-based image classification, called low-rank multi-view matrix completion (lrMMC), which first seeks a low-dimensional common representation of all views by utilizing the proposed low-rank multi-view learning (lrMVL) algorithm. In lrMVL, the common subspace is constrained to be low rank so that it is suitable for MC. In addition, combination weights are learned to explore complementarity between different views. An efficient solver based on fixed-point continuation (FPC) is developed for optimization, and the learned low-rank representation is then incorporated into MC-based image classification. Extensive experimentation on the challenging PASCAL VOC' 07 dataset demonstrates the superiority of lrMMC compared to other multi-label image classification approaches. Meng Liu 0003, Yong Luo 0002, Dacheng Tao, Chao Xu 0006, Yonggang Wen 0001 |
AAAI | 5 |
| 2015 | Manifold Regularized Transfer Distance Metric LearningabstractThe performance of many computer vision and machine learning algorithms are heavily depend on the distance metric between samples. It is necessary to exploit abundant of side information like pairwise constraints to learn a robust and reliable distance metric[2, 3]. Let D = {(xl i ,xj,yi j)} l i, j=1 denotes the labeled training set for the target task, wherein xi, x j ∈ Rd and yi j = ±1 indicates xl i and xl i are similar/dissimilar to each other. Then, a metric is usually learned to minimize the distance between the data from the same class and maximize their distance otherwise. This leads to the following loss function for learning the metric A: Haibo Shi, Yong Luo 0002, Chao Xu 0006, Yonggang Wen 0001 |
BMVC | 4 |
| 2015 | Cost-efficient and QoS-aware content management in media cloud: Implementation and evaluationabstractAdaptive bitrate streaming has been proposed to encode video contents into multiple versions for device heterogeneity and changing network conditions. This solution, however, could consume enormous computing and storage resource. In fact, only a small fraction of videos are frequently requested. Thus, caching multiple versions for unpopular contents is not cost efficient. In this paper, we design a cost-efficient and QoS-aware content management system for video streaming. The system consists of a set of streaming servers and a computing cluster, where streaming servers can cache video contents or transcode them in real time, and the computing cluster can perform transcoding tasks on behalf of streaming servers. Based on this architecture, to provide cost-efficient and QoS-aware video service, first, we design a cost-efficient content cache management module to minimize the operational cost, by dynamically determining whether a segment should be cached or transcoded on fly according to their popularity. Second, to reduce transcoding latency, we design a QoS-aware transcoding task delegation module to determine whether a transcoding task in streaming server should be delegated to the computing cluster according to the streaming server's workload. We implement the system and evaluate the performance in a real environment. The results demonstrate that our method can greatly reduce the operational cost and guarantee the QoS in providing video services. Guanyu Gao, Yonggang Wen 0001, Han Hu 0003 |
ICC | 2 |
| 2015 | Cloud-assisted collaborative execution for mobile applications with general task topologyabstractMobile cloud computing has been touted as an effective solution to extend the capabilities of resource-poor mobile devices for executing computation intensive applications. In this paper, we investigate cloud-assisted collaborative execution for mobile applications with general task topology to reduce the energy consumption on mobile devices. A mobile application consists of fine-grained tasks organized in general topology. Each task can be executed either on the mobile device or offloaded to the cloud for execution, which is referred to as collaborative task execution. We aim to minimize the energy consumption on the mobile device while meeting a time deadline, by strategically mapping the task execution to the mobile device or to the cloud. We formulate the collaborative task execution as a delay-constrained workflow scheduling problem. For the workflow scheduling, we first leverage partial critical path analysis (PCP) to find out the critical path formed by a set of critical parents, in which the critical parent is defined as the parent node of a task that results in the maximum value of the earliest start time of the task. Then, for each path, we find its sub-deadline and apply one-climb policy to schedule the tasks on the path, in which there exists at most one migration from the mobile device to the cloud if ever for the minimum energy consumption. Simulation results show that the proposed collaborative task execution can save energy consumption compared to the local execution and is more flexible than the remote execution. Yonggang Wen 0001 |
ICC | 2 |
| 2015 | Reducing Vector I/O for Faster GPU Sparse Matrix-Vector MultiplicationabstractSparse matrix-vector multiplication (Spiv) is an important kernel used in solving many scientific and engineering problems. The massive parallelism of graphics processing units (GPUs) makes them well suited for Spiv computations. However, fully utilizing the power of GPUs is challenging because Spiv makes a large number of scattered memory accesses which saturate the Gnu's memory bandwidth. Most previous works sought to address the bandwidth limitation by using efficient storage formats for the matrix. However, we show that for most matrices, a majority of the bandwidth is consumed by accesses to the vector. In this paper, we introduce two techniques to significantly decrease the I/O for vector accesses, by making novel use of the Gnu's fast shared memory. A key advantage of our vector optimizations is that they are complementary to existing matrix I/O optimizations, so that it is possible to use both techniques in conjunction. Furthermore, combining the optimizations requires only minor code changes. We demonstrate how to combine our techniques with the widely used CUSP Spiv algorithm and the currently highest performing yaSpMV algorithm to significantly improve both algorithms' performance. We experimented with a wide range of matrices, and show that the modified version of CUSP on average reduces vector I/O by 37% and reduces the total I/O by 31%, while the modified version of yaSpMV reduces the vector and total I/O by 36% and 31%, resp. We improve CUSP's total throughput by 14% on average and up to 77% for certain matrices, and improve yaSpMV's throughput by 12% on average and 35% for some matrices. Pham Nguyen Quang Anh, Rui Fan 0004, Yonggang Wen 0001 |
IPDPS | 3 |
| 2015 | Adaptive configuration of cloud video transcodingabstractCloud computing is emerging as a new paradigm which enables big data computing, including high quality media processing. However, considering the media dynamics on resource consumption and the QoS criteria, dynamically providing the cloud computing resource to meet the QoS requirements of media processing is not easy. The current cloud computing infrastructure usually employs auto-scaling to dynamically adjust the computing resource allocation, which is typically performed at relatively long time scale and cannot adapt to the dynamic changes of video arrivals or content changes at relatively short time scale. In this paper, we propose to adaptively configure the video transcoding mode to deal with the short-term transcoding QoS and computing resource mismatch problem. We formulate the problem as the one to minimize the output bit-rate with the queue stability constraint, for which we use the Lyapunov optimization framework to solve it. Simulation results show that, compared with the static configuration strategy, the proposed adaptive method achieves smooth transcoding QoS degradation when system load becomes heavier and much better transcoding delay performance. Ming Yang 0018, Jianfei Cai 0001, Yonggang Wen 0001, Chuan Heng Foh |
ISCAS | 4 |
| 2015 | Opportunities and Challenges of Global Network CamerasabstractSince the introduction of consumer digital cameras, user-created multimedia content has become increasingly popular. Digital cameras, together with inexpensive editing tools, and free hosting sites have made multimedia an integral part of everyday life. Today, hundreds of hours video are uploaded to hosting sites every minute. Video-on-demand through wireless networks and smartphones have profoundly changed how people consume multimedia content. Meanwhile, the widely deployed network cameras can provide live views of many parts of the world. These cameras can provide rich sources creating multimedia content. This panel will explore the opportunities and discuss the challenges using global network cameras for creating multimedia contents and understanding the world. Every year, millions of network cameras are deployed. The data from some of these network cameras are publicly available, continuously streaming live views of national parks, city halls, streets, highways, and shopping malls. A person may see multiple tourist attractions through these cameras, without leaving home. Researchers may observe the weather in different cities. Using the data from the cameras, it is possible to observe natural disasters, such as volcano eruption or tsunami, at a safe distance. News reporters may obtain instant views of an unfolding riot without risking their lives. A spectator may watch a celebration parade from multiple locations using the street cameras. Despite the many promising applications, the opportunities of using global network cameras for creating multimedia content have not been fully exploited. Joanna Batstone, Touradj Ebrahimi, Tiejun Huang 0001, Yung-Hsiang Lu, Yonggang Wen 0001 |
ACM Multimedia | 5 |
| 2015 | Towards joint resource allocation and routing to optimize video distribution over future internetabstractGiven the exploding growth of video traffic, efficient video distribution is essential to the future Internet. Therefore, how to optimize its networking cost is a critical research problem. In this paper, we introduce Network Function Virtualization (NFV) in conjunction with Software-Defined Networking (SDN) to minimize the cost via joint orchestration of caching, transcoding and routing functions. Specifically, we propose a two-step iterative approach. First, in NFV-based resource allocation phase, we maximize total cache hits by optimally allocating storage and computing resources for a giving routing policy. Second, in SDN-based routing phase, we minimize the networking cost by optimally configuring the routing matrix for a given resource placement. Finally, we analytically prove their iterative repeat converges to the joint optimum. Through extensive simulations, we verify its convergence, and performance gains compared with the optimal solution of either phase alone. By examining numerical results, we obtain some operational guidelines. From the resource allocation aspect, we should allocate more resources to the node with heavier request rate. From the routing aspect, for each node-server pair, the node should split the traffic across multiple paths with identical shortest hops if there are many, or use the shortest path alone if there is only one. Yichao Jin 0002, Yonggang Wen 0001, Cédric Westphal |
Networking | 2 |
| 2015 | Adaptive and scalable load balancing for metadata server cluster in cloud-scale file systems
Quanqing Xu, Rajesh Vellore Arumugam, Khai Leong Yong, Yonggang Wen 0001, Yew-Soon Ong, Weiya Xi |
Frontiers Comput. Sci. | 4 |
| 2015 | Preface
Wenwu Zhu 0001, Yonggang Wen 0001, Zhi Wang 0001 |
J. Comput. Sci. Technol. | 2 |
| 2015 | Improving Energy Efficiency for Mobile Media Cloud via Virtual Machine Consolidation
Liang Zhou 0002, Yichao Jin 0002, Yonggang Wen 0001 |
Mob. Networks Appl. | 4 |
| 2015 | SmartGW: Enabling Bandwidth-Efficient Group Watching in Cloud Social TV Systems
Zheng Xue, Di Wu 0001, Xueyan Xie, Yonggang Wen 0001 |
Mob. Networks Appl. | 4 |
| 2015 | Mobile cloud computing based privacy protection in location-based information survey applicationsabstractAbstract Nowadays, location‐based service (LBS) has become pervasive. Given its high utility value, LBS, however, presents serious privacy concerns for cautious users. In this paper, we investigate privacy preserving for location‐based information survey application, which calculates the geographic distribution of user's information. The design objective is twofold: (i) calculate an information distribution for a pool of mobile users and (ii) protect the location and value privacy of individual user, in the presence of malicious servers and possible corrupted users. Our proposed solution leverages a mobile cloud computing paradigm, in which each mobile device is replicated with a system‐level clone in cloud. The computing of distribution function is distributed among the set of cloud clones via a P2P protocol. We further enhance our basic scheme with the multiple aggregation mechanism, aiming to protect the correctness of the aggregate result from the active attacker. Compared to the approaches based on centralized server or aggregate proxy, our proposed scheme and its enhanced version are advantageous in avoiding single point of failure/attack, load balancing, and overhead reduction. Simulation results verify these advantages and the protection to the correctness of aggregate result and suggest that our proposed scheme is suitable for large‐scale applications. Copyright © 2014 John Wiley & Sons, Ltd. Hao Zhang 0016, Nenghai Yu, Yonggang Wen 0001 |
Secur. Commun. Networks | 3 |
| 2015 | Optimal Transcoding and Caching for Adaptive Streaming in Media Cloud: an Analytical ApproachabstractNowadays, large-scale video distribution feeds a significant fraction of the global Internet traffic. However, existing content delivery networks may not be cost efficient enough to distribute adaptive video streaming, mainly due to the lack of orchestration on storage, computing, and bandwidth resources. In this paper, we leverage Media Cloud to deliver on-demand adaptive video streaming services, where those resources can be dynamically scheduled in an on-demand fashion. Our objective is to minimize the total operational cost by optimally orchestrating multiple resources. Specifically, we formulate an optimization problem, by examining a three-way tradeoff between the caching, transcoding, and bandwidth costs, at each edge server. Then, we adopt a two-step approach to analytically derive the closed-form solution of the optimal transcoding configuration and caching space allocation, respectively, for every edge server. Finally, we verify our solution throughout extensive simulations. The results indicate that our approach achieves significant cost savings compared with the existing methods used in content delivery networks. In addition, we also find the optimal strategy and its benefits can be affected by a list of system parameters, including the unit cost of different resources, the hop distance to the origin server, the Zipf parameter of users' request patterns, and the settings of different bitrate versions for one segment. Yichao Jin 0002, Yonggang Wen 0001, Cédric Westphal |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2015 | Tensor Canonical Correlation Analysis for Multi-View Dimension ReductionabstractCanonical correlation analysis (CCA) has proven an effective tool for two-view dimension reduction due to its profound theoretical foundation and success in practical applications. In respect of multi-view learning, however, it is limited by its capability of only handling data represented by two-view features, while in many real-world applications, the number of views is frequently many more. Although the ad hoc way of simultaneously exploring all possible pairs of features can numerically deal with multi-view data, it ignores the high order statistics (correlation information) which can only be discovered by simultaneously exploring all features. Therefore, in this work, we develop tensor CCA (TCCA) which straightforwardly yet naturally generalizes CCA to handle the data of an arbitrary number of views by analyzing the covariance tensor of the different views. TCCA aims to directly maximize the canonical correlation of multiple (more than two) views. Crucially, we prove that the main problem of multi-view canonical correlation maximization is equivalent to finding the best rank-1 approximation of the data covariance tensor, which can be solved efficiently using the well-known alternating least squares (ALS) algorithm. As a consequence, the high order correlation information contained in the different views is explored and thus a more reliable common subspace shared by all features can be obtained. In addition, a non-linear extension of TCCA is presented. Experiments on various challenge tasks, including large scale biometric structure prediction, internet advertisement classification, and web image annotation, demonstrate the effectiveness of the proposed method. Yong Luo 0002, Dacheng Tao, Kotagiri Ramamohanarao, Chao Xu 0006, Yonggang Wen 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2015 | Towards Cost-Efficient Video Transcoding in Media Cloud: Insights Learned From User Viewing PatternsabstractVideo transcoding in an adaptive bitrate streaming (ABR) system is demanded to support video streaming over heterogenous devices and varying networks. However, it could incur a tremendous cost. Meanwhile, most viewers terminate viewing sessions within 20% of their durations; only a small fraction of each video is consumed. Built upon this user viewing pattern, we propose a Partial Transcoding Scheme for content management in media clouds. Particularly, each content is encoded into different bitrates and split into segments. Some of the segments are stored in cache, resulting in storage cost; others are transcoded online in the case of cache miss, resulting in computing cost. We aim to minimize the long-term overall cost by determining whether a segment should be cached or transcoded online. We formulate it as a constrained stochastic optimization problem. Leveraging Lyapunov optimization framework and Lagrangian relaxation, we design an online algorithm which can achieve the optimal solution within provable upper bounds. Experiments demonstrate that our proposed method can reduce 30% of operational cost, compared with the scheme of caching all the segments. Guanyu Gao, Yonggang Wen 0001, Zhi Wang 0001, Wenwu Zhu 0001 |
IEEE Trans. Multim. | 3 |
| 2015 | How Much to Coordinate? Optimizing In-Network Caching in Content-Centric NetworksabstractIn content-centric networks, it is challenging how to optimally provision in-network storage to cache contents, to balance the tradeoffs between the network performance and the provisioning cost. To address this problem, we first propose a holistic model for intradomain networks to characterize the network performance of routing contents to clients and the network cost incurred by globally coordinating the in-network storage capability. We then derive the optimal strategy for provisioning the storage capability that optimizes the overall network performance and cost, and analyze the performance gains via numerical evaluations on real network topologies. Our results reveal interesting phenomena; for instance, different ranges of the Zipf exponent can lead to opposite optimal strategies, and the tradeoffs between the network performance and the provisioning cost have great impacts on the stability of the optimal strategy. We also demonstrate that the optimal strategy can achieve significant gain on both the load reduction at origin servers and the improvement on the routing performance. Moreover, given an optimal coordination level ℓ*, we design a routing-aware content placement (RACP) algorithm that runs on a centralized server. The algorithm computes and assigns contents to each CCN router to store, which can minimize the overall routing cost, e.g., transmission delay or hop counts, to deliver contents to clients. By conducting extensive simulations using a large-scale trace dataset collected from a commercial 3G network in China, our results demonstrate that our caching scheme can achieve 4% to 22% latency reduction on average over the state-of-the-art caching mechanisms. Haiyong Xie 0001, Yonggang Wen 0001, Chi-Yin Chow, Zhi-Li Zhang |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2015 | Algorithms and Applications for Community Detection in Weighted NetworksabstractCommunity detection is an important issue due to its wide use in designing network protocols such as data forwarding in Delay Tolerant Networks (DTN) and worm containment in Online Social Networks (OSN). However, most of the existing community detection algorithms focus on binary networks. Since most networks are naturally weighted such as DTN or OSN, in this article, we address the problems of community detection in weighted networks, exploit community for data forwarding in DTN and worm containment in OSN, and demonstrate how community can facilitate these network designs. Specifically, we propose a novel community detection algorithm, and introduce two metrics: intra-centrality and inter-centrality, to characterize nodes in communities, based on which we propose an efficient data forwarding algorithm for DTN and a worm containment strategy for OSN. Extensive trace-driven simulation results show that the proposed community detection algorithm, the data forwarding algorithm, and the worm containment strategy significantly outperform existing works. Zongqing Lu 0002, Xiao Sun 0010, Yonggang Wen 0001, Guohong Cao, Thomas La Porta |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2015 | Supporting Seamless Virtual Machine Migration via Named Data Networking in Cloud Data CenterabstractVirtual machine migration has been touted as one of the crucial technologies in improving data center efficiency, such as reducing energy cost and maintaining load balance. However, traditional approaches could not avoid the service interruption completely. Moreover, they often result in longer delay and are prone to failures. In this paper, we leverage the emerging named data networking (NDN) to design an efficient and robust protocol to support seamless virtual machine migration in cloud data center. Specifically, virtual machines (VMs) are named with the services they provide. Request routing is based on service names instead of IP addresses that are normally bounded with physical machines. As such, services would not be interrupted when migrating supported VMs to different physical machines. We further analyze the performance of our proposed NDN-based VM migration protocol, and optimize its performance via a load balancing algorithm. Our extensive evaluations verify the effectiveness and the efficiency of our approach and demonstrate that it is interruption-free. Ruitao Xie, Yonggang Wen 0001, Xiaohua Jia, Haiyong Xie 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2015 | Collaborative Task Execution in Mobile Cloud Computing Under a Stochastic Wireless ChannelabstractThis paper investigates collaborative task execution between a mobile device and a cloud clone for mobile applications under a stochastic wireless channel. A mobile application is modeled as a sequence of tasks that can be executed on the mobile device or on the cloud clone. We aim to minimize the energy consumption on the mobile device while meeting a time deadline, by strategically offloading tasks to the cloud. We formulate the collaborative task execution as a constrained shortest path problem. We derive a one-climb policy by characterizing the optimal solution and then propose an enumeration algorithm for the collaborative task execution in polynomial time. Further, we apply the LARAC algorithm to solving the optimization problem approximately, which has lower complexity than the enumeration algorithm. Simulation results show that the approximate solution of the LARAC algorithm is close to the optimal solution of the enumeration algorithm. In addition, we consider a probabilistic time deadline, which is transformed to hard deadline by Markov inequality. Moreover, compared to the local execution and the remote execution, the collaborative task execution can significantly save the energy consumption on the mobile device, prolonging its battery life. Yonggang Wen 0001, Dapeng Oliver Wu |
IEEE Trans. Wirel. Commun. | 2 |
| 2014 | Virt Cache: Managing Virtual Disk Performance Variation in Distributed File Systems for the CloudabstractAs Applications are moved from physical servers to virtual machines sharing storage resources, they experience large variation in I/O latencies. While maintaining average performance in such virtualized environments is important to conform to service level agreements (SLA), cloud users also expect their applications to have minimum variation in tail end latencies like 90th percentile latency for predictable performance. This becomes a challenging problem as the deviation in the application's 90th percentile I/O latency from average latency under storage resource sharing (VM consolidation) can be very high. We show through experiments under VM consolidation that during peak loads this latency variation from average can be as much as 5 times compared to when the application has exclusive access to the storage devices. This variation in performance exists for both Hard drives (HDD) and Solid state drives (SSD). To minimize this large latency variation, we propose a dynamic I/O redirection and caching mechanism called Virt Cache. Virt Cache can pro-actively detect storage device contention at the storage server and temporarily redirect the peaking virtual disk workload to a dynamically instantiated distributed read-write cache. We have implemented our system in Gluster FS, a commonly used distributed file system deployed as a backing store in the cloud. Our system can achieve from 50% to 83% reduction in the 90th percentile latency deviation from average compared to previous work as we move from low load conditions to peak non uniform consolidated VM workloads. With our Virt Cache system, Cloud providers can guarantee predictable performance for the cloud users as if their application has exclusive access to the storage resources. Rajesh Vellore Arumugam, Quanqing Xu, Haixiang Shi, Qingchao Cai, Yonggang Wen 0001 |
CloudCom | 5 |
| 2014 | Joint virtual machine and bandwidth allocation in software defined network (SDN) and cloud computing environmentsabstractCloud computing provides users with great flexibility when provisioning resources, with cloud providers offering a choice of reservation and on-demand purchasing options. Reservation plans offer cheaper prices, but must be chosen in advance, and therefore must be appropriate to users' requirements. If demand is uncertain, the reservation plan may not be sufficient and on-demand resources have to be provisioned. Previous work focused on optimally placing virtual machines with cloud providers to minimize total cost. However, many applications require large amounts of network bandwidth. Therefore, considering only virtual machines offers an incomplete view of the system. Exploiting recent developments in software defined networking (SDN), we propose a unified approach that integrates virtual machine and network bandwidth provisioning. We solve a stochastic integer programming problem to obtain an optimal provisioning of both virtual machines and network bandwidth, when demand is uncertain. Numerical results clearly show that our proposed solution minimizes users' costs and provides superior performance to alternative methods. We believe that this integrated approach is the way forward for cloud computing to support network intensive applications. Jonathan Chase, Rakpong Kaewpuang, Yonggang Wen 0001, Dusit Niyato |
ICC | 3 |
| 2014 | Cost optimal video transcoding in media cloud: Insights from user viewing patternabstractVideo transcoding has been touted as an enabling technology to support growing media consumption over heterogenous devices. However, on-line transcoding could incur tremendous, if not prohibitive, cost in deploying or renting resources. In this research, we leverage an insight into the viewing pattern of video consumers to reduce the operating cost of video transcoding services. Specifically, it has been reported that viewers tend to terminate their session before the whole video is watched. As such, it is not cost-efficient for service providers to store or transcode all segments of the videos. Built upon this insight, we propose a partial transcoding scheme for content management in a media cloud to reduce the operating cost. Particularly, each content is split into multiple segments and stored in different files of varying playback rates. Some of the segments are stored in cache, resulting in storage cost; while some are transcoded in real-time in case of cache miss, resulting in computing cost. We aim to minimize the long-term operational cost by determining the number of segments for each playback rate to be cached or transcoded in real-time. We formulate this partial transcoding scheme as a constrained integer optimization problem. Leveraging Lagrangian relaxation and a subgradient method, we obtain the approximate solution to the integer program. Numerical results indicate that our proposed partial transcoding scheme can save more than 30% of operational cost, compared with a brute-force scheme of caching all the segments. Guanyu Gao, Yonggang Wen 0001, Zhi Wang 0001, Wenwu Zhu 0001, Yap-Peng Tan |
ICME | 3 |
| 2014 | Community based effective social video contents placement in cloud centric CDN networkabstractThe increasing popularity of online social networks (OSNs) has been transforming the dissemination pattern of social video contents. Considering the unique features of social videos, e.g., huge volume, long-tailed, and short length, how to utilize the information propagation pattern to improve the efficiency of content distribution for social videos attracts more and more attention. In this paper, we first conduct a large scale measurement to explore the social video viewing behavior under the community classification. Based on the measurement, we investigate the community driven sharing video distribution problem under the cloud-centric content delivery network (CDN) architecture. In particular, we formulate it as a constrained optimization problem with the objective to minimize the operational cost. The constraint is the averaged transmission delay. Following that, we propose a dynamic algorithm to seek the optimal solution. Our trace-driven experiments further demonstrate our algorithm can make a better tradeoff between monetary cost and QoS, and outperforms the traditional method with less operational cost while satisfying the QoS requirement. Han Hu 0003, Yonggang Wen 0001, Tat-Seng Chua, Zhi Wang 0001, Wenwu Zhu 0001, Di Wu 0001 |
ICME | 2 |
| 2014 | Toward profit-seeking virtual network embedding algorithm via global resource capacityabstractIn this paper, after proposing a novel metric, i.e., global resource capacity (GRC), to quantify the embedding potential of each substrate node, we propose an efficient heuristic virtual network embedding (VNE) algorithm, called as GRC-VNE. The proposed algorithm aims to maximize the revenue and to minimize the cost of the infrastructure provider (InP). Based on GRC, the proposed algorithm applies a greedy load-balance manner to embed each virtual node sequentially, and then adopts the shortest path routing to embed each virtual link. Simulation results demonstrate that our proposed GRC-VNE algorithm achieves lower request blocking probability and higher revenue due to the more appropriate consideration of the resource distribution of the entire network, when compared to the two lastest VNE algorithms that also consider the resources of entire substrate network. Then, we introduce a classical reserved cloud revenue model, which consists of fixed revenue and variable one. Based on this revenue model, we design a novel admission control policy selectively accepting the VNR with high revenue-to-cost ratio to maximize the InP's profit based on an empirical threshold. Through extensive simulations, we observe that the optimal empirical threshold is proportional to the ratio of variable revenue to the fixed one. Long Gong, Yonggang Wen 0001, Zuqing Zhu |
INFOCOM | 2 |
| 2014 | Information diffusion in mobile social networks: The speed perspectiveabstractThe emerging of mobile social networks opens opportunities for viral marketing. However, before fully utilizing mobile social networks as a platform for viral marketing, many challenges have to be addressed. In this paper, we address the problem of identifying a small number of individuals through whom the information can be diffused to the network as soon as possible, referred to as the diffusion minimization problem. Diffusion minimization under the probabilistic diffusion model can be formulated as an asymmetric k-center problem which is NP-hard, and the best known approximation algorithm for the asymmetric k-center problem has approximation ratio of log*n and time complexity O(n5). Clearly, the performance and the time complexity of the approximation algorithm are not satisfiable in large-scale mobile social networks. To deal with this problem, we propose a community based algorithm and a distributed set-cover algorithm. The performance of the proposed algorithms is evaluated by extensive experiments on both synthetic networks and a real trace. The results show that the community based algorithm has the best performance in both synthetic networks and the real trace, and the distributed setcover algorithm outperforms the approximation algorithm in the real trace in terms of diffusion time. Zongqing Lu 0002, Yonggang Wen 0001, Guohong Cao |
INFOCOM | 2 |
| 2014 | PAINT: Partial in-network transcoding for adaptive streaming in information centric networkabstractInformation centric network (ICN) has emerged as a promising architecture to efficiently distribute content over the future Internet. However, ICN proposals may still not be cost efficient enough for adaptive video streaming. The problem is, each ICN node caches duplicated copies of the same content for each bitrate version in its limited storage space. Thus the cache hit ratio drops, and the bandwidth cost of serving the cache missed requests increases. This paper proposes PAINT (Partial In-Network Transcoding) scheme to reduce the operational cost of delivering adaptive video streaming over ICN. Specifically, we consider both the in-network caching and transcoding services at each ICN node, where the storage and transcoding resources can be dynamically scheduled. Then we formulate an optimization problem to balance the trade-off between the transcoding and bandwidth costs. Next we analytically derive the optimal strategy, and quantify cost savings compared with existing schemes. Finally, we verify our solution by intensive numerical evaluations. The results indicate PAINT can achieve significant cost savings (e.g., up to 50% in typical scenarios). Besides, we find the optimal strategy and the cost savings can be affected by the cache capacity, the unit price ratio, the hop distance to origin server, and the Zipf parameter of users' request patterns. Yichao Jin 0002, Yonggang Wen 0001 |
IWQoS | 2 |
| 2014 | Social TV analytics: a novel paradigm to transform TV watching experienceabstractThe blooming online social networks have revolutionized the way information is created, disseminated and consumed, positing significant challenges to the conventional information propagation carriers, especially for the television land-scape. In this paper, we design and develop a multi-screen cloud social TV integrated with social media via a second screen as a novel paradigm in response to this trend. Our system comprises three building blocks, including a cloud based social TV system, a social TV analytics system, and a multi-screen orchestration system. In particular, we leverage the cloud infrastructure to improve the system scalability, and design intelligent social media collection & analysis mechanisms to mine deeper social perception. Furthermore, we demonstrate two key features of our system based on a real user case. Han Hu 0003, Yonggang Wen 0001, Chang Wen Chen, Tat-Seng Chua |
MMSys | 4 |
| 2014 | CREATE: Correlation enhanced traffic matrix estimation in Data Center NetworksabstractUnderstanding the pattern of end-to-end traffic flows in Data Center Networks (DCNs) is essential to many DCN designs and operations (e.g., traffic engineering and load balancing). However, little research work has been done to obtain traffic information efficiently and yet accurately. Researchers often assume the availability of traffic tracing tools (e.g., OpenFlow) when their proposals require traffic information as input, but these tools may generate high monitoring overhead and consume significant switch resources even if they are available in a DCN. Although estimating the traffic matrix between origin-destination pairs using only basic switch SNMP counters is a mature practice in IP networks, traffic flows in DCNs are notoriously more irregular and volatile, while the large number of redundant routes in a DCN further complicates the situation. To this end, we propose to utilize the service placement logs for deducing the correlations among top-of-rack switches, and to leverage the uneven traffic distribution in DCNs for reducing the number of routes potentially used by a flow. These allow us to develop an efficient CoRrelation Enhanced trAffic maTrix Estimation (CREATE) method that achieves high accuracy. We compare CREATE with two existing representative methods through both experiments and simulations; the results strongly confirm the promising performance of CREATE. Zhiming Hu 0001, Jun Luo 0001, Peng Sun 0006, Yonggang Wen 0001 |
Networking | 5 |
| 2014 | MUTAS: Multi-screen TV experience as a service through cloud centric media networkabstractRecently, the TV landscape is rapidly shifting from the traditional “laid-back” experience to a “lean-forward” multiscreen experience. In this paper, we propose MUTAS (MUltiscreen TV experience As a Service), a novel cloud-based service delivery model, in response to this trend. The design objective is to facilitate the development process of new multi-screen features, and improve user experiences by offering an all-in-one solution. The enabling technology is to encapsulate basic functions into a unified cloud platform, and expose divergent multi-screen services through a cloud clone per user. Based on MUTAS, we will use one system to demonstrate four different multi-screen experiences (i.e., synchronized social TV watching, video teleportation, social networking integration, and advertising re-distribution). This demo provides a reference to build cloud-based frameworks for flexible, extensible, and scalable multi-screen TV experience. Yichao Jin 0002, Han Hu 0003, Yonggang Wen 0001 |
SECON | 4 |
| 2014 | Skeleton construction in mobile social networks: Algorithms and applicationsabstractMobile social networks have emerged as a new frontier in the mobile computing research society, and the commonly used social structure (i.e., community) has been exploited to facilitate the design of network protocols and applications, such as data forwarding and worm containment. However, community based approaches may not be accurate when applied for predicting node contacts and may separate two frequently contacted nodes into different communities. In this paper, to address these problems, we propose skeleton, a tree structure specially designed for organizing network nodes, as the underlying structure in mobile social networks. We address the challenges on how to uncover skeleton from network, how to adapt skeleton with dynamic network and how to leverage skeleton for network protocol designs. Skeleton is constructed based on best friendship and skeleton construction is simple and efficient (e.g., less computational complexity than community detection). Algorithms are also designed to adapt skeleton construction to dynamic network. Moreover, a data forwarding algorithm and a worm containment strategy are designed based on skeleton. Trace-driven simulation results show that the skeleton based data forwarding algorithm and worm containment strategy outperform existing schemes based on community. Zongqing Lu 0002, Xiao Sun 0010, Yonggang Wen 0001, Guohong Cao |
SECON | 3 |
| 2014 | Toward a biometric-aware cloud service engine for multi-screen video applicationsabstractThe emergence of portable devices and online social networks (OSNs) has changed the traditional video consumption paradigm by simultaneously providing multi-screen video watching, social networking engagement, etc. One challenge is to design a unified solution to support ever-growing features while guarantee system performance. In this demo, we design and implement a multi-screen technology to provide multi-screen interactions over wide area network (WAN). Furthermore, we incorporate face-detection technology into our system to identify users' bio-features and employ a machine learning based traffic scheduling mechanism to improve the system performance. Han Hu 0003, Yichao Jin 0002, Yonggang Wen 0001, Tat-Seng Chua, Xuelong Li 0001 |
SIGCOMM | 3 |
| 2014 | A QoS-aware routing algorithm based on ant-cluster in wireless multimedia sensor networks
Haiping Huang, Xiao Cao, Ruchuan Wang 0001, Yonggang Wen 0001 |
Sci. China Inf. Sci. | 4 |
| 2014 | Towards optimal noise distribution for privacy preserving in data aggregation
Hao Zhang 0016, Nenghai Yu, Yonggang Wen 0001, Weiming Zhang 0001 |
Comput. Secur. | 3 |
| 2014 | On the Cost-QoE Tradeoff for Cloud-Based Video Streaming Under Amazon EC2's Pricing ModelsabstractThe emergence of cloud computing provides a cost-effective approach to deliver video streams to a large number of end users with the desired user quality of experience (QoE). Under such a paradigm, a video service provider (VSP) can launch its own video streaming services virtually by renting the distribution infrastructure from one or more cloud service providers (CSPs). However, CSPs such as Amazon EC2 normally offer multiple pricing options for virtual machine (VM) instances that they can provide, such as on-demand instances, reserved instances, and spot instances. Such diverse pricing models make it challenging for a VSP to determine how to optimally procure the required number of VM instances in different types to satisfy dynamic user demands. Given the limited budget, a VSP needs to carefully balance the procurement cost and the achieved QoE for end users. In this paper, we investigate the tradeoff between the cost incurred by VM instance procurement and the achieved QoE of end users under Amazon EC2's pricing models, and formulate the VM instance provisioning and procurement problem into a constrained stochastic optimization problem. By applying the Lyapunov optimization framework, we design an online procurement algorithm, which approaches the optimal solution with explicitly provable upper bounds. We also conduct extensive trace-driven simulations and our results show that our proposed algorithm (OPT-ORS) achieves a good balance between the procurement cost and the user QoE for cloud-based VSPs. In the achieved near-optimal situation, our algorithm guarantees that reserved VM instances are fully utilized to satisfy the baseline user demand, on-demand VM instances are only rented to handle flash crowds, while more spot VM instances are rented than on-demand VM instances to serve user demand over the baseline due to their low prices. Jian He 0002, Yonggang Wen 0001, Jianwei Huang 0001, Di Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2014 | Distributed Wireless Video Scheduling With Delayed Control InformationabstractTraditional distributed wireless video scheduling is based on perfect control channels in which instantaneous control information from the neighbors is available. However, it is difficult, sometimes even impossible, to obtain this information in practice, especially for dynamic wireless networks. Thus, neither the distortion-minimum scheduling approaches aiming to meet the longterm video quality demands nor the solutions that focus on minimum delay can be applied directly. This motivates us to investigate the distributed wireless video scheduling with delayed control information (DCI). First, to exploit in a tractable framework, we translate this scheduling problem into a stochastic optimization rather than a convex optimization problem. Next, we consider two classes of DCI distributions: 1) the class with finite mean and variance and 2) a general class that does not employ any parametric representation. In each case, we study the relationship between the DCI and scheduling performance, and provide a general performance property bound for any distributed scheduling. Subsequently, a class of distributed scheduling scheme is proposed to achieve the performance bound by making use of the correlation among the time-scale control information. Finally, we provide simulation results to demonstrate the correctness of the theoretical analysis and the efficiency of the proposed scheme. Liang Zhou 0002, Zhen Yang 0001, Yonggang Wen 0001, Joel J. P. C. Rodrigues |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2014 | μ DC2: unified data collection for data centers
Wenfeng Xia 0002, Yonggang Wen 0001, Haiyong Xie 0001, Bin Liu 0022 |
J. Supercomput. | 2 |
| 2014 | CBM: Online Strategies on Cost-Aware Buffer Management for Mobile Video StreamingabstractMobile video traffic, owing to the rapid adoption of smartphones and tablets, has been growing exponentially in recent years and started to dominate the mobile Internet. In reality, mobile video applications commonly adopt buffering techniques to handle bandwidth fluctuation and minimize the impact of stochastic wireless channels on user experiences. However, recent measurement work reveals that mobile users tend to abort more frequently than PC users during viewing videos. Such a high abortion rate results in a significant wastage of buffered video data, which is directly translated into monetary and energy cost for mobile users. In this paper, we propose an intelligent buffer management strategy called CBM (Cost-aware Buffer Management), for mobile video streaming applications. Our purpose is to minimize cost induced by un-consumed video data while respecting certain user experience requirements. To this objective, we formulate the problem into a constrained stochastic optimization problem, and apply the Lyapunov optimization theory to derive the corresponding online strategy for cost minimization. Different from conventional heuristic-based strategies, our proposed CBM strategy can provide provably performance guarantee with explicit bounds. We also conduct extensive simulations to validate the effectiveness of our proposed strategy and our experimental results show that CBM achieves significant gains over existing schemes. Jian He 0002, Zheng Xue, Di Wu 0001, Dapeng Oliver Wu, Yonggang Wen 0001 |
IEEE Trans. Multim. | 5 |
| 2014 | Reducing Operational Costs in Cloud Social TV: An Opportunity for Cloud CloningabstractThe emergence of social TV has transformed TV experiences, providing a unified media experience across different devices. In response to this trend, we have implemented a multi-screen social TV system, offering video teleportation as an attractive feature. The enabling technology is instantiating a cloud clone to support all media outlets of each user. As the user shifts his attention from one device to the other, the cloud clone might migrate to a better location to reduce its operational cost. This paper investigates this cloud clone migration problem, aiming to minimize the monetary cost on operating video teleportation. Specifically, we formulate it into a Markov Decision Problem, to balance the trade-off between the migration cost and the content transmission cost. Under this framework, four algorithms are proposed to solve this optimization problem. We first characterize an upper and a lower bound for the optimal cost, by considering a random fixed placement and an offline algorithm. We then present a semi-online and a more practical Q-learning approach to make online decisions. Their performances are evaluated based on both simulated and real user traces. The results show that the Q-learning method achieves up to 25% cost compared to random fixed placement in typical scenarios. The savings are affected by the delivery path length, the migration size, and the user behavior pattern. Moreover, our investigations reveal the optimal cloud clone location is either at the nearest or the furthest node to the user along the content delivery path for a single user scenario. Yichao Jin 0002, Yonggang Wen 0001, Han Hu 0003, Marie-José Montpetit |
IEEE Trans. Multim. | 2 |
| 2014 | Dynamic Request Redirection and Elastic Service Scaling in Cloud-Centric Media NetworksabstractWe consider the problem of optimally redirecting user requests in a cloud-centric media network (CCMN) to multiple destination Virtual Machines (VMs), which elastically scale their service capacities in order to minimize a cost function that includes service response times, computing costs, and routing costs. We also allow the request arrival process to switch between normal and flash crowd modes to model user requests to a CCMN. We quantify the trade-offs in flash crowd detection delay and false alarm frequency, request allocation rates, and service capacities at the VMs. We show that under each request arrival mode (normal or flash crowd), the optimal redirection policy can be found in terms of a price for each VM, which is a function of the VM's service cost, with requests redirected to VMs in order of nondecreasing prices, and no redirection to VMs with prices above a threshold price. Applying our proposed strategy to a YouTube request trace data set shows that our strategy outperforms various benchmark strategies. We also present simulation results when various arrival traffic characteristics are varied, which again suggest that our proposed strategy performs well under these conditions. Jianhua Tang, Wee-Peng Tay, Yonggang Wen 0001 |
IEEE Trans. Multim. | 3 |
| 2014 | Cloud Mobile Media: Reflections and OutlookabstractThis paper surveys the emerging paradigm of cloud mobile media. We start with two alternative perspectives for cloud mobile media networks: an end-to-end view and a layered view. Summaries of existing research in this area are organized according to the layered service framework: i) cloud resource management and control in infrastructure-as-a-service (IaaS), ii) cloud-based media services in platform-as-a-service (PaaS), and iii) novel cloud-based systems and applications in software-as-a-service (SaaS). We further substantiate our proposed design principles for cloud-based mobile media using a concrete case study: a cloud-centric media platform (CCMP) developed at Nanyang Technological University. Finally, this paper concludes with an outlook of open research problems for realizing the vision of cloud-based mobile media. Yonggang Wen 0001, Joel J. P. C. Rodrigues, Chang Wen Chen |
IEEE Trans. Multim. | 1 |
| 2013 | Revenue-driven virtual network embedding based on global resource informationabstractVirtual network embedding (VNE), working as a key step for network virtualization, has recently gained intensive attentions from the research community. In this paper, we propose a novel VNE algorithm that aims to maximize the infrastructure provider's revenue from serving virtual network (VN) requests, with the help of the global resource information. The proposed algorithm, named as revenue-driven VNE (RD-VNE), adopts a node-ranking approach that takes the global resource information into account in a recursive manner to assist the greedy node mapping, and leverages the shortest-path routing for link mapping. Our simulation results suggest that the proposed VNE algorithm outperforms two existing VNE algorithms that also take global resource information into consideration, in terms of request blocking probability, and brings higher time-average revenue to the infrastructure provider (InP). Long Gong, Yonggang Wen 0001, Zuqing Zhu |
GLOBECOM | 2 |
| 2013 | Power-efficient collaborative distribution of social videos over wireless community cloudabstractThe prevalence of social networking services dramatically changes the landscape of video distribution, in which social video contents spread much faster than traditional video-sharing portals. The pervasive wireless connectivity further enables users to view and generate videos from anywhere at any time. In this paper, we focus on the problem of collaborative distribution of social videos in a wireless community cloud. We aim to minimize the total power consumption of all participants in the community. To this purpose, we first analyze the distribution problem using a Markovian model and study how the soft deadline threshold impacts the total power consumption. We derive the closed-form expression to reveal the relationship between the optimal power allocation strategy and the soft deadline threshold. Our numerical results show that the minimum power consumption increases convexly as the soft deadline threshold approaches one. Moreover, we also observe that when more paths are used for parallel transmission, the total power consumption increases in spite that the power consumption of each individual path is reduced. Jian He 0002, Yonggang Wen 0001, Di Wu 0001 |
GLOBECOM | 2 |
| 2013 | Minimizing monetary cost via cloud clone migration in multi-screen cloud social TV systemabstractThe emergence of multi-screen cloud social TV has the potential to transform TV experience, providing a unified media experience across a diverse set of devices at an affordable cost. One key technology to support unified media experience across multiple screens is to instantiate a virtual machine (VM) as a cloud clone of the user, to manage all his/her media outlets (e.g., TV and smartphone), as implemented in our Cloud-Centric Media Network (CCMN). In this case, as the user shifts his attention from one device to another, the cloud clone can migrate to another location for better quality of experience. In this paper, we investigates the problem of cloud-clone migration for the multi-screen social TV application, minimizing its monetary cost. This problem can be cast into the Markov Decision Process (MDP) framework, to balance a trade-off between the migration cost and the transmission cost. Under this framework, we first derive an upper and lower bound for the optimal monetary cost, by considering a fixed placement policy and an offline policy. We then follow up with an online policy using a dynamic programming approach. Our numerical results indicate, up to 10% monetary cost can be saved, by optimally migrating the cloud clone. Moreover, the cost reduction depends on the length of content-delivery path, the data size associated with VM migration, and the user behavior pattern. These insights would offer operational guidelines to deliver cost effective multi-screen social TV services over CCMN, potentially easing its adoption. Yichao Jin 0002, Yonggang Wen 0001, Han Hu 0003 |
GLOBECOM | 2 |
| 2013 | Dynamic transparent virtual network embedding over elastic optical infrastructuresabstractWe propose a novel dynamic transparent virtual network embedding (VNE) algorithm, which considers node mapping and link mapping jointly, for network virtualization over optical orthogonal frequency-division multiplexing (O-OFDM) based elastic optical infrastructures. For each virtual optical network (VON) request, the algorithm first transfers the substrate optical network into a layered-auxiliary-graph according to the spectrum usage of each fiber link, then applies a node mapping approach that considers the local information of all substrate nodes, and accomplishes the link mapping, in a single layer of the auxiliary graph. The simulation results verify that the proposed algorithm considers the uniqueness of O-OFDM networks and outperforms two reference algorithms that directly apply the VNE schemes developed for Layer 2/3 or WDM network virtualization, by providing lower VON blocking probability. The simulations with a realistic topology also demonstrate that the average lengths of embedded substrate paths are well-controlled within the typical transmission reaches of O-OFDM signals. To the best of our knowledge, this is the first proposal that includes both link mapping and node mapping to address dynamic transparent VNE over elastic optical infrastructures. Long Gong, Wenwen Zhao, Yonggang Wen 0001, Zuqing Zhu |
ICC | 3 |
| 2013 | Coordinating In-Network Caching in Content-Centric Networks: Model and AnalysisabstractIn-network content storage has become an inherent capability of routers in the content-centric networking architecture. This raises new challenges in utilizing and provisioning the in-network caching capability, namely, how to optimally provision individual routers' storage to cache contents, so as to balance the trade-offs between the network performance and the provisioning cost. To address this problem, we first propose a holistic model to characterize the network performance of routing contents to clients and the network cost incurred by globally coordinating the in-network storage capability. We then derive the optimal strategy for provisioning the storage capability that optimizes the overall network performance and cost, and analyze the performance gains via numerical evaluations on real network topologies. Our results reveal interesting phenomena; for instance, different ranges of the Zipf exponent can lead to opposite optimal strategies, and the trade-offs between the network performance and the provisioning cost have great impacts on the stability of the optimal strategy. We also demonstrate that the optimal strategy can achieve significant gain on both the load reduction at origin servers and the improvement on the routing performance. Haiyong Xie 0001, Yonggang Wen 0001, Zhi-Li Zhang |
ICDCS | 3 |
| 2013 | Toward monetary cost effective content placement in cloud centric media networkabstractIn recent years, technical challenges are emerging on how to efficiently distribute the rapid growing user-generated contents (UGCs) with long-tailed nature. To address this issue, we have previously proposed cloud centric media network (CCMN) for cost-efficient UGCs delivery. In this paper, we further study the content placement problem in CCMN. Our objective is to minimize the monetary cost incurred by using cloud resources to orchestrate an elastic and global content delivery network (CDN) service. In particular, this objective is achieved via a two-step method. First, for a single content, we map it into a k-center problem, and find a logarithmic relationship between the mean hop distance from users to contents, and the reciprocal of replica number. Second, for multiple contents, we formulate a convex optimization with storage and bandwidth capacity constraints, which can be solved by our proposed algorithm. Finally, we verify the algorithm based on real-world traces collected from a popular video website in China. Our numerical results suggest that, the optimal number of replica for each content follows a power law in respect to its popularity, under feasible storage and bandwidth constraints, in a set of deployed backbone networks. Yichao Jin 0002, Yonggang Wen 0001, Kyle Guan, Daniel C. Kilper, Haiyong Xie 0001 |
ICME | 2 |
| 2013 | vRGW: Towards network function virtualization enabled by software defined networkingabstractIt has been a significant challenge for network carriers to deploy and provision a large number of Customer-Premises Equipment (CPE) devices located at subscribers' premises and connected to a carrier's network infrastructure. In this paper, we make a first systematic attempt to fundamentally re-shape the access networks into a software defined networking architecture by virtualizing the network functionality of residential gateways (vRGW). Our approach can be generalized to other CPE such as set-top boxes. Our analysis suggests that vRGW can achieve significant economic benefits ranging from up to 90% reduction on the call center cost and up to 46% reduction on the product return cost. Haiyong Xie 0001, Diego R. López, Tina Tsou, Yonggang Wen 0001 |
ICNP | 6 |
| 2013 | Energy-efficient scheduling policy for collaborative execution in mobile cloud computingabstractIn this paper, we investigate the scheduling policy for collaborative execution in mobile cloud computing. A mobile application is represented by a sequence of fine-grained tasks formulating a linear topology, and each of them is executed either on the mobile device or offloaded onto the cloud side for execution. The design objective is to minimize the energy consumed by the mobile device, while meeting a time deadline. We formulate this minimum-energy task scheduling problem as a constrained shortest path problem on a directed acyclic graph, and adapt the canonical “LARAC” algorithm to solving this problem approximately. Numerical simulation suggests that a one-climb offloading policy is energy efficient for the Markovian stochastic channel, in which at most one migration from mobile device to the cloud is taken place for the collaborative task execution. Moreover, compared to standalone mobile execution and cloud execution, the optimal collaborative execution strategy can significantly save the energy consumed on the mobile device. Yonggang Wen 0001, Dapeng Oliver Wu |
INFOCOM | 2 |
| 2013 | Mobile media communication, processing, and analysis: A review of recent advancesabstractIn this paper, we review recent advances in mobile media communication, processing, and analysis. To identify the opportunities and challenges in fast growing mobile media computing, we discuss several emerging topics including mobile visual search, retargeting, mobile video streaming, and cloud based mobile media computing. According to the infrastructure of mobile devices vs. servers, we come up with essential concerns in mobile media computing such as wireless bandwidth consumption, mobile energy saving, media adaptation for better quality of services, the computational load shift from mobiles to servers, etc. With booming mobile Apps on diverse media consumption, it is envisioned that mobile media research and development is bringing about significant achievements in traditional topics of communication, processing, and analytics. Wen Gao 0001, Ling-Yu Duan, Jun Sun 0007, Junsong Yuan 0001, Yonggang Wen 0001, Yap-Peng Tan, Jianfei Cai 0001, Alex Chichung Kot |
ISCAS | 5 |
| 2013 | Inter-screen interaction for session recognition and transfer based on cloud centric media networkabstractRecently, there is a growing trend that people tend to consume media over multi-screens simultaneously. This paper proposes an efficient and convenient inter-screen interaction approach based on our cloud centric media network, where all the ongoing sessions on the client side are always synchronized with the media cloud. This approach realizes the session recognition and transfer over different devices in a three-step manner. First, the users are required to use the mobile device camera to scan the main screen. Second, after the screen edge detection and image correction, users are allowed to choose one or more ongoing session on the corrected image via touch screen to request the transfer. Finally, the selected sessions are identified by the cloud, and those sessions are delivered to the mobile devices to complete the session transfer. The algorithms and strategies involved are discussed in detail. We also implement a testbed on top of a private cloud at NTU. The results prove our proposed method is robust and easy to use. Yichao Jin 0002, Yonggang Wen 0001, Jianfei Cai 0001 |
ISCAS | 3 |
| 2013 | Multi-screen cloud social TV: transforming TV experience into 21st centuryabstractNowadays, TV experience has been transformed from the traditional "laid-back" video watching experience to a "lean-forward" social and multi-screen experience. In this demo, we design and develop a multi-screen cloud social TV system in response to this trend. Our system is built upon two enabling technologies, including a cloud based back-end infrastructure and a multi-screen front-end application. We demonstrate two key features of our system based on real user scenarios, including a living-room video watching experience with remote viewers, and the video teleportation as an enhanced multi-screen experience. Yichao Jin 0002, Yonggang Wen 0001, Haiyong Xie 0001 |
ACM Multimedia | 3 |
| 2013 | Multi-screen social TV over cloud-centric media platformabstractMulti-screen social TV is an innovative application for transforming the traditional "laid-back" video watching behavior with the emerging "lean-forward" social network experience. In this video, we demonstrate a set of use cases in which the multi-screen social TV, developed over our patent-pending cloud-centric media platform, is leveraged to change the way we work, study and play in future. Yonggang Wen 0001 |
MobiSys | 3 |
| 2013 | Community detection in weighted networks: Algorithms and applicationsabstractCommunity detection is an important issue due to its wide use in designing network protocols such as data forwarding in Delay Tolerant Networks (DTN) and worm containment in Online Social Networks (OSN). However, most of the existing community detection algorithms focus on binary networks. Since most networks are weighted such as social networks, DTN or OSN, in this paper, we address the problems of community detection in weighted networks and exploit community for data forwarding in DTN and worm containment in OSN. We propose a novel community detection algorithm, and then introduce two metrics called intra-centrality and inter-centrality, to characterize nodes in communities. Based on these metrics, we propose an efficient data forwarding algorithm for DTN and an efficient worm containment strategy for OSN. Extensive trace-driven simulation results show that the data forwarding algorithm and the worm containment strategy significantly outperform existing works. Zongqing Lu 0002, Yonggang Wen 0001, Guohong Cao |
PerCom | 2 |
| 2013 | Cloud3DView: an interactive tool for cloud data center operationsabstractThe emergence of cloud computing has promoted growing demand and rapid deployment of data centers. However, data center operations require a set of sophisticated skills (e.g., command-line-interface), resulting in a high operational cost. In this demo, to reduce the data center operational cost, we design and build a novel cloud data center management system, based on the concept of 3D gamification. In particular, we apply data visualization techniques to overlay operational status upon a data center 3D model, allowing the operators to monitor the real-time situation and control the data center from a friendly user interface. This demo highlights: (1)a data center 3D view from a First Person Shooter (FPS) camera, (2)a run-time presentation of visualized infrastructures information. Moreover, to improve the user experience, we employ cutting-edge HCI technologies from multi-touch, for remote access to Cloud3DView. Jianxiong Yin, Peng Sun 0006, Yonggang Wen 0001, Hai-gang Gong, Ming Liu 0002, Xuelong Li 0001, Haipeng You, Jinqi Gao, Cynthia Lin |
SIGCOMM | 3 |
| 2013 | An Empirical Investigation of the Impact of Server Virtualization on Energy Efficiency for Green Data CenterabstractThe swift adoption of cloud services is accelerating the deployment of data centers. These data centers are consuming a large amount of energy, which is expected to grow dramatically under the existing technological trends. Therefore, research efforts are in great need to architect green data centers with better energy efficiency. The most prominent approach is the consolidation enabled by virtualization. However, little effort has been paid to the potential overhead in energy usage and the throughput reduction for virtualized servers. Clear understanding of energy usage on virtualized servers lays out a solid foundation for green data-center architecture. This paper investigates how virtualization affects the energy usage in servers under different task loads, aiming to understand a fundamental trade-off between the energy saving from consolidation and the detrimental effects from virtualization. We adopt an empirical approach to measure the server energy usage with different configurations, including a benchmark case and two alternative hypervisors. Based on the collected data, we report a few findings on the impact of virtualization on server energy usage and their implications to green data-center architecture. We envision that these technical insights would bring tremendous value propositions to green data-center architecture and operations. Yichao Jin 0002, Yonggang Wen 0001, Zuqing Zhu |
Comput. J. | 2 |
| 2013 | Toward Efficient Distributed Algorithms for In-Network Binary Operator Tree Placement in Wireless Sensor NetworksabstractIn-network processing is touted as a key technology to eliminate data redundancy and minimize data transmission, which are crucial to saving energy in wireless sensor networks (WSNs). Specifically, operators participating in in-network processing are mapped to nodes in a sensor network. They receive data from downstream operators, process them and route the output to either the upstream operator or the sink node. The objective of operator tree placement is to minimize the total energy consumed in performing in-network processing. Two types of placement algorithms, centralized and distributed, have been proposed. A problem with the centralized algorithm is that it does not scale to large WSN's, because each sensor node is required to know the complete topology of the network. A problem with the distributed algorithm is their high message complexity. In this paper, we propose a heuristic algorithm to place a treestructured operator graph, and present a distributed implementation to optimize in-network processing cost and reduce the communication overhead. We prove a tight upper bound on the minimum in-network processing cost, and show that the heuristic algorithm has better performance than a canonical greedy algorithm. Simulation-based evaluations demonstrate the superior performance of our heuristic algorithm. We also give an improved distributed implementation of our algorithm that has a message overhead of O(M) per node, which is much less than the O(√NM log2M) and O(√NM) complexities for two previously proposed algorithms, Sync and MCFA, respectively. Here, N is the number of network nodes and M is the size of the operator tree. Zongqing Lu 0002, Yonggang Wen 0001, Rui Fan 0004, Su-Lim Tan, Jit Biswas |
IEEE J. Sel. Areas Commun. | 2 |
| 2013 | Toward Optimal Deployment of Cloud-Assisted Video Distribution ServicesabstractFor Internet video services, the high fluctuation of user demands in geographically distributed regions results in low resource utilizations of traditional content distribution network systems. Due to the capability of rapid and elastic resource provisioning, cloud computing emerges as a new paradigm to reshape the model of video distribution over the Internet, in which resources (such as bandwidth, storage) can be rented on demand from cloud data centers to meet volatile user demands. However, it is challenging for a video service provider (VSP) to optimally deploy its distribution infrastructure over multiple geo-distributed cloud data centers. A VSP needs to minimize the operational cost induced by the rentals of cloud resources without sacrificing user experience in all regions. The geographical diversity of cloud resource prices further makes the problem complicated. In this paper, we investigate the optimal deployment problem of cloud-assisted video distribution services and explore the best tradeoff between the operational cost and the user experience. We aim to pave the way for building the next-generation video cloud. Toward this objective, we first formulate the deployment problem into a min-cost network flow problem, which takes both the operational cost and the user experience into account. Then, we apply the Nash bargaining solution to solve the joint optimization problem efficiently and derive the optimal bandwidth provisioning strategy and optimal video placement strategy. In addition, we extend the algorithms to the online case and consider the scenario when peers participate into video distribution. Finally, we conduct extensive simulations to evaluate our algorithms in the realistic settings. Our results show that our proposed algorithms can achieve a good balance among multiple objectives and effectively optimize both operational cost and user experience. Jian He 0002, Di Wu 0001, Yupeng Zeng, Xiaojun Hei, Yonggang Wen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2013 | Guest Editorial - Special section on cloud-based mobile media: Infrastructure, services, and applicationsabstractIt is the aim of this special section to report on the latest research that explores various aspects of cloud-based mobile media. Through an open call for papers, we received 40 submissions. Fourteen (14) papers were accepted for final publication after two rounds of highly competitive reviews. The final papers were selected on the basis of originality and significance of the technical work, as well as relevance to the theme topic. The papers in this special section span a wide range of novel algorithms, techniques, applications, and system-level solutions. They naturally fall into two categories: addressing existing challenges or exploring emerging opportunities. Chang Wen Chen, Yonggang Wen 0001, Joel J. P. C. Rodrigues |
IEEE Trans. Multim. | 3 |
| 2013 | QoE-Driven Cache Management for HTTP Adaptive Bit Rate Streaming Over Wireless NetworksabstractIn this paper, we investigate the problem of optimal content cache management for HTTP adaptive bit rate (ABR) streaming over wireless networks. Specifically, in the media cloud, each content is transcoded into a set of media files with diverse playback rates, and appropriate files will be dynamically chosen in response to channel conditions and screen forms. Our design objective is to maximize the quality of experience (QoE) of an individual content for the end users, under a limited storage budget. Deriving a logarithmic QoE model from our experimental results, we formulate the individual content cache management for HTTP ABR streaming over wireless network as a constrained convex optimization problem. We adopt a two-step process to solve the snapshot problem. First, using the Lagrange multiplier method, we obtain the numerical solution of the set of playback rates for a fixed number of cache copies and characterize the optimal solution analytically. Our investigation reveals a fundamental phase change in the optimal solution as the number of cached files increases. Second, we develop three alternative search algorithms to find the optimal number of cached files, and compare their scalability under average and worst complexity metrics. Our numerical results suggest that, under optimal cache schemes, the maximum QoE measurement, i.e., mean-opinion-score (MOS), is a concave function of the allowable storage size. Our cache management can provide high expected QoE with low complexity, shedding light on the design of HTTP ABR streaming services over wireless networks. Yonggang Wen 0001, Ashish Khisti |
IEEE Trans. Multim. | 2 |
| 2013 | Multiview Vector-Valued Manifold Regularization for Multilabel Image ClassificationabstractIn computer vision, image datasets used for classification are naturally associated with multiple labels and comprised of multiple views, because each image may contain several objects (e.g., pedestrian, bicycle, and tree) and is properly characterized by multiple visual features (e.g., color, texture, and shape). Currently, available tools ignore either the label relationship or the view complementarily. Motivated by the success of the vector-valued function that constructs matrix-valued kernels to explore the multilabel structure in the output space, we introduce multiview vector-valued manifold regularization (MV(3)MR) to integrate multiple features. MV(3)MR exploits the complementary property of different features and discovers the intrinsic local geometry of the compact support shared by different features under the theme of manifold regularization. We conduct extensive experiments on two challenging, but popular, datasets, PASCAL VOC' 07 and MIR Flickr, and validate the effectiveness of the proposed MV(3)MR for image classification. Yong Luo 0002, Dacheng Tao, Chang Xu 0002, Chao Xu 0006, Hong Liu 0008, Yonggang Wen 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2013 | Energy-Optimal Mobile Cloud Computing under Stochastic Wireless ChannelabstractThis paper provides a theoretical framework of energy-optimal mobile cloud computing under stochastic wireless channel. Our objective is to conserve energy for the mobile device, by optimally executing mobile applications in the mobile device (i.e., mobile execution) or offloading to the cloud (i.e., cloud execution). One can, in the former case sequentially reconfigure the CPU frequency; or in the latter case dynamically vary the data transmission rate to the cloud, in response to the stochastic channel condition. We formulate both scheduling problems as constrained optimization problems, and obtain closed-form solutions for optimal scheduling policies. Furthermore, for the energy-optimal execution strategy of applications with small output data (e.g., CloudAV), we derive a threshold policy, which states that the data consumption rate, defined as the ratio between the data size (L) and the delay constraint (T), is compared to a threshold which depends on both the energy consumption model and the wireless channel model. Finally, numerical results suggest that a significant amount of energy can be saved for the mobile device by optimally offloading mobile applications to the cloud in some cases. Our theoretical framework and numerical investigations will shed lights on system implementation of mobile cloud computing under stochastic wireless channel. Yonggang Wen 0001, Kyle Guan, Daniel C. Kilper, Haiyun Luo, Dapeng Oliver Wu |
IEEE Trans. Wirel. Commun. | 2 |
| 2013 | Resource Allocation with Incomplete Information for QoE-Driven Multimedia CommunicationsabstractMost existing Quality of Experience (QoE)-driven multimedia resource allocation methods assume that the QoE model of each user is known to the controller before the start of the multimedia playout. However, this assumption may be invalid in many practical scenarios. In this paper, we address the resource allocation problem with incomplete information where the realized mean opinion score (MOS) can only be observed over time, but the underlying QoE model and playout time are unknown. We consider two variants of this problem: 1) the form of the QoE model is known but the parameters are unknown; 2) both the form and the parameters of the QoE model are unknown. For both cases, we develop dynamic resource allocation schemes based on online test-optimization strategy. Simply speaking, one first spends appropriate time on testing the QoE model, then optimizes the sum of the MOS in the remaining playout time. The highlight of this paper lies in resolving the inherent tension between the test and optimization by jointly considering the uncertainties of QoE model and playout time. Furthermore, we derive tight bounds on the MOS loss incurred by the proposed schemes in comparison with the optimal scheme that knows the QoE model a priori and prove that the performance gap, as the playout time tends to infinity, asymptotically shrinks to zero. Liang Zhou 0002, Zhen Yang 0001, Yonggang Wen 0001, Haohong Wang, Mohsen Guizani |
IEEE Trans. Wirel. Commun. | 3 |
| 2012 | On P2P mechanisms for VM image distribution in cloud data centers: Modeling, analysis and improvementabstractTo provide elastic cloud services with QoS guarantee, it is essential for cloud data centers to provision virtual machines rapidly according to user requests. Due to bandwidth bottleneck of centralized model, P2P model is recently adopted in data centers to relieve server workload by enabling sharing among VM instances. In this paper, we develop a simple theoretic model to analyze two typical P2P models for VM image distribution, namely, isolated-image P2P distribution model and cross-image P2P distribution model. We compare their efficiency under different parameter settings and derive their corresponding optimal server bandwidth allocation strategies. In addition, we also propose a practical optimal server bandwidth provisioning algorithm for chunk-level cross-image P2P distribution mechanism to further improve its efficiency. Extensive simulations are conducted to validate the effectiveness of our proposed algorithm. Di Wu 0001, Yupeng Zeng, Jian He 0002, Yonggang Wen 0001 |
CloudCom | 5 |
| 2012 | Towards end-to-end secure content storage and delivery with public cloudabstractRecent years have witnessed the trend of leveraging cloud-based services for large scale content storage, processing, and distribution. Security and privacy are among top concerns for the public cloud environments. Towards end-to-end content security, we propose and implement CloudSeal, a scheme for securely sharing and distributing content via the public cloud. CloudSeal ensures the confidentiality of content in the public cloud environments with flexible access control policies for subscribers and efficient content distribution via content delivery network. Huijun Xiong, Xinwen Zhang, Danfeng Yao, Xiaoxin Wu 0001, Yonggang Wen 0001 |
CODASPY | 5 |
| 2012 | Energy optimizations for data center network: Formulation and its solutionabstractData center consumes increasing amount of power nowadays, together with expanding number of data centers and upgrading data center scale, its power consumption becomes a knotty issue. While main efforts of this research focus on server and storage power reduction, network devices as part of the key components of data centers, also contribute to the overall power consumption as data centers expand. In this paper, we address this problem with two perspectives. First, in a macro level, we attempt to reduce redundant energy usage incurred by network redundancies for load balancing. Second, in the micro level, we design algorithm to limit port rate in order to reduce unnecessary power consumption. Given the guidelines we obtained from problem formulation, we propose a solution based on greedy approach with integration of network traffic and minimization of switch link rate. We also present results from a simulation-based performance evaluation which shows that expected power saving is achieved with tolerable delay. Shuo Fang, Chuan Heng Foh, Yonggang Wen 0001, Khin Mi Mi Aung |
GLOBECOM | 4 |
| 2012 | Content routing and lookup schemes using global bloom filter for content-delivery-as-a-serviceabstractLeveraging cloud computing technology, we have proposed content-delivery-as-a-service (CoDaaS) to distribute user generated content (UGC) in an efficient and economical fashion. However, due to the exponential increases of Internet traffic, traditional hashing-based content routing and lookup scheme suffers from high delay. This paper introduces a global compressed counting bloom filter (CCBF) into CoDaaS to address this issue. The global CCBF adds our system with the capability to early check the existence of any specific content among all the peering surrogates, before any local checking on each cache node. Using this global CCBF, we propose two content routing and lookup mechanisms (parallel and cut-through schemes) to reduce the delay for better user experience. We verify the comparative performance of those approaches via both mathematical modeling and experimental simulation. The results show that for light traffic load, the mean response time can be saved by up to 65.2%. Besides, the impacts and overheads of different synchronization schemes for the CCBF are quantified to provide valuable insights for further optimizations. Yichao Jin 0002, Yonggang Wen 0001 |
GLOBECOM | 2 |
| 2012 | Optimizing content retrieval delay for LT-based distributed cloud storage systemsabstractAmong different setups of cloud storage systems, fountain-codes based distributed cloud storage system provides reliable online storage solution through placing coded content fragments into multiple storage nodes. Luby Transform (LT) code is one of the popular fountain codes for storage systems due to its efficient recovery. However, to ensure high success decoding of fountain codes based storage, retrieval of additional fragments is required, and this requirement introduces additional delay, which is critical for content retrieval or downloading applications. In this paper, we show that multiple-stage retrieval of fragments is effective to reduce the content-retrieval delay. We first develop a delay model for various multiple-stage retrieval schemes applicable to our considered system. With the developed model, we study optimal retrieval schemes given the success decodability requirement. Our numerical results demonstrate that the content-retrieval delay can be significantly reduced by optimally scheduling packet requests in a multi-stage fashion. Haifeng Lu, Chuan Heng Foh, Yonggang Wen 0001, Jianfei Cai 0001 |
GLOBECOM | 3 |
| 2012 | QoE-driven cache management for HTTP adaptive bit rate (ABR) streaming over wireless networksabstractIn this paper, we investigate the problem of how to cache a set of media files with optimal streaming rates, under HTTP adaptive bit rate streaming over wireless networks. The design objective is to achieve the optimal expected QoE under a limited storage budget, which is measured by the logarithmic relation between the required bit rate and the actual streaming bit rate. We formulate the content cache management of streaming files as a constrained optimization problem. Lagrange multiplier method is employed, and we obtain the numerical solution of the optimal streaming files. Particularly, we characterize the properties of the solution, and find there is a fundamental phase change in the optimal solution as the number of cached files grows. Moreover, the simulation results indicate that with the increase of cache size, more copies of different bit rate should be cached for a better QoE. Our comprehensive investigation reveals insightful guidelines to provide HTTP ABR streaming services over wireless networks. Yonggang Wen 0001, Ashish Khisti |
GLOBECOM | 2 |
| 2012 | Energy minimization via dynamic voltage scaling for real-time video encoding on mobile devicesabstractThis paper investigates the problem of minimizing energy consumption for real-time video encoding on mobile devices, by dynamically configuring the clock frequency in the CPU via the dynamic voltage scaling (DVS) technology. The problem can be formulated as a constrained optimization problem, whose objective is to minimize the total energy consumption of encoding video contents while respecting a real-time delay constraint. Under a probabilistic workload model, we obtain closed-form solutions for both the optimal clock frequency configuration and the resulted minimum energy. We also compare the optimal solution with a brute force flat frequency configuration. Numerical results indicate that our derived optimal solution outperforms the brute-force approach significantly. Moreover, we apply the optimal solution for real-time H.264/AVC video encoding application. Our numerical results suggest that an energy saving of 10%-20% can be achieved, compared to the flat clock frequency scheduling. Ming Yang 0018, Yonggang Wen 0001, Jianfei Cai 0001, Chuan Heng Foh |
ICC | 2 |
| 2012 | Energy-optimal mobile application execution: Taming resource-poor mobile devices with cloud clonesabstractIn this paper, we propose to leverage cloud computing to tame resource-poor mobile devices. Specifically, mobile applications can be executed in the mobile device (known as mobile execution) or offloaded to the cloud clone for execution (known as cloud execution), with an objective to conserve energy for mobile device. The energy-optimal execution policy is obtained by solving two constrained optimization problems, i.e., how to optimally configure the clock frequency to complete CPU cycles for mobile execution, and how to optimally schedule the data transmission for cloud execution in order to achieve the minimal energy within time delay. Closed-form solutions are obtained for both cases and applied to decide the optimal condition under whether the local execution or the remote execution is more energy-efficient for the mobile device. Moreover, numerical results illustrate that a significant amount of energy (e.g., up to 13 times for a typical mobile application profile) can be saved by optimally offloading the mobile application to the cloud clone. Yonggang Wen 0001, Haiyun Luo |
INFOCOM | 1 |
| 2012 | Credit routing for source-location privacy protection in wireless sensor networksabstractSource-location privacy became one of major issues due to the open nature of wireless sensor networks. The adversary can eavesdrop and trace the message movements so as to capture the source. In the paper, first we propose Credit routing to provide the source-location privacy protection. Credit routing is able to route the message within the assigned credit at each message and randomize the routing path. Unlike other location privacy protection schemes in WSN, Credit routing not only can provide strong protection but also precisely control the transmission cost of each message. Then, we propose Hybrid credit routing, which routes the message to the receiver through three phases: totally random walk, forwarding random walk and credit random walk. These phases provide tri-fold protection to prevent the source from being captured by the adversary. We evaluate our proposed schemes based on several metrics including safe period, latency and protection efficiency. The simulation results show that Credit is able to provide the strong and efficient protection compared with other schemes including Phantom, LPR and RRIN. It is also shown Hybrid improves the protection strength and efficiency even further. The performance of Credit and Hybrid can be tuned by the assigned credit. For real application, the credit can be the real power consumption for forwarding the message from the source to the sink. So both Credit and Hybrid can be used to precisely control the power consumption for source-location protection. Zongqing Lu 0002, Yonggang Wen 0001 |
MASS | 2 |
| 2012 | Distributed and Asynchronous Solution to Operator Placement in Large Wireless Sensor NetworksabstractDue to energy limitation of wireless sensor networks, in-network aggregation and distributed data fusion are proposed to perform the desired aggregation (or fusion) operators en route-eliminating data redundancy, minimizing transmissions and thus saving energy. An operator involved with in-network processing will be placed on a network node, which receives the data from sources, process them and send the output to either the next operator or sink node. As transmitting data from one operator to other imposes a cost, which is dependent on the placement of operators, the placement of operators can greatly affect the energy cost of in-network processing. In this research work, we propose a minimum-cost forwarding based asynchronous algorithm (MCFA) to find the optimal placement for operator tree with minimized energy cost of in-network processing. It is shown that minimum-cost forwarding can dramatically reduce message overhead of asynchronous algorithm. It is also shown that MCFA has less message overhead than synchronous algorithm by both mathematical analysis and simulation-based evaluation. For a regular grid network and a complete binary operator tree, the messages sent at each node are O(√NM) for MCFA, meanwhile O(√NM log2M) for synchronous algorithm, where N is the number of network nodes and M is the number of data objects in operator tree. Zongqing Lu 0002, Yonggang Wen 0001 |
MSN | 2 |
| 2012 | IPAD: Intelligent Parking-Assisted Ads Dissemination over Urban VANETsabstractAdvertisement dissemination via vehicular ad hoc networks (VANETs), fuelled by commercial interests, has attracted tremendous research efforts lately. However, the existing ads dissemination schemes could either incur a high operational overhead in inter-vehicle communication or lead to a substantial capital overhead of constructing roadside infrastructure. To reduce the aforementioned costs, we propose IPAD (Intelligent Parking-assisted Ads Dissemination), which taps into the unused resources (e.g., wireless device, rechargeable battery, storage capability, on-board computer chip) offered by roadside parking in urban areas to facilitate ads dissemination to mobile vehicles. Our proposed IPAD scheme is substantiated with a novel architecture, which is cost saving and could support large scale ads dissemination. To realize efficient ads dissemination over this artitecure, we put forward an effective routing scheme to distribute each ad to appropriate roadside parking and introduce the pub/sub scheme into the last stage of ads dissemination. Finally, we investigate IPAD through theoretic analysis and simulation. The numerical results obtained verify that our scheme offers an enhanced ad delivery ratio and reduces the network traffic overhead in ads dissemination via VANETs. Hai-gang Gong, Ming Liu 0002, Yonggang Wen 0001 |
MSN | 5 |
| 2012 | Depth-color based 3D image transmission over wireless networks with QoE provisions
Honggang Wang 0001, Yonggang Wen 0001, Dalei Wu, Ken C. K. Lee |
Comput. Commun. | 3 |
| 2011 | Towards name-based trust and security for content-centric networkabstractTrust and security have been considered as built-in properties for future Internet architecture. Leveraging the concept of named content in recently proposed information centric network, we propose a name-based trust and security protection mechanism. Our scheme is built with identity-based cryptography (IBC), where the identity of a user or device can act as a public key string. Uniquely, in named content network such as content-centric network (CCN), a content name or its prefixes can be used as public identities, with which content integrity and authenticity can be achieved with IBC algorithms. The trust of a content is seamlessly integrated with the verification of the content's integrity and authenticity with its name or prefix, instead of the public key certificate of its publisher. In addition, flexible confidentiality protection is enabled between content publishers and consumers. For scalable deployment purpose, we further propose to use a hybrid scheme combined with traditional public-key infrastructure (PKI) and IBC. We have implemented this scheme with CCNx open source project on Android. Xinwen Zhang, Katharine Chang, Huijun Xiong, Yonggang Wen 0001, Guangyu Shi, Guoqiang Wang 0001 |
ICNP | 4 |
| 2011 | The last minute: Efficient Data Evacuation strategy for sensor networks in post-disaster applicationsabstractDisasters (e.g., earthquakes, flooding, tornadoes, oil spilling and mining accidents) often result in tremendous cost to our society. Previously, wireless sensor networks (WSNs) have been proposed and deployed to provide information for decision making in post-disaster relief operations. The existing WSN solutions for post-disaster operations normally assume that the deployed sensor network can tolerate the damage caused by disasters and maintain its connectivity and coverage, even though a significant portion of nodes have been physically destroyed. In reality, however, this assumption is often invalid for disastrous events like earthquakes in large scale, limiting the relief capability of the existing solutions. Inspired by the “blackbox” technique in flight industry, we propose that preserving “the last snapshot” of the whole network and transferring those data to a safe zone would be the most logical approach to provide necessary information for rescuing lives and control damages. In this paper, we introduce Data Evacuation (DE), an original idea that takes advantage of the survival time of the WSN, i.e., the gap from the time when the disaster hits and the time when the WSN is paralyzed, to transmit critical data to sensor nodes in the safe area. Mathematically, the problem can be formulated as a nonlinear programming problem with multiple minimums in its support. We propose a gradient-based DE algorithm (GRAD-DE) to verify our DE strategy. Numerical investigations reveal the effectiveness of GRAD-DE algorithm. Ming Liu 0002, Hai-gang Gong, Yonggang Wen 0001, Guihai Chen, Jiannong Cao 0001 |
INFOCOM | 3 |
| 2011 | On file-based content distribution over wireless networks via multiple paths: Coding and delay trade-offabstractWith the emergence of the adaptive bit rate (ABR) streaming technology, the video/content streaming technology is shifting toward a file-based content distribution. That is, video content is encoded into a set of smaller media files containing video of 2–10 seconds before transmission. This file-based content distribution, coupled with increasingly rapid adoption of smartphones, requires an efficient file-based distribution algorithm to satisfy the QoS demand in wireless networks. In this paper, we study the transmission of a finite-sized file over wireless networks using multipath routing, with the objective to minimize file transmission delay instead of average packet delay. The file transmission delay is defined as the time interval from the instant that a file is first transmitted to the time at which the file can be reconstructed in the destination node. We observe that file transmission delay depends not only on the mean of the packet delay but also on its distribution, especially the tail. This observation leads to a better understanding of the file transfer delay in wireless networks and a minimum delay file transmission strategy. In a wireless multipath communication scenario, we propose to use packet level erasure code (e.g., digital fountain code) to transmit data file with redundancy. Given that a file with k packets is encoded into n packets for transmission, the use of digital fountain code allows the file to be received when only k out of n packets are received. By adding redundant packets, the destination node does not have to wait for the packet to arrive late, hence reducing the delay of the file transmission. We characterize the tradeoff between the code rate (i.e., the ratio of the number of transmitted packets to the number of the original packets) and the file delay reduction. As a rule of thumb, we provide practical guidelines in determining an appropriate code rate for a fixed file to achieve a reasonable transmission delay. We show that only a few redundant packets are needed to achieve a significant reduction in file transmission delay. Jun Sun 0007, Yonggang Wen 0001, Lizhong Zheng |
INFOCOM | 2 |
| 2011 | Buffer and Switch: An Efficient Road-to-Road Routing Scheme for VANETsabstractVehicular Ad Hoc Networks (VANETs) are getting increasing attention from academic researchers and automotive industries. Timely and cost-efficient multi-hop data delivery among vehicles is essential for VANETs, and various routing protocols are envisioned for infrastructure-less vehicle-to vehicle (V2V) communications. Due to the road-constrained data delivery and highly dynamic topology of vehicle nodes, it's better to construct routing based on the road-to-road pattern than the traditional node-to-node routing pattern in MANETs. However, the challenging issue for the road-to-road routing in VANETs is the opportunistic forwarding at intersections. Therefore, we propose a novel routing scheme, called Buffer and Switch (BAS). In BAS, each road buffers the data packets with multiple duplicates propagation in order to provide more opportunities for packet switching at intersections. Different from conventional protocols in VANETs, the propagation of duplicates in BAS is bidirectional along the routing path. Moreover, BAS's cost is much lower than other flooding-based protocols due to its spatio-temporally controlled duplicates propagation. We conduct the extensive simulations to evaluate the performance of BAS based on the road map of a real city collected from Google Earth. The simulation results show that BAS can outperform the existing protocols, especially when the network resources are limited. Chao Song 0002, Ming Liu 0002, Yonggang Wen 0001, Jiannong Cao 0001, Guihai Chen |
MSN | 3 |
| 2008 | Cost-Efficient Transmitter/Receiver Deployment for Proactive Fault Diagnosis in All-Optical NetworksabstractA scalable fault management system, including fault detection and localization capability, is crucial for future all-optical networks. In our previous work [5-7], we have proposed adaptive fault diagnosis schemes that deploy proactive lightpath probes to identify network failures, and have developed an asymptotically optimal run-length probing scheme to minimize the diagnostic effort (i.e., the number of lightpath probes). In this research, we aim to investigate the diagnostic hardware cost, i.e., the cost resulted from transmitter/receiver (Tx/Rx) pairs for probe transmission and detection. As a benchmark, we first show that, in order to identify all possible network failures, all the network nodes have to be equipped with diagnostic Tx/Rx pairs. We then develop a probabilistic analysis framework to characterize the trade-off between hardware cost (i.e., the number of nodes equipped with Tx/Rx pairs) and diagnosis capability (i.e., the probability of successful failure detection and localization). Our results suggest that, for practical situations, the hardware cost can be reduced significantly by accepting a small amount of uncertainty about the failure status. Yonggang Wen 0001, Vincent W. S. Chan, Eric A. Swanson |
ICC | 1 |
| 2008 | Optimizing Joint Erasure- and Error-Correction Coding for Wireless Packet TransmissionsabstractTo achieve reliable packet transmission over a wireless link without feedback, we propose a layered coding approach that uses error-correction coding within each packet and erasure-correction coding across the packets. This layered approach is also applicable to an end-to-end data transport over a network where a wireless link is the performance bottleneck. We investigate how to optimally combine the strengths of error- and erasure-correction coding to optimize the system performance with a given resource constraint, or to maximize the resource utilization efficiency subject to a prescribed performance. Our results determine the optimum tradeoff in splitting redundancy between error-correction coding and erasure-correction codes, which depends on the fading statistics and the average signal to noise ratio (SNR) of the wireless channel. For severe fading channels, such as Rayleigh fading channels, the tradeoff leans towards more redundancy on erasure-correction coding across packets, and less so on error-correction coding within each packet. For channels with better fading conditions, more redundancy can be spent on error-correction coding. The analysis has been extended to a limiting case with a large number of packets, and a scenario where only discrete rates are available via a finite number of transmission modes. Christian R. Berger, Shengli Zhou 0001, Yonggang Wen 0001, Peter Willett 0001, Krishna R. Pattipati |
IEEE Trans. Wirel. Commun. | 3 |
| 2007 | Non-Adaptive Fault Diagnosis for All-Optical Networks via Combinatorial Group Testing on GraphsabstractWe consider the problem of detecting failures for all-optical networks, with the objective of keeping the diagnosis cost low. Compared to the passive paradigm based on parity check in SONET, optical probing signals are sent proactively along lightpaths to probe their state of health and failure pattern is identified through the set of test results (i.e., probe syndromes). As an alternative to our previous adaptive approach where all the probes are sent sequentially, we consider in this work a non-adaptive approach where all the probes are sent in parallel. The design objective is to minimize the number of parallel probes, so as to keep network cost low. The non-adaptive fault diagnosis approach motivates a new technical framework that we introduce: combinatorial group testing with graph-based constraints. Using this framework, we develop several new probing schemes to detect network faults for all-optical networks with different topologies. The efficiency of our schemes often depends on the network topology; in many cases we can show that our schemes are optimal in minimizing the number of probes. Nicholas J. A. Harvey, Mihai Patrascu, Yonggang Wen 0001, Sergey Yekhanin, Vincent W. S. Chan |
INFOCOM | 3 |
| 2007 | On Minimum-Delay Data Block Transport over Two-Connected Mesh NetworksabstractIn this paper we investigate the problem of sending data blocks containing finite number of packets over two independent routing paths in mesh networks. The objective is to minimize the average block delay by allocating packets into the two routing paths optimally. Previous researchers have shown that using more paths can reduce packet delay and the rate-based allocation policy is optimal, based on the delay metric solely depending on delay mean of each path. In our research, generalizing the multi-path routing scheme to accommodate data block transport, we first establish an upper bound for the average block delay, depending on both delay mean and variance of each path, and then solve a non-linear optimization problem to obtain the optimal packet allocation policy, which either allocates all the packets to the faster path for blocks of small size or allocate all the packets to both paths in proportion to their service rates for blocks of large size. Contrary to conventional results, our analysis suggests that using additional slower routing path could increase the average block delay in some cases. We further characterize the whole spectrum of optimal packet allocation policies as a function of block sizes, and conclude that the existing rate-based allocation policy is a special case for large data block. Yonggang Wen 0001, Jun Sun 0007 |
WCNC | 1 |
| 2006 | Efficient Fault Detection and Localization for All-Optical NetworksabstractWe investigate the fault diagnosis problem for all- optical networks with probabilistic link and node failures in this paper. Our major contribution is the development of diagnosis algorithms that minimize the operating effort to identify failures. We achieve this by employing the fault diagnosis approach based on proactive probing with perfect feedback: knowledge of the network state is progressively refined through a sequence of optical probe signals, each of which is determined upon the results of previous probe signals (i.e. probe syndromes). To detect and localize failures in all-optical networks with probabilistic node and link failures, we introduce a network transformation that maps both link and node failures in an undirected graph into arc failures in a directed graph and apply our previously developed run-length probing scheme [1] to the directed graph. Our analytical and numerical investigation verifies our previously established guideline for efficient fault diagnosis algorithms: each probe should provide approximately 1-bit of state information, and thus the total number of probes required is approximately equal to the entropy of the network state. Hence the complexity of optical network fault management functionality is fundamentally related to the information entropy of the network state. Yonggang Wen 0001, Vincent W. S. Chan, Lizhong Zheng |
GLOBECOM | 1 |
| 2006 | Efficient fault diagnosis for all-optical networks: an information theoretic approachabstractNetwork management and control contribute to at least half of the operating cost of current optical networks. All-optical networks with end-to-end transparent lightpaths promise significant cost savings using optical switching at network nodes. However, this cost saving cannot be realized unless the cost of network management is also reduced. In this paper we explore a promising technique towards that goal. The fault diagnosis problem for all-optical networks is investigated via an information theoretic approach, with the objective to minimize the operating 'cost' of failure detection and localization in the optical layer. Under a probabilistic link failure model, we first interpret the run-length probing scheme previously developed for Eulerian graphs as a constrained source-coding algorithm, and characterize its performance via the code rate of its corresponding run-length code. We then extend the run-length probing scheme to non-Eulerian graphs via two alternative approaches: the disjoint-trail decomposition approach and the path-augmentation approach, and obtain their performance analytically. The analytical and numerical results indicate that the run-length probing scheme is asymptotically optimum for both Eulerian and non-Eulerian graphs of large size. The property of the run-length probing scheme also suggests that each probe in an efficient probing scheme should provide approximately one bit of network state information and thus the total number of probes (or equivalently, the operating cost of failure identification) is lower-bounded and approximated by the entropy of the network states. We believe that our approach using information theory in an inter-disciplinary effort can provide new insights on network management, and substantial cost-reduction for all-optical networks can be realized Yonggang Wen 0001, Vincent W. S. Chan, Lizhong Zheng |
ISIT | 1 |
| 2005 | Network monitoring in multicast networks using network codingabstractIn this paper we show how information contained in robust network codes can be used for passive inference of possible locations of link failures or losses in a network. For distributed randomized network coding, we bound the probability of being able to distinguish among a given set of failure events, and give some experimental results for one and two link failures in randomly generated networks. We also bound the required field size and complexity for designing a robust network code that distinguishes among a given set of failure events Tracey Ho, Ben Leong, Yu-Han Chang, Yonggang Wen 0001, Ralf Koetter |
ISIT | 4 |
| 2005 | Ultra-Reliable Communication Over Vulnerable All-Optical Networks Via Lightpath DiversityabstractIn this paper, we propose using spatial diversity via multiple node-disjointed lightpaths at the optical layer to achieve ultra-reliable communication with low delay between any source-destination pair in all-optical networks. Using a doubly stochastic point process model and a "genie-aided" receiver, we obtain an exponentially tight error probability bound for the lightpath diversity scheme under an independent lightpath failure model. Error probability of the proposed scheme can be designed to be significantly lower than that of a system without lightpath diversity, and system parameters (e.g., the number of lightpaths) can be optimized to achieve efficient utilization of a limited amount of transmitted optical energy. In particular, at the optimum operating point, each lightpath is allocated an optimum average number of signal photons per bit and is biased to have an effective error probability 2f if the decision is based on that path alone, where f is the lightpath failure probability. We also investigate the tradeoff between the error probability and the implementation complexity within the class of all "structured" receivers. We derive receiver architectures for both the optimal receiver, which has the best error performance but complicated receiver architecture, and the equal-gain-combining (EGC) receiver, which has suboptimum error performance but simpler receiver architecture. Closed-form error bounds for both receivers are obtained and compared with the "genie-aided" limit of the lightpath diversity scheme. Performance comparison shows that the simpler equal-gain-combing receiver provides similar performance as the optimal receiver in the regime of high signal-to-noise photon rate ratio (/spl Omega/=/spl lambda//sub s//M/spl lambda//sub n/, where /spl lambda//sub s//M is the signal photon rate per path, /spl lambda//sub n/ is the noise photon rate per path), and performs slightly worse than the optimal receiver in the low and medium signal-to-noise photon rate ratio regimes. It indicates that the simpler EGC receiver is preferred over the complicated optimum receiver in practical receiver design. Yonggang Wen 0001, Vincent W. S. Chan |
IEEE J. Sel. Areas Commun. | 1 |
| 2004 | Ultrareliable communication over vulnerable optical networks via lightpath diversity: receiver architectures and performanceabstractWe develop a class of structured receivers for a lightpath diversity scheme, which was introduced to provide ultrareliable communication with low delay in vulnerable all-optical networks. We explore the trade-off between implementation complexity and error probability to achieve optimum and near-optimum performance within a class of structured receivers. Using a doubly-stochastic point process model, we develop receiver architectures for both the optimal receiver with respect to error performance, and the equal-gain-combining receiver with suboptimum error performance but simpler receiver architecture. Closed-form error bounds for both receivers are obtained and compared with the 'Genie-aided' limit of the lightpath diversity transmission scheme. The comparison shows that the error performance of both receivers approaches the 'Genie-aided' limit when the signal is strong. Numerical results also demonstrate that additional power over what is required for the optimum receiver is needed to be transmitted in order for the equal-gain-combining receiver to achieve the same target error probability, and the power penalty decreases with decreasing noise level. These results suggest that the simpler equal-gain-combing receiver provides similar performance as the optimal receiver in the high signal-to-noise ratio (SNR) regime, but the optimal receiver should be used in the low SNR regime for significantly better performance. Yonggang Wen 0001, Vincent W. S. Chan |
ICC | 1 |
| 2003 | Ultra-reliable communication over unreliable optical networks via lightpath diversity: system characterization and optimizationabstractWe propose using diversity via multiple disjoint lightpaths at the optical layer to achieve ultra-reliable communication with low delay between any source-destination pair of all-optical networks. A doubly-stochastic point process model is used to characterize the photo-events of a direct detection receiver. The error probability can be designed to be significantly lower than that of a system without lightpath diversity. System parameters, such as the number of lightpaths used, are optimized to achieve efficient utilization of the limited optical transmitter power. Yonggang Wen 0001, Vincent W. S. Chan |
GLOBECOM | 1 |