VLDB 2026 Research / reviewers in the wild / expert
Zhi Wang 0001
dblp:95/6543-1
· DBLP profile ↗
191ranked-venue papers
25as first author
112since 2021 · last 2026
0000-0002-5462-6178ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 87 · 7 first-author · 54 since 2021Artificial intelligence and machine learning · 59 · 7 first-author · 54 since 2021Computer networks · 50 · 8 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 6 since 2021Systems, architecture and hardware · 6 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 5 · 2 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EMS-GL: Adaptive Evict-then-Merge Strategy for KV Cache Compression Based on Global-Local Importance
Yingxin Li, Ye Li 0016, Xinzhu Ma, Zihan Geng, Shutao Xia, Zhi Wang 0001 |
KSEM (1) | 7 |
| 2026 | PRANCE: Joint Token-Optimization and Structural Channel-Pruning for Adaptive ViT InferenceabstractThe troublesome model size and quadratic computational complexity associated with token quantity pose significant deployment challenges for Vision Transformers (ViTs) in practical applications. Despite recent advancements in model pruning and token reduction techniques speed up the inference speed of ViTs, these approaches either adopt a fixed sparsity ratio or overlook the meaningful interplay between architectural optimization and token selection. Consequently, this static and single-dimension compression often leads to pronounced accuracy degradation under aggressive compression rates, as they fail to fully explore redundancies across these two orthogonal dimensions. Therefore, we introduce PRANCE, a framework which can jointly optimize activated channels and tokens on a per-sample basis, aiming to accelerate ViTs' inference process from a unified data and architectural perspective. However, the joint framework poses challenges to both architectural and decision-making aspects. First, while ViTs inherently support variable-token inference, they do not facilitate dynamic computations for variable channels. To overcome this limitation, we propose a meta-network using weight-sharing techniques to support arbitrary channels of the Multi-Head Self-Attention (MHSA) and Multi-Layer Perceptron (MLP) layers, serving as a foundational model for architectural decision-making. Second, simultaneously optimizing the model structure and input data constitutes a combinatorial optimization problem with an extremely large decision space, reaching up to around $10^{14}$1014, making supervised learning infeasible. To this end, we design a lightweight selector employing Proximal Policy Optimization algorithm (PPO) for efficient decision-making. Furthermore, we introduce a novel "Result-to-Go" training mechanism that models ViTs' inference process as a Markov decision process, significantly reducing action space and mitigating delayed-reward issues during training. Additionally, our framework simultaneously supports different kinds of token optimization methods such as pruning, merging, and sequential pruning-merging strategies. Extensive experiments demonstrate the effectiveness of PRANCE in reducing FLOPs by approximately 50%, retaining only about 10% of tokens while achieving lossless Top-1 accuracy. Ye Li 0016, Jiajun Fan, Zenghao Chai, Xinzhu Ma, Zhi Wang 0001, Wenwu Zhu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2026 | Learning gated experts for segment anything in the wild
Yizhen Guo, Hang Guo 0002, Tao Dai 0001, Zhi Wang 0001, Bin Chen 0011, Shutao Xia |
Pattern Recognit. | 4 |
| 2026 | Universal image restoration via task-adaptive diffusion degradation oriented model
Junxi Wu, Sicheng Pan, Naiqi Li, Bin Chen 0011, Baoyi An 0002, Zhi Wang 0001, Yaowei Wang 0001, Shutao Xia |
Pattern Recognit. | 6 |
| 2026 | Leveraging Neural Architecture Search for improved downstream-agnostic adversarial attack
Haodong Xiao, Bin Chen 0011, Hao Fang 0011, Yulin Wu 0001, Xuan Wang 0002, Zhi Wang 0001, Shutao Xia |
Pattern Recognit. | 8 |
| 2026 | Multi-Beholder: Biomarker Prediction for Low-Grade Glioma With Multiple Instance Learning and One-Class ClassificationabstractBiomarker detection is an indispensable part of the diagnosis and treatment of low-grade glioma (LGG). However, current LGG biomarker detection methods rely on expensive and complex molecular genetic testing, for which professionals are required to analyze the results, and intra-rater variability is often reported. To overcome these challenges, we propose an interpretable deep learning pipeline, named Multi-Biomarker Histomorphology Discoverer (Multi-Beholder), to predict the status of five biomarkers in LGG using only hematoxylin and eosin-stained whole slide images. Specifically, Multi-Beholder incorporates one-class classification into the multiple instance learning framework to achieve accurate instance-level pseudo-labeling, thereby complementing slide-level labels and improving prediction performance. Multi-Beholder demonstrates high performance on two LGG cohorts with diverse races and scanning protocols, with area under the receiver operating characteristic curve up to 0.973 on the internal-validated TCGA-LGG dataset and 0.820 on the external-validated Xiangya cohort. Moreover, the interpretability of Multi-Beholder allows for discovering quantitative and qualitative correlations between biomarker status and histomorphology characteristics. Our pipeline not only provides a novel approach for biomarker prediction, enhancing the applicability of molecular treatments for LGG patients but also facilitates the discovery of new mechanisms in molecular functionality and LGG progression. Code can be accessed athttps://github.com/Vison307/Multi-Beholder. Zijie Fang, Yifeng Wang 0001, Yang Chen 0036, Changjing Cai, Yiyang Lin, Zhi Wang 0001, Shan Zeng, Yongbing Zhang 0002 |
IEEE Trans. Comput. Biol. Bioinform. | 9 |
| 2026 | Breaking the Curse of Knowledge: Towards Effective Multimodal Recommendation Using Knowledge Soft IntegrationabstractA critical challenge in contemporary recommendation systems lies in effectively leveraging multimodal content to enhance recommendation personalization. Although various solutions have been proposed, most fail to account for discrepancies between knowledge extracted through isolated feature extraction and its application in recommendation tasks. Specifically, multimodal feature extraction does not incorporate task-specific prior knowledge, while downstream recommendation tasks typically use these features as auxiliary information. This misalignment often introduces biases in model fitting and degrades performance, a phenomenon we refer to as the curse of knowledge. To address this challenge, we propose a knowledge soft integration framework designed to balance the utilization of multimodal features with the biases they may introduce. The framework, namedKnowledgeSoftIntegration (KSI), comprises two key components: the Structure Efficient Injection (SEI) module and the Semantic Soft Integration (SSI) module. The SEI module employs a Refined Graph Neural Network (RGNN) to model inter-modal correlations among items while introducing a regularization term to minimize redundancy in user and item representations. In parallel, the SSI module utilizes a self-supervised retrieval task to implicitly integrate multimodal semantic knowledge, thereby enhancing the semantic distinctiveness of item representations. We conduct comprehensive experiments on three benchmark datasets, demonstrating KSI's effectiveness. Furthermore, these results underscore the ability of the SEI and SSI modules to reduce representation redundancy and mitigate the curse of knowledge in multimodal recommendation systems. Kai Ouyang, Zenghao Chai, Wenhao Zheng 0001, Xiangjin Xie, Xuanji Xiao, Zhi Wang 0001 |
IEEE Trans. Multim. | 7 |
| 2025 | JAQ: Joint Efficient Architecture Design and Low-Bit Quantization with Hardware-Software Co-ExplorationabstractThe co-design of neural network architectures, quantization precisions, and hardware accelerators offers a promising approach to achieving an optimal balance between performance and efficiency, particularly for model deployment on resource-constrained edge devices. In this work, we propose the JAQ Framework, which jointly optimizes the three critical dimensions. However, effectively automating the design process across the vast search space of those three dimensions poses significant challenges, especially when pursuing extremely low-bit quantization. Specifical, the primary challenges include: (1) Memory overhead in software-side: Low-precision quantization-aware training can lead to significant memory usage due to storing large intermediate features and latent weights for backpropagation, potentially causing memory exhaustion. (2) Search time-consuming in hardware-side: The discrete nature of hardware parameters and the complex interplay between compiler optimizations and individual operators make the accelerator search time-consuming. To address these issues, JAQ mitigates the memory overhead through a channel-wise sparse quantization (CSQ) scheme, selectively applying quantization to the most sensitive components of the model during optimization. Additionally, JAQ designs BatchTile, which employs a hardware generation network to encode all possible tiling modes, thereby speeding up the search for the optimal compiler mapping strategy. Extensive experiments demonstrate the effectiveness of JAQ, achieving approximately 7% higher Top-1 accuracy on ImageNet compared to previous methods and reducing the hardware search time per iteration to 0.15 seconds. Mingzi Wang, Weixiang Zhang, Yijian Qin, Yang Yao 0003, Yingxin Li, Tongtong Feng, Xin Wang 0019, Xun Guan, Zhi Wang 0001, Wenwu Zhu 0001 |
AAAI | 11 |
| 2025 | Q-DiT: Accurate Post-Training Quantization for Diffusion TransformersabstractRecent advancements in diffusion models, particularly the architectural transformation from UNet-based models to Diffusion Transformers (DiTs), significantly improve the quality and scalability of image and video generation. However, despite their impressive capabilities, the substantial computational costs of these large-scale models pose significant challenges for real-world deployment. Post-Training Quantization (PTQ) emerges as a promising solution, enabling model compression and accelerated inference for pretrained models, without the costly retraining. However, research on DiT quantization remains sparse, and existing PTQ frameworks, primarily designed for traditional diffusion models, tend to suffer from biased quantization, leading to notable performance degradation. In this work, we identify that DiTs typically exhibit significant spatial variance in both weights and activations, along with temporal variance in activations. To address these issues, we propose Q-DiT, a novel approach that seamlessly integrates two key techniques: automatic quantization granularity allocation to handle the significant variance of weights and activations across input channels, and sample-wise dynamic activation quantization to adaptively capture activation changes across both timesteps and samples. Extensive experiments conducted on ImageNet and VBench demonstrate the effectiveness of the proposed Q-DiT. Specifically, when quantizing DiT-XL/2 to W6A8 on ImageNet (256 × 256), Q-DiT achieves a remarkable reduction in FID by 1.09 compared to the baseline. Under the more challenging W4A8 setting, it maintains high fidelity in image and video generation, establishing a new benchmark for efficient, high-quality quantization in DiTs. Xinzhu Ma, Jingyan Jiang, Xin Wang 0019, Zhi Wang 0001, Wenwu Zhu 0001 |
CVPR | 7 |
| 2025 | Dynamic Model Fusion for Multi-Source Test-Time AdaptationabstractDeep Neural Networks suffer significant performance degradation when faced with distribution shifts between training and test data. Test-time adaptation (TTA) has emerged as a practical solution that enables models to adapt to the shifted test distribution. Currently, most existing TTA methods are designed around a single model, which incorporate limited information from a singular data distribution. In practice, pre-trained models derived from diverse source domains are readily accessible, each capturing a distinct data distribution and containing complementary information. To exploit this diversity, we propose Model Fusion-based multi-source Test-Time Adaptation (MFTTA), which constructs a target model by fusing the parameters of multiple source models. Drawing inspiration from deep model fusion, we introduce a fine-grained fusion mechanism governed by an off-policy reinforcement learning agent, which dynamically assigns fusion weights based on the current data distribution. Furthermore, we design a correlation-aware model update strategy that prioritizes the source model most relevant to the incoming test data. Extensive experiments on standard out-of-distribution benchmarks demonstrate that our method effectively integrates knowledge from multiple source models, adapts robustly to dynamic distribution shifts, and alleviates the problem of forgetting in long-term adaptation. Yuan Xue 0013, Qinting Jiang, Xingxuan Zhang, Jingyan Jiang, Zhi Wang 0001 |
ECAI | 7 |
| 2025 | Beyond A Single AI Cluster: A Survey of Decentralized LLM TrainingabstractThe emergence of large language models (LLMs) has revolutionized AI development, yet their resource demands beyond a single cluster or even datacenter, limiting accessibility to well-resourced organizations.Decentralized training has emerged as a promising paradigm to leverage dispersed resources across clusters, datacenters and even regions, offering the potential to democratize LLM development for broader communities.As the first comprehensive exploration of this emerging field, we present decentralized LLM training as a resource-driven paradigm and categorize existing efforts into community-driven and organizational approaches.We further clarify this through: (1) a comparison with related paradigms, (2) characterization of decentralized resources, and (3) a taxonomy of recent advancements.We also provide up-to-date case studies and outline future directions to advance research in decentralized LLM training. Haotian Dong, Jingyan Jiang, Rongwei Lu, Jiajun Luo, Jiajun Song, Zhi Wang 0001 |
EMNLP | 8 |
| 2025 | DICE: Staleness-Centric Optimizations for Parallel Diffusion MoE Inference
Jiajun Luo, Lizhuo Luo, Jianru Xu, Jiajun Song, Rongwei Lu, Zhi Wang 0001 |
ICCV | 7 |
| 2025 | CATP-LLM: Empowering Large Language Models for Cost-Aware Tool PlanningabstractUtilizing large language models (LLMs) for tool planning has emerged as a promising avenue for developing general AI systems, where LLMs automatically schedule external tools (e.g., vision models) to tackle complex tasks based on task descriptions. To push this paradigm toward practical applications, it is crucial for LLMs to consider tool execution costs (e.g., execution time) for tool planning. Unfortunately, prior studies overlook the tool execution costs, leading to the generation of expensive plans whose costs outweigh their benefits in terms of task performance. To fill this gap, we propose the Cost-Aware Tool Planning with LLMs (CATP-LLM) framework, which for the first time provides a coherent design to empower LLMs for cost-aware tool planning. Specifically, To facilitate efficient concurrent tool execution and cost reduction, we design a tool planning language to enhance the LLM for creating multi-branch non-sequential plans. Moreover, we propose a cost-aware offline reinforcement learning algorithm to fine-tune the LLM to optimize the performance-cost trade-off in tool planning. In the lack of public cost-related datasets, we further present OpenCATP, the first dataset for cost-aware planning, which comprises 11,100 evaluation samples from diverse tasks. Extensive experiments show that CATP-LLM outperforms GPT-4 even when using Llama2-7B as its backbone, with the average improvement of 1.5%-93.9% in terms of plan quality. Codes and dataset are available at: https://github.com/duowuyms/OpenCATP-LLM. Duo Wu, Jinghe Wang, Zhi Wang 0001 |
ICCV | 6 |
| 2025 | Anti-FT: Towards Practical Deep Leakage From GradientsabstractFederated learning is usually regarded as a privacy-preserving training paradigm for it enables multiple clients to participate in a training task without sharing their private data. However, recent studies revealed that a malicious server can still recover private data from the victim clients based on the shared gradients via deep leakage from gradients (DLG). Currently, almost all DLG attacks are designed based on the average loss, leading to a significant decrease in attack efficiency when the batch size is greater than 1. In this paper, we revisit DLG attacks from the perspective of the loss function. We reveal that not all samples in the target batch are equally susceptible to DLG attacks: the sample with the highest loss value tends to be easily recovered by DLG attacks. Based on these observations, we propose a simple yet effective DLG method under practical FL settings. Specifically, the adversaries can enhance the effectiveness of DLG by perturbing the global model through finetuning it with a few mislabeled samples (dubbed ‘Anti-FT’). Extensive experiments are conducted on benchmark datasets, which verify the effectiveness of our method and its resistance to potential defenses. The codes are available at https://github.com/zlh-thu/anti-finetune. Linghui Zhu, Yiming Li 0004, Haiqin Weng, Shutao Xia, Zhi Wang 0001 |
ICIP | 6 |
| 2025 | Expansive Supervision for Neural Radiance FieldsabstractNeural Radiance Field (NeRF) has achieved remarkable success in creating immersive media representations through its exceptional reconstruction capabilities. However, the computational demands of dense forward passes and volume rendering during training continue to challenge its real-world applications. In this paper, we introduce Expansive Supervision to reduce time and memory costs during NeRF training from the perspective of partial ray selection for supervision. Specifically, we observe that training errors exhibit a long-tail distribution correlated with image content. Based on this observation, our method selectively renders a small but crucial subset of pixels and expands their values to estimate errors across the entire area for each iteration. Compared to conventional supervision, our approach effectively bypasses redundant rendering processes, resulting in substantial reductions in both time and memory consumption. Experimental results demonstrate that integrating Expansive Supervision within existing state-of-the-art acceleration frameworks achieves 52% memory savings and 16% time savings while maintaining comparable visual quality. Our code is available at this link. Weixiang Zhang, Shuzhao Xie, Shijia Ge, Zhi Wang 0001 |
ICME | 6 |
| 2025 | Divide and Conquer: Grounding LLMs as Efficient Decision-Making Agents via Offline Hierarchical Reinforcement LearningabstractWhile showing sophisticated reasoning abilities, large language models (LLMs) still struggle with long-horizon decision-making tasks due to deficient exploration and long-term credit assignment, especially in sparse-reward scenarios. Inspired by the divide-and-conquer principle, we propose an innovative framework GLIDER (Grounding Language Models as EffIcient Decision-Making Agents via Offline HiErarchical Reinforcement Learning) that introduces a parameter-efficient and generally applicable hierarchy to LLM policies. We develop a scheme where the low-level controller is supervised with abstract, step-by-step plans that are learned and instructed by the high-level policy. This design decomposes complicated problems into a series of coherent chain-of-thought reasoning sub-tasks, providing flexible temporal abstraction to significantly enhance exploration and learning for long-horizon tasks. Furthermore, GLIDER facilitates fast online adaptation to non-stationary environments owing to the strong transferability of its task-agnostic low-level skills. Experiments on ScienceWorld and ALFWorld benchmarks show that GLIDER achieves consistent performance gains, along with enhanced generalization capabilities. Zican Hu, Wei Liu 0131, Xiaoye Qu, Xiangyu Yue 0001, Chunlin Chen 0001, Zhi Wang 0001, Yu Cheng 0001 |
ICML | 6 |
| 2025 | γ-FedHT: Stepsize-Aware Hard-Threshold Gradient Compression in Federated Learning
Rongwei Lu, Yifei Zhu 0001, Bin Chen 0011, Zhi Wang 0001 |
INFOCOM | 7 |
| 2025 | Quantization Meets OOD: Generalizable Quantization-aware Training from a Flatness PerspectiveabstractCurrent quantization-aware training (QAT) methods primarily focus on enhancing the performance of quantized models on in-distribution (I.D) data, while overlooking the potential performance degradation on out-of-distribution (OOD) data. In this paper, we first substantiate this problem through rigorous experiment, showing that QAT can lead to a significant OOD generalization performance degradation. Further, we find the contradiction between the perspective that flatness of loss landscape gives rise to superior OOD generalization and the phenomenon that QAT lead to a sharp loss landscape, can cause the above problem. Therefore, we propose a flatness-oriented QAT method, FQAT, to achieve generalizable QAT. Specifically, i) FQAT introduces a layer-wise freezing mechanism to mitigate the gradient conflict issue between dual optimization objectives (i.e., vanilla QAT and flatness). ii) FQAT proposes an disorder-guided adaptive freezing algorithm to dynamically determines which layers to freeze at each training step, effectively addressing the challenges caused by interference between layers. A gradient disorder metric is designed to help the algorithm identify unstable layers during training. Extensive experiments on influential OOD benchmark demonstrate the superiority of our method over state-of-the-art baselines under both I.D and OOD image classification tasks. Han Yu 0009, Zhi Wang 0001, Wenwu Zhu 0001 |
ACM Multimedia | 6 |
| 2025 | SizeGS: Size-aware Compression of 3D Gaussian Splatting via Mixed Integer ProgrammingabstractRecent advances in 3D Gaussian Splatting (3DGS) have greatly improved 3D reconstruction. However, its substantial data size poses a significant challenge for transmission and storage. While many compression techniques have been proposed, they fail to efficiently adapt to fluctuating network bandwidth, leading to resource wastage. We address this issue from the perspective of size-aware compression, where we aim to compress 3DGS to a desired size by quickly searching for suitable hyperparameters. Through a measurement study, we identify key hyperparameters that affect the size - namely, the reserve ratio of Gaussians and bit-width settings for Gaussian attributes. Then, we formulate this hyperparameter optimization problem as a mixed-integer nonlinear programming (MINLP) problem, with the goal of maximizing visual quality while respecting the size budget constraint. To solve the MINLP, we decouple this problem into two parts: discretely sampling the reserve ratio and determining the bit-width settings using integer linear programming (ILP). To solve the ILP more quickly and accurately, we design a quality loss estimator and a calibrated size estimator, as well as implement a CUDA kernel. Extensive experiments on multiple 3DGS variants demonstrate that our method achieves state-of-the-art performance in post-training compression. Furthermore, our method can achieve comparable quality to leading training-required methods after fine-tuning. Shuzhao Xie, Weixiang Zhang, Shijia Ge, Sicheng Pan, Yunpeng Bai, Cong Zhang 0002, Xiaoyi Fan 0001, Zhi Wang 0001 |
ACM Multimedia | 10 |
| 2025 | Accelerating Parallel Diffusion Model Serving with Residual CompressionabstractDiffusion models produce realistic images and videos but require substantial computational resources, necessitating multi-accelerator parallelism for real-time deployment. However, parallel inference introduces significant communication overhead from exchanging large activations between devices, limiting efficiency and scalability. We present CompactFusion, a compression framework that significantly reduces communication while preserving generation quality. Our key observation is that diffusion activations exhibit strong temporal redundancy—adjacent steps produce highly similar activations, saturating bandwidth with near-duplicate data carrying little new information. To address this inefficiency, we seek a more compact representation that encodes only the essential information. CompactFusion achieves this via Residual Compression that transmits only compressed residuals (step-wise activation differences). Based on empirical analysis and theoretical justification, we show that it effectively removes redundant data, enabling substantial data reduction while maintaining high fidelity. We also integrate lightweight error feedback to prevent error accumulation. CompactFusion establishes a new paradigm for parallel diffusion inference, delivering lower latency and significantly higher generation quality than prior methods. On 4$\times$L20, it achieves $3.0\times$ speedup while greatly improving fidelity. It also uniquely supports communication-heavy strategies like sequence parallelism on slow networks, achieving $6.7\times$ speedup over prior overlap-based method. CompactFusion applies broadly across diffusion models and parallel settings, and integrates easily without requiring pipeline rework. Portable implementation demonstrated on xDiT is publicly available at https://github.com/Cobalt-27/CompactFusion Jiajun Luo, Yicheng Xiao, Jianru Xu, Yangxiu You, Rongwei Lu, Jingyan Jiang, Zhi Wang 0001 |
NeurIPS | 8 |
| 2025 | Mixture-of-Experts Meets In-Context Reinforcement LearningabstractIn-context reinforcement learning (ICRL) has emerged as a promising paradigm for adapting RL agents to downstream tasks through prompt conditioning. However, two notable challenges remain in fully harnessing in-context learning within RL domains: the intrinsic multi-modality of the state-action-reward data and the diverse, heterogeneous nature of decision tasks. To tackle these challenges, we propose **T2MIR** (**T**oken- and **T**ask-wise **M**oE for **I**n-context **R**L), an innovative framework that introduces architectural advances of mixture-of-experts (MoE) into transformer-based decision models. T2MIR substitutes the feedforward layer with two parallel layers: a token-wise MoE that captures distinct semantics of input tokens across multiple modalities, and a task-wise MoE that routes diverse tasks to specialized experts for managing a broad task distribution with alleviated gradient conflicts. To enhance task-wise routing, we introduce a contrastive learning method that maximizes the mutual information between the task and its router representation, enabling more precise capture of task-relevant information. The outputs of two MoE components are concatenated and fed into the next layer. Comprehensive experiments show that T2MIR significantly facilitates in-context learning capacity and outperforms various types of baselines. We bring the potential and promise of MoE to ICRL, offering a simple and scalable architectural enhancement to advance ICRL one step closer toward achievements in language and vision communities. Our code is available at [https://github.com/NJU-RL/T2MIR](https://github.com/NJU-RL/T2MIR). Fuhong Liu, Haoru Li, Zican Hu, Daoyi Dong, Chunlin Chen 0001, Zhi Wang 0001 |
NeurIPS | 7 |
| 2025 | Learning to Reason under Off-Policy GuidanceabstractRecent advances in large reasoning models (LRMs) demonstrate that sophisticated behaviors such as multi-step reasoning and self-reflection can emerge via reinforcement learning with verifiable rewards~(RLVR).
However, existing RLVR approaches are inherently ``on-policy'', limiting learning to a model's own outputs and failing to acquire reasoning abilities beyond its initial capabilities.
To address this issue, we introduce LUFFY (Learning to reason Under oFF-policY guidance), a framework that augments RLVR with off-policy reasoning traces.
LUFFY dynamically balances imitation and exploration by combining off-policy demonstrations with on-policy rollouts during training.
Specifically, LUFFY combines the Mixed-Policy GRPO framework, which has a theoretically guaranteed convergence rate, alongside policy shaping via regularized importance sampling to avoid superficial and rigid imitation during mixed-policy training.
Compared with previous RLVR methods, LUFFY achieves an over +6.4 average gain across six math benchmarks and an advantage of over +6.2 points in out-of-distribution tasks.
Most significantly, we show that LUFFY successfully trains weak models in scenarios where on-policy RLVR completely fails. These results provide compelling evidence that LUFFY transcends the fundamental limitations of on-policy RLVR and demonstrates the great potential of utilizing off-policy guidance in RLVR. Jianhao Yan, Yafu Li, Zican Hu, Zhi Wang 0001, Ganqu Cui, Xiaoye Qu, Yu Cheng 0001, Yue Zhang 0004 |
NeurIPS | 4 |
| 2025 | Text-to-Decision Agent: Offline Meta-Reinforcement Learning from Natural Language SupervisionabstractOffline meta-RL usually tackles generalization by inferring task beliefs from high-quality samples or warmup explorations. The restricted form limits their generality and usability since these supervision signals are expensive and even infeasible to acquire in advance for unseen tasks. Learning directly from the raw text about decision tasks is a promising alternative to leverage a much broader source of supervision. In the paper, we propose **T**ext-to-**D**ecision **A**gent (**T2DA**), a simple and scalable framework that supervises offline meta-RL with natural language. We first introduce a generalized world model to encode multi-task decision data into a dynamics-aware embedding space. Then, inspired by CLIP, we predict which textual description goes with which decision embedding, effectively bridging their semantic gap via contrastive language-decision pre-training and aligning the text embeddings to comprehend the environment dynamics. After training the text-conditioned generalist policy, the agent can directly realize zero-shot text-to-decision generation in response to language instructions. Comprehensive experiments on MuJoCo and Meta-World benchmarks show that T2DA facilitates high-capacity zero-shot generalization and outperforms various types of baselines. Our code is available at [https://github.com/NJU-RL/T2DA](https://github.com/NJU-RL/T2DA). Zican Hu, Jianxiang Tang, Chunlin Chen 0001, Daoyi Dong, Yu Cheng 0001, Zhenhong Sun, Zhi Wang 0001 |
NeurIPS | 10 |
| 2025 | Understanding Bias Terms in Neural RepresentationsabstractIn this paper, we examine the impact and significance of bias terms in Implicit Neural Representations (INRs). While bias terms are known to enhance nonlinear capacity by shifting activations in typical neural networks, we discover their functionality differs markedly in neural representation networks.
Our analysis reveals that INR performance neither scales with increased number of bias terms nor shows substantial improvement through bias term gradient propagation. We demonstrate that bias terms in INRs primarily serve to eliminate \textit{spatial aliasing} caused by symmetry from both coordinates and activation functions, with input-layer bias terms yielding the most significant benefits.
These findings challenge the conventional practice of implementing full-bias INR architecture.
We propose using freezing bias terms exclusively in input layers, which consistently outperforms fully biased networks in signal fitting tasks.
Furthermore, we introduce Feature-Biased INRs~(Feat-Bias), which initialize input-layer bias with high-level features extracted from pre-trained models. This feature-biasing approach effectively addresses the limited performance in INR post-processing tasks due to neural parameter uninterpretability, achieving superior accuracy while reducing parameter count and improving reconstruction quality. Weixiang Zhang, Boxi Li, Shuzhao Xie, Chengwei Ren, Yuan Xue 0013, Zhi Wang 0001 |
NeurIPS | 6 |
| 2025 | Harnessing WebRTC for Large-Scale Live StreamingabstractLive streaming that supports real-time interaction has become increasingly popular. To support the ensuing requirements on low end-to-end latency, RTM, the state-of-the-art live streaming system at Douyin, replaces the HTTP-FLV streaming protocol with WebRTC. To tailor the WebRTC stack to the live streaming scenario, we focus on optimizing first-frame delay, startup video rebuffering, audio-to-video drift, and per-session CPU usage. Those are the top-priority metrics identified from an importance analysis with respect to two user engagement metrics, i.e., viewer penetration and viewing time. To date, WebRTC-based streaming in RTM has been in operation for 4 years, and serves billions of viewer sessions every day. It dramatically optimizes QoE metrics (e.g., end-to-end latency reduced by 54.5%), and delivers statistically significant user engagement gains (e.g., number of paid orders increased by 0.8%). In this paper, we report our deployment experiences comprehensively. Wei Zhang 0074, Tong Meng, Changqing Yan, Feng Qian 0001, Lei Zhang 0066, Zhi Wang 0001 |
SIGCOMM | 11 |
| 2025 | PDD: Planning Offline Meta-RL with Prompt Decision DiffuserabstractInspired by diffusion models that revolutionized image generation through language conditioning, offline reinforcement learning (RL) has been reformulated as a sequence modeling problem. Similar to how language descriptions guide image generation, RL sequence models require task-specific conditioning information to generalize to new tasks. We propose Prompt Decision Diffuser (PDD), addressing the generalization challenge in offline meta reinforcement learning (OMRL) by treating it as a sequence modeling problem, with demonstration-based prompting for few-shot adaptation. These prompts, encoded from few-shot demonstrations, effectively capture task-specific information to guide cross-task policy generation. Experimental results on Mujoco and Point-Robot benchmarks demonstrate PDD's superior few-shot generalization capabilities compared to baseline approaches. Zican Hu, Jiangxiang Tang, Zhi Wang 0001 |
SMC | 6 |
| 2025 | MIXRTs: Toward Interpretable Multi-Agent Reinforcement Learning via Mixing Recurrent Soft Decision TreesabstractWhile achieving tremendous success in various fields, existing multi-agent reinforcement learning (MARL) with a black-box neural network makes decisions in an opaque manner that hinders humans from understanding the learned knowledge and how input observations influence decisions. In contrast, existing interpretable approaches usually suffer from weak expressivity and low performance. To bridge this gap, we propose MIXing Recurrent soft decision Trees (MIXRTs), a novel interpretable architecture that can represent explicit decision processes via the root-to-leaf path and reflect each agent's contribution to the team. Specifically, we construct a novel soft decision tree using a recurrent structure and demonstrate which features influence the decision-making process. Then, based on the value decomposition framework, we linearly assign credit to each agent by explicitly mixing individual action values to estimate the joint action value using only local observations, providing new insights into interpreting the cooperation mechanism. Theoretical analysis confirms that MIXRTs guarantee additivity and monotonicity in the factorization of joint action values. Evaluations on complex tasks like Spread and StarCraft II demonstrate that MIXRTs compete with existing methods while providing clear explanations, paving the way for interpretable and high-performing MARL systems. Zichuan Liu, Yuanyang Zhu, Zhi Wang 0001, Yang Gao 0001, Chunlin Chen 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Multi-View Clustering With Incremental Instances and ViewsabstractMulti-view clustering (MVC) has attracted increasing attention with the emergence of various data collected from multiple sources. In real-world dynamic environment, instances are continually gathered, and the number of views expands as new data sources become available. Learning for such simultaneous increment of instances and views, particularly in unsupervised scenarios, is crucial yet underexplored. In this paper, we address this problem by proposing a novel MVC method with Incremental Instances and Views, MVC-IIV for short. MVC-IIV contains two stages, an initial stage and an incremental stage. In the initial stage, a basic latent multi-view subspace clustering model is constructed to handle existing data, which can be viewed as traditional static MVC. In the incremental stage, the previously trained model is reused to guide learning for newly arriving instances with new views, transferring historical knowledge while avoiding redundant computations. In specific, we design and reuse two modules, i.e., multi-view embedding module for low-dimensional representation learning, and consensus centroids module for cluster probability learning. By adding consistency regularization on the two modules, the knowledge acquired from previous data is used, which not only enhances the exploration within current data batch, but also extracts the between-batch data correlations. The proposed model can be efficiently solved with linear space and time complexity. Extensive experiments demonstrate the effectiveness and efficiency of our method compared with the state-of-the-art approaches. Chao Zhang 0078, Zhi Wang 0001, Xiuyi Jia, Zechao Li, Chunlin Chen 0001, Huaxiong Li |
IEEE Trans. Image Process. | 2 |
| 2025 | Data-Aware Gradient Compression for FL in Communication-Constrained Mobile ComputingabstractFederated Learning (FL) in mobile environments faces significant communication bottlenecks. Gradient compression has proven as an effective solution to this issue, offering substantial benefits in environments with limited bandwidth and metered data. Yet, it encounters severe performance drops in non-IID environments due to a one-size-fits-all compression approach, which does not account for the varying data volumes across workers. Assigning varying compression ratios to workers with distinct data distributions and volumes is therefore a promising solution. This work derives the convergence rate of distributed SGD with non-uniform compression, which reveals the intricate relationship between model convergence and the compression ratios applied to individual workers. Accordingly, we frame the relative compression ratio assignment as an$n$-variable chi-squared nonlinear optimization problem, constrained by a limited communication budget. We propose DAGC-R, which assigns conservative compression to workers handling larger data volumes. Recognizing the computational limitations of mobile devices, we propose the DAGC-A, which is computationally less demanding and enhances the robustness of compression in non-IID scenarios. Our experiments confirm that the DAGC-R and DAGC-A can speed up the training speed by up to 25.43% and 16.65% compared to the uniform compression respectively, when dealing with highly imbalanced data volume distribution and restricted communication. Rongwei Lu, Yinan Mao, Bin Chen 0011, Laizhong Cui, Zhi Wang 0001 |
IEEE Trans. Mob. Comput. | 7 |
| 2025 | HisynSeg: Weakly-Supervised Histopathological Image Segmentation via Image-Mixing Synthesis and Consistency RegularizationabstractTissue semantic segmentation is one of the key tasks in computational pathology. To avoid the expensive and laborious acquisition of pixel-level annotations, a wide range of studies attempt to adopt the class activation map (CAM), a weakly-supervised learning scheme, to achieve pixel-level tissue segmentation. However, CAM-based methods are prone to suffer from under-activation and over-activation issues, leading to poor segmentation performance. To address this problem, we propose a novel weakly-supervised semantic segmentation framework for histopathological images based on image-mixing synthesis and consistency regularization, dubbed HisynSeg. Specifically, synthesized histopathological images with pixel-level masks are generated for fully-supervised model training, where two synthesis strategies are proposed based on Mosaic transformation and Bézier mask generation. Besides, an image filtering module is developed to guarantee the authenticity of the synthesized images. In order to further avoid the model overfitting to the occasional synthesis artifacts, we additionally propose a novel self-supervised consistency regularization, which enables the real images without segmentation masks to supervise the training of the segmentation model. By integrating the proposed techniques, the HisynSeg framework successfully transforms the weakly-supervised semantic segmentation problem into a fully-supervised one, greatly improving the segmentation accuracy. Experimental results on three datasets prove that the proposed method achieves a state-of-the-art performance. Code is available at https://github.com/Vison307/HisynSeg. Zijie Fang, Yifeng Wang 0001, Peizhang Xie, Zhi Wang 0001, Yongbing Zhang 0002 |
IEEE Trans. Medical Imaging | 4 |
| 2025 | An Efficient Implicit Neural Representation Image Codec Based on Mixed Autoregressive Model for Low-Complexity DecodingabstractDisplaying high-quality images on edge devices, such as augmented reality devices, is essential for enhancing the user experience. However, these devices often face power consumption and computing resource limitations, making it challenging to apply many deep learning-based image compression algorithms in this field. Implicit Neural Representation (INR) for image compression is an emerging technology that offers two key benefits compared to cutting-edge autoencoder models: low computational complexity and parameter-free decoding. It also outperforms many traditional and early neural compression methods in terms of quality. In this study, we introduce a new Mixed AutoRegressive Model (MARM) to significantly reduce the decoding time for the current INR codec, along with a new synthesis network to enhance reconstruction quality. MARM includes our proposed AutoRegressive Upsampler (ARU) blocks, which are highly computationally efficient, and ARM from previous work to balance decoding time and reconstruction quality. We also propose enhancing ARU's performance using a checkerboard two-stage decoding strategy. Moreover, the ratio of different modules can be adjusted to maintain a balance between quality and speed. Comprehensive experiments demonstrate that our method significantly improves computational efficiency while preserving image quality. With different parameter settings, our method can achieve over a magnitude acceleration in decoding time without industrial level optimization or achieve state-of-the-art reconstruction quality compared with other INR codecs. To the best of our knowledge, our method is the first INR-based codec comparable with Ballé et al. [1] in both decoding speed and quality while maintaining low complexity. Jiahong Chen, Bin Chen 0011, Zimo Liu, Baoyi An 0002, Shutao Xia, Zhi Wang 0001 |
IEEE Trans. Multim. | 7 |
| 2025 | SkyML: A MLaaS Federation Design for Multicloud-Based Multimedia AnalyticsabstractThe advent of deep learning has precipitated a surge in public machine learning as a service (MLaaS) for multimedia analysis. However, reliance on a single MLaaS can result in product dependency and a loss of better performance offered by multiple MLaaSes. Consequently, many enterprises opt for an intercloud broker capable of managing jobs across various clouds. Though existing works explore the efficient utilization of inter-cloud computational resources and the enhancement of inter-cloud data transfer throughput, they disregard improving the overall accuracy of multiple MLaaSes. In response, we conduct a measurement study on object detection services, which are designed to identify and locate various objects within an image. We discover that combining predictions from multiple MLaaSes can improve analytical performance. However, more MLaaSes do not necessarily equate to better performance. Therefore, we propose SkyML, a user-side MLaaS federation broker that selects a subset of MLaaSes based on the characteristics of the request to achieve optimal multimedia analytical performance. Initially, we design a combinatorial reinforcement learning approach to select the sound MLaaS combination, thereby maximizing user experience. We also present an ingenious, automated taxonomy unification algorithm to minimize human efforts in merging MLaaS-specific labels into a user-preferred label space. Moreover, we devise an optimized ensemble strategy to aggregate predictions from the selected MLaaSes. Evaluations indicate that our similarity-based taxonomy unification approach can reduce annotation costs by 90%. Moreover, real-world trace-driven evaluations further prove that our MLaaS selection method can achieve similar levels of accuracy with a 67% reduction in inference fees. Shuzhao Xie, Yuan Xue 0013, Yifei Zhu 0001, Zhi Wang 0001 |
IEEE Trans. Multim. | 4 |
| 2025 | Discretizing Continuous Action Space With Unimodal Probability Distributions for On-Policy Reinforcement LearningabstractFor on-policy reinforcement learning (RL), discretizing action space for continuous control can easily express multiple modes and is straightforward to optimize. However, without considering the inherent ordering between the discrete atomic actions, the explosion in the number of discrete actions can possess undesired properties and induce a higher variance for the policy gradient (PG) estimator. In this article, we introduce a straightforward architecture that addresses this issue by constraining the discrete policy to be unimodal using Poisson probability distributions. This unimodal architecture can better leverage the continuity in the underlying continuous action space using explicit unimodal probability distributions. We conduct extensive experiments to show that the discrete policy with the unimodal probability distribution provides significantly faster convergence and higher performance for on-policy RL algorithms in challenging control tasks, especially in highly complex tasks such as Humanoid. We provide theoretical analysis on the variance of the PG estimator, which suggests that our attentively designed unimodal discrete policy can retain a lower variance and yield a stable learning process. Yuanyang Zhu, Zhi Wang 0001, Yuanheng Zhu, Chunlin Chen 0001, Dongbin Zhao |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Understanding 5G Performance for Real-World Services: A Content Provider's PerspectiveabstractRecent years have witnessed a rapid growth of both 5G coverage and 5G users, attracting several measurement studies on its coverage, reliability and quality of service. However, the capabilities and potential impacts of 5G, especially Standalone (SA) 5G, still remain to be fully understood from a content provider (CP)’s perspective. This paper fills this gap by studying 5G networks used by over 23 million users in one year in Kuaishou, a popular crowdsourced live streaming platform. With passive and active measurements, we have the following key findings: i) SA 5G generally provides end-to-end performance improvements compared to 4G or Non-Standalone (NSA) 5G, but its advantage depends on both the number of cellular users and CP-level configurations. ii) In the radio access networks, SA 5G is more sensitive to access density, but has better handover tolerance. iii) Controlled experiments with 29 mobile device models on energy consumption refute some “conventional wisdom,” including the notion that 5G always consumes more power. iv) Traceroute-based active experiments in over 300 cities show that although users are “closer” to the internet in SA 5G due to the control and user plane separation, whether end-to-end latency benefits from that partly depends on the routing policy at the gateways. Furthermore, we propose a 5G-aware rebuffer strategy tested by 9 million viewers in Kuaishou, showing a 7% reduction in rebuffer proportion. Finally, we also provide new design space for other 5G participants. Xinjie Yuan, Mingzhou Wu, Zhi Wang 0001, Yifei Zhu 0001, Junjian Guo, Zhi-Li Zhang, Wenwu Zhu 0001 |
IEEE Trans. Netw. | 3 |
| 2025 | Role Assignment for Agent Evaluation Under Uncertainty: A Distributionally Robust ApproachabstractRole-based collaboration (RBC) is an emerging and advanced methodology for problem-solving. A critical aspect of RBC theory is agent evaluation, which aims to assess agents’ abilities through a qualification value derived from a comprehensive analysis of their characteristics. This evaluation directly impacts the quality of role assignments. Existing research typically assumes that the qualification value is either predetermined, based on multiscale criteria, or following a predefined distribution. These assumptions, however, are overly idealistic and difficult to generalize, failing to capture the inherent volatility of the qualification value. To address this challenge, this article introduces a Wasserstein-based ambiguity set to model potential fluctuations in the qualification value, drawing on empirical distributions derived from historical sample data. Building upon the RBC framework and its abstract model environments, classes, agents, roles, groups, and objects (E-CARGO), we propose two data-driven models: distributionally robust group role assignment (DRGRA) and group multirole assignment (DRGMRA). These models aim to achieve more robust and optimal role assignments under uncertainty in agent evaluation. Leveraging strong duality, we reformulate DRGRA and DRGMRA as tractable finite mixed 0–1 convex problems, providing an approximation framework that reduces computational complexity. Notably, these models are adaptable to other problems with no uncertainty in agent evaluation, highlighting their modeling scalability. Experimental results demonstrate the effectiveness and robustness of the proposed models. Zhihang Yu, Bo Wang 0027, Libo Zhang 0006, Zhi Wang 0001, Haibin Zhu 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2025 | TextIR: A Simple Framework for Text-Based Editable Image RestorationabstractMany current image restoration approaches utilize neural networks to acquire robust image-level priors from extensive datasets, aiming to reconstruct missing details. Nevertheless, these methods often falter with images that exhibit significant information gaps. While incorporating external priors or leveraging reference images can provide supplemental information, these strategies are limited in their practical scope. Alternatively, textual inputs offer greater accessibility and adaptability. In this study, we develop a sophisticated framework enabling users to guide the restoration of deteriorated images via textual descriptions. Utilizing the text-image compatibility feature of CLIP enhances the integration of textual and visual data. Our versatile framework supports multiple restoration activities such as image inpainting, super-resolution, and colorization. Comprehensive testing validates our technique's efficacy. Yunpeng Bai, Cairong Wang, Shuzhao Xie, Chao Dong 0005, Chun Yuan 0003, Zhi Wang 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2024 | OVD-Explorer: Optimism Should Not Be the Sole Pursuit of Exploration in Noisy EnvironmentsabstractIn reinforcement learning, the optimism in the face of uncertainty (OFU) is a mainstream principle for directing exploration towards less explored areas, characterized by higher uncertainty. However, in the presence of environmental stochasticity (noise), purely optimistic exploration may lead to excessive probing of high-noise areas, consequently impeding exploration efficiency. Hence, in exploring noisy environments, while optimism-driven exploration serves as a foundation, prudent attention to alleviating unnecessary over-exploration in high-noise areas becomes beneficial. In this work, we propose Optimistic Value Distribution Explorer (OVD-Explorer) to achieve a noise-aware optimistic exploration for continuous control. OVD-Explorer proposes a new measurement of the policy's exploration ability considering noise in optimistic perspectives, and leverages gradient ascent to drive exploration. Practically, OVD-Explorer can be easily integrated with continuous control RL algorithms. Extensive evaluations on the MuJoCo and GridChaos tasks demonstrate the superiority of OVD-Explorer in achieving noise-aware optimistic exploration. Jinyi Liu 0002, Zhi Wang 0001, Yan Zheng 0002, Jianye Hao, Chenjia Bai, Junjie Ye 0002, Zhen Wang 0004, Haiyin Piao |
AAAI | 2 |
| 2024 | Procedural Level Generation with Diffusion Models from a Single ExampleabstractLevel generation is a central focus of Procedural Content Generation (PCG), yet deep learning-based approaches are limited by scarce training data, i.e., human-designed levels. Despite being a dominant framework, Generative Adversarial Networks (GANs) exhibit a substantial quality gap between generated and human-authored levels, alongside rising training costs, particularly with increasing token complexity. In this paper, we introduce a diffusion-based generative model that learns from just one example. Our approach involves two core components: 1) an efficient yet expressive level representation, and 2) a latent denoising network with constrained receptive fields. To start with, our method utilizes token semantic labels, similar to word embeddings, to provide dense representations. This strategy not only surpasses one-hot encoding in representing larger game levels but also improves stability and accelerates convergence in latent diffusion. In addition, we adapt the denoising network architecture to confine the receptive field to localized patches of the data, aiming to facilitate single-example learning. Extensive experiments demonstrate that our model is capable of generating stylistically congruent samples of arbitrary sizes compared to manually designed levels. It suits a wide range of level structures with fewer artifacts than GAN-based approaches. The source code is available at https://github.com/shiqi-dai/diffusioncraft. Shiqi Dai, Naiqi Li, Tao Dai 0001, Zhi Wang 0001 |
AAAI | 5 |
| 2024 | Towards Compact 3D Representations via Point Feature Enhancement Masked AutoencodersabstractLearning 3D representation plays a critical role in masked autoencoder (MAE) based pre-training methods for point cloud, including single-modal and cross-modal based MAE. Specifically, although cross-modal MAE methods learn strong 3D representations via the auxiliary of other modal knowledge, they often suffer from heavy computational burdens and heavily rely on massive cross-modal data pairs that are often unavailable, which hinders their applications in practice. Instead, single-modal methods with solely point clouds as input are preferred in real applications due to their simplicity and efficiency. However, such methods easily suffer from limited 3D representations with global random mask input. To learn compact 3D representations, we propose a simple yet effective Point Feature Enhancement Masked Autoencoders (Point-FEMAE), which mainly consists of a global branch and a local branch to capture latent semantic features. Specifically, to learn more compact features, a share-parameter Transformer encoder is introduced to extract point features from the global and local unmasked patches obtained by global random and local block mask strategies, followed by a specific decoder to reconstruct. Meanwhile, to further enhance features in the local branch, we propose a Local Enhancement Module with local patch convolution to perceive fine-grained local context at larger scales. Our method significantly improves the pre-training efficiency compared to cross-modal alternatives, and extensive downstream experiments underscore the state-of-the-art effectiveness, particularly outperforming our baseline (Point-MAE) by 5.16%, 5.00%, and 5.04% in three variants of ScanObjectNN, respectively. Code is available at https://github.com/zyh16143998882/AAAI24-PointFEMAE. Yaohua Zha, Huizhen Ji, Jinmin Li, Rongsheng Li, Tao Dai 0001, Bin Chen 0011, Zhi Wang 0001, Shutao Xia |
AAAI | 7 |
| 2024 | Vision-Language Pre-training with Object Contrastive Learning for 3D Scene UnderstandingabstractIn recent years, vision language pre-training frameworks have made significant progress in natural language processing and computer vision, achieving remarkable performance improvement on various downstream tasks. However, when extended to point cloud data, existing works mainly focus on building task-specific models, and fail to extract universal 3D vision-language embedding that generalize well. We carefully investigate three common tasks in semantic 3D scene understanding, and derive key insights into the development of a pre-training model. Motivated by these observations, we propose a vision-language pre-training framework 3DVLP (3D vision-language pre-training with object contrastive learning), which transfers flexibly on 3D vision-language downstream tasks. 3DVLP takes visual grounding as the proxy task and introduces Object-level IoU-guided Detection (OID) loss to obtain high-quality proposals in the scene. Moreover, we design Object-level Cross-Contrastive alignment (OCC) task and Object-level Self-Contrastive learning (OSC) task to align the objects with descriptions and distinguish different objects in the scene, respectively. Extensive experiments verify the excellent performance of 3DVLP on three 3D vision-language tasks, reflecting its superiority in semantic 3D scene understanding. Code is available at https://github.com/iridescentttt/3DVLP. Taolin Zhang 0003, Sunan He, Tao Dai 0001, Zhi Wang 0001, Bin Chen 0011, Shutao Xia |
AAAI | 4 |
| 2024 | MamMIL: Multiple Instance Learning for Whole Slide Images with State Space ModelsabstractRecently, pathological diagnosis has achieved superior performance by combining deep learning models with the multiple instance learning (MIL) framework using whole slide images (WSIs). However, the giga-pixeled nature of WSIs poses a great challenge for efficient MIL. Existing studies either do not consider global dependencies among instances, or use approximations such as linear attentions to model the pair-to-pair instance interactions, which inevitably brings performance bottlenecks. To tackle this challenge, we propose a framework named MamMIL for WSI analysis by cooperating the selective structured state space model (i.e., Mamba) with MIL, enabling the modeling of global instance dependencies while maintaining linear complexity. Specifically, considering the irregularity of the tissue regions in WSIs, we represent each WSI as an undirected graph. To address the problem that Mamba can only process 1D sequences, we further propose a topology-aware scanning mechanism to serialize the WSI graphs while preserving the topological relationships among the instances. Finally, in order to further perceive the topological structures among the instances and incorporate short-range feature interactions, we propose an instance aggregation block based on graph neural networks. Experiments show that MamMIL can achieve advanced performance than the state-of-the-art frameworks. The code can be accessed at https://github.com/Vison307/MamMIL. Zijie Fang, Yifeng Wang 0001, Ye Zhang 0043, Zhi Wang 0001, Jian Zhang 0018, Xiangyang Ji, Yongbing Zhang 0002 |
BIBM | 4 |
| 2024 | Retraining-free Model Quantization via One-Shot Weight-Coupling LearningabstractQuantization is of significance for compressing the over-parameterized deep neural models and deploying them on resource-limited devices. Fixed-precision quantization suf-fers from performance drop due to the limited numerical representation ability. Conversely, mixed-precision quan-tization (MPQ) is advocated to compress the model ef-fectively by allocating heterogeneous bit-width for layers. MPQ is typically organized into a searching-retraining two-stage process. Previous works only focus on determining the optimal bit-width configuration in the first stage effi-ciently, while ignoring the considerable time costs in the second stage and thus hindering deployment efficiency sig-nificantly. In this paper, we devise a one-shot training-searching paradigm for mixed-precision model compression. Specifically, in the first stage, all potential bit-width configurations are coupled and thus optimized simultane-ously within a set of shared weights. However, our ob-servations reveal a previously unseen and severe bit-width interference phenomenon among highly coupled weights during optimization, leading to considerable performance degradation under a high compression ratio. To tackle this problem, we first design a bit-width scheduler to dy-namically freeze the most turbulent bit-width of layers during training, to ensure the rest bit-widths converged prop-erly. Then, taking inspiration from information theory, we present an information distortion mitigation technique to align the behaviour of the bad-performing bit-widths to the well-performing ones. In the second stage, an inference-only greedy search scheme is devised to evaluate the good-ness of configurations without introducing any additional training costs. Extensive experiments on three representative models and three datasets demonstrate the effective-ness of the proposed method. Code can be available on https://github.com/1hunters/retraining-free-quantization. Shuzhao Xie, Rongwei Lu, Xinzhu Ma, Zhi Wang 0001, Wenwu Zhu 0001 |
CVPR | 7 |
| 2024 | MesonGS: Post-training Compression of 3D Gaussians via Efficient Attribute Transformation
Shuzhao Xie, Weixiang Zhang, Yunpeng Bai, Rongwei Lu, Shijia Ge, Zhi Wang 0001 |
ECCV (33) | 7 |
| 2024 | A Joint Approach to Local Updating and Gradient Compression for Efficient Asynchronous Federated Learning
Jiajun Song, Jiajun Luo, Rongwei Lu, Shuzhao Xie, Bin Chen 0011, Zhi Wang 0001 |
Euro-Par (3) | 6 |
| 2024 | Attention-Guided Contrastive Role Representations for Multi-agent Reinforcement LearningabstractReal-world multi-agent tasks usually involve dynamic team composition with the emergence of roles, which should also be a key to efficient cooperation in multi-agent reinforcement learning (MARL). Drawing inspiration from the correlation between roles and agent's behavior patterns, we propose a novel framework of **A**ttention-guided **CO**ntrastive **R**ole representation learning for **M**ARL (**ACORM**) to promote behavior heterogeneity, knowledge transfer, and skillful coordination across agents. First, we introduce mutual information maximization to formalize role representation learning, derive a contrastive learning objective, and concisely approximate the distribution of negative pairs. Second, we leverage an attention mechanism to prompt the global state to attend to learned role representations in value decomposition, implicitly guiding agent coordination in a skillful role space to yield more expressive credit assignment. Experiments on challenging StarCraft II micromanagement and Google research football tasks demonstrate the state-of-the-art performance of our method and its advantages over existing approaches. Our code is available at [https://github.com/NJU-RL/ACORM](https://github.com/NJU-RL/ACORM). Zican Hu, Zongzhang Zhang, Huaxiong Li, Chunlin Chen 0001, Hongyu Ding, Zhi Wang 0001 |
ICLR | 6 |
| 2024 | Collaborative Edge Caching in LEO Satellites Networks: A MAPPO Based ApproachabstractLow Earth Orbit satellite networks, as a crucial component of global low-latency internet access, are expected to carry significant user traffic in the future. Caching frequently requested content, e.g., popular videos on short-video platforms, in satellite networks can significantly alleviate traffic congestion. However, the satellite’s brief overhead passing time, which is less than ten minutes, makes it difficult for satellites to capture the content popularity distribution. And the changing relative position between satellites poses challenges for cooperation. To address the challenges, we propose a method called SEC_MAPPO for deploying cooperative edge caching in satellite networks. First, we model this novel scenario and transform it into a Partially Observable Markov Decision Process (POMDP). Then, we design a multi-agent reinforcement learning algorithm specifically tailored for this scenario. Trace-driven simulation using a real-world LEO satellite constellation and video request dataset demonstrated that our proposed algorithm could achieve a reduction in average video request latency ranging from 4.53% to 9.31% compared to the baseline solutions. Mingzhou Wu, Shiqi Dai, Han Hu 0003, Zhi Wang 0001 |
ICME | 4 |
| 2024 | Invertible Residual Rescaling Models
Jinmin Li, Tao Dai 0001, Yaohua Zha, Yilu Luo, Longfei Lu, Bin Chen 0011, Zhi Wang 0001, Shutao Xia |
IJCAI | 7 |
| 2024 | Meta-DT: Offline Meta-RL as Conditional Sequence Modeling with World Model DisentanglementabstractA longstanding goal of artificial general intelligence is highly capable generalists that can learn from diverse experiences and generalize to unseen tasks. The language and vision communities have seen remarkable progress toward this trend by scaling up transformer-based models trained on massive datasets, while reinforcement learning (RL) agents still suffer from poor generalization capacity under such paradigms. To tackle this challenge, we propose Meta Decision Transformer (Meta-DT), which leverages the sequential modeling ability of the transformer architecture and robust task representation learning via world model disentanglement to achieve efficient generalization in offline meta-RL. We pretrain a context-aware world model to learn a compact task representation, and inject it as a contextual condition to the causal transformer to guide task-oriented sequence generation. Then, we subtly utilize history trajectories generated by the meta-policy as a self-guided prompt to exploit the architectural inductive bias. We select the trajectory segment that yields the largest prediction error on the pretrained world model to construct the prompt, aiming to encode task-specific information complementary to the world model maximally. Notably, the proposed framework eliminates the requirement of any expert demonstration or domain knowledge at test time. Experimental results on MuJoCo and Meta-World benchmarks across various dataset types show that Meta-DT exhibits superior few and zero-shot generalization capacity compared to strong baselines while being more practical with fewer prerequisites. Our code is available at https://github.com/NJU-RL/Meta-DT. Zhi Wang 0001, Yuanheng Zhu, Dongbin Zhao, Chunlin Chen 0001 |
NeurIPS | 1 |
| 2024 | LCM: Locally Constrained Compact Point Cloud Model for Masked Point ModelingabstractThe pre-trained point cloud model based on Masked Point Modeling (MPM) has exhibited substantial improvements across various tasks. However, these models heavily rely on the Transformer, leading to quadratic complexity and limited decoder, hindering their practice application. To address this limitation, we first conduct a comprehensive analysis of existing Transformer-based MPM, emphasizing the idea that redundancy reduction is crucial for point cloud analysis. To this end, we propose a Locally constrained Compact point cloud Model (LCM) consisting of a locally constrained compact encoder and a locally constrained Mamba-based decoder. Our encoder replaces self-attention with our local aggregation layers to achieve an elegant balance between performance and efficiency. Considering the varying information density between masked and unmasked patches in the decoder inputs of MPM, we introduce a locally constrained Mamba-based decoder. This decoder ensures linear complexity while maximizing the perception of point cloud geometry information from unmasked patches with higher information density. Extensive experimental results show that our compact model significantly surpasses existing Transformer-based models in both performance and efficiency, especially our LCM-based Point-MAE model, compared to the Transformer-based model, achieved an improvement of 1.84%, 0.67%, and 0.60% in performance on the three variants of ScanObjectNN while reducing parameters by 88% and computation by 73%. The code is available at https://github.com/zyh16143998882/LCM. Yaohua Zha, Naiqi Li, Yanzi Wang, Tao Dai 0001, Hang Guo 0002, Bin Chen 0011, Zhi Wang 0001, Zhihao Ouyang, Shutao Xia |
NeurIPS | 7 |
| 2024 | NetLLM: Adapting Large Language Models for NetworkingabstractMany networking tasks now employ deep learning (DL) to solve complex prediction and optimization problems. However, current design philosophy of DL-based algorithms entails intensive engineering overhead due to the manual design of deep neural networks (DNNs) for different networking tasks. Besides, DNNs tend to achieve poor generalization performance on unseen data distributions/environments. Duo Wu, Xianda Wang, Yaqi Qiao, Zhi Wang 0001, Junchen Jiang, Shuguang Cui, Fangxin Wang 0001 |
SIGCOMM | 4 |
| 2024 | Pyramid hybrid pooling quantization for efficient fine-grained image retrieval
Ziyun Zeng, Jinpeng Wang 0002, Bin Chen 0011, Tao Dai 0001, Shutao Xia, Zhi Wang 0001 |
Pattern Recognit. Lett. | 6 |
| 2024 | Enhancing Unsupervised Semantic Segmentation Through Context-Aware ClusteringabstractDespite the great progress of semantic segmentation with supervised learning, annotating large amounts of pixel-wise labels is, however, very expensive and time-consuming. To this end, Unsupervised Semantic Segmentation(USS) has been proposed to learn semantic segmentation, without any form of annotations. This approach involves dense prediction of semantics which is however challenging due to the unreliable nature of local representations. To solve this problem, we propose a newly context-aware unsupervised semantic segmentation framework, which aims to enhance the unsupervised semantic segmentation by leveraging contextual knowledge within and across images. In particular, we introduce a training strategy based on our Pyramid Semantic Guidance (PSG), which utilizes holistic semantics on pyramid views to guide pixel clustering with a siamese network-based framework. Additionally, we introduce a Context-Aware Embedding (CAE) module to fuse global features with low-level geometrical and appearance representations. We evaluate our method on the COCO-Stuff dataset and achieved competitive results compared to both the convolutional and ViT-based USS methods. Specifically, we attain significant improvements of +4.5% and +5% mIoU for Stuff and all class segmentation respectively, compared to previous approaches that employ unsupervised convolutional backbones. Yuan Wang 0083, Junliang Chen 0002, Songhe Deng, Zhi Wang 0001, LinLin Shen, Wenwu Zhu 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | Efficient Bayesian Policy Reuse With a Scalable Observation Model in Deep Reinforcement LearningabstractBayesian policy reuse (BPR) is a general policy transfer framework for selecting a source policy from an offline library by inferring the task belief based on some observation signals and a trained observation model. In this article, we propose an improved BPR method to achieve more efficient policy transfer in deep reinforcement learning (DRL). First, most BPR algorithms use the episodic return as the observation signal that contains limited information and cannot be obtained until the end of an episode. Instead, we employ the state transition sample, which is informative and instantaneous, as the observation signal for faster and more accurate task inference. Second, BPR algorithms usually require numerous samples to estimate the probability distribution of the tabular-based observation model, which may be expensive and even infeasible to learn and maintain, especially when using the state transition sample as the signal. Hence, we propose a scalable observation model based on fitting state transition functions of source tasks from only a small number of samples, which can generalize to any signals observed in the target task. Moreover, we extend the offline-mode BPR to the continual learning setting by expanding the scalable observation model in a plug-and-play fashion, which can avoid negative transfer when faced with new unknown tasks. Experimental results show that our method can consistently facilitate faster and more efficient policy transfer. Jinmei Liu, Zhi Wang 0001, Chunlin Chen 0001, Daoyi Dong |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Depthwise Convolution for Multi-Agent Communication With Enhanced Mean-Field ApproximationabstractMulti-Agent settings remain a fundamental challenge in the reinforcement learning (RL) domain due to the partial observability and the lack of accurate real-time interactions across agents. In this article, we propose a new method based on local communication learning to tackle the multi-agent RL (MARL) challenge within a large number of agents coexisting. First, we design a new communication protocol that exploits the ability of depthwise convolution to efficiently extract local relations and learn local communication between neighboring agents. To facilitate multi-agent coordination, we explicitly learn the effect of joint actions by taking the policies of neighboring agents as inputs. Second, we introduce the mean-field approximation into our method to reduce the scale of agent interactions. To more effectively coordinate behaviors of neighboring agents, we enhance the mean-field approximation by a supervised policy rectification network (PRN) for rectifying real-time agent interactions and by a learnable compensation term for correcting the approximation bias. The proposed method enables efficient coordination as well as outperforms several baseline approaches on the adaptive traffic signal control (ATSC) task and the StarCraft II multi-agent challenge (SMAC). Donghan Xie, Zhi Wang 0001, Chunlin Chen 0001, Daoyi Dong |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Weakly-Supervised Semantic Segmentation for Histopathology Images Based on Dataset Synthesis and Feature Consistency ConstraintabstractTissue segmentation is a critical task in computational pathology due to its desirable ability to indicate the prognosis of cancer patients. Currently, numerous studies attempt to use image-level labels to achieve pixel-level segmentation to reduce the need for fine annotations. However, most of these methods are based on class activation map, which suffers from inaccurate segmentation boundaries. To address this problem, we propose a novel weakly-supervised tissue segmentation framework named PistoSeg, which is implemented under a fully-supervised manner by transferring tissue category labels to pixel-level masks. Firstly, a dataset synthesis method is proposed based on Mosaic transformation to generate synthesized images with pixel-level masks. Next, considering the difference between synthesized and real images, this paper devises an attention-based feature consistency, which directs the training process of a proposed pseudo-mask refining module. Finally, the refined pseudo-masks are used to train a precise segmentation model for testing. Experiments based on WSSS4LUAD and BCSS-WSSS validate that PistoSeg outperforms the state-of-the-art methods. The code is released at https://github.com/Vison307/PistoSeg. Zijie Fang, Yang Chen 0036, Yifeng Wang 0001, Zhi Wang 0001, Xiangyang Ji, Yongbing Zhang 0002 |
AAAI | 4 |
| 2023 | Curriculum Multi-Negative Augmentation for Debiased Video GroundingabstractVideo Grounding (VG) aims to locate the desired segment from a video given a sentence query. Recent studies have found that current VG models are prone to over-rely the groundtruth moment annotation distribution biases in the training set. To discourage the standard VG model's behavior of exploiting such temporal annotation biases and improve the model generalization ability, we propose multiple negative augmentations in a hierarchical way, including cross-video augmentations from clip-/video-level, and self-shuffled augmentations with masks. These augmentations can effectively diversify the data distribution so that the model can make more reasonable predictions instead of merely fitting the temporal biases. However, directly adopting such data augmentation strategy may inevitably carry some noise shown in our cases, since not all of the handcrafted augmentations are semantically irrelevant to the groundtruth video. To further denoise and improve the grounding accuracy, we design a multi-stage curriculum strategy to adaptively train the standard VG model from easy to hard negative augmentations. Experiments on newly collected Charades-CD and ActivityNet-CD datasets demonstrate our proposed strategy can improve the performance of the base model on both i.i.d and o.o.d scenarios. Xiaohan Lan, Yitian Yuan, Hong Chen 0011, Xin Wang 0019, Zequn Jie, Lin Ma 0002, Zhi Wang 0001, Wenwu Zhu 0001 |
AAAI | 7 |
| 2023 | FSR: A General Frequency-Oriented Framework to Accelerate Image Super-resolution NetworksabstractDeep neural networks (DNNs) have witnessed remarkable achievement in image super-resolution (SR), and plenty of DNN-based SR models with elaborated network designs have recently been proposed. However, existing methods usually require substantial computations by operating in spatial domain. To address this issue, we propose a general frequency-oriented framework (FSR) to accelerate SR networks by considering data characteristics in frequency domain. Our FSR mainly contains dual feature aggregation module (DFAM) to extract informative features in both spatial and transform domains, followed by a four-path SR-Module with different capacities to super-resolve in the frequency domain. Specifically, DFAM further consists of a transform attention block (TABlock) and a spatial context block (SCBlock) to extract global spectral information and local spatial information, respectively, while SR-Module is a parallel network container that contains four to-be-accelerated branches. Furthermore, we propose an adaptive weight strategy for a trade-off between image details recovery and visual quality. Extensive experiments show that our FSR can save FLOPs by almost 40% while reducing inference time by 50% for other SR methods (e.g., FSRCNN, CARN, SRResNet and RCAN). Code is available at https://github.com/THU-Kingmin/FSR. Jinmin Li, Tao Dai 0001, Mingyan Zhu 0001, Bin Chen 0011, Zhi Wang 0001, Shutao Xia |
AAAI | 5 |
| 2023 | BiERL: A Meta Evolutionary Reinforcement Learning Framework via Bilevel OptimizationabstractEvolutionary reinforcement learning (ERL) algorithms recently raise attention in tackling complex reinforcement learning (RL) problems due to high parallelism, while they are prone to insufficient exploration or model collapse without carefully tuning hyperparameters (aka meta-parameters). In the paper, we propose a general meta ERL framework via bilevel optimization (BiERL) to jointly update hyperparameters in parallel to training the ERL model within a single agent, which relieves the need for prior domain knowledge or costly optimization procedure before model deployment. We design an elegant meta-level architecture that embeds the inner-level’s evolving experience into an informative population representation and introduce a simple and feasible evaluation of the meta-level fitness function to facilitate learning efficiency. We perform extensive experiments in MuJoCo and Box2D tasks to verify that as a general framework, BiERL outperforms various baselines and consistently improves the learning performance for a diversity of ERL algorithms. Yuanyang Zhu, Zhi Wang 0001, Yan Zheng 0002, Jianye Hao, Chunlin Chen 0001 |
ECAI | 3 |
| 2023 | GIFD: A Generative Gradient Inversion Method with Feature Domain OptimizationabstractFederated Learning (FL) has recently emerged as a promising distributed machine learning framework to preserve clients' privacy, by allowing multiple clients to upload the gradients calculated from their local data to a central server. Recent studies find that the exchanged gradients also take the risk of privacy leakage, e.g., an attacker can invert the shared gradients and recover sensitive data against an FL system by leveraging pre-trained generative adversarial networks (GAN) as prior knowledge. However, performing gradient inversion attacks in the latent space of the GAN model limits their expression ability and generalizability. To tackle these challenges, we propose Gradient Inversion over Feature Domains (GIFD), which disassembles the GAN model and searches the feature domains of the intermediate layers. Instead of optimizing only over the initial latent code, we progressively change the optimized layer, from the initial latent space to intermediate layers closer to the output images. In addition, we design a regularizer to avoid unreal image generation by adding a small l1ball constraint to the searching range. We also extend GIFD to the out-of-distribution (OOD) setting, which weakens the assumption that the training sets of GANs and FL tasks obey the same data distribution. Extensive experiments demonstrate that our method can achieve pixel-level reconstruction and is superior to the existing methods. Notably, GIFD also shows great generalizability under different defense strategy settings and batch sizes. Hao Fang 0011, Bin Chen 0011, Xuan Wang 0002, Zhi Wang 0001, Shutao Xia |
ICCV | 4 |
| 2023 | Instance-aware Dynamic Prompt Tuning for Pre-trained Point Cloud ModelsabstractPre-trained point cloud models have found extensive applications in 3D understanding tasks like object classification and part segmentation. However, the prevailing strategy of full fine-tuning in downstream tasks leads to large per-task storage overhead for model parameters, which limits the efficiency when applying large-scale pre-trained models. Inspired by the recent success of visual prompt tuning (VPT), this paper attempts to explore prompt tuning on pre-trained point cloud models, to pursue an elegant balance between performance and parameter efficiency. We find while instance-agnostic static prompting, e.g. VPT, shows some efficacy in downstream transfer, it is vulnerable to the distribution diversity caused by various types of noises in real-world point cloud data. To conquer this limitation, we propose a novel Instance-aware Dynamic Prompt Tuning (IDPT) strategy for pre-trained point cloud models. The essence of IDPT is to develop a dynamic prompt generation module to perceive semantic prior features of each point cloud instance and generate adaptive prompt tokens to enhance the model's robustness. Notably, extensive experiments demonstrate that IDPT outperforms full finetuning in most tasks with a mere 7% of the trainable parameters, providing a promising solution to parameter-efficient learning for pre-trained point cloud models. Code is available at https://github.com/zyh16143998882/ICCV23-IDPT. Yaohua Zha, Jinpeng Wang 0002, Tao Dai 0001, Bin Chen 0011, Zhi Wang 0001, Shutao Xia |
ICCV | 5 |
| 2023 | Unsupervised Anomaly Detection with Local-Sensitive VQVAE and Global-Sensitive TransformersabstractUnsupervised anomaly detection (UAD) has been widely implemented in industrial and medical applications, which reduces the cost of manual annotation and improves efficiency in disease diagnosis. Recently, deep auto-encoder with its variants has demonstrated its advantages in many UAD scenarios. Training on the normal data, these models are expected to locate anomalies by producing higher reconstruction error for the abnormal areas than the normal ones. However, this assumption does not always hold because of the uncontrollable generalization capability. To solve this problem, we present LSGS, a method that builds on Vector Quantised-Variational Autoencoder (VQVAE) with a novel aggregated codebook and transformers with global attention. In this work, the VQVAE focus on feature extraction and reconstruction of images, and the transformers fit the manifold and locate anomalies in the latent space. Then, leveraging the generated encoding sequences that conform to a normal distribution, we can reconstruct a more accurate image for locating the anomalies. Experiments on various datasets demonstrate the effectiveness of the proposed method. Mingqing Wang, Jiawei Li 0006, Chengxiao Luo, Bin Chen 0011, Shutao Xia, Zhi Wang 0001 |
ICIP | 7 |
| 2023 | Multi-stream Adaptive Offloading of Joint Compressed Video Streams, Feature Streams, and Semantic Streams in Edge Computing SystemsabstractEdge computing (EC) is a promising paradigm for serving latency-sensitive video applications. However, massive compressed video transmission and analysis require considerable bandwidth and computing resources, posing enormous challenges for current multimedia frameworks. Novel multi-stream frameworks that incorporate feature streams are more practical. The reason is that feature streams containing compact video frame feature data have a lower bitrate and better serve machine vision tasks. Nevertheless, feature extraction by devices increases the latency and energy consumption of local computing. Therefore, how to offload suitable streams according to video task requirements and system resources is a challenging issue. This paper studies EC-based multi-stream adaptive offloading. We model the multi-stream offloading and computation problem to maximize system utility by jointly optimizing offloading decisions, computation resource allocation, and video frame sampling rates. Frame sampling rates, processing latency, and energy consumption are considered in system utility modeling. The formulated optimization problem is a mixed-integer programming (MIP) problem. We propose an efficient algorithm to address this MIP problem. The proposed algorithm relies on the Hungarian algorithm and improved greedy Markov approximation. The simulation results validate our proposed algorithm’s superior performance. Dieli Hu 0001, Wen Ji 0003, Zhi Wang 0001 |
ICME | 3 |
| 2023 | Collaborative Edge Caching: a Meta Reinforcement Learning Approach with Edge SamplingabstractCurrent learning-based edge caching schemes usually suffer from dynamic content popularity, e.g., in the emerging short video platforms, users’ request patterns shift significantly over time and across different edges. An intuitive solution for a specific local edge cache is to collect more request histories from other edge caches. However, uniformly merging these request histories may not perform satisfactorily due to heterogeneous content distributions on different edges. To solve this problem, we propose a collaborative edge caching framework. First, we design a meta-learning-based collaborative strategy to guarantee that the local model can timely meet the continually changing content popularity. Then, we design an edge sampling method to select more "valuable" neighbor edges to participate in the local training. To evaluate the proposed framework, we conduct trace-driven experiments to demonstrate the effectiveness of our design: it improves the average cache hit rate by up to 10.12% (normalized) compared with other baselines. Yinan Mao, Bowei He, Shiji Zhou, Chen Ma 0001, Zhi Wang 0001 |
ICME | 5 |
| 2023 | Dynamic Edge Caching via Online Meta-RLabstractThe content request patterns perceived by edge devices are becoming highly dynamic, especially for emerging short video platforms compared to traditional video platforms. This calls for caching policies that can continuously adapt to dynamic environments, challenging previously popular reinforcement learning (RL)-based policies. A straightforward solution, i.e., repeatedly restarting and training RL agents, would fail to converge timely while meeting the observed adaptation process. Offering transferable knowledge is considered a possible method to speed up the adaptation process. Unfortunately, it fails to outperform the RL-based approach as an alternative solution in these scenarios. To alleviate this drawback, we 1) design a sequential-pair meta-learning for edge caching that captures the meta-knowledge of dynamic changes from sequential-pair-wise intervals, which are segmentations from the whole dynamic episode, and 2) develop an online meta-RL-based solution called Online Meta Actor-Critic (OMAC), which updates the meta-knowledge in an online manner. To evaluate the proposed framework, we conduct trace-driven experiments to demonstrate the effectiveness of our design: it improves the average cache hit rate by up to 37.4% (normalized) compared with other baselines. Yinan Mao, Shiji Zhou, Zhi Wang 0001, Wenwu Zhu 0001 |
IJCNN | 4 |
| 2023 | DAGC: Data-Aware Adaptive Gradient CompressionabstractGradient compression algorithms are widely used to alleviate the communication bottleneck in distributed ML. However, existing gradient compression algorithms suffer from accuracy degradation in Non-IID scenarios, because a uniform compression scheme is used to compress gradients at workers with different data distributions and volumes, since workers with larger volumes of data are forced to adapt to the same aggressive compression ratios as others. Assigning different compression ratios to workers with different data distributions and volumes is thus a promising solution. In this study, we first derive a function from capturing the correlation between the number of training iterations for a model to converge to the same accuracy, and the compression ratios at different workers; This function particularly shows that workers with larger data volumes should be assigned with higher compression ratios1to guarantee better accuracy. Then, we formulate the assignment of compression ratios to the workers as an n-variables chi-square nonlinear optimization problem under fixed and limited total communication constrain. We propose an adaptive gradient compression strategy called DAGC, which assigns each worker a different compression ratio according to their data volumes. Our experiments confirm that DAGC can achieve better performance facing highly imbalanced data volume distribution and restricted communication. Rongwei Lu, Jiajun Song, Bin Chen 0011, Laizhong Cui, Zhi Wang 0001 |
INFOCOM | 5 |
| 2023 | One-stage Low-resolution Text Recognition with High-resolution Knowledge TransferabstractRecognizing characters from low-resolution (LR) text images poses a significant challenge due to the information deficiency as well as the noise and blur in low-quality images. Current solutions for low-resolution text recognition (LTR) typically rely on a two-stage pipeline that involves super-resolution as the first stage followed by the second-stage recognition. Although this pipeline is straightforward and intuitive, it has to use an additional super-resolution network, which causes inefficiencies during training and testing. Moreover, the recognition accuracy of the second stage heavily depends on the reconstruction quality of the first stage, causing ineffectiveness.In this work, we attempt to address these challenges from a novel perspective: adapting the recognizer to low-resolution inputs by transferring the knowledge from the high-resolution. Guided by this idea, we propose an efficient and effective knowledge distillation framework to achieve multi-level knowledge transfer.Specifically, the visual focus loss is proposed to extract the character position knowledge with resolution gap reduction and character region focus, the semantic contrastive loss is employed to exploit the contextual semantic knowledge with contrastive learning, and the soft logits loss facilitates both local word-level and global sequence-level learning from the soft teacher label.Extensive experiments show that the proposed one-stage pipeline significantly outperforms super-resolution based two-stage frameworks in terms of effectiveness and efficiency, accompanied by favorable robustness.Code is available at https://github.com/csguoh/KD-LTR. Hang Guo 0002, Tao Dai 0001, Mingyan Zhu 0001, Guanghao Meng, Bin Chen 0011, Zhi Wang 0001, Shutao Xia |
ACM Multimedia | 6 |
| 2023 | SEAM: Searching Transferable Mixed-Precision Quantization Policy through Large Margin RegularizationabstractMixed-precision quantization (MPQ) suffers from the time-consuming process of searching the optimal bit-width allocation (i.e., the policy) for each layer, especially when using large-scale datasets such as ISLVRC-2012. This limits the practicality of MPQ in real-world deployment scenarios. To address this issue, this paper proposes a novel method for efficiently searching for effective MPQ policies using a small proxy dataset instead of the large-scale dataset used for training the model. Deviating from the established norm of employing a consistent dataset for both model training and MPQ policy search stages, our approach, therefore, yields a substantial enhancement in the efficiency of MPQ exploration. Nonetheless, using discrepant datasets poses challenges in searching for a transferable MPQ policy. Driven by the observation that quantization noise of sub-optimal policy exerts a detrimental influence on the discriminability of feature representations---manifesting as diminished class margins and ambiguous decision boundaries---our method aims to identify policies that uphold the discriminative nature of feature representations, i.e., intra-class compactness and inter-class separation. This general and dataset-independent property makes us search for the MPQ policy over a rather small-scale proxy dataset and then the policy can be directly used to quantize the model trained on a large-scale dataset. Our method offers several advantages, including high proxy data utilization, no excessive hyper-parameter tuning, and high searching efficiency. We search high-quality MPQ policies with the proxy dataset that has only 4% of the data scale compared to the large-scale target dataset, achieving the same accuracy as searching directly on the latter, improving MPQ searching efficiency by up to 300×. Kai Ouyang, Zenghao Chai, Yunpeng Bai, Zhi Wang 0001, Wenwu Zhu 0001 |
ACM Multimedia | 6 |
| 2023 | JAVP: Joint-Aware Video Processing with Edge-Cloud Collaboration for DNN InferenceabstractCurrently, massive video inference tasks are processed through edge-cloud collaboration. However, the diverse scenarios make it difficult to allocate the inference tasks efficiently, resulting in many wasted resources. In this paper, we propose a joint-aware video processing (JAVP) architecture for edge-cloud collaboration. First, we develop a multiscale complexity-aware model for predicting task complexity and determining its suitability for edge or cloud servers. The task is subsequently efficiently scheduled to the appropriate servers by integrating complexity with an adaptive resource-aware optimization algorithm. For input tasks, JAVP can dynamically and intelligently select the most appropriate server. The evaluation results on public datasets show that JAVP can improve the through-put by more than 70% compared to traditional cloud-only solutions while meeting accuracy requirements. And JAVP can improve the accuracy by 3%-5% and reduce delay and energy consumption by 16%-50% compared to state-of-the-art edge-cloud solutions. Zheming Yang, Wen Ji 0003, Qi Guo 0009, Zhi Wang 0001 |
ACM Multimedia | 4 |
| 2023 | Breaking Filter Bubble: A Reinforcement Learning Framework of Controllable Recommender SystemabstractIn the information-overloaded era of the Web, recommender systems that provide personalized content filtering are now the mainstream portal for users to access Web information. Recommender systems deploy machine learning models to learn users’ preferences from collected historical data, leading to more centralized recommendation results due to the feedback loop. As a result, it will harm the ranking of content outside the narrowed scope and limit the options seen by users. In this work, we first conduct data analysis from a graph view to observe that the users’ feedback is restricted to limited items, verifying the phenomenon of centralized recommendation. We further develop a general simulation framework to derive the procedure of the recommender system, including data collection, model learning, and item exposure, which forms a loop. To address the filter bubble issue under the feedback loop, we then propose a general and easy-to-use reinforcement learning-based method, which can adaptively select few but effective connections between nodes from different communities as the exposure list. We conduct extensive experiments in the simulation framework based on large-scale real-world datasets. The results demonstrate that our proposed reinforcement learning-based control method can serve as an effective solution to alleviate the filter bubble and the separated communities induced by it. We believe the proposed framework of controllable recommendation in this work can inspire not only the researchers of recommender systems, but also a broader community concerned with artificial intelligence algorithms’ impact on humanity, especially for those vulnerable populations on the Web. Yancheng Dong, Chen Gao 0001, Dong Li 0016, Jianye Hao, Kai Zhang 0012, Yong Li 0008, Zhi Wang 0001 |
WWW | 9 |
| 2023 | Rolling horizon wind-thermal unit commitment optimization based on deep reinforcement learning
Jinhao Shi, Bo Wang 0027, Ran Yuan, Zhi Wang 0001, Chunlin Chen 0001, Junzo Watada |
Appl. Intell. | 4 |
| 2023 | A Dirichlet Process Mixture of Robust Task Models for Scalable Lifelong Reinforcement LearningabstractWhile reinforcement learning (RL) algorithms are achieving state-of-the-art performance in various challenging tasks, they can easily encounter catastrophic forgetting or interference when faced with lifelong streaming information. In this article, we propose a scalable lifelong RL method that dynamically expands the network capacity to accommodate new knowledge while preventing past memories from being perturbed. We use a Dirichlet process mixture to model the nonstationary task distribution, which captures task relatedness by estimating the likelihood of task-to-cluster assignments and clusters the task models in a latent space. We formulate the prior distribution of the mixture as a Chinese restaurant process (CRP) that instantiates new mixture components as needed. The update and expansion of the mixture are governed by the Bayesian nonparametric framework with an expectation maximization (EM) procedure, which dynamically adapts the model complexity without explicit task boundaries or heuristics. Moreover, we use the domain randomization technique to train robust prior parameters for the initialization of each task model in the mixture; thus, the resulting model can better generalize and adapt to unseen tasks. With extensive experiments conducted on robot navigation and locomotion domains, we show that our method successfully facilitates scalable lifelong RL and outperforms relevant existing methods. Zhi Wang 0001, Chunlin Chen 0001, Daoyi Dong |
IEEE Trans. Cybern. | 1 |
| 2023 | Interference-Aware Mobile Backscatter Communication: A PHY-Assisted Rate Adaptive ApproachabstractOver the past decade, backscatter nodes have received booming interest for many emerging mobile applications, such as sports analytics and interactive gaming. However, backscatter networks are not ready to provide a high-throughput and stable communication platform for billions of such mobile nodes due to two main factors in rate adaptation. First, the common mapping paradigm that chooses the optimal rate based on RSSIs is hardly adaptable to hardware diversity. Second, the current probing processes are not optimized for mobile scenarios due to inefficient probing trigger, inaccurate channel estimation, and unique self-interference. To address those issues, we propose MobiRate, a mobility-aware rate adaptation link-layer that fully exploits the mobility hints from PHY information to deliver a high-throughput link-layer for mobile backscatter networks. The key insight is that mobility-hints can greatly benefit link-layer design, including rate selection and channel probing. Specifically, we introduce a novel velocity-based loss-rate estimation module, a mobility-assisted probing trigger, a selective probing module, and a robust self-interference detection module, significantly saving probing time and improving probing accuracy. As MobiRate is fully compatible with the current standard, we prototype it using COTS RFID readers and commercial tags. Our extensive experiments demonstrate that MobiRate can successfully identify self-interference with detection accuracy over 90% for tags of different velocities. Moreover, it achieves up to 3.8x throughput gain over the state-of-the-art methods across a wide range of mobility, channel conditions, and tag types. Si Chen 0003, Wei Gong 0001, Jiangchuan Liu, Zhi Wang 0001, Jia Zhao 0006 |
IEEE Trans. Mob. Comput. | 4 |
| 2023 | Towards Real-Time Video Caching at Edge Servers: A Cost-Aware Deep Q-Learning SolutionabstractGiven the rapid growth of user-generated videos, internet traffic has been heavily dominated by online video streaming. Caching videos on edge servers in close proximity to users has been an effective approach to reduce the backbone traffic and the request response time, as well as to improve the video quality on the user side. Video popularity, however, can be highly dynamic over time. The cost of cache replacement at edge servers, particularly that related to service interruption during replacement, is not yet well understood. This paper presents a novel lightweight video caching algorithm for edge servers, seeking to optimize the hit rate with real-time decisions and minimized cost. Inspired by recent advances in deep Q-learning, our DQN-based online video caching (DQN-OVC) makes effective use of the rich and readily available information from users and networks. We decompose the Q-value function as a product of the video value function and the action function, which significantly reduces the state space. We instantiate the action function for cost-aware caching decisions with low complexity so that the cached videos can be updated continuously and instantly with dynamic video popularity. We used video traces from Tencent, one of the largest online video providers in China, to evaluate the performance of our DQN-OVC and to compare it with state-of-the-art solutions. The results demonstrate that DQN-OVC significantly outperforms the baseline algorithms in the edge caching context. Laizhong Cui, Erchao Ni, Yipeng Zhou, Zhi Wang 0001, Lei Zhang 0066, Jiangchuan Liu, Yuedong Xu 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | Caching in Dynamic Environments: A Near-Optimal Online Learning ApproachabstractThe rapid growth of rich multimedia data in today’s Internet, especially video traffic, has challenged the content delivery networks (CDNs). Caching serves as an important means to reduce user access latency so as to enable faster content downloads. Motivated by the dynamic nature of the real-world edge traces, this paper introduces aprovably wellonline caching policy in dynamic environments where: 1) the popularity is highly dynamic; 2) no regular stochastic pattern can model this dynamic evaluation process. First, we design an online optimization framework, which aims to minimize thedynamic regretthat finds the distance between an online caching policy and the best dynamic policy in hindsight. Second, we propose a dynamic online learning method to solve the non-stationary caching problem formulated in the previous framework. Compared to the linear dynamic regret of previous methods, our proposal is proved to achieve asublinear dynamic regret, from which it is guaranteed to be nearly optimal. We verify the design using both synthetic and real-world traces: the proposed policy achieves the best performance in the synthetic traces with different levels of dynamicity, which verifies the dynamic adaptation; our proposal consistently achieves at least 9.4% improvement than the baselines, including LRU, LFU, Static Online Learning based replacement, and Deep Reinforcement Learning based replacement, in random edge areas from real-world traces (from iQIYI), further verifying the effectiveness and robustness on the edge. Shiji Zhou, Zhi Wang 0001, Chenghao Hu, Yinan Mao, Haopeng Yan, Shanghang Zhang, Chuan Wu 0001, Wenwu Zhu 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | Instance Weighted Incremental Evolution Strategies for Reinforcement Learning in Dynamic EnvironmentsabstractEvolution strategies (ESs), as a family of black-box optimization algorithms, recently emerge as a scalable alternative to reinforcement learning (RL) approaches such as Q-learning or policy gradient and are much faster when many central processing units (CPUs) are available due to better parallelization. In this article, we propose a systematic incremental learning method for ES in dynamic environments. The goal is to adjust previously learned policy to a new one incrementally whenever the environment changes. We incorporate an instance weighting mechanism with ES to facilitate its learning adaptation while retaining scalability of ES. During parameter updating, higher weights are assigned to instances that contain more new knowledge, thus encouraging the search distribution to move toward new promising areas of parameter space. We propose two easy-to-implement metrics to calculate the weights: instance novelty and instance quality. Instance novelty measures an instance's difference from the previous optimum in the original environment, while instance quality corresponds to how well an instance performs in the new environment. The resulting algorithm, instance weighted incremental evolution strategies (IW-IESs), is verified to achieve significantly improved performance on challenging RL tasks ranging from robot navigation to locomotion. This article thus introduces a family of scalable ES algorithms for RL domains that enables rapid learning adaptation to dynamic environments. Zhi Wang 0001, Chunlin Chen 0001, Daoyi Dong |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | A Survey on Temporal Sentence Grounding in VideosabstractTemporal sentence grounding in videos (TSGV), which aims at localizing one target segment from an untrimmed video with respect to a given sentence query, has drawn increasing attentions in the research community over the past few years. Different from the task of temporal action localization, TSGV is more flexible since it can locate complicated activities via natural languages, without restrictions from predefined action categories. Meanwhile, TSGV is more challenging since it requires both textual and visual understanding for semantic alignment between two modalities (i.e., text and video). In this survey, we give a comprehensive overview for TSGV, which (i) summarizes the taxonomy of existing methods, (ii) provides a detailed description of the evaluation protocols (i.e., datasets and metrics) to be used in TSGV, and (iii) in-depth discusses potential problems of current benchmarking designs and research directions for further investigations. To the best of our knowledge, this is the first systematic survey on temporal sentence grounding. More specifically, we first discuss existing TSGV approaches by grouping them into four categories, i.e., two-stage methods, single-stage methods, reinforcement learning-based methods, and weakly supervised methods. Then we present the benchmark datasets and evaluation metrics to assess current research progress. Finally, we discuss some limitations in TSGV through pointing out potential problems improperly resolved in the current evaluation protocols, which may push forwards more cutting-edge research in TSGV. Besides, we also share our insights on several promising directions, including four typical tasks with new and practical settings based on TSGV. Xiaohan Lan, Yitian Yuan, Xin Wang 0019, Zhi Wang 0001, Wenwu Zhu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | A Closer Look at Debiased Temporal Sentence Grounding in Videos: Dataset, Metric, and ApproachabstractTemporal Sentence Grounding in Videos (TSGV) , which aims to ground a natural language sentence that indicates complex human activities in an untrimmed video, has drawn widespread attention over the past few years. However, recent studies have found that current benchmark datasets may have obvious moment annotation biases, enabling several simple baselines even without training to achieve state-of-the-art (SOTA) performance. In this paper, we take a closer look at existing evaluation protocols for TSGV, and find that both the prevailing dataset splits and evaluation metrics are the devils that lead to untrustworthy benchmarking. Therefore, we propose to re-organize the two widely-used datasets, making the ground-truth moment distributions different in the training and test splits, i.e., out-of-distribution (OOD) test. Meanwhile, we introduce a new evaluation metric “dR@ n ,IoU= m ” that discounts the basic recall scores especially with small IoU thresholds, so as to alleviate the inflating evaluation caused by biased datasets with a large proportion of long ground-truth moments. New benchmarking results indicate that our proposed evaluation protocols can better monitor the research progress in TSGV. Furthermore, we propose a novel causality-based Multi-branch Deconfounding Debiasing (MDD) framework for unbiased moment prediction. Specifically, we design a multi-branch deconfounder to eliminate the effects caused by multiple confounders with causal intervention. In order to help the model better align the semantics between sentence queries and video moments, we enhance the representations during feature encoding. Specifically, for textual information, the query is parsed into several verb-centered phrases to obtain a more fine-grained textual feature. For visual information, the positional information has been decomposed from the moment features to enhance the representations of moments with diverse locations. Extensive experiments demonstrate that our proposed approach can achieve competitive results among existing SOTA approaches and outperform the base model with great gains. Xiaohan Lan, Yitian Yuan, Xin Wang 0019, Long Chen 0016, Zhi Wang 0001, Lin Ma 0002, Wenwu Zhu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2022 | bert2BERT: Towards Reusable Pretrained Language ModelsabstractCheng Chen, Yichun Yin, Lifeng Shang, Xin Jiang, Yujia Qin, Fengyu Wang, Zhi Wang, Xiao Chen, Zhiyuan Liu, Qun Liu. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Yichun Yin, Lifeng Shang, Xin Jiang 0002, Yujia Qin, Zhi Wang 0001, Xiao Chen 0012, Zhiyuan Liu 0001, Qun Liu 0001 |
ACL (1) | 7 |
| 2022 | Online Continual Adaptation with Active Self-TrainingabstractModels trained with offline data often suffer from continual distribution shifts and expensive labeling in changing environments. This calls for a new online learning paradigm where the learner can continually adapt to changing environments with limited labels. In this paper, we propose a new online setting – Online Active Continual Adaptation, where the learner aims to continually adapt to changing distributions using both unlabeled samples and active queries of limited labels. To this end, we propose Online Self-Adaptive Mirror Descent (OSAMD), which adopts an online teacher-student structure to enable online self-training from unlabeled data, and a margin-based criterion that decides whether to query the labels to track changing distributions. Theoretically, we show that, in the separable case, OSAMD has an $O({T}^{2/3})$ dynamic regret bound under mild assumptions, which is aligned with the $\Omega(T^{2/3})$ lower bound of online learning algorithms with full labels. In the general case, we show a regret bound of $O({T}^{2/3} + \alpha^* T)$, where $\alpha^*$ denotes the separability of domains and is usually small. Our theoretical results show that OSAMD can fast adapt to changing environments with active queries. Empirically, we demonstrate that OSAMD achieves favorable regrets under changing environments with limited labels on both simulated and real-world data, which corroborates our theoretical findings. Shiji Zhou, Han Zhao 0002, Shanghang Zhang, Lianzhe Wang, Heng Chang, Zhi Wang 0001, Wenwu Zhu 0001 |
AISTATS | 6 |
| 2022 | Mixed-Precision Neural Network Quantization via Learned Layer-Wise Importance
Kai Ouyang, Zhi Wang 0001, Yifei Zhu 0001, Wen Ji 0003, Yaowei Wang 0001, Wenwu Zhu 0001 |
ECCV (11) | 3 |
| 2022 | Improved DC Estimation for JPEG Compression Via Convex RelaxationabstractMass image transmission has undergone an explosion of growth with the development of the internet, DCT-based lossy image compression like JPEG is pervasively conducted to save the transmission bandwidth. Recently, DCT-domain coefficient estimation approaches have been proposed to further improve the compression ratio by discarding DC coefficients at the sender’s end while recovering them at the receiver’s end via DC estimation. However, known DC estimation needs to enumerate all possible DC coefficients. Consequently, they are limited and resource-consuming due to the low delay requirements in real-time transmission. In this paper, we propose an improved DC estimation method via convex relaxation, which achieves state-of-the-art performance in terms of both recovery image quality and time complexity. Extensive experiments across various data sets demonstrate the advantages of our method. Jianghui Zhang, Bin Chen 0011, Yujun Huang, Han Qiu 0001, Zhi Wang 0001, Shutao Xia |
ICIP | 5 |
| 2022 | Maxim: DRL-Based Cross-Camera Streaming Configuration for Real-Time Video AnalyticsabstractReal-time video analytics (VA) with requirement of high accuracy necessitates intensive bandwidth resource consumption, calling for an adaptive streaming configuration strategy to strike a balance in VA pipelines. Existing works however suffer from a key limitation: the profiling-based strategy not only wastes unnecessary resources with the golden configuration transmission but also is trapped into only coarse-grained adaptation due the contradiction between profiling granularity and the profiling cost. In this paper, we for the first time reveal the correlation between the video dynamics and motion degree and highlight the limitation of traditional profiling-based strategy. We then propose Maxim, the first learning-based framework that solves the fine-grained VA configuration adaptation problem employing a novel deep reinforcement learning-based methodology without using any golden configuration in inference stage. Maxim could optimize the trade-off between resources cost and inference accuracy. Besides, Maxim also employs an enhanced cross-camera collaboration based on spatial and temporal correlation among cameras, which further improves robustness and performance in a large camera network. Extensive experiments confirm the superiority of our work compared with SOTA works, with a 76.6% improvement in comprehensive. Yutao Zhou, Fangxin Wang 0001, Zhi Wang 0001 |
ICME | 4 |
| 2022 | Fast Probabilistic Policy Reuse via Reward Function FittingabstractTransfer learning has shown great potential to accelerate reinforcement learning (RL) by utilizing prior knowledge of relevant task that has been learned in the past. Policy Reuse Q-learning (PRQL) is a general policy transfer framework, which speeds up the learning process of the target task by probabilistically reusing source policies from the policy library. In this paper, we propose an improved PRQL method to achieve more fast probabilistic policy reuse in deep reinforcement learning (DRL). First, we extend the basic PRQL algorithm to DRL, proposing a probability policy reuse algorithm that builds on DRL to solve more complex problems. Second, PRQL algorithms usually use a metric based on the average gain to measure the similarity between tasks. However, it contains very limited information and must be delayed until the end of an episode to update, which is inefficient. Instead, we propose a new metric based on fitting the reward function, which can make the agent converge to the most suitable reuse policy more quickly and accurately. We demonstrate the detection accuracy, received cumulative reward, and speed of convergence of our method in three complex Markov tasks. Experimental results show that our method can consistently achieve efficient policy transfer in these tasks. Jinmei Liu, Zhi Wang 0001, Chunlin Chen 0001 |
IJCNN | 2 |
| 2022 | HRL2E: Hierarchical Reinforcement Learning with Low-level EnsembleabstractGoal-conditioned hierarchical reinforcement learning (HRL) is a promising approach to solve challenging tasks with sparse rewards and long horizons. However, it suffers from the non-stationary problem due to the updating and unstable low level. To stabilize the low level more quickly and accelerate the non-stationary stage, we propose a novel HRL method: Hierarchical Reinforcement Learning with Low-level Ensemble (HRL2E). In HRL2E, the high level generates goals as high-level actions based on current states. Then the low level made up of several homogeneous policies attempts to complete these goals within a specific timestep budget. The improvement of our approach to the general goal-conditioned HRL algorithms can be summarized in two aspects. First, we estimate the target value function with the ensemble, stabilizing the training process. Second, we propose the Gates module composed of several scoring machines to score each low-level policy and judge which one has the most success potential to execute a specific goal. We adopt Twin Delayed Deep Deterministic Policy Gradient (TD3) in each level. Experimental comparison between our method and state-of-the-art goal-conditioned HRL methods on challenging continuous control tasks in MuJoCo domains shows our method can significantly accelerate training. You Qin, Zhi Wang 0001, Chunlin Chen 0001 |
IJCNN | 2 |
| 2022 | Cost Effective MLaaS Federation: A Combinatorial Reinforcement Learning ApproachabstractWith the advancement of deep learning techniques, major cloud providers and niche machine learning service providers start to offer their cloud-based machine learning tools, also known as machine learning as a service (MLaaS), to the public. According to our measurement, for the same task, these MLaaSes from different providers have varying performance due to the proprietary datasets, models, etc. Federating different MLaaSes together allows us to improve the analytic performance further. However, naively aggregating results from different MLaaSes not only incurs significant momentary cost but also may lead to sub-optimal performance gain due to the introduction of possible false-positive results. In this paper, we propose Armol, a framework to federate the right selection of MLaaS providers to achieve the best possible analytic performance. We first design a word grouping algorithm to unify the output labels across different providers. We then present a deep combinatorial reinforcement learning based-approach to maximize the accuracy while minimizing the cost. The predictions from the selected providers are then aggregated together using carefully chosen ensemble strategies. The real-world trace-driven evaluation further demonstrates that Armol is able to achieve the same accuracy results with 67% less inference cost. Shuzhao Xie, Yuan Xue 0013, Yifei Zhu 0001, Zhi Wang 0001 |
INFOCOM | 4 |
| 2022 | Batch Adaptative Streaming for Video AnalyticsabstractVideo streaming plays a critical role in the video analytics pipeline and thus its adaptation scheme has been a focus of optimization. As machine learning algorithms have become main consumers of video contents, the streaming adaptation decision should be made to optimize their inference performance. Existing video streaming adaptation schemes for video analytics are usually designed to adapt to bandwidth and content variations separately, which fail to consider the coordination between transmission and computation. Given the nature of batch transmission in video streaming and batch processing in deep learning-based inference, we observe that the choices of the batch sizes directly affects the bandwidth efficiency, the response delay and the accuracy of the deep learning inference in video analytics. In this work, we investigate the effect of the batch size in transmission and processing, formulate the optimal batch size adaptation problem, and further develop the deep reinforcement learning-based solution. Practical issues are further addressed for Implementation. Extensive simulations are conducted for performance evaluation, whose results demonstrate the superiority of our proposed batch adaptive streaming approach over the baseline streaming approaches. Lei Zhang 0066, Ximing Wu, Fangxin Wang 0001, Laizhong Cui, Zhi Wang 0001, Jiangchuan Liu |
INFOCOM | 6 |
| 2022 | Target-oriented Semi-supervised Domain Adaptation for WiFi-based HARabstractIncorporating domain adaptation is a promising solution to mitigate the domain shift problem of WiFi-based human activity recognition (HAR). The state-of-the-art solutions, however, do not fully exploit all the data, only focusing either on unlabeled samples or labeled samples in the target WiFi environment. Moreover, they largely fail to carefully consider the discrepancy between the source and target WiFi environments, making the adaptation of models to the target environment with few samples become much less effective. To cope with those issues, we propose a Target-Oriented Semi-Supervised (TOSS) domain adaptation method for WiFi-based HAR that can effectively leverage both labeled and unlabeled target samples. We further design a dynamic pseudo label strategy and an uncertainty-based selection method to learn the knowledge from both source and target environments. We implement TOSS with a typical meta learning model and conduct extensive evaluations. The results show that TOSS greatly outperforms state-of-the-art methods under comprehensive 1 on 1 and multi-source one-shot domain adaptation experiments across multiple real-world scenarios. Feng Wang 0001, Jihong Yu, Ju Ren 0001, Zhi Wang 0001, Wei Gong 0001 |
INFOCOM | 5 |
| 2022 | Personalized 360-Degree Video Streaming: A Meta-Learning ApproachabstractOver the past decades, 360-degree videos have attracted wide interest for the immersive experience they bring to viewers. The rising of high-resolution 360-degree videos greatly challenges the traditional video streaming systems in limited network environments. Given the limited bandwidth, tile-based video streaming with adaptive bitrate selection has been widely studied to improve the Quality of Experience (QoE) of viewers by tiling the video frames and allocating different bitrates for tiles inside and outside viewers' viewports. Existing solutions for viewport prediction and bitrate selection train general models without catering to the intrinsic need for personalization. In this paper, we present the first meta-learning-based personalized 360-degree video streaming framework. The commonality among viewers of different viewing patterns and QoE preferences is captured by efficient meta-network designs. Specifically, we design a meta-based long-short term memory model for viewport prediction and a meta-based reinforcement learning model for bitrate selection. Extensive experiments on real-world datasets demonstrate that our framework not only outperforms the state-of-the-art data-driven approaches in prediction accuracy by 11% on average and improves QoE by 27% on average, but also quickly adapts to users with new preferences with on average 67%-88% less training epochs. Yiyun Lu, Yifei Zhu 0001, Zhi Wang 0001 |
ACM Multimedia | 3 |
| 2022 | Arbitrary Bit-width Network: A Joint Layer-Wise Quantization and Adaptive Inference ApproachabstractConventional model quantization methods use a fixed quantization scheme to different data samples, which ignores the inherent"recognition difficulty" differences between various samples. We propose to feed different data samples with varying quantization schemes to achieve a data-dependent dynamic inference, at a fine-grained layer level. However, enabling this adaptive inference with changeable layer-wise quantization schemes is challenging because the combination of bit-widths and layers is growing exponentially, making it extremely difficult to train a single model in such a vast searching space and use it in practice. To solve this problem, we present the Arbitrary Bit-width Network (ABN), where the bit-widths of a single deep network can change at runtime for different data samples, with a layer-wise granularity. Specifically, first we build a weight-shared layer-wise quantizable "super-network" in which each layer can be allocated with multiple bit-widths and thus quantized differently on demand. The super-network provides a considerably large number of combinations of bit-widths and layers, each of which can be used during inference without retraining or storing myriad models. Second, based on the well-trained super-network, each layer's runtime bit-width selection decision is modeled as a Markov Decision Process (MDP) and solved by an adaptive inference strategy accordingly. Experiments show that the super-network can be built without accuracy degradation, and the bit-widths allocation of each layer can be adjusted to deal with various inputs on the fly. On ImageNet classification, we achieve 1.1% top1 accuracy improvement while saving 36.2% BitOps. Haoyu Zhai, Kai Ouyang, Zhi Wang 0001, Yifei Zhu 0001, Wenwu Zhu 0001 |
ACM Multimedia | 4 |
| 2022 | AdaConfigure: Reinforcement Learning-Based Adaptive Configuration for Video Analytics Services
Zhaoliang He, Yuan Wang 0083, Zhi Wang 0001, Wenwu Zhu 0001, Chenyang Guo, Zhibo Chen 0006 |
MMM (1) | 4 |
| 2022 | Understanding 5G performance for real-world services: a content provider's perspectiveabstract5G has seen rapid growth recently, attracting several measurement studies on its coverage, connectivity and quality of service. However, there is still a lack of understanding of 5G's capabilities and potential impacts from a content provider (CP)'s perspective. This paper fills in this gap by studying 5G networks used by over 23 million users in one year in Kuaishou, a popular crowdsourced live streaming platform. Our measurements provide the following discoveries. i) Standalone (SA) 5G generally provides end-to-end performance improvement as compared with 4G or non-SA (NSA) 5G, but its advantage depends on both the number of cellular users and CP-level configurations. ii) In the radio access network, SA 5G is more sensitive to access density but has better handover tolerance. iii) Controlled experiments with 29 mobile device models on energy consumption refute some "conventional wisdom," including that 5G always consumes more power. iv) Traceroute-based active experiments in over 300 cities show that although users are "closer" to the internet in SA 5G, their end-to-end latency may not benefit from that. Furthermore, we show new design space for 5G participants and provide a 5G-aware rebuffer strategy tested by 9 million viewers in Kuaishou, with a 7% reduction in rebuffer proportion. Xinjie Yuan, Mingzhou Wu, Zhi Wang 0001, Yifei Zhu 0001, Junjian Guo, Zhi-Li Zhang, Wenwu Zhu 0001 |
SIGCOMM | 3 |
| 2022 | HEBO: An Empirical Study of Assumptions in Bayesian OptimisationabstractIn this work we rigorously analyse assumptions inherent to black-box optimisation hyper-parameter tuning tasks. Our results on the Bayesmark benchmark indicate that heteroscedasticity and non-stationarity pose significant challenges for black-box optimisers. Based on these findings, we propose a Heteroscedastic and Evolutionary Bayesian Optimisation solver (HEBO). HEBO performs non-linear input and output warping, admits exact marginal log-likelihood optimisation and is robust to the values of learned parameters. We demonstrate HEBO’s empirical efficacy on the NeurIPS 2020 Black-Box Optimisation challenge, where HEBO placed first. Upon further analysis, we observe that HEBO significantly outperforms existing black-box optimisers on 108 machine learning hyperparameter tuning tasks comprising the Bayesmark benchmark. Our findings indicate that the majority of hyper-parameter tuning tasks exhibit heteroscedasticity and non-stationarity, multiobjective acquisition ensembles with Pareto front solutions improve queried configurations, and robust acquisition maximisers afford empirical advantages relative to their non-robust counterparts. We hope these findings may serve as guiding principles for practitioners of Bayesian optimisation. Alexander I. Cowen-Rivers, Wenlong Lyu, Rasul Tutunov, Zhi Wang 0001, Antoine Grosnit, Ryan-Rhys Griffiths, Alexandre Maraval, Jianye Hao, Jun Wang 0012, Jan Peters 0001, Haitham Bou-Ammar |
J. Artif. Intell. Res. | 4 |
| 2022 | CDFKD-MFS: Collaborative Data-Free Knowledge Distillation via Multi-Level Feature SharingabstractRecently, the compression and deployment of powerful deep neural networks (DNNs) on resource-limited edge devices to provide intelligent services have become attractive tasks. Although knowledge distillation (KD) is a feasible solution for compression, its requirement on the original dataset raises privacy concerns. In addition, it is common to integrate multiple pretrained models to achieve satisfactory performance. How to compress multiple models into a tiny model is challenging, especially when the original data are unavailable. To tackle this challenge, we propose a framework termed collaborative data-free knowledge distillation via multi-level feature sharing (CDFKD-MFS), which consists of a multi-header student module, an asymmetric adversarial data-free KD module, and an attention-based aggregation module. In this framework, the student model equipped with a multi-level feature-sharing structure learns from multiple teacher models and is trained together with a generator in an asymmetric adversarial manner. When some real samples are available, the attention module adaptively aggregates predictions of the student headers, which can further improve performance. We conduct extensive experiments on three popular computer visual datasets. In particular, compared with the most competitive alternative, the accuracy of the proposed framework is 1.18% higher on the CIFAR-100 dataset, 1.67% higher on the Caltech-101 dataset, and 2.99% higher on the mini-ImageNet dataset. Zhiwei Hao 0001, Yong Luo 0002, Zhi Wang 0001, Han Hu 0003, Jianping An |
IEEE Trans. Multim. | 3 |
| 2022 | Active Gradual Domain Adaptation: Dataset and ApproachabstractAdapting deep neural networks to the changing environments is critical in practical utility, especially for online web applications, where the data distribution changes gradually due to the evolving environments. For instance, the web photos of cellphones change gradually over years due to appearance changes. This paper deals with such a problem via active gradual domain adaptation, where the learner continually and actively selects the most informative labels from the target to enhance labeling efficiency and utilizes both labeled and unlabeled samples to improve the model adaptation under gradual domain drift. We propose the active gradual self-training (AGST) algorithm with novel designs of active pseudolabeling and gradual semi-supervised domain adaptation. Specifically, AGST pseudolabels the samples with high confidence, and selects the most informative labels from the unconfident samples based on both uncertainty and diversity, and then gradually self-trains itself by confident pseudolabels and queried labels. To study the gradual domain shift problem in the web data and verify the proposed algorithm, we create a new dataset -- Evolving-Image-Search (EVIS), collected from the web search engine and covers a 12-years range. Since the appearance of the products evolves over these years, such dataset naturally contains gradual domain drift. We extensively evaluate AGST on the synthetic dataset, real-world dataset, and EVIS dataset. AGST achieves up to 62% accuracy improvement (absolute value) against unsupervised gradual self-training with only 5% additional labels, and 19% accuracy improvement against directly applying CLUE, demonstrating the effectiveness of the designs of active pseudolabel and gradual semi-supervised domain adaptation. Shiji Zhou, Lianzhe Wang, Shanghang Zhang, Zhi Wang 0001, Wenwu Zhu 0001 |
IEEE Trans. Multim. | 4 |
| 2022 | Lifelong Incremental Reinforcement Learning With Online Bayesian InferenceabstractA central capability of a long-lived reinforcement learning (RL) agent is to incrementally adapt its behavior as its environment changes and to incrementally build upon previous experiences to facilitate future learning in real-world scenarios. In this article, we propose lifelong incremental reinforcement learning (LLIRL), a new incremental algorithm for efficient lifelong adaptation to dynamic environments. We develop and maintain a library that contains an infinite mixture of parameterized environment models, which is equivalent to clustering environment parameters in a latent space. The prior distribution over the mixture is formulated as a Chinese restaurant process (CRP), which incrementally instantiates new environment models without any external information to signal environmental changes in advance. During lifelong learning, we employ the expectation-maximization (EM) algorithm with online Bayesian inference to update the mixture in a fully incremental manner. In EM, the E-step involves estimating the posterior expectation of environment-to-cluster assignments, whereas the M-step updates the environment parameters for future learning. This method allows for all environment models to be adapted as necessary, with new models instantiated for environmental changes and old models retrieved when previously seen environments are encountered again. Simulation experiments demonstrate that LLIRL outperforms relevant existing methods and enables effective incremental adaptation to various dynamic environments for lifelong learning. Zhi Wang 0001, Chunlin Chen 0001, Daoyi Dong |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2021 | Extract then Distill: Efficient and Effective Task-Agnostic BERT Distillation
Yichun Yin, Lifeng Shang, Zhi Wang 0001, Xin Jiang 0002, Xiao Chen 0012, Qun Liu 0001 |
ICANN (3) | 4 |
| 2021 | A Spherical Mixture Model Approach for 360 Video Virtual Cinematographyabstract360 video virtual cinematography attempts to direct a virtual camera and capture the most salient regions of 360 videos. In this paper, we propose a data-drive solution to achieve high-quality and diversified 360 cinematography based on crowd-sourced viewing histories. Specifically, we try to address two problems: 1) how to locate the semantically important regions of interest (RoI) from raw data, 2) how to generate virtual camera paths that follow chronological narratives. We first design a dynamic spherical mixture model based algorithm to locate variable number of RoIs on each video frame. We then model the camera transition and chronological orders with a Bayesian network and conditional probabilities. With the above two designs, we can generate “optimal” cinematography paths based on a dynamic programming algorithm. By modeling the RoIs as spherical mixture model, we are also able to provide diversified cinematography results. We show its effectiveness through extensive experiments. Chenglei Wu, Zhi Wang 0001, Lifeng Sun |
ICIP | 2 |
| 2021 | Model Compression via Collaborative Data-Free Knowledge Distillation for Edge IntelligenceabstractModel compression without the original data for fine-tuning is challenging for deploying large-size models on resource constrained edge devices. To this end, we propose a novel data-free model compression framework based on knowledge distillation (KD), where multiple teachers are utilized in a collaborative manner to enable reliable distillation. It mainly consists of three components: adversarial data generation, multi-teacher KD, and adaptive outputs aggregation. In particular, some synthesized data are generated in an adversarial manner to mimic the original data for model compression. Then a multi-header module is developed to simultaneously leverage diverse knowledge from multiple teachers. The distillation outputs are adaptively aggregated for final prediction. The experimental results demonstrate that our framework outperforms the data-free counterpart significantly (4.48% on MNIST and 2.96% on CIFAR-10). Effectiveness of different components of our method is also verified via carefully designed ablation study. Zhiwei Hao 0001, Yong Luo 0002, Zhi Wang 0001, Han Hu 0003, Jianping An |
ICME | 3 |
| 2021 | DRL-Based Collaborative Edge Content Replication with Popularity DistillationabstractThe infrastructure for multimedia content delivery has been using more and more edge infrastructure (e.g., base stations, smart routers, etc.), which not only alleviates the centralized servers but also improves the quality of service by letting users access content nearby. Algorithms based on deep reinforcement learning (DRL) have been widely adopted by such edge cache replacement strategies due to their capability to adapt to changing request patterns. However, a DRL cache replacement agent learns extremely slow at an edge cache because of the sparse requests. In this paper, we propose a popularity distillation framework that allows edge caches to refer to content replication strategies of other edge caches. First, we design a collaborative edge cache framework that lets edge caches learn their strategies by handling the local requests using deep reinforcement learning and learn from others by exchanging the "soft" popularity distributions experienced by different edge caches. Second, we design a neighbor maintenance mechanism in which an agent iteratively selects only a small number of neighboring edge caches to perform the collaboration. Experiments driven by a real-world mobile video dataset show that our design can improve the cache hit rate by 3.0% compared with a non-popularity distillation baseline with only a small overhead of transmission data during distillation. Haopeng Yan, Zeming Chen 0002, Zhi Wang 0001, Wenwu Zhu 0001 |
ICME | 3 |
| 2021 | An Adaptive Logarithm Quantization Method for DNN Compression
Yuan Wang 0083, Zhaoliang He, Zhi Wang 0001, Wenwu Zhu 0001 |
ICONIP (5) | 4 |
| 2021 | STG-Meta: Spatial-Temporal Graph Meta-Learning for Traffic ForecastingabstractFor current traffic datasets, data scarcity is common in many districts. This limits the performance of existing spatial-temporal models. Current works tackle this problem by transferring knowledge from other cities. They mainly focus on grid-like data, thus not sharing graph structure information across cities. However, for graph-structured data like highway traffic flow, the spatial dependency between nodes is much more obvious. Furthermore, existing research focuses solely on data inconsistency between cities, ignoring data inconsistency in different periods. Towards these, we propose STG-Meta, a meta-learning-based framework for graph-based traffic prediction tasks with only limited training samples. Specifically, STG-Meta adopts the cross-city-cross-period task construction method to reflect the variation in data across periods. STG-Meta includes the structure memory to store the embedding of the structure patterns. Additionally, the optimization-based meta-learning method is utilized to extract knowledge such as the memory and the initialization parameters of spatial-temporal graph (STG) networks, from other cities. Experiments on realworld datasets of two traffic data types demonstrate that our method outperforms state-of-the-art approaches for data-scarce cities. Jiadong Li, Wang Pan, Qipu Deng, Zhi Wang 0001, Wenwu Zhu 0001 |
IJCNN | 4 |
| 2021 | STSIR: A Spatial Temporal Pandemic Model with Mobility Data - A COVID-19 StudyabstractWith the outbreak of COVID-19, how to mitigate and suppress its spread is a big issue to the government. Department of public health need powerful models to analyze and predict the trend and scale of such pandemic. And models that could evaluate the effect of the public policy are also essential to the fight with COVID-19. A main limitation of existing models is that they can only evaluate the policy by calculating R0after infection happens instead of giving observable index. To tackle this, based on the transmission characteristics of the COVID-19, we propose a novel framework Spatial-Temporal-Susceptible-Infected-Removed (STSIR) model. In particular, we combine both intra-city and inter-city mobility indices with the traditional SIR dynamics and make it a dynamic system. And we prove that the STSIR system is a closed system which makes the system self-consistent. And finally we proposed a Multi-Stage Simulated Annealing (MSSA) algorithm to find the optimal parameters of the system. In our experiments, based on Baidu Mobility dataset [1], and China pandemic dataset provided by Dingxiangyuan [2], our model can effectively predict the total scale of the pandemic and also give clear policy analysis with the observable index. Wang Pan, Qipu Deng, Jiadong Li, Zhi Wang 0001, Wenwu Zhu 0001 |
IJCNN | 4 |
| 2021 | Joint Cache Size Scaling and Replacement Adaptation for Small Content ProvidersabstractElastic Content Delivery Networks (Elastic CDNs) have been introduced to support explosive Internet traffic growth by providing small Content Providers (CPs) with just-in-time services. Due to the diverse requirements of small CPs, they need customized adaptive caching modules to help them adjust the cached contents to maximize their long-term utility. The traditional adaptive caching module is usually a built-in service in a cloud CDN. They adaptively change cache contents using size-scaling-only methods or strategy-adaptation-only methods. A natural question is: can we jointly optimize size and strategy to achieve tradeoff and better performance for small CPs when renting services from elastic CDNs? The problem is challenging because the two decision variables could involve both discrete and categorical variables, where discrete variables have an intrinsic order while categorical variables do not. In this paper, we propose a distribution-guided reinforcement learning framework JEANA to learn the joint cache size scaling and strategy adaptation policy. We design a distribution-guided regularizer to keep the intrinsic order of discrete variables. More importantly, we prove that our algorithm has a theoretical guarantee of performance improvement. Trace-driven experimental results demonstrate our method can improve the hit ratio while reducing the rental cost. Jiahui Ye, Zichun Li, Zhi Wang 0001, Zhuobin Zheng, Han Hu 0003, Wenwu Zhu 0001 |
INFOCOM | 3 |
| 2021 | Mix-order Attention Networks for Image RestorationabstractConvolutional neural networks (CNNs) have obtained great success in image restoration tasks, like single image denoising, demosaicing, and super-resolution. However, most existing CNN-based methods neglect the diversity of image contents and degradations in the corrupted images and treat channel-wise features equally, thus hindering the representation ability of CNNs. To address this issue, we propose deep mix-order attention networks (MAN) to extract features that capture rich feature statistics within networks. Our MAN is mainly built on simple residual blocks and our mix-order channel attention (MOCA) module, which further consists of feature gating and feature pooling blocks to capture different types of semantic information. With our MOCA, our MAN can be flexible to handle various types of image contents and degradations. Besides, our MAN can be generalized to different image restoration tasks, like image denoising, super-resolution, and demosaicing. Extensive experiments demonstrate that our method obtains favorably against state-of-the-art methods in terms of quantitative and qualitative metrics. Tao Dai 0001, Yalei Lv, Bin Chen 0011, Zhi Wang 0001, Zexuan Zhu 0001, Shutao Xia |
ACM Multimedia | 4 |
| 2021 | Towards cloud-edge collaborative online video analytics with fine-grained serverless pipelinesabstractThe ever-growing deployment scale of surveillance cameras and the users' increasing appetite for real-time queries have urged online video analytics. Synergizing the virtually unlimited cloud resources with agile edge processing would deliver an ideal online video analytics system; yet, given the complex interaction and dependency within and across video query pipelines, it is easier said than done. This paper starts with a measurement study to acquire a deep understanding of video query pipelines on real-world camera streams. We identify the potentials and practical challenges towards cloud-edge collaborative video analytics. We then argue that the newly emerged serverless computing paradigm is the key to achieve fine-grained resource partitioning with minimum dependency. We accordingly propose CEVAS, a Cloud-Edge collaborative Video Analytics system empowered by fine-grained Serverless pipelines. It builds flexible serverless-based infrastructures to facilitate fine-grained and adaptive partitioning of cloud-edge workloads for multiple concurrent query pipelines. With the optimized design of individual modules and their integration, CEVAS achieves real-time responses to highly dynamic input workloads. We have developed a prototype of CEVAS over Amazon Web Services (AWS) and conducted extensive experiments with real-world video streams and queries. The results show that by judiciously coordinating the fine-grained serverless resources in the cloud and at the edge, CEVAS reduces 86.9% cloud expenditure and 74.4% data transfer overhead of a pure cloud scheme and improves the analysis throughput of a pure edge scheme by up to 20.6%. Thanks to the fine-grained video content-aware forecasting, CEVAS is also more adaptive than the state-of-the-art cloud-edge collaborative scheme. Miao Zhang 0003, Fangxin Wang 0001, Yifei Zhu 0001, Jiangchuan Liu, Zhi Wang 0001 |
MMSys | 5 |
| 2021 | PAAS: a preference-aware deep reinforcement learning approach for 360° video streamingabstractConventional tile-based 360° video streaming methods, including deep reinforcement learning (DRL) based, ignore the interactive nature of 360° video streaming and download tiles following fixed sequential orders, thus failing to respond to the user's head motion changes. We show that these existing solutions suffer from either the prefetch accuracy or the playback stability drop. Furthermore, these methods are constrained to serve only one fixed streaming preference, causing extra training overhead and the lack of generalization on unseen preferences. In this paper, we propose a dual-queue streaming framework, with accuracy and stability purposes respectively, to enable the DRL agent to determine and change the tile download order without incurring overhead. We also design a preference-aware DRL algorithm to incentivize the agent to learn preference-dependent ABR decisions efficiently. Compared with state-of-the-art DRL baselines, our method not only significantly improves the streaming quality, e.g., increasing the average streaming quality by 13.6% on a public dataset, but also demonstrates better performance and generalization under dynamic preferences, e.g., an average quality improvement of 19.9% on unseen preferences. Chenglei Wu, Zhi Wang 0001, Lifeng Sun |
NOSSDAV | 2 |
| 2021 | Joint Model and Data Adaptation for Cloud Inference ServingabstractReal-time deep learning inference serving systems often require prohibitive resources and diverse user requirements. The existing design of inference serving systems mainly focusing on computation resource efficiency, largely ignoring the trade-off between computation and bandwidth resources in need. Sub-optimal resource utilization usually leads to huge serving cost waste. In this paper, we tackle the dual challenge of computation-bandwidth trade-off and cost-effectiveness by proposing A2, an efficient joint Adaptive model, and Adaptive data deep learning serving solution across the geo-datacenters. Inspired by the insight that a trade-off between computational cost and bandwidth cost in achieving the same accuracy, we design a real-time inference serving framework, which selectively places different "versions" of the deep learning models at different geo-locations, and schedules different data sample versions to be sent to those model versions for inference. The goal is to minimize the total serving cost while meeting latency and accuracy demand for the serving requests. We formulate a joint placement and serving problem and propose an efficient approximation algorithm to solve it with a theoretical performance guarantee. We deploy A2on Amazon EC2 for experiments, which shows that A2achieves 30%-50% serving cost reduction under the same required latency and accuracy as compared to baselines. Jingyan Jiang, Ziyue Luo, Chenghao Hu, Zhaoliang He, Zhi Wang 0001, Shutao Xia, Chuan Wu 0001 |
RTSS | 5 |
| 2021 | Multi-UAV-Enabled Mobile-Edge Computing for Time-Constrained IoT ApplicationsabstractUnmanned-aerial-vehicle (UAV)-enabled mobile-edge computing (MEC) has emerged as a promising paradigm to extend the coverage of computation service for Internet of Things (IoT) applications, which are usually time sensitive and computation intensive. In this article, a novel design framework is proposed for a multi-UAV-enabled MEC system, where edge servers are equipped on multiple UAVs to provide flexible computation assistance to IoT devices with hard deadlines. The aim is to maximize the number of served IoT devices through jointly optimizing UAV trajectory and service indicator as well as resource allocation and computation offloading, where the chosen IoT devices will complete their computation tasks on time under given energy budgets and co-channel interference is taken into account. We formulate the optimization problem as a mixed integer nonlinear programming (MINLP), which is challenging to solve directly. The problem is first reformulated to a more mathematically tractable form by adding a penalty term to the objective function. We then decouple the problem into two subproblems and develop an iterative algorithm by solving the two subproblems with alternating optimization and successive convex approximation techniques, where the proposed algorithm converges to a Karush–Kuhn–Tucker (KKT) solution. In addition, an efficient initialization scheme is proposed based on multiple traveling salesman problem with time windows (m-TSPTWs) method. Finally, simulation results are provided to demonstrate that the proposed joint design achieves significant performance gains over baseline schemes. Cheng Zhan, Han Hu 0003, Zhi Liu 0002, Zhi Wang 0001, Shiwen Mao |
IEEE Internet Things J. | 4 |
| 2021 | Unsupervised text-to-image synthesis
Yanlong Dong, Ying Zhang 0021, Lin Ma 0002, Zhi Wang 0001, Jiebo Luo 0001 |
Pattern Recognit. | 4 |
| 2021 | TCLiVi: Transmission Control in Live Video Streaming Based on Deep Reinforcement LearningabstractCurrently, video content accounts for the majority of network traffic. With increased live streaming, rigorous requirements have been introduced for better Quality of Experience (QoE). It is challenging to meet satisfactory QoE in live streaming, where the aim is to achieve a balance between 1) enhancing the video quality and stability and 2) reducing the rebuffering time and end-to-end delay, under different scenarios with various network conditions and user preferences, where the fluctuation in the network throughput degrades the QoE severely. In this paper, we propose an approach to improve the QoE for live video streaming based on Deep Reinforcement Learning (DRL). The new approach jointly adjusts the streaming parameters, including the video bitrate and target buffer size. With the basic DRL framework, TCLiVi can automatically generate the inference model based on the playback information, to achieve the joint optimization of the video quality, stability, rebuffering time and latency parameters. We evaluate our framework on real-world data in different live streaming broadcast scenarios, such as a talent show and a sports competition under different network conditions. We compare TCLiVi with other algorithms, such as the Double DQN, MPC and Buffer-based algorithms. The simulation results show that TCLiVi significantly improves the video quality and decreases the rebuffering time, consequently increasing the QoE score by 40.84% in average. We also show that TCLiVi is self-adaptive in different scenarios. Laizhong Cui, Dongyuan Su, Shu Yang 0002, Zhi Wang 0001, Zhong Ming 0001 |
IEEE Trans. Multim. | 4 |
| 2021 | Adaptive Compression for Online Computer Vision: An Edge Reinforcement Learning ApproachabstractWith the growth of computer vision-based applications, an explosive amount of images have been uploaded to cloud servers that host such online computer vision algorithms, usually in the form of deep learning models. JPEG has been used as the de facto compression and encapsulation method for images. However, standard JPEG configuration does not always perform well for compressing images that are to be processed by a deep learning model—for example, the standard quality level of JPEG leads to 50% of size overhead (compared with the best quality level selection) on ImageNet under the same inference accuracy in popular computer vision models (e.g., InceptionNet and ResNet). Knowing this, designing a better JPEG configuration for online computer vision-based services is still extremely challenging. First, cloud-based computer vision models are usually a black box to end-users; thus, it is challenging to design JPEG configuration without knowing their model structures. Second, the “optimal” JPEG configuration is not fixed; instead, it is determined by confounding factors, including the characteristics of the input images and the model, the expected accuracy and image size, and so forth. In this article, we propose a reinforcement learning (RL)-based adaptive JPEG configuration framework, AdaCompress. In particular, we design an edge (i.e., user-side) RL agent that learns the optimal compression quality level to achieve an expected inference accuracy and upload image size, only from the online inference results, without knowing details of the model structures. Furthermore, we design an explore-exploit mechanism to let the framework fast switch an agent when it detects a performance degradation, mainly due to the input change (e.g., images captured across daytime and night). Our evaluation experiments using real-world online computer vision-based APIs from Amazon Rekognition, Face++, and Baidu Vision show that our approach outperforms existing baselines by reducing the size of images by one-half to one-third while the overall classification accuracy only decreases slightly. Meanwhile, AdaCompress adaptively re-trains or re-loads the RL agent promptly to maintain the performance. Zhaoliang He, Hongshan Li, Zhi Wang 0001, Shutao Xia, Wenwu Zhu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2021 | Dissimilarity Analysis-Based Multimode Modeling for Complex Distributed Parameter SystemsabstractFor complex distributed parameter systems (DPSs) with strong nonlinearities and time-varying dynamics, the conventional spatiotemporal modeling methods become ill-suited since the elementary assumption that the process data follow a unimodal Gaussian distribution usually becomes invalid. In this paper, a multimode method is proposed for modeling of such systems. First, the original operating space is partitioned along the time dimension into several subspaces via modified dissimilarity analysis. Each subspace represents the local spatiotemporal characteristics of the original system. Second, the Karhunen-Loève decomposition (KLD)-based spatiotemporal modeling approach is applied to approximate the local dynamics of each subspace. Finally, an ensemble model is obtained using the soft weighting sum of the local ones, where the corresponding weights are calculated by principal component regression. By properly decomposing the original space into several local parts, the ensemble model is capable of handling the strong nonlinearities and time-varying dynamics of the system. The validity and efficiency of the proposed method are verified on two representative applications: 1) a one-dimensional parabolic catalytic rod and 2) a two-dimensional curing thermal process. The experimental results show that the proposed method provides a superior performance regarding modeling accuracy compared to several baselines. Zhi Wang 0001, Han-Xiong Li |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2020 | Texture and Shape Biased Two-Stream Networks for Clothing Classification and Attribute RecognitionabstractClothes category classification and attribute recognition have achieved distinguished success with the development of deep learning. People have found that landmark detection plays a positive role in these tasks. However, little research is committed to analyzing these tasks from the perspective of clothing attributes. In our work, we explore the usefulness of landmarks and find that landmarks can assist in extracting shape features; and using landmarks for joint learning can increase classification and recognition accuracy effectively. We also find that texture features have an impelling effect on these tasks and that the pre-trained ImageNet model has good performance in extracting texture features. To this end, we propose to use two streams to enhance the extraction of shape and texture, respectively. In particular, this paper proposes a simple implementation, Texture and Shape biased Fashion Networks (TS-FashionNet). Comprehensive and rich experiments demonstrate our discoveries and the effectiveness of our model. We improve the top-3 classification accuracy by 0.83% and improve the top-3 attribute recognition recall rate by 1.39% compared to the state-of-the-art models. Chun Yuan 0003, Zhi Wang 0001 |
CVPR | 4 |
| 2020 | Fine-Grained Garment Parsing: A Body Generation ApproachabstractCurrent human parsing methods segment an image into different semantic parts including background, body parts and garments. A major limitation of today’s human parsing methodologies is that they are not able to provide fine-grained garment segmentation (e.g., left and right sleeves), and it is mainly due to the lack of a dataset with such fine-grained semantic garment part labels. To tackle this, we propose a body generation approach for fine-grained garment parsing. In particular, we first use a body generation module based on image inpainting, to locate the fine-grained garment parts corresponding to where the generated body parts are, e.g., the left sleeve is assumed to be associated with the left arm; we then extract the garment parts from the original whole garment based on the positions above. In our experiments based on a public dataset focusing on top clothing images, our solution can effectively separate a top garment into a left sleeve, a right sleeve and front, as compared to state-of-the-art solutions that parse it as a whole. Zhi Wang 0001 |
ICME | 4 |
| 2020 | IEDQN: Information Exchange DQN with a Centralized Coordinator for Traffic Signal ControlabstractFinding the optimal control strategy for traffic signals, especially for multi-intersection traffic signals, is still a difficult task. The use of reinforcement learning (RL) algorithms to this problem is greatly limited because of the partially observable and nonstationary environment. In this paper, we study how to eliminate the above influence from the environment through communication among agents. The proposed method, called Information Exchange Deep Q-Network (IEDQN), has a learning communication protocol, which makes each local agent pay unbalanced and asymmetric attention to other agents' information. Besides the protocol, each agent has the ability to abstract local information from its own history data for interacting, which means that the communication can avoid the dependent instant information and it is robust to the potential time delay of communication. Specifically, by alleviating the effects of partial observation, experience replay can recover to good performance. We evaluate IEDQN via simulation experiments in the simulation of urban mobility (SUMO) in a traffic grid, and it outperforms the comparative multi-agent RL (MARL) methods in both efficiency and effectiveness. Donghan Xie, Zhi Wang 0001, Chunlin Chen 0001, Daoyi Dong |
IJCNN | 2 |
| 2020 | Look Ahead at the First-mile in Livecast with Crowdsourced Highlight PredictionabstractRecently, data-driven prediction strategies have shown the potential of shepherding the optimization strategies for end viewer's Quality-of-Experience in practical streaming applications. The current prediction-based designs have largely focused on optimizing the last-mile, i.e., viewer-side, which 1) need the real-time feedback from viewers to improve the prediction accuracy; and 2) need quick responses to guarantee the effectiveness of optimization strategies in the future. Thanks to the emerged crowdsourced livecast services, e.g., Twitch.tv, we for the first time exploit the opportunity to realize the long-term prediction and optimization with the assistance derived from the first-mile, i.e., source broadcasters.In this paper, we propose a novel framework CastFlag, which analyzes the broadcasters' operations and interactions, predicts the key events (i.e., highlights), and optimizes the transcoding stage in the corresponding live streams, even before the encoding stage. Taking the most popular eSports gamecast as an example, we illustrate the effectiveness of this framework in the game highlight prediction and transcoding workload allocation. The trace-driven evaluation shows the superiority of CastFlag as it: (1) improves the prediction accuracy over other learning-based approaches by up to 30%; (2) achieves an average of 10% saving of the transcoding latency at less cost. Cong Zhang 0002, Jiangchuan Liu, Zhi Wang 0001, Lifeng Sun |
INFOCOM | 3 |
| 2020 | A dataset for exploring gaze behaviors in text summarizationabstractAutomatic text summarization has been a hot research topic for years. Though most of the existing studies only use the content itself to generate the summaries, researchers believe that an individual's reading behaviors have much to do with the summaries s/he generates, usually regarded as the ground truth. However, such research is limited by the lack of a dataset that provides the connection between people's reading behaviors and the summaries provided by them. This paper fills in this gap by providing a dataset covering 50 individuals' gaze behaviors collected by a high-accurate eye tracking device (that generates 100 gaze points per second) when they are reading 100 articles (from 10 popular categories) and composing the corresponding summaries for each article. Collected in a controlled environment, our dataset with 157 million gaze points in total, provides not only the basic gaze behaviors when different people read an article and compose its corresponding summary, but also the connections between different behavior patterns and the summaries they will provide. We believe such a dataset will be valuable for a wide range of studies, and we also provide sample use cases of the dataset. Weifeng Jiang, Zhi Wang 0001, Lifeng Sun |
MMSys | 4 |
| 2020 | Reinforcement Learning-Based Optimal Sensor Placement for Spatiotemporal ModelingabstractA reinforcement learning-based method is proposed for optimal sensor placement in the spatial domain for modeling distributed parameter systems (DPSs). First, a low-dimensional subspace, derived by Karhunen-Loève decomposition, is identified to capture the dominant dynamic features of the DPS. Second, a spatial objective function is proposed for the sensor placement. This function is defined in the obtained low-dimensional subspace by exploiting the time-space separation property of distributed processes, and in turn aims at minimizing the modeling error over the entire time and space domain. Third, the sensor placement configuration is mathematically formulated as a Markov decision process (MDP) with specified elements. Finally, the sensor locations are optimized through learning the optimal policies of the MDP according to the spatial objective function. The experimental results of a simulated catalytic rod and a real snap curing oven system are provided to demonstrate the feasibility and efficiency of the proposed method in solving the combinatorial optimization problems, such as optimal sensor placement. Zhi Wang 0001, Han-Xiong Li, Chunlin Chen 0001 |
IEEE Trans. Cybern. | 1 |
| 2020 | Unmanned Aircraft System Aided Adaptive Video Streaming: A Joint Optimization ApproachabstractDue to the coverage constraint of a wireless base station, mobile users suffer from the unstable network connection and poor service quality, especially for the prevalent video services. As an alternative solution, an unmanned aerial vehicle (UAV) is able to reach the cell edge and serve ground users (GUs). In this paper, we extend the UAV applications to the more challenging adaptive streaming service over fading channel. First, we decompose the system into different modules, and present mathematical models for each of them, including a trajectory model of the UAV, fading channels between the UAV and GUs, and video streaming utility. Second, we formulate the problem as a non-convex optimization problem by optimizing the UAV trajectory and transmit power allocation, jointly with transmission schedule and rate allocation for multiple users. The objective is to maximize the overall utility while guaranteeing the fairness among multiple users under the UAV energy budget and rate-outage probability constraints. Third, to tackle this problem, we first analyze the relationship between transmission rate and rate-outage probability over the fading channel, and then divide the original problem into three subproblems, which can be solved by leveraging the successive convex approximation technique. Furthermore, an overall iterative algorithm over the three subproblems is proposed to obtain a locally optimal solution by applying the block coordinate descent technique. Finally, through extensive experiments, we demonstrate that the proposed design can achieve almost 30% performance gain in terms of max-min streaming utility for all users, compared with other benchmark schemes. Cheng Zhan, Han Hu 0003, Zhi Wang 0001, Rongfei Fan, Dusit Niyato |
IEEE Trans. Multim. | 3 |
| 2020 | Incremental Reinforcement Learning in Continuous Spaces via Policy Relaxation and Importance WeightingabstractIn this paper, a systematic incremental learning method is presented for reinforcement learning in continuous spaces where the learning environment is dynamic. The goal is to adjust the previously learned policy in the original environment to a new one incrementally whenever the environment changes. To improve the adaptability to the ever-changing environment, we propose a two-step solution incorporated with the incremental learning procedure: policy relaxation and importance weighting. First, the behavior policy is relaxed to a random one in the initial learning episodes to encourage a proper exploration in the new environment. It alleviates the conflict between the new information and the existing knowledge for a better adaptation in the long term. Second, it is observed that episodes receiving higher returns are more in line with the new environment, and hence contain more new information. During parameter updating, we assign higher importance weights to the learning episodes that contain more new information, thus encouraging the previous optimal policy to be faster adapted to a new one that fits in the new environment. Empirical studies on continuous controlling tasks with varying configurations verify that the proposed method achieves a significantly faster adaptation to various dynamic environments than the baselines. Zhi Wang 0001, Han-Xiong Li, Chunlin Chen 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2019 | Towards QoS-Aware Cloud Live Transcoding: A Deep Reinforcement Learning ApproachabstractVideo transcoding is widely adopted in live streaming services to bridge the format and resolution gap between content producers and consumers (i.e., broadcasters and viewers). Meanwhile, the cloud has been recognized as one of the most reliable and cost-effective ways for video transcoding. However, due to the dynamic and uncertainty of the transcoding workloads in live streaming, it is very challenging for cloud service providers to provision computing resources and schedule transcoding tasks while guaranteeing the Service Level Agreement (SLA). To this end, we propose a joint resource provisioning and task scheduling approach for transcoding live streams in the cloud. We adopt Deep Reinforcement Learning (DRL) to train a neural network model for resource provisioning under dynamic workloads. Moreover, we design a QoS-aware task scheduling algorithm that maps transcoding tasks to Virtual Machines (VMs) by considering the real-time QoS requirement. We evaluate our approach with trace-driven experiments and the results demonstrate that our approach outperforms heuristic baselines by up to 89% improvements on average QoS with 4% extra resource overhead at most. Zhengyuan Pang, Lifeng Sun, Tianchi Huang, Zhi Wang 0001, Shiqiang Yang |
ICME | 4 |
| 2019 | Incremental Learning Based Subspace Modeling for Distributed Parameter SystemsabstractIn this paper, a novel incremental learning based subspace modeling method is developed for spatiotemporal modeling of distributed parameter systems (DPSs). First, the streaming snapshots are collected into small batches at a preset time interval in an online mode. The initial batch belongs to the first nominal subspace. Second, the dissimilarity analysis is further utilized to assign each new batch to one of the existing subspaces or a new subspace. Third, the local basis functions corresponding to the assigned subspace is updated or generated through incremental learning of the new batch data. Finally, all the local models are ensembled to approximate the system's dynamics over the whole time-space domain in real-time. The proposed method is tested on a hyperbolic advection system and a one-dimensional diffusion-reaction system. Results demonstrate that the proposed method is superior to the conventional global modeling, and achieves higher modeling accuracy for DPSs. Zhi Wang 0001, Han-Xiong Li |
IJCNN | 1 |
| 2019 | Fine-grained Fitting Experience Prediction: A 3D-slicing Attention ApproachabstractThe comfortableness of fashion items (e.g., footwear) when people actually wear them has become an increasingly important factor in today's fashion experience. However, existing solutions usually only provide general metrics, e.g., a size of a pair of shoes, for people to roughly infer the fitness possibility, failing to tell the details about how much it fits or why it does not fit a person. In this paper, we propose a fine-grained fitting experience prediction framework based on 3D shapes of both fashion items and people's bodies. First, we propose a 3D-slicing sampling method, by extracting a series of parallel slices from an object, to represent the spatial details of the object with a much smaller amount of features. Second, we propose a spatial self-attention based fitness prediction model including a sub-region attention method and a sequence attention method, which can capture users' comfortable preferences for fine-grained regions divided from slices. Our design can capture users' try-on preferences and landmark positions that may or may not fit (e.g., too tight or too loose). Then, we design a multi-position experience module to predict users' fitting experiences, which can help to explore the spatial differences among slices better. Finally, we use subjective experiments over 500 people trying 32 pairs of fashion shoes with detailed places' comfortableness reported in questionnaires to verify our design, which has accuracies of $77.7%$ and $80.9%$ in reporting the comfortableness of tightness and length respectively, and an overall fitness accuracy of $83.6%$. Zhi Wang 0001, Laizhong Cui, Yong Jiang 0001 |
ACM Multimedia | 2 |
| 2019 | AdaCompress: Adaptive Compression for Online Computer Vision ServicesabstractWith the growth of computer vision based applications and services, an explosive amount of images have been uploaded to cloud servers which host such computer vision algorithms, usually in the form of deep learning models. JPEG has been used as the \em de facto compression and encapsulation method before one uploads the images, due to its wide adaptation. However, standard JPEG configuration does not always perform well for compressing images that are to be processed by a deep learning model, e.g., the standard quality level of JPEG leads to 50% of size overhead (compared with the best quality level selection) on ImageNet under the same inference accuracy in popular computer vision models including InceptionNet, ResNet, etc. Knowing this, designing a better JPEG configuration for online computer vision services is still extremely challenging: 1) Cloud-based computer vision models are usually a black box to end-users; thus it is difficult to design JPEG configuration without knowing their model structures. 2) JPEG configuration has to change when different users use it. In this paper, we propose a reinforcement learning based JPEG configuration framework. In particular, we design an agent that adaptively chooses the compression level according to the input image's features and backend deep learning models. Then we train the agent in a reinforcement learning way to adapt it for different deep learning cloud services that act as the \em interactive training environment and feeding a reward with comprehensive consideration of accuracy and data size. In our real-world evaluation on Amazon Rekognition, Face++ and Baidu Vision, our approach can reduce the size of images by 1/2 -- 1/3 while the overall classification accuracy only decreases slightly. Hongshan Li, Zhi Wang 0001, Shutao Xia, Wenwu Zhu 0001 |
ACM Multimedia | 3 |
| 2019 | Content Harvest Network: Optimizing First Mile for Crowdsourced Live StreamingabstractCrowdsourced live streaming (CLS), such as Twitch.tv and Inke.tv, has emerged as an important multimedia application in recent years. Video delivery in such CLS service involves two steps: 1) video uploading-video streaming (i.e., a live channel) generated from a broadcaster is uploaded to the server, which we call the “first mile” network and 2) video distribution-the video streaming is then delivered to viewers in the channel. Today's CLS services usually use conventional content delivery network solutions to address the video distribution problem, while little attention has been paid to improve the video uploading quality. Our measurement study shows that the first mile network causes 17% viewer rebuffers, and some viewers quit the channel once encountering rebuffer. In this paper, we propose a content harvest network (CHN) architecture to address the uploading problem in the CLS service. Specifically, the CHN architecture employs edge devices in the network as relays to receive the streaming uploaded by broadcasters and then forward to the central servers. On one hand, we need to reduce the latency since it is live streaming; on the other hand, we must provide sustainable upload bandwidth. It is challenging to achieve both at the same time, especially in such high dynamic system as CLS. In order to provide global optimal and real-time assignment, we propose a hybrid solution, i.e., centralized and distributed assignment. Specifically, we formulate the centralized relay assignment problem as an optimization problem to achieve both low latency and sustainable bandwidth. To cope with the frequent channel establishments, we use a multi-armed bandit method to characterize the time-variant network condition. Experiment results on a large-scale trace provided by Inke.tv show that our solution can reduce the overall viewer cost by 40% compared to state-of-the-art solutions. The viewers' rebuffer can also be reduced by 50%. Haitian Pang, Zhi Wang 0001, Qinghua Ding, Jiangchuan Liu, Lifeng Sun |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Incremental Spatiotemporal Learning for Online Modeling of Distributed Parameter SystemsabstractAn incremental spatiotemporal learning scheme is proposed for online modeling of distributed parameter systems (DPSs). A novel incremental learning method is developed to recursively update the spatial basis functions and the corresponding temporal model based on the Karhunen-Loève decomposition for time-space separation. The time-space synthesis continually evolves by adding new increment data with more updated information and revising the existing parameters of the dynamic system. In this way, the spatiotemporal structure is inherited and updated efficiently as output data increases over time. The adaptive nature of this evolving structure makes it promising for online modeling of DPSs under streaming data environment. The proposed incremental modeling scheme is evaluated on the classical benchmark of a catalytic rod problem. The simulation results demonstrate the viability and efficiency of the proposed method for online modeling of DPSs. Zhi Wang 0001, Han-Xiong Li |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2018 | Delay-Aware Upload Balancing Cross APs Based on User RelaysabstractRecent years have witnessed the rapid growth of user-generated content, produced by end users and uploaded via edge networks, in which 802.11 wireless LANs (Wi-Fi) serve a dominant fraction of such traffic. However, the user experience for content uploading via Wi-Fi networks, especially in today's metropolises, remains unsatisfactory. In this paper, we address this problem using a delay-aware wireless Access Point (AP) upload balancing solution: content items can be carried by users from one overloaded AP to another AP for uploading before their deadlines, to balance the load between Wi-Fi APs and achieve a high overall upload capacity. Our contributions are as follows. First, we conduct measurement studies to validate that the AP upload balancing framework is feasible in practice. Second, we formulate the delay-aware AP upload balancing problem as a Cache Utility Maximization (CUM) model and propose an efficient algorithm to guarantee a high success upload rate and average AP load utilization. Third, trace-driven experiments are conducted to evaluate our solution. The results show that our solution can achieve significant performance improvements. Miao Zhang 0003, Zhi Wang 0001, Yong Jiang 0001 |
ICC | 3 |
| 2018 | JALAD: Joint Accuracy-And Latency-Aware Deep Structure Decoupling for Edge-Cloud ExecutionabstractRecent years have witnessed a rapid growth of deep-network based services and applications. A practical and critical problem thus has emerged: how to effectively deploy the deep neural network models such that they can be executed efficiently. Conventional cloud-based approaches usually run the deep models in data center servers, causing large latency because a significant amount of data has to be transferred from the edge of network to the data center. In this paper, we propose JALAD, a joint accuracy- and latency-aware execution framework, which decouples a deep neural network so that a part of it will run at edge devices and the other part inside the conventional cloud, while only a minimum amount of data has to be transferred between them. Though the idea seems straightforward, we are facing challenges including i)how to find the best partition of a deep structure; ii)how to deploy the component at an edge device that only has limited computation power; and iii)how to minimize the overall execution latency. Our answers to these questions are a set of strategies in JALAD, including 1)A normalization based in-layer data compression strategy by jointly considering compression rate and model accuracy; 2)A latency-aware deep decoupling strategy to minimize the overall execution latency; and 3)An edge-cloud structure adaptation strategy that dynamically changes the decoupling for different network conditions. Experiments demonstrate that our solution can significantly reduce the execution latency: it speeds up the overall inference execution with a guaranteed model accuracy loss. Hongshan Li, Chenghao Hu, Jingyan Jiang, Zhi Wang 0001, Yonggang Wen 0001, Wenwu Zhu 0001 |
ICPADS | 4 |
| 2018 | MobiRate: Mobility-Aware Rate Adaptation Using PHY Information for Backscatter NetworksabstractIn the past few years, various backscatter nodes have been invented for many emerging mobile applications, such as sports analytics, interactive gaming, and mobile healthcare. Backscatter networks are expected to provide a high-throughput and stable communication platform for those interconnected mobile nodes. Yet, through experiments, we find state-of-the-art rate adaptation methods for backscatter networks share a fundamental limitation of accommodating the hardware diversity of nodes because the common mapping paradigm that chooses the optimal rate based on the radio signal strength indicator (RSSI) or the like is hardly adaptable to hardware-dependent RSSIs. To address this issue, we propose MobiRate (Mobility-aware Rate adaptation) that fully exploits the mobility hints from PHY information and the characteristics of backscatter systems. The key insight is that mobility-hints, like velocity and position, can greatly benefit rate selection and channel probing. Specifically, we introduce a novel velocity-based loss rate estimation method that dynamically re-weighs packets based on time and mobility. In addition, we design a mobility-assisted probing trigger and a new selective-probing mechanism, significantly saving probing time. As MobiRate is fully compatible with the current standard, it is prototyped using a COTS RFID reader and a variety of commercial tags. Our extensive experiments demonstrate that MobiRate achieves up to 3.8x throughput gain over the state-of-the-art methods across a wide range of mobility, channel conditions, and tag types. Wei Gong 0001, Si Chen 0003, Jiangchuan Liu, Zhi Wang 0001 |
INFOCOM | 4 |
| 2018 | Optimizing Personalized Interaction Experience in Crowd-Interactive Livecast: A Cloud-Edge ApproachabstractEnabling users to interact with broadcasters and audience, the crowd-interactive livecast greatly improves viewer's quality of experience (QoE) and attracts millions of daily active users recently. In addition to striking the balance between resource utilization and viewers' QoE met in the traditional video streaming service, this novel service needs to take supererogatory efforts to improve the interaction QoE, which reflects the viewer interaction experience. To tackle this issue, we conduct measurement studies over a large-scale dataset crawled from a representative livecast service provider. We observe that the individual's interaction pattern is quite heterogeneous: only 10% viewers proactively participate in the interaction, and the rest viewers usually watch passively. Incorporating the insight into the emerging cloud-edge architecture, we propose a framework PIECE, which optimizes the Personalized Interaction Experience with Cloud-Edge architecture (PIECE) for intelligent user access control and livecast distribution. In particular, we first devise a novel deep neural network based algorithm to predict users' interaction intensity using the historical viewer pattern. We then design an algorithm to maximize the individual's QoE, by strategically matching viewer sessions and transcoding-delivery paths over cloud-edge infrastructure. Finally, we use trace-driven experiments to verify the effectiveness of PIECE. Our results show that our prediction algorithm outperforms the state-of-the-art algorithms with a much smaller mean absolute error (40% reduction). Furthermore, in comparison with the cloud-based video delivery strategy, the proposed framework can simultaneously improve the average viewers QoE (26% improvement) and interaction QoE (21% improvement), while maintaining a high streaming bitrate. Haitian Pang, Cong Zhang 0002, Fangxin Wang 0001, Han Hu 0003, Zhi Wang 0001, Jiangchuan Liu, Lifeng Sun |
ACM Multimedia | 5 |
| 2018 | Wireless Caching in Large-Scale Edge Access Points: A Local Distributed ApproachabstractToday's mobile users achieve unsatisfactory quality of experience mainly due to the large network distance to the centralized infrastructure. To improve users' experiences, caching at the wireless access points (APs) has been proposed for bringing the contents closer to users. However, the wireless content placement is challenging as the placement is affected by many realistic constraints, such as a large number of APs, interaction among neighboring APs, various local content popularities. In this paper, we study the wireless caching problem, i.e., which contents should be stored by which APs. First, we fulfil these constraints to formulate our problem and introduce an objective function that maximizes the total cache hit rate of all APs. Next, we prove the NP-hardness of the problem and propose a local distributed caching algorithm to address it. Furthermore, we provide a game theoretic perspective on the problem and prove that the proposed algorithm can converge to the Nash Equilibrium in polynomial time. Finally, we perform simulations on a real-world dataset to demonstrate the effectiveness of our algorithm. Ge Ma, Zhi Wang 0001, Jiahui Ye, Wenwu Zhu 0001 |
MobiCom | 2 |
| 2018 | Understanding Gaming Experience in Mobile Multiplayer Online Battle Arena GamesabstractOnline mobile game (OMG) is booming recently, which has driven user expectations for high-quality game service. Therefore, it is crucial for game operators to understand if and how system factors (i.e. network quality metrics such as delay and mobile phone's rendering performance such as frame rate) affect gaming experience and how to optimize resource allocation to improve it. This paper is a first step towards addressing these problems. Despite the rich literature on multimedia services and Quality of Experience (QoE) measurement, the understanding of gaming experience is limited because the game platform shifts from traditional PC end to the mobile end, which has yet to be explored in depth. Based on a large-scale dataset collected from Honour of Kings, the world's top grossing mobile game, we carry out elaborate studies to explore gaming experience from the aspects of user behavior and game quality. Our key findings are as follows. First, user behavior is mainly limited to the game logic itself, such as I of a game is restricted by game rules. Therefore, it cannot well represent gaming experience. Second, among all system factors, the biggest impact on game quality comes from network performance rather than the frame rate of device, which is different from the influence mode in PC games. Third, some context factors, such as AP/BaseStations used by mobile phones, can have an indirect but huge impact on gaming experience. Based on the above observations, we provide insights that can enhance OMG's system from the perspective of resource allocation and real-time gaming experience monitoring. To the best of our knowledge, we are the first to conduct large-scale measurement to study gaming experience of OMGs in the wild. We believe our study is not only crucial for the understanding of gaming experience, but also helpful for game operators to optimize their systems. Chou Mo, Guowei Zhu, Zhi Wang 0001, Wenwu Zhu 0001 |
NOSSDAV | 3 |
| 2018 | Guess your size: A hybrid model for footwear size recommendation
Zhi Wang 0001, Yong Jiang 0001 |
Adv. Eng. Informatics | 2 |
| 2018 | Toward Wi-Fi AP-Assisted Content Prefetching for an On-Demand TV Series: A Learning-Based ApproachabstractThe emergence of smart Wi-Fi access points (AP), which are equipped with huge storage space, opens a new research area on how to utilize these resources at the edge network to improve users' quality of experience (e.g., a short startup delay and smooth playback). One important research interest in this area is content prefetching which predicts and accurately fetches contents ahead of users' requests to shift the traffic away during peak periods. However, in practice, the different video watching patterns among users and the varying network connection status lead to the time-varying server load, which eventually makes the content prefetching problem challenging. To understand this challenge, this paper first performs a large-scale measurement study on users' AP connection and TV series watching patterns using real traces. Then, based on the obtained insights, we formulate the content prefetching problem as a Markov decision process. The objective is to strike a balance between the increased prefetching and storage cost incurred by incorrect prediction and the reduced content download delay because of successful prediction. A learning-based approach is proposed to solve this problem and another three algorithms are adopted as baselines. In particular, first we investigate the performance lower bound by using a random algorithm and the upper bound by using an ideal offline approach. Then, we present a heuristic algorithm as another baseline. Finally, we design a reinforcement learning algorithm that is more practical to work in the online manner. Through extensive trace-based experiments, we demonstrate the performance gain of our design. Remarkably, our learning-based algorithm achieves a better precision and hit ratio (e.g., 80%) with about 70% (resp. 50%) cost saving compared to the random (resp. heuristic) algorithm Wen Hu 0003, Yichao Jin 0002, Yonggang Wen 0001, Zhi Wang 0001, Lifeng Sun |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2018 | Characterizing User Behaviors in Mobile Personal Livecast: Towards an Edge Computing-assisted ParadigmabstractMobile personal livecast (MPL) services are emerging and have received great attention recently. In MPL, numerous and geo-distributed ordinary people broadcast their video contents to worldwide viewers. Different from conventional social networking services like Twitter and Facebook, which have a tolerance for interaction delay, the interactions (e.g., chat messages) in a personal livecast must be in real-time with low feedback latency. These unique characteristics inspire us to: (1) investigate how the relationships (e.g., social links and geo-locations) between viewers and broadcasters influence the user behaviors, which has yet to be explored in depth; and (2) explore insights to benefit the improvement of system performance. In this article, we carry out extensive measurements of a representative MPL system, with a large-scale dataset containing 11M users. In the current costly and limited cloud-based MPL system, which is faced with scalability problem, we find: (1) the long content uploading distances between broadcasters and cloud ingesting servers result in an impaired system QoS, including a high broadcast latency and a frequently buffering events; and (2) most of the broadcasters in MPL are geographically locally popular (the majority of the views come from the same region of the broadcaster), which consume vast computation and bandwidth resources of the clouds and Content Delivery Networks. Fortunately, the emergence of edge computing, which provides cloud-computing capabilities at the edge of the mobile network, naturally sheds new light on the MPL system; i.e., localized ingesting, transcoding, and delivering locally popular live content is possible. Based on these critical observations, we propose an edge-assisted MPL system that collaboratively utilizes the core-cloud and abundant edge computing resources to improve the system efficiency and scalability. In our framework, we consider a dynamic broadcaster assignment to minimize the broadcast latency while keeping the resource lease cost low. We formulate the broadcaster scheduling as a stable matching with migration problem to solve it effectively. Compared with the current pure cloud-based system, our edge-assisted delivery approach reduces the broadcast latency by about 35%. Lei Zhang 0066, Jiangchuan Liu, Zhi Wang 0001, Haitian Pang, Lifeng Sun, Guangling Hou, Kaiyan Chu |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2017 | Joint Request Balancing and Content Aggregation in Crowdsourced CDNabstractRecent years have witnessed a new content delivery paradigm named crowdsourced CDN, in which devices deployed at edge network can prefetch contents and provide content delivery service. Crowdsourced CDN offers high-quality experience to end-users by reducing their content access latency and alleviates the load of network backbone by making use of network and storage resources at millions of edge devices. In such paradigm, redirecting content requests to proper devices is critical for user experience. The uniqueness of request redirection in such crowdsourced CDN lies that: on one hand, the bandwidth capacity of the crowdsourced CDN devices is limit, hence devices located at a crowded place can be easily overwhelmed when serving nearby user requests; on the other hand, contents requested in one device can be significantly different from another one, making request redirection strategies used in conventional CDNs which only aim to balance request loads ineffective. In this paper, we explore request redirection strategies that take both workload balance of devices and content requested by users into consideration. Our contributions are as follows. First, we conduct measurement studies, coving 1.8M users watching 0.4M videos, to understand request patterns in crowdsourced CDN. We observe that the loads of nearby devices can be very different and the contents requested at nearby devices can also be significantly different. These observations lead to our design for request balancing at nearby devices. Second, we formulate the request redirection problem by taking both the content access latency and the content replication cost into consideration, and propose a request balancing and content aggregation solution. Finally, we evaluate the performance of our design using trace-driven simulations, and observe our scheme outperforms the traditional strategy in terms of many metrics, e.g., we observe a content access latency reduction by 50% over traditional mechanisms such as the Nearest/Random request routing scheme. Zhi Wang 0001, Jiangchuan Liu, Lifeng Sun |
ICDCS | 2 |
| 2017 | APRank: Joint mobility and preference-based mobile video prefetchingabstractToday's internet has witnessed a fast growth of mobile video streaming. Different from traditional PC/laptop-based video streaming, mobile video streaming relies on the usage of mobile devices and wireless networks, allowing people to receive video content on the move. The change has challenged traditional video content delivery, which uses centralized infrastructure (e.g., CDN) inside the network for content distribution, in a sense that mobile users (connected to Wi-Fi or cellular networks) encounter large delay and small download speed. One promising solution is to prefetch content in the edge of the network, e.g., on access points (APs). However, it faces the great challenges: 1) It is difficult to prefetch content in such edge APs with limited storage capacity; 2) Users' mobility cross APs affects the content delivery; 3) Popularity of content may change significantly across APs. Previous approaches make mobile video content delivery inefficient without jointly considering these problems. In this paper, we propose an AP-assisted mobile video delivery framework to solve these problems. First, using large-scale measurement studies of users' trajectories and preferences of videos, we reveal that both users' mobility patterns and their intrinsic preferences are important for AP-assisted content delivery. Second, we formulate the AP content prefetching as an optimization problem, and develop an online solution, APRank, to solve it. Third, we evaluate the effectiveness of our design, compared with four baselines, random-based, popularity-based, preference-based and offline prefetching. Ge Ma, Zhi Wang 0001, Minghua Chen 0001, Wenwu Zhu 0001 |
ICME | 2 |
| 2017 | CP-operated dash caching via reinforcement learningabstractIn recent years, Dynamic Adaptive Streaming over HTTP (DASH) has gained momentum as an effective solution for delivering videos on the Internet. This trend is further driven by the deployment of existing HTTP cache infrastructures in DASH systems to reduce the traffic load as well as to serve clients better. However, deploying conventional cache servers in DASH systems still suffers from low cache hit ratio and bitrate oscillations, which makes it challenging for content providers (CPs) to balance the user-perceived quality-of-experience (QoE) and the operating cost in cache-enabled DASH systems. To address this challenge, we propose a CP-operated DASH caching framework to provide good user QoE with low cost. In particular, we first formulate the caching decision problem as a stochastic optimization problem over a finite time horizon. The objective of this problem is to maximize a weighted sum of the user QoE and the operating cost, termed as the utility. Then we design a reinforcement learning based online algorithm which can obtain approximately optimal solution of this problem. Through extensive trace-driven experiments, we show that our approach not only achieves 40% average improvement of the overall utility compared to baseline approaches, but also adapts to the server load. Zhengyuan Pang, Lifeng Sun, Zhi Wang 0001, Wen Hu 0003, Shiqiang Yang |
ICME | 3 |
| 2017 | Beyond the touch: Interaction-aware mobile gamecasting with gazing pattern predictionabstractRecent years have witnessed an explosion of gamecasting applications in the market, in which game players (or gamers in short) broadcast their game scenes in real-time. Such pioneer applications as YouTube Gaming, Twitch, and Mobcrush have attracted a massive number of online broadcasters, and each of them can attract hundreds or thousands of fellow viewers. The growing number however has created significant challenges to the network and end-devices, particularly considering bandwidth- and battery-limited smartphones or tablets are becoming dominating for both gamers and viewers. Yet the unique touch operations of the mobile interface offer opportunities, too. In this paper, our crowdsourced measurement reveals that strong associations exist between the gamers' touch interactions and the viewers' gazing patterns. Motivated by this, we present a novel interaction-aware optimization framework to improve the energy utilization and stream quality for mobile gamecasting (MGC). Our framework incorporates a touch-assisted prediction module to extract association rules for gazing pattern prediction and a tile-based optimization module to utilize energy on mobile devices efficiently. Trace-driven simulations illustrate the effectiveness of our framework in terms of energy consumption and streaming quality. Our user study experiments also demonstrate much improved (3%-13%) quality satisfaction than the state-of-the-art solution with similar network resources. Cong Zhang 0002, Qiyun He, Jiangchuan Liu, Zhi Wang 0001 |
INFOCOM | 4 |
| 2017 | Understanding viewing engagement and video quality in a large-scale mobile video systemabstractWith the advances in wireless communication and the growing popularity of mobile devices, it has become rather normal to watch videos using mobile devices such as mobile phones and tablets. In order to improve user engagement in mobile video sessions, it is important for content providers and CDN operators to understand the correlations between video quality and Quality of Experience (QoE) of mobile users. In this paper, we measure how video quality metrics affect user engagement, and identify several counter-intuitive relationships, based on the large-scale trace data from a famous IPTV provider. We discover the internal causes of these relationships later, which give us new insights about how we improve user engagement. Based on our new insights, we further design a user engagement prediction framework to let content providers predict how long viewers will remain in video sessions with specific video quality metrics. Finally, we verify the effectiveness of the prediction framework by the real-world data. Zitao Chen 0003, Laizhong Cui, Yong Jiang 0001, Zhi Wang 0001 |
ISCC | 4 |
| 2017 | MUSA: Wi-Fi AP-assisted video prefetching via Tensor LearningabstractDriven by the exponentially increasing amount of mobile video traffic, caching videos closer to the end users has become an appealing solution to reduce the traffic through the backbone network while improving users' perceived quality-of-experience (e.g., better video quality and reduced service delay). This research interest has been gaining lots of momentums due to the emergence of smart Access Points (APs), which are equipped with large storage space (several GBs). To address the “small population” problem involved in the prefetching at the edge, we propose to prefetch videos to APs ahead of users' requests via tensor learning: We first adopt the weighted tensor model to mine the hidden semantic pattern to characterize both users' preference for different types of videos and the dynamic video popularity over time; Then, based on the resulting low-dimension matrixes generated by the tensor factorization, we adopt an exponential smoothing model to capture the temporal pattern to predict users' propensity to unwatched videos; Finally, based on the predicted video popularity, we proactively replicate videos from the original CDN server to the APs at the edge. Through trace-driven simulations, we show that the proposed prefetching solution can outperform the baseline algorithms: compared with the SVD-based prefetching strategy, our design achieves a better hit ratio (e.g., surpassing about 10%) and accuracy (e.g., surpassing about 15%); compared with the history based strategy, our design also have about 40% (resp. 20%) improvement in terms of hit ratio (resp. accuracy). Wen Hu 0003, Zhi Wang 0001, Peng Wang 0012, Yonggang Wen 0001, Kaiyan Chu, Lifeng Sun |
IWQoS | 3 |
| 2017 | When Cloud Meets Uncertain Crowd: An Auction Approach for Crowdsourced Livecast TranscodingabstractIn the emerging crowd sourced live cast services, numerous amateur broadcasters live stream their video contents to worldwide viewers and constantly interact with them through chat messages. Live video contents are transcoded into multiple quality versions to better service viewers with different network and device configurations. Cloud computing becomes a natural choice to handle these computational intensive tasks due to its elasticity and the "pay-as-you-go" billing model. However, given the significantly large number of concurrent channel numbers and the diverse viewer geo-distributions in this new crowd sourced live cast service, even the cloud becomes significantly expensive to cover the whole community and inadequate in fulfilling the latency requirement. In this paper, after observing the abundant computational resources residing in end viewers, we propose a Cloud-Crowd collaborative system, C2, which combines end viewers with cloud to perform video transcoding in a cost-efficient way. To quantify the heterogeneity and uncertainty of viewers and pass the asymmetric information barrier, we incorporate statistical descriptions into our bidding language and design truthful auctions to recruit stable viewers with appropriate incentives. We further tailor redundancy strategies for workloads with different Quality of Service requirements to improve the stability of our system. Desirable economic properties, like social efficiency, ex-post incentive compatibility, individual rationality, are proved to be guaranteed in our studied scenarios. Using traces captured from the popular Twitch platform, we show that C2 achieves up to 93% more cost saving than a pure cloud-based solution, and significantly outperforms other baseline approaches in both social welfare and system stability. Yifei Zhu 0001, Jiangchuan Liu, Zhi Wang 0001, Cong Zhang 0002 |
ACM Multimedia | 3 |
| 2017 | Understanding Performance of Edge Prefetching
Zhengyuan Pang, Lifeng Sun, Zhi Wang 0001, Yuan Xie 0005, Shiqiang Yang |
MMM (1) | 3 |
| 2017 | A Dataset for Exploring User Behaviors in VR Spherical Video StreamingabstractWith Virtual Reality (VR) devices and content getting increasingly popular, understanding user behaviors in virtual environment is important for not only VR product design but also user experience improvement. In VR applications, the head movement is one of the most important user behaviors, which can reflect a user's visual attention, preference, and even unique motion pattern. However, to the best of our knowledge, no dataset containing this information is publicly available. In this paper, we present a head tracking dataset composed of 48 users (24 males and 24 females) watching 18 sphere videos from 5 categories. We carefully record how users watch the videos, how their heads move in each session, what directions they focus, and what content they can remember after each session. Based on this dataset, we show that people share certain common patterns in VR spherical video streaming, which are different from conventional video streaming. We believe the dataset can serve good resource for exploring user behavior patterns in VR applications. Chenglei Wu, Zhihao Tan, Zhi Wang 0001, Shiqiang Yang |
MMSys | 3 |
| 2017 | Characterizing User Behaviors in Mobile Personal LivecastabstractMobile personal livecast (MPL) services are emerging and have received great attention recently. Unlike traditional livecast services with commercial content providers (e.g., live TV), the live contents in MPL are crowdsourced from and consumed among geo-distributed individuals. Although there exist typical social relationships in MPL (i.e., follower-followee), different from conventional social networking services like Twitter and Facebook, which have much of a tolerance for interaction delay, the interactions in MPL must be in real-time. These unique characteristics intrigue us to investigate how the relationships (e.g., social links and geo-locations) between viewers and broadcasters influence the user behaviors, which has yet to be explored in depth. In this paper, we carry out extensive measurements of Inke, one of the most popular MPL providers, with a large-scale dataset containing 11M users. Our key findings are as follows. First, compared with traditional livecast services, the user interests shift much more frequently and the average viewing duration is considerably shorter in MPL. Second, the existence of social relationships significantly strengthens viewer stickiness---followers dedicating longer viewing time (contributing 81% of the total viewing time) and being 2x more patient when suffering poor network connectivity than non-followers. Third, most of the broadcasts in MPL are geographically local-popular (the majority of the views come from the same region of the broadcaster). Based on these critical observations, we provide insights that can enhance the MPL system design from the perspectives of efficient resource allocation and envision a future MPL framework that collaboratively utilizes the cloud and edge computing resources to improve efficiency and scalability for Inke-like services. Lei Zhang 0066, Jiangchuan Liu, Zhi Wang 0001, Guangling Hou, Lifeng Sun |
NOSSDAV | 4 |
| 2017 | An Exponential Time-Aware Recommendation Model for Mobile Notification Services
Chenglin Zeng, Laizhong Cui, Zhi Wang 0001 |
PAKDD (2) | 3 |
| 2017 | Understanding Performance of Edge Content Caching for Mobile Video StreamingabstractToday's Internet has witnessed an increase in the popularity of mobile video streaming, which is expected to exceed 3/4 of the global mobile data traffic by 2019. To satisfy the considerable amount of mobile video requests, video service providers have been pushing their content delivery infrastructure to edge networks-from regional content delivery network (CDN) servers to peer CDN servers (e.g., smartrouters in users' homes)-to cache content and serve users with storage and network resources nearby. Among the edge network content caching paradigms, Wi-Fi access point caching and cellular base station caching have become two mainstream solutions. Thus, understanding the effectiveness and performance of these solutions for large-scale mobile video delivery is important. However, the characteristics and request patterns of mobile video streaming are unclear in practical wireless network. In this paper, we use real-world data sets containing 50 million trace items of nearly 2 million users viewing more than 0.3 million unique videos using mobile devices in a metropolis in China over two weeks, not only to understand the request patterns and user behaviors in mobile video streaming, but also to evaluate the effectiveness of Wi-Fi and cellular-based edge content caching solutions. To understand the performance of edge content caching for mobile video streaming, we first present temporal and spatial video request patterns, and we analyze their impacts on caching performance using frequency-domain and entropy analysis approaches. We then study the behaviors of mobile video users, including their mobility and geographical migration behaviors, which determine the request patterns. Using trace-driven experiments, we compare strategies for edge content caching, including least recently used (LRU) and least frequently used (LFU), in terms of supporting mobile video requests. We reveal that content, location, and mobility factors all affect edge content caching performance. Moreover, we design an efficient caching strategy based on the measurement insights and experimentally evaluate its performance. The results show that our design significantly improves the cache hit rate by up to 30% compared with LRU/LFU. Ge Ma, Zhi Wang 0001, Miao Zhang 0003, Jiahui Ye, Minghua Chen 0001, Wenwu Zhu 0001 |
IEEE J. Sel. Areas Commun. | 2 |
| 2017 | CrowdNavi: Demystifying Last Mile Navigation With Crowdsourced Driving InformationabstractWith detailed digital map of the transport network and even real-time traffic, today's navigation services provide good quality routes in the major route level. Once entering the last mile near the destination, they unfortunately can be ineffective and, instead, local drivers often have a better understanding of the routes there. With the deep penetration of 3G/4G mobile networks, drivers today are well connected anytime and anywhere; they can readily access information from the Internet and share information to the driver's community. This motivates our design of CrowdNavi, a complementary service to existing navigation systems, seeking to combat the last mile puzzle. CrowdNavi collects the crowdsourced driving information from users to identify their local driving patterns, and recommend the best local routes for users to reach their destinations. In this paper, we present the architectural design of CrowdNavi and identifies the unique challenges therein, particularly on identifying the last segment in a route from the crowdsourced driving information and navigate drivers through the last segment. We offer a complete set of algorithms to identify the last segment from the drivers' trajectories, scoring the landmark, and locating best routes with user preferences. We then present effective navigation algorithm to locate the best route along the landmarks for the last segment. We further realize the potential risks of attacks in crowdsourced systems and develop a multisensor cross-validation method against them. We have implemented the CrowdNavi app on Android mobile OS, and have examined its performance under various circumstances. The experimental results demonstrate its superiority in navigating drivers in the last segment toward the destination. Xiaoyi Fan 0001, Jiangchuan Liu, Zhi Wang 0001, Yong Jiang 0001, Xue (Steve) Liu |
IEEE Trans. Ind. Informatics | 3 |
| 2017 | ACM TIST Special Issue on Data-Driven Intelligence for Wireless NetworkingabstractNo abstract available. Wenwu Zhu 0001, Jean C. Walrand, Yike Guo, Zhi Wang 0001 |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2017 | Propagation- and Mobility-Aware D2D Social Content ReplicationabstractMobile online social network services have seen rapid expansion; thus, the corresponding huge amounts of user-generated social media contents propagating between users via social connections have significantly challenged the traditional content delivery paradigm. First, replicating all the contents generated by users to edge servers that well “fit” the receivers becomes difficult due to limited bandwidth and storage capacities. Motivated by device-to-device (D2D) communication, which allows users with smart devices to transfer content directly, we propose replicating bandwidth-intensive social contents in a device-to-device manner. Based on large-scale measurement studies on social content propagation and user mobility patterns in edge-network regions, we observe the following: (1) Device-to-device replication can significantly help users download social contents from neighboring peers. (2) Both social propagation and mobility patterns affect how contents should be replicated. (3) The replication strategies depend on regional characteristics (e.g., how users move across regions). Using these measurement insights, we propose a propagationand mobility-aware content replication strategy for edge-network regions, in which social contents are assigned to users in edge-network regions according to a joint consideration of social graphs, content propagation, and user mobility. We formulate the replication scheduling as an optimization problem and design a distributed algorithm using only historical, local, and partial information to solve it. Trace-driven experiments further verify the superiority of our proposal: compared with conventional pure-movement-based and popularity-based approaches, our design can significantly improve (2 - 4-fold improvement) the amount of social content successfully delivered via device-to-device replication. Zhi Wang 0001, Lifeng Sun, Miao Zhang 0003, Haitian Pang, Erfang Tian, Wenwu Zhu 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2017 | Cost-Effective Low-Delay Design for Multiparty Cloud Video ConferencingabstractMultiparty cloud video conferencing architecture has been recently advocated to exploit rich computing and bandwidth resources in the cloud to effectively improve video conferencing performance. As a typical design in this architecture, multiple agents, i.e., virtual machines, are deployed in different cloud sites, and users are assigned to the agents. Then, the users communicate through the agents, and the agents might transcode the recorded videos given the heterogeneities among devices in terms of hardware specification and connectivity. In this architecture, two critical and nontrivial challenges are: 1) assigning users to agents to reduce the operational cost and the user-to-user conferencing delay and 2) identifying best agents to perform transcoding tasks, taking into account the heterogeneous bandwidth and processing availabilities. To address these challenges, we cast a joint problem of user-to-agent assignment and transcoding-agent selection. The ultimate objective is to simultaneously minimize the cost of the service provider and the conferencing delay. The problem is combinatorial in nature, which belongs to the NP-hard node assignment problems. We leverage the Markov approximation framework and devise an adaptive parallel algorithm that finds a close-to-optimal solution to our problem with a bounded performance guarantee. To evaluate the performance of our solution, we implement a prototype video conferencing system and carry out trace-driven experiments. In a set of largescale experiments using PlanetLab traces, our solution decreases the operational cost by 77% and simultaneously yields lower conferencing delay compared with an existing alternative. Mohammad Hajiesmaili, Lok To Mak, Zhi Wang 0001, Chuan Wu 0001, Minghua Chen 0001, Ahmad Khonsari |
IEEE Trans. Multim. | 3 |
| 2017 | Social-Aware Video Recommendation for Online Social GroupsabstractGroup recommendation plays a significant role in today's social media systems, where users form social groups to receive multimedia content together and interact with each other, instead of consuming the online content individually. Limitations of traditional group recommendation approaches are as follows. First, they usually infer group members’ preferences by their historical behaviors, failing to captureinactiveusers’ preferences from the sparse historical data. Second, relationships between group members are not studied by these approaches, which fail to capture the inherent personality of members in a group. To address these issues, we propose a social-aware group recommendation framework that jointly utilizes both social relationships and social behaviors to not only infer a group's preference, but also model thetoleranceandaltruismcharacteristics of group members. Based on the observation that thefollowingrelationship in the online social network reflects common interests of users, we propose a group preference model based on external experts of group members. Furthermore, we model users’ tolerance (willingness to receive content not preferred) and altruism (willingness to receive content preferred by friends). Finally, based on the group preference model, we design recommendation algorithms for users under different social contexts. Experimental results demonstrate the effectiveness of our approach, which significantly improves the recommendation accuracy against traditional approaches, especially in the cases of inactive group members. Lifeng Sun, Zhi Wang 0001, H. Vicky Zhao, Wenwu Zhu 0001 |
IEEE Trans. Multim. | 3 |
| 2017 | Exploring Viewer Gazing Patterns for Touch-Based Mobile GamecastingabstractRecent years have witnessed an explosion of gamecasting applications, in which game players (or gamers in short) broadcast game playthroughs by their personal devices in real time. Such pioneer platforms, such as YouTube Gaming, Twitch, and Mobcrush, have attracted a massive number of online broadcasters, and each of them can have hundreds or thousands of fellow viewers. The growing number, however, has created significant challenges to the network and end-devices, particularly considering that bandwidth- and battery-limited smartphones or tablets are becoming dominating for both gamers and viewers. Yet the unique touch operations of the mobile interface offer opportunities, too. In this paper, our measurements based on the real traces from gamers and viewers reveal that strong associations exist between the gamers' touch interactions and the viewers' gazing patterns. Motivated by this, we present a novel interaction-aware optimization framework to improve the energy utilization and stream quality for mobile gamecasting. Our framework incorporates a touch-assisted prediction module to extract association rules for gazing pattern prediction and a tilebased optimization module to utilize energy on mobile devices efficiently. Trace-driven simulations illustrate the effectiveness of our framework in terms of energy consumption and stream quality. Our user study experiments also demonstrate much improved (3%-13%) quality satisfaction over the state-of-the-art solution with similar network resources. Cong Zhang 0002, Qiyun He, Jiangchuan Liu, Zhi Wang 0001 |
IEEE Trans. Multim. | 4 |
| 2016 | Towards Network-Failure-Tolerant Web Content Delivery: A Path-Aware Peer-Assisted ApproachabstractPopularly used to distribute a variety of multimedia content items in today's Internet, HTTP-based web content delivery still suffers from various content delivery failures, including server failures [1], network failures [2] and routing failures [3]. Hindered by the expensive deployment cost, the conventional CDN can not deploy as many edge servers as possible to successfully deliver content items to all users under these delivery failures. In this paper, we propose a joint CDN and peer-assisted web content delivery framework to address the delivery failure problem. Different from conventional peer- assisted approaches for web content delivery, which mainly focus on alleviating the CDN servers' bandwidth load, we study how to use a browser-based peer-assisted scheme, namely WebRTC, to resolve content delivery failures. To this end, we carry out large-scale measurement studies on how users access and view webpages. Our measurement results demonstrate the challenges (e.g., peers stay on a webpage extremely short) that can not be directly solved by conventional P2P strategies, and some important webpage viewing patterns (e.g., predictability of users' webpage viewing time). Due to these unique characteristics, WebRTC peers open up new possibilities for helping the web content delivery, coming with the problem of how to utilize the dynamic resources efficiently. We formulate the peer selection that is the critical strategy in our framework, as an optimization problem, and design a heuristic algorithm based on the mea- surement insights to solve it. Our simulation experiments driven by the traces from Tencent QZone, one of the most popular social service platforms in China, demonstrate the effectiveness of our design: compared with non-peer-assisted strategy and random peer selection strategy, our design significantly improves the successful relay ratio of web content items under network failures, e.g., our design improves the content download ratio up to 60% even when users located in a particular region (e.g., city) where none can connect to the regional CDN server. Wen Hu 0003, Zhi Wang 0001, Lifeng Sun |
GLOBECOM | 2 |
| 2016 | Crowdsourced Live Streaming over Aggregated Edge NetworksabstractRecent years have witnessed a dramatic increase of user-generated video services. In such user-generated video services, crowdsourced live streaming (e.g., Periscope, Twitch) has significantly challenged today's content delivery infrastructure: today's edge networks (e.g., 4G, Wi-Fi) have limited uplink capacity support, making high-bitrate live streaming over such links fundamentally impossible. In this paper, we propose to let broadcasters (i.e., users who generate the video) upload crowdsourced video streams using aggregated network resources from multiple edge networks. There are several challenges in the proposal: First, how to design a framework that aggregates bandwidth from multiple edge networks? Second, how to make this framework transparent to today's crowdsourced live stream- ing services? Third, how to maximize the streaming quality for the whole system? We design a multi-objective and deployable bandwidth aggregation system BASS to address these challenges: (1) We propose an aggregation framework transparent to today's crowdsourced live streaming services, using an edge proxy box and aggregation cloud paradigm; (2) We dynamically allocate geo- distributed cloud aggregation servers to enable MPTCP (i.e., multi- path TCP), according to location and network characteristics of both broadcasters and the original streaming servers; (3) We maximize the overall performance gain for the whole system, by matching streams with the best aggregation paths. Chenglei Wu, Zhi Wang 0001, Jiangchuan Liu, Shiqiang Yang |
GLOBECOM | 2 |
| 2016 | User Mapping Strategies in Multi-Cloud Streaming: A Data-Driven ApproachabstractUsing content delivery networks (CDNs) for video distribution has become a de facto approach for today's video streaming, due to the easy usage and good scalability. Today, it has become a norm rather than an exception for video providers to hire multiple cloud CDNs for their video services in a pay-per-use manner, to not only serve users at different locations, but also reduce the operation costs. Given the multiple CDNs and their peering servers at many different locations, mapping a user to an edge CDN server has become a critical decision that can affect the quality of experience (QoE) of users. Conventional user mapping strategies are generally rule-based, e.g., assigning users to CDN servers according to only their locations or ISPs, which cannot guarantee any QoE. In this paper, we first propose to use a data-driven approach to study factors determining the streaming QoE in the multi-cloud CDN paradigm. Our findings suggest that the streaming QoE is affected by a combination of not only network factors but also user factors including their preference of video content. Then, we design a machine learning based predictive model to capture the QoE given the network conditions and user preference. Finally, we formulate the user mapping problem as an optimization problem and design algorithms to solve it: our algorithms identify users whose QoE are mostly affected by QoS and assign users to CDN servers so that the overall QoE can be maximized. Trace-driven experiments further verify the effectiveness of our design. Guowei Zhu, Chou Mo, Zhi Wang 0001, Wenwu Zhu 0001 |
GLOBECOM | 3 |
| 2016 | Understanding the Power of Smartrouter-Based Peer CDN for Video StreamingabstractRecent years have witnessed a new video delivery paradigm: smartrouter-based video delivery network, which is enabled by smartrouters deployed at users' homes, together with the conventional video servers deployed in the datacenters. Recently, ChinaCache (a largest CDN provider) and Youku (a video service provider using smartrouters to assist video delivery) announced their cooperation to create a new paradigm of content delivery based on householders' network resources. This new paradigm is different from the conventional peer-to-peer (P2P) approach, because such dedicated smartrouters are inherently operated by the centralized video service provides in a coordinative manner. It is intriguing to study the strategies, performance and potential impact on the content delivery ecosystem of such peer CDN systems. In this paper, we study the Youku peer CDN, which has deployed over 300K smartrouter devices for its video streaming. In our measurement, 78K videos were investigated and 3TB traffic has been analyzed, over controlled peer nodes and players. Our contributions are the following measurement insights. First, a global replication and caching strategy is essential for the peer CDN systems, and proactively scheduling replication and caching on a daily basis can guarantee their performance. Second, such peer CDN deployment can itself form an effective QoS monitoring sub-system, which can be used for fine-grained user request redirection. We also provide our analysis on the performance issues and potential improvement to the peer CDN systems. Zhi Wang 0001, Lifeng Sun |
ICCCN | 2 |
| 2016 | Online influence maximization in non-stationary Social NetworksabstractSocial networks have been popular platforms for information propagation. An important use case is viral marketing: given a promotion budget, an advertiser can choose some influential users as the seed set and provide them free or discounted sample products; in this way, the advertiser hopes to increase the popularity of the product in the users' friend circles by the world-of-mouth effect, and thus maximizes the number of users that information of the production can reach. There has been a body of literature studying the influence maximization problem. Nevertheless, the existing studies mostly investigate the problem on a one-off basis, assuming fixed known influence probabilities among users, or the knowledge of the exact social network topology. In practice, the social network topology and the influence probabilities are typically unknown to the advertiser, which can be varying over time, i.e., in cases of newly established, strengthened or weakened social ties. In this paper, we focus on a dynamic non-stationary social network and design a randomized algorithm, RSB, based on multi-armed bandit optimization, to maximize influence propagation over time. The algorithm produces a sequence of online decisions and calibrates its explore-exploit strategy utilizing outcomes of previous decisions. It is rigorously proven to achieve an upper-bounded regret in reward and applicable to large-scale social networks. Practical effectiveness of the algorithm is evaluated using real-world datasets, which demonstrates that our algorithm outperforms previous stationary methods under non-stationary conditions. Yixin Bao, Zhi Wang 0001, Chuan Wu 0001, Francis C. M. Lau 0001 |
IWQoS | 3 |
| 2016 | More is Better? Measurement of MPTCP Based Cellular Bandwidth Aggregation in the Wildabstract4G/3G Networks have been widely deployed around the world to provide high wireless bandwidth for mobile users. However, the achievable 3G/4G bandwidth is still much lower than their theoretic maximum. Signal strengths and available backhaul capacities may vary significantly at different locations and times, often leading to unsatisfactory performance. Band-width aggregation, which uses multiple interfaces concurrently for data transfer, is a readily deployable solution. Specifically, Multi-Path TCP (MPTCP) has been advocated as a promising approach for leveraging multiple source-destination paths simultaneously in the transport layer. In this paper, we investigate the efficiency of an MPTCP-based bandwidth aggregation frame-work based on extensive measurements. In particular, we evaluate the gain for bandwidth aggregation across up to 4 cellular operators' networks, with respect to factors such as time, user location, data size, aggregation proxy location and congestion control algorithm. Our measurement studies reveal that (1) bandwidth aggregation in general improves the cellular network bandwidth experienced by mobile users, but the performance gain is significant only for bandwidth-intensive delay-tolerant flows, (2) the effectiveness of aggregation depends on many network factors, including QoS of individual cellular interfaces and the location of aggregation proxy, (3) contextual factors, including the time of day and the mobility of a user, also affect the aggregation performance. Zhixiong Niu, Zhi Wang 0001, Hong Xu 0001, Chuan Wu 0001, Francis C. M. Lau 0001 |
MASS | 2 |
| 2016 | Understanding content placement strategies in smartrouter-based peer video CDNabstractRecent years have witnessed a new video delivery paradigm: smartrouter-based peer video content delivery network, which is enabled by smartrouters deployed at users' homes. ChinaCache (one of the largest CDN providers in China) and Youku (a video provider using smartrouters to assist video delivery) announced their cooperation in 2015, to create a new paradigm of content delivery based on householders' network resources [2]. This new paradigm is different from the conventional peer-to-peer (P2P) approach, because millions of dedicated smartrouters are operated by the centralized video service providers in a coordinative manner. Thus it is intriguing to study the content placement strategies used in a smartrouter-based content delivery system, as well as its potential impact on the content delivery ecosystem. In this paper, we carry out measurement studies of Youku's peer video CDN, who has deployed over 300K smartrouter devices for its video delivery. In our measurement studies, 104K videos were investigated and 4TB traffic has been analyzed, over controlled smartrouter nodes and players. Our measurement insights are as follows. First, a global content replication strategy is essential for the peer CDN systems. Second, such peer CDN deployment itself can form an effective sub-system for end-to-end QoS monitoring, which can be used for fine-grained request redirection (e.g., user-level) and content replication. We also show our analysis on the performance limitations and propose potential improvements to the peer CDN systems. Zhi Wang 0001, Lifeng Sun |
NOSSDAV | 2 |
| 2016 | Edge Video CDN: A Wi-Fi Content Hotspot Solution
Wen Hu 0003, Zhi Wang 0001, Lifeng Sun |
J. Comput. Sci. Technol. | 2 |
| 2016 | Dispersing Instant Social Video Service Across Multiple CloudsabstractInstant social video sharing which combines the online social network and user-generated short video streaming services, has become popular in today’s Internet. Cloud-based hosting of such instant social video contents has become a norm to serve the increasing users with user-generated contents. A fundamental problem of cloud-based social video sharing service is that users are located globally, who cannot be served with good service quality with a single cloud provider. In this paper, we investigate the feasibility of dispersing instant social video contents to multiple cloud providers. The challenge is that inter-cloud socialpropagationis indispensable with such multi-cloud social video hosting, yet such inter-cloud traffic incurs substantial operational cost. We analyze and formulate the multi-cloud hosting of an instant social video system as an optimization problem. We conduct large-scale measurement studies to show the characteristics of instant social video deployment, and demonstrate the trade-off between satisfying users with their ideal cloud providers, and reducing the inter-cloud data propagation. Our measurement insights of the social propagation allow us to propose a heuristic algorithm with acceptable complexity to solve the optimization problem, by partitioning a propagation-weighted social graph in two phases: a preference-aware initial cloud provider selection and a propagation-aware re-hosting. Our simulation experiments driven by real-world social network traces show the superiority of our design. Zhi Wang 0001, Baochun Li, Lifeng Sun, Wenwu Zhu 0001, Shiqiang Yang |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2015 | Responsive multipath TCP in SDN-based datacentersabstractA basic need in datacenter networks is to provide high throughput for large flows such as the massive shuffle traffic flows in a MapReduce application. Multipath TCP (MPTCP) has been investigated as an effective approach toward this goal, by spreading one TCP flow onto multiple paths. However, the current MPTCP implementation has two major limitations: (1) a fixed number of subflows are used without reacting to the actual traffic condition; (2) the routing of subflows of a multipath TCP connection relies heavily on the ECMP-based random hashing. The former may lead to a waste of both the server and network resources, while the latter can cause throughput degradation when multiple subflows collide on the same path. This paper proposes a responsive MPTCP system to resolve the two limitations simultaneously. Our system employs a centralized controller for intelligent subflow route calculation and a monitor running on each server for actively adjusting the number of subflows. Working in synergy, the two modules enable MPTCP flows to respond to the traffic conditions and pursue high throughput on the fly, at very low computation and messaging overhead. NS3-based experiments show that our system achieves satisfactory throughput with less resource overhead, or better throughput at similar amounts of overhead, as compared to common alternatives. Jingpu Duan, Zhi Wang 0001, Chuan Wu 0001 |
ICC | 2 |
| 2015 | Guyot: a hybrid learning- and model-based RTT predictive approachabstractKnowing the Round-Trip Time (RTT) between a client and a server is important for an online interactive multi-media service to provide satisfactory quality of user experience. Due to the intrinsic dynamics of the topology and routing strategies in the Internet, it is however challenging to predict the RTT accurately from limited information, e.g., only the IP pair of the client and server. To address this challenge, we propose Guyot, a hybrid learning- and model-based approach to predict RTT, which requires significantly smaller amount of data to be collected than traditional approaches, while achieving a similar prediction accuracy. Our design is based on a large-scale measurement study from a content provider's perspective. Based on an information gain analysis, we design a hybrid RTT prediction approach involving two types of predictions: (1) Learning-based prediction: We train a decision tree to predict RTT between IP pairs with large geographic distance, requiring only a small set of features to be collected. (2) Model-based prediction: We use a model-based framework to predict RTT between IP pairs with small distance, providing an accurate RTT prediction over time. By strategically dividing RTT prediction tasks to these two types according to the distance of the inferred geo-locations of the IPs, our prediction approach can scale with satisfactory accuracy. Our experiments further confirm the superiority of our design. Wen Hu 0003, Zhi Wang 0001, Lifeng Sun |
ICC | 2 |
| 2015 | Cost-Effective Low-Delay Cloud Video ConferencingabstractThe cloud computing paradigm has been advocated in recent video conferencing system design, which exploits the rich on-demand resources spanning multiple geographic regions of a distributed cloud, for better conferencing experience. A typical architectural design in cloud environment is to create video conferencing agents, i.e., Virtual machines, in each cloud site, assign users to the agents, and enable inter-user communication through the agents. Given the diversity of devices and network connectivities of the users, the agents may also transcode the conferencing streams to the best formats and bitrates. In this architecture, two key issues exist on how to effectively assign users to agents and how to identify the best agent to perform a Transco ding task, which are nontrivial due to the following: (1) the existing proximity-based assignment may not be optimal in terms of inter-user delay, which fails to consider the whereabouts of the other users in a conferencing session, (2) the agents may have heterogeneous bandwidth and processing availability, such that the best Transco ding agents should be carefully identified, for cost minimization while best serving all the users requiring the transcoded streams. To address these challenges, we formulate the user-to-agent assignment and Transco ding-agent selection problems, which targets at minimizing the operational cost of the conferencing provider while keeping the conferencing delay low. The optimization problem is combinatorial in nature and difficult to solve. Using Markov approximation framework, we design a decentralized algorithm that provably converges to a bounded neighborhood of the optimal solution. An agent ranking scheme is also proposed to properly initialize our algorithm so as to improve its convergence. The results from a prototype system implementation show that our design in a set of Internet-scale scenarios reduces the operational cost by 77% as compared to a commonly-adopted alternative, while simultaneously yielding lower conferencing delays. Mohammad Hajiesmaili, Lok To Mak, Zhi Wang 0001, Chuan Wu 0001, Minghua Chen 0001, Ahmad Khonsari |
ICDCS | 3 |
| 2015 | WINET: Indoor white space network designabstractThe Federal Communications Commission (FCC) released the final rule to approve of TV white spaces (TVWS), i.e., locally vacant TV channels, for unlicensed use in 2010. This TV spectrum will mitigate the shortage of wireless spectrum resources and provide opportunities for new applications. TVWS differ from the conventional Wi-Fi spectrum in three aspects: spectrum fragmentation, spatial variation, and temporal variation. These differences make the network design over TVWS challenging and fundamentally different from Wi-Fi networks. While most prior works on TVWS network design focused on outdoor large-area scenario, the important indoor scenario is largely open for investigation. In this paper, we present WINET (for White-space Indoor NETwork), the first design framework for indoor multi-AP white space network. We optimize AP placement, spectrum allocation, and AP association. Spectrum fragmentation, spatial variation, and temporal variation are all tackled in our network design. We build a test-bed and conduct extensive measurements inside an office building across four months to obtain real-world traces. Experimental results show that WINET can increase AP coverage area by an average of 62.2% and obtain 67.9% higher system throughput while achieving fairness among users as compared to alternative approaches. Minghua Chen 0001, Zhi Wang 0001 |
INFOCOM | 4 |
| 2015 | Path-aware peer-assisted web content delivery against network failuresabstractPopularly used to distribute a variety of multimedia contents in today's Internet, HTTP-based web content delivery still suffers from failures occurring both in the network and at the servers. Limited by the deployment of edge servers, it is hard for a CDN to successfully deliver contents to all users under these network failures. Different from conventional peer-assisted approaches for web content delivery, which mainly focus on alleviating the CDN servers' bandwidth load, we study how to use a browser-based peer-assisted scheme to resolve content delivery failures. Based on large-scale measurement studies on how users access and view webpages, we observe the challenges (e.g., peers stay on a webpage extremely short) that cannot be directly solved by conventional P2P strategies, and some important webpage viewing patterns. Due to these unique characteristics, WebRTC peers give us a novel way to deliver web content, coming with the problem of how to utilize the dynamic resources efficiently. We formulate the peer selection as an optimization problem, and design a heuristic algorithm based on the measurement insights to strategically select the most appropriate peer. Our prototype implementation on Tencent QZone and simulation experiments further demonstrate the effectiveness of our design: compared with non-peer-assisted strategy and random peer selection strategy, our design significantly improves the successful relay ratio of web contents under network failures, e.g., our design achieves a successful content download ratio of 60% even when users located in a particular region are not able to connect to the CDN servers. Wen Hu 0003, Zhi Wang 0001, Lifeng Sun |
IWQoS | 2 |
| 2015 | MAP: Microblogging Assisted Profiling of TV Shows
Xiahong Lin, Zhi Wang 0001, Lifeng Sun |
MMM (1) | 2 |
| 2015 | Power-Efficient Resource Utilization in Cellular Multimedia MulticastabstractWith advancements in wireless communication technologies, broadband wireless services will be prevalent in the near future. Meanwhile, the capability of mobile devices is drastically increasing the mobile data usage, which is far in excess of mobile network capacities. Therefore, despite the high availability of these networks, the large scale of users they support, and their improved spectral efficiencies, effective utilization of wireless resources is still required to keep up with the ever increasing user demands for mobile services. This paper targets high-throughput data transmission in advanced cellular wireless networks that have been widely used for broadband access and are constantly enhanced for future applications. We present a power-efficient resource allocation solution to meet the transmission requirements for bandwidth-intensive applications -- video streaming. In particular, our design strategically groups mobile users into multicast groups with different video quality requirement, and utilizes cellular resource to meet their video requirements. We compare our proposed methods with state-of-the-art solutions and prove their effectiveness: our design achieves 5% to 18% improvement in base station power consumption, and 13% to 25% improvement in user device power conservation. Ouldooz Baghban Karimi, Jiangchuan Liu, Zhi Wang 0001 |
MSN | 3 |
| 2015 | Preface
Wenwu Zhu 0001, Yonggang Wen 0001, Zhi Wang 0001 |
J. Comput. Sci. Technol. | 3 |
| 2015 | Towards Cost-Efficient Video Transcoding in Media Cloud: Insights Learned From User Viewing PatternsabstractVideo transcoding in an adaptive bitrate streaming (ABR) system is demanded to support video streaming over heterogenous devices and varying networks. However, it could incur a tremendous cost. Meanwhile, most viewers terminate viewing sessions within 20% of their durations; only a small fraction of each video is consumed. Built upon this user viewing pattern, we propose a Partial Transcoding Scheme for content management in media clouds. Particularly, each content is encoded into different bitrates and split into segments. Some of the segments are stored in cache, resulting in storage cost; others are transcoded online in the case of cache miss, resulting in computing cost. We aim to minimize the long-term overall cost by determining whether a segment should be cached or transcoded online. We formulate it as a constrained stochastic optimization problem. Leveraging Lyapunov optimization framework and Lagrangian relaxation, we design an online algorithm which can achieve the optimal solution within provable upper bounds. Experiments demonstrate that our proposed method can reduce 30% of operational cost, compared with the scheme of caching all the segments. Guanyu Gao, Yonggang Wen 0001, Zhi Wang 0001, Wenwu Zhu 0001 |
IEEE Trans. Multim. | 4 |
| 2015 | A Joint Online Transcoding and Delivery Approach for Dynamic Adaptive StreamingabstractDynamic adaptive streaming has emerged as a popular approach for video services in today's Internet. To date, the two important components in dynamic adaptive streaming, video transcoding that generates the adaptive bitrates of a video and video delivery that streams the videos to users, have been separately studied, resulting in a huge waste of computation and storage resource due to producing and caching different versions of videos regardless of their demands. We conduct extensive measurement studies of video sharing systems, including an IPTV service which streams regular, professionally made videos and an instant video clip sharing service which provides extremely short user-generated videos, as well as the availability of computation resource in conventional content delivery networks (CDNs). Based on the measurement insights, we propose an online joint transcoding and delivery approach for adaptive video streaming. We formulate optimization problems to enable high streaming quality for the users, and low computation and replication costs for the system. In particular, our strategy connects video transcoding and video delivery based on users' preferences of CDN regions and regional preferences of video versions. We analyze hardness of these problems and design distributed solutions. Extensive trace-driven experiments further demonstrate the superiority of our design. Zhi Wang 0001, Lifeng Sun, Chuan Wu 0001, Wenwu Zhu 0001, Qidong Zhuang, Shiqiang Yang |
IEEE Trans. Multim. | 1 |
| 2015 | CPCDN: Content Delivery Powered by Context and User IntelligenceabstractThere is an unprecedented trend that content providers (CPs) are building their own content delivery networks (CDNs) to provide a variety of content services to their users. By exploiting powerful CP-level information in content distribution, these CP-built CDNs open up a whole new design space and are changing the content delivery landscape. In this paper, we adopt a measurement-based approach to understanding why, how, and how much CP-level intelligences can help content delivery. We first present a measurement study of the CDN built by Tencent, a largest content provider based in China. We observe new characteristics and trends in content delivery which pose great challenges to the conventional content delivery paradigm and motivate the proposal of CPCDN, a CDN powered by CP-aware information. We then reveal the benefits obtained by exploiting two indispensable CP-level intelligences, namely context intelligence and user intelligence, in content delivery. Inspired by the insights learnt from the measurement studies, we systematically explore the design space of CPCDN and present the novel architecture and algorithms to address the new content delivery challenges that have arisen. Our results not only demonstrate the potential of CPCDN in pushing content delivery performance to the next level, but also identify new research problems calling for further investigation. Zhi Wang 0001, Wenwu Zhu 0001, Minghua Chen 0001, Lifeng Sun, Shiqiang Yang |
IEEE Trans. Multim. | 1 |
| 2015 | Enhancing Internet-Scale Video Service Deployment Using Microblog-Based PredictionabstractOnline microblogging has been very popular in today's Internet, where users follow other people they are interested in and exchange information between themselves. Among these exchanges, video links are a representative type on a microblogging site. The impact is fundamental-not only are viewers in a video service directly coming from the microblog sharing and recommendation, but also are the users in the microblogging site representing a promising sample to all the viewers. It is intriguing to study a proactive service deployment for such videos, using the propagation patterns of microblogs. Based on extensive traces from Youku and Tencent Weibo, a popular video sharing site and a favored microblogging system, we explore how video propagation patterns in the microblogging system are correlated with video popularity on the video sharing site. Using influential factors summarized from the measurement studies, we further design a neural network-based learning framework to predict the number of potential viewers and their geographic distribution. We then design proactive video deployment algorithms based on the prediction framework, which not only determines the upload capacities of servers in different regions, but also strategically replicates videos to these regions to serve users. Our PlanetLab-based experiments verify the effectiveness of our design. Zhi Wang 0001, Lifeng Sun, Chuan Wu 0001, Shiqiang Yang |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2014 | Characterizing cascade dynamics in a microblogging systemabstractOnline microblogging sites have become increasingly important platforms for information diffusion in today's world, where users post short messages and follow various messages posted by people that they are interested in. It is intriguing to qualitatively study the temporal dynamics of an information cascade in a microblogging system, in terms of the number of users influenced at any given time, which may provide valuable input to facilitate emerging applications such as online advertising and content distribution. In this paper, we model information diffusion in a microblogging network as an age-dependent branching process, based on practical observations from Tencent Weibo, a popular microblogging site in China. This model enables careful characterization of the diffusion topology, the different delays for users to respond to new information, and the evolution of the size of the information cascade over time. We derive the expected cascade size at any time. We validate our model based on Tencent Weibo traces, and demonstrate its effectiveness in capturing information diffusion dynamics in the real world. Shengkai Shi, Zhi Wang 0001, Chuan Wu 0001, Xiaojun Lin 0001 |
ICC | 2 |
| 2014 | Cost optimal video transcoding in media cloud: Insights from user viewing patternabstractVideo transcoding has been touted as an enabling technology to support growing media consumption over heterogenous devices. However, on-line transcoding could incur tremendous, if not prohibitive, cost in deploying or renting resources. In this research, we leverage an insight into the viewing pattern of video consumers to reduce the operating cost of video transcoding services. Specifically, it has been reported that viewers tend to terminate their session before the whole video is watched. As such, it is not cost-efficient for service providers to store or transcode all segments of the videos. Built upon this insight, we propose a partial transcoding scheme for content management in a media cloud to reduce the operating cost. Particularly, each content is split into multiple segments and stored in different files of varying playback rates. Some of the segments are stored in cache, resulting in storage cost; while some are transcoded in real-time in case of cache miss, resulting in computing cost. We aim to minimize the long-term operational cost by determining the number of segments for each playback rate to be cached or transcoded in real-time. We formulate this partial transcoding scheme as a constrained integer optimization problem. Leveraging Lagrangian relaxation and a subgradient method, we obtain the approximate solution to the integer program. Numerical results indicate that our proposed partial transcoding scheme can save more than 30% of operational cost, compared with a brute-force scheme of caching all the segments. Guanyu Gao, Yonggang Wen 0001, Zhi Wang 0001, Wenwu Zhu 0001, Yap-Peng Tan |
ICME | 4 |
| 2014 | Community based effective social video contents placement in cloud centric CDN networkabstractThe increasing popularity of online social networks (OSNs) has been transforming the dissemination pattern of social video contents. Considering the unique features of social videos, e.g., huge volume, long-tailed, and short length, how to utilize the information propagation pattern to improve the efficiency of content distribution for social videos attracts more and more attention. In this paper, we first conduct a large scale measurement to explore the social video viewing behavior under the community classification. Based on the measurement, we investigate the community driven sharing video distribution problem under the cloud-centric content delivery network (CDN) architecture. In particular, we formulate it as a constrained optimization problem with the objective to minimize the operational cost. The constraint is the averaged transmission delay. Following that, we propose a dynamic algorithm to seek the optimal solution. Our trace-driven experiments further demonstrate our algorithm can make a better tradeoff between monetary cost and QoS, and outperforms the traditional method with less operational cost while satisfying the QoS requirement. Han Hu 0003, Yonggang Wen 0001, Tat-Seng Chua, Zhi Wang 0001, Wenwu Zhu 0001, Di Wu 0001 |
ICME | 4 |
| 2014 | Joint online transcoding and geo-distributed delivery for dynamic adaptive streamingabstractDynamic adaptive video streaming has emerged as a popular approach for video streaming in today's Internet. To date the two important components in dynamic adaptive streaming, video transcoding which generates the adaptive bitrates of a video and video delivery which streams the videos to users, have been separately studied, resulting in a huge waste of computation and storage resource due to transcoding useless videos and suboptimal streaming quality due to homogeneous video replication. In this paper, we propose to jointly perform video transcoding and video delivery for adaptive streaming in an online manner. We conduct extensive measurement studies of a video sharing system and a CDN to motivate our design. We formulate and solve optimization problems to enable high streaming quality for the users, and low computation and replication costs for the system. In particular, our design connects video transcoding and video delivery based on users' preferences of CDN regions and regional preferences of video versions. Extensive trace-driven experiments further confirm the superiority of our design. Zhi Wang 0001, Lifeng Sun, Chuan Wu 0001, Wenwu Zhu 0001, Shiqiang Yang |
INFOCOM | 1 |
| 2013 | Joint Social and Content Recommendation for User-Generated Videos in Online Social NetworkabstractOnline social network is emerging as a promising alternative for users to directly access video contents. By allowing users to import videos and re-share them through the social connections, a large number of videos are available to users in the online social network. The rapid growth of the user-generated videos provides enormous potential for users to find the ones that interest them; while the convergence of online social network service and online video sharing service makes it possible to perform recommendation using social factors and content factors jointly. In this paper, we design a joint social-content recommendation framework to suggest users which videos to import or re-share in the online social network. In this framework, we first propose a user-content matrix update approach which updates and fills in cold user-video entries to provide the foundations for the recommendation. Then, based on the updated user-content matrix, we construct a joint social-content space to measure the relevance between users and videos, which can provide a high accuracy for video importing and re-sharing recommendation. We conduct experiments using real traces from Tencent Weibo and Youku to verify our algorithm and evaluate its performance. The results demonstrate the effectiveness of our approach and show that our approach can substantially improve the recommendation accuracy. Zhi Wang 0001, Lifeng Sun, Wenwu Zhu 0001, Shiqiang Yang, Dapeng Oliver Wu |
IEEE Trans. Multim. | 1 |
| 2013 | Peer-Assisted Social Media Streaming with Social ReciprocityabstractOnline video sharing and social networking are cross-pollinating rapidly in today's Internet: Online social network users are sharing more and more media contents among each other, while online video sharing sites are leveraging social connections among users to promote their videos. An intriguing development as it is, the operational challenge in previous video sharing systems persists, em i.e., the large server cost demanded for scaling of the systems. Peer-to-peer video sharing could be a rescue, only if the video viewers' mutual resource contribution has been fully incentivized and efficiently scheduled. Exploring the unique advantages of a social network based video sharing system, we advocate to utilize social reciprocities among peers with social relationships for efficient contribution incentivization and scheduling, so as to enable high-quality video streaming with low server cost. We exploit social reciprocity with two give-and-take ratios at each peer: (1) peer contribution ratio (em PCR), which evaluates the reciprocity level between a pair of social friends, and (2) system contribution ratio (em SCR), which records the give-and-take level of the user to and from the entire system. We design efficient peer-to-peer mechanisms for video streaming using the two ratios, where each user optimally decides which other users to seek relay help from and help in relaying video streams, respectively, based on combined evaluations of their social relationship and historical reciprocity levels. Our design achieves effective incentives for resource contribution, load balancing among relay peers, as well as efficient social-aware resource scheduling. We also discuss practical implementation and implement our design in a prototype social media sharing system. Our extensive evaluations based on PlanetLab experiments verify that high-quality large-scale social media sharing can be achieved with conservative server costs. Zhi Wang 0001, Chuan Wu 0001, Lifeng Sun, Shiqiang Yang |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2013 | Two decades of internet video streaming: A retrospective viewabstractFor over two decades, video streaming over the Internet has received a substantial amount of attention from both academia and industry. Starting from the design of transport protocols for streaming video, research interests have later shifted to the peer-to-peer paradigm of designing streaming protocols at the application layer. More recent research has focused on building more practical and scalable systems, using Dynamic Adaptive Streaming over HTTP. In this article, we provide a retrospective view of the research results over the past two decades, with a focus on peer-to-peer streaming protocols and the effects of cloud computing and social media. Baochun Li, Zhi Wang 0001, Jiangchuan Liu, Wenwu Zhu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2013 | Propagation-based social-aware multimedia content distributionabstractOnline social networks have reshaped how multimedia contents are generated, distributed, and consumed on today's Internet. Given the massive number of user-generated contents shared in online social networks, users are moving to directly access these contents in their preferred social network services. It is intriguing to study the service provision of social contents for global users with satisfactory quality of experience. In this article, we conduct large-scale measurement of a real-world online social network system to study the social content propagation. We have observed important propagation patterns, including social locality, geographical locality, and temporal locality. Motivated by the measurement insights, we propose a propagation-based social-aware delivery framework using a hybrid edge-cloud and peer-assisted architecture. We also design replication strategies for the architecture based on three propagation predictors designed by jointly considering user, content, and context information. In particular, we design a propagation region predictor and a global audience predictor to guide how the edge-cloud servers backup the contents, and a local audience predictor to guide how peers cache the contents for their friends. Our trace-driven experiments further demonstrate the effectiveness and superiority of our design. Zhi Wang 0001, Wenwu Zhu 0001, Xiangwen Chen, Lifeng Sun, Jiangchuan Liu, Minghua Chen 0001, Peng Cui 0001, Shiqiang Yang |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2012 | Cloud-based social application deployment using local processing and global distributionabstractSocial applications represent a paradigm shift on how the Internet is to be used, and have already changed the way we work, live, and play. When it comes to deploying social applications, cloud computing platforms are used to meet the Internet-scale, self-propagating, and fast-growing demands from these applications. Yet, to deploy social media applications in the most effective and economic fashion, we need to strategically design and follow a set of theoretical and practical principles. In this paper, we seek to design a set of new principles to guide social application deployment. Learning from large-scale measurement-based observations using a real-world social application, the gist of our principles is to detach the typically integrated "collection → processing → distribution" work ows in social applications into separate local processing and global distribution procedures, which can be effectively deployed using different cloud services. Moreover, based on a predictive model of regional propagation, we formulate the resource allocation problems in the processes of collecting/processing and distributing content as two optimization problems, which can be solved by efficient algorithms. Finally, based on our theoretical design, we have implemented an example social application on Amazon EC2 and Google AppEngine, where IaaS-based computation instances perform content collection and processing, and the PaaS-based platform is employed to distribute the contents that are widely propagating. Our PlanetLab-based trace-driven experiments have further confirmed the superiority of our design. Zhi Wang 0001, Baochun Li, Lifeng Sun, Shiqiang Yang |
CoNEXT | 1 |
| 2012 | Group Recommendation Using External Followee for Social TVabstractGroup recommendation plays a significant role in Social TV systems, where online friends form into temporary groups to enjoy watching video together and interact with each other. Online microblogging systems introduce the "following" relationship that reflects the common interests between users in a group and external representative followees outside the group. Traditional group recommendation only considers internal group members' preferences and their relationship. In our study, we measure the external followees' impact on group interest and establish group preference model based on external experts' guidance for group recommendation. In addition, we take advantage of the current watching video to improve context-aware recommendations. Experimental results show that our solution works much better in situations of high group dynamic and inactive group members than traditional approaches. Lifeng Sun, Zhi Wang 0001, Da Meng |
ICME | 3 |
| 2012 | Guiding internet-scale video service deployment using microblog-based predictionabstractOnline microblogging has been very popular in today's Internet, where users exchange short messages and follow various contents shared by people that they are interested in. Among the variety of exchanges, video links are a representative type on a microblogging site. More and more viewers of an Internet video service are coming from microblog recommendations. It is intriguing research to explore the connections between the patterns of microblog exchanges and the popularity of videos, in order to potentially use the propagation patterns of microblogs to guide proactive service deployment of a video sharing system. Based on extensive traces from Youku and Tencent Weibo, a popular video sharing site and a favored microblogging system in China, we explore how patterns of video link propagation in the microblogging system are correlated with video popularity on the video sharing site, at different times and in different geographic regions. Using influential factors summarized from the measurement studies, we further design neural network-based learning frameworks to predict the number of potential viewers of different videos and the geographic distribution of viewers. Experiments show that our neural network-based frameworks achieve better prediction accuracy, as compared to a classical approach that relies on historical numbers of views. We also briefly discuss how proactive video service deployment can be effectively enabled by our prediction frameworks. Zhi Wang 0001, Lifeng Sun, Chuan Wu 0001, Shiqiang Yang |
INFOCOM | 1 |
| 2012 | Propagation-based social-aware replication for social video contentsabstractOnline social network has reshaped the way how video contents are generated, distributed and consumed on today's Internet. Given the massive number of videos generated and shared in online social networks, it has been popular for users to directly access video contents in their preferred social network services. It is intriguing to study the service provision of social video contents for global users with satisfactory quality-of-experience. In this paper, we conduct large-scale measurement of a real-world online social network system to study the propagation of the social video contents. We have summarized important characteristics from the video propagation patterns, including social locality, geographical locality and temporal locality. Motivated by the measurement insights, we propose a propagation-based social-aware replication framework using a hybrid edge-cloud and peer-assisted architecture, namely PSAR, to serve the social video contents. Our replication strategies in PSAR are based on the design of three propagation-based replication indices, including a geographic influence index and a content propagation index to guide how the edge-cloud servers backup the videos, and a social influence index to guide how peers cache the videos for their friends. By incorporating these replication indices into our system design, PSAR has significantly improved the replication performance and the video service quality. Our trace-driven experiments further demonstrate the effectiveness and superiority of PSAR, which improves the local download ratio in the edge-cloud replication by 30%, and the local cache hit ratio in the peer-assisted replication by 40%, against traditional approaches. Zhi Wang 0001, Lifeng Sun, Xiangwen Chen, Wenwu Zhu 0001, Jiangchuan Liu, Minghua Chen 0001, Shiqiang Yang |
ACM Multimedia | 1 |
| 2011 | Peer-assisted online games with social reciprocityabstractOnline games and social networks are cross-pollinating rapidly in today's Internet: Online social network sites are deploying more and more games in their systems, while online game providers are leveraging social networks to power their games. An intriguing development as it is, the operational challenge in the previous game persists, i.e., the large server operational cost remains a non-negligible obstacle for deploying high-quality multi-player games. Peer-to-peer based game network design could be a rescue, only if the game players' mutual resource contribution has been fully incentivized and efficiently scheduled. Exploring the unique advantage of social network based games (social games), we advocate to utilize social reciprocities among peers with social relationships for efficient contribution incentivization and scheduling, so as to power a high-quality online game with low server cost. In this paper, social reciprocity is exploited with two give-and-take ratios at each peer: (1) peer contribution ratio (PCR), which evaluates the reciprocity level between a pair of social friends, and (2) system contribution ratio (SCR), which records the give-and-take level of the player to and from the entire network. We design efficient peer-to-peer mechanisms for game state distribution using the two ratios, where each player optimally decides which other players to seek relay help from and help in relaying game states, respectively, based on combined evaluations of their social relationship and historical reciprocity levels. Our design achieves effective incentives for resource contribution, load balancing among relay peers, as well as efficient social-aware resource scheduling. We also discuss practical implementation concerns and implement our design in a prototype online social game. Our extensive evaluations based on experiments on PlanetLab verify that high-quality large-scale social games can be achieved with conservative server costs. Zhi Wang 0001, Chuan Wu 0001, Lifeng Sun, Shiqiang Yang |
IWQoS | 1 |
| 2011 | Prefetching strategy in peer-assisted social video streamingabstractOnline social network has emerged as the most popular approach for people to directly access multimedia contents. Among these contents, video sharing is a challenging task due to the demand on a large amount of uplink bandwidth at the dedicated server. We leverage a P2P paradigm to alleviate the server to distribute shared videos. By investigating traces obtained from a popular online social network in China, we observe that users' preferences can be predicted. We design a user preference guided prefetching strategy to reduce video startup delays, enabling smooth playback. Simulation experiments show that our design achieves high prefetch accuracy and short startup delay with conservative storage and bandwidth capacities at peers. Zhi Wang 0001, Lifeng Sun, Shiqiang Yang, Wenwu Zhu 0001 |
ACM Multimedia | 1 |
| 2010 | Strategies of Collaboration in Multi-Channel P2P VoD StreamingabstractAs compared to live peer-to-peer (P2P) streaming, modern P2P video-on-demand (VoD) systems have brought much larger volumes of videos and more interactive controls to the Internet users. Nevertheless, the larger number of available videos and the flexibility of allowing users to jump back and forth in a video, have led to much fewer numbers of concurrent peers watching at a similar pace, that reduces the chance for collaborative chunk supply among peers and thus significantly increases the server bandwidth cost. Towards the ultimate goal of maximizing peer resource utilization, in this paper, we design effective strategies for both cross-channel and intra-channel collaborations in multi- channel P2P VoD systems, such that individual peer's resources, including download/upload bandwidths and the cache capacity, are effectively utilized to maximize the streaming qualities in all the channels. In particular, each peer actively and strategically determines the supply-and-demand imbalance in different channels, as well as that among different chunks within each video, makes use of its surplus download capacity to fetch chunks with the most need, and then serves those chunks using its idle upload bandwidth, all without impairing its own streaming quality. Our extensive trace-driven simulations show the effectiveness of our strategies in reducing the server cost while guaranteeing high streaming qualities in the entire system, even during extreme scenarios such as unexpected flash crowds. Zhi Wang 0001, Chuan Wu 0001, Lifeng Sun, Shiqiang Yang |
GLOBECOM | 1 |
| 2010 | Strategies of buffering schedule in P2P VoD streamingabstractAs compared to live peer-to-peer (P2P) streaming, modern P2P video-on-demand (VoD) systems have brought much larger volumes of videos and more interactive controls to the Internet users. As the increase of bitrate of the videos and the full VCR controls of P2P VoD, the behavior “buffering” motivates us to design different schedule and service strategies for peers, to improve the playback performance, and the alleviation of the dedicated streaming server, by making best use of the bandwidth and cache capacities of these buffering peers. In our design, peers strategically decide which segments in the video to download first, and which requests to serve first. We conduct extended simulations to evaluate the performance of the strategies, and the results show our design outperforms the conventional sequential scheme, with respect to improving the playback quality and reducing the server load. Zhi Wang 0001, Lifeng Sun, Shiqiang Yang |
MMSP | 1 |
| 2007 | A Novel Event-Oriented Segment-of-Interest Discovery Method for Surveillance VideoabstractDuring recent years, the quick development of computer techniques has witnessed the ever-increasing surveillance video data, which essentially pose great challenge on the data storage, management, analysis and even retrieval. Considering that most of the high volume of data is with no interest, we mainly investigate the problem of effectively and efficiently discovering segments-of-interest (SoI) in this paper. To do so, we propose a novel event-oriented Sol discovery method in two steps: first, we represent an event by modeling pixels' change in temporal-spatial space, aiming to unify both the inter-frames and frames-background changes; second, with the benefit of unsupervised learning, the prototype-event models could be learned from these detected events and in turn exploited to measure the interest factor of each prototype-event. The experiment results demonstrate that the proposed method precisely discriminate different events and effectively discover Sols. Peng Cui 0001, Lifeng Sun, Zhi Wang 0001, Shiqiang Yang |
ICME | 3 |