Yuan Zhang 0013

dblp:48/2168-13 · DBLP profile ↗
← Back
48ranked-venue papers
6as first author
38since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 27 · 2 first-author · 21 since 2021Computer networks · 13 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 8 · 8 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multi-Task Representation Alignment on Language Understanding: A Mutual Information Perspective
abstract
Multi-task learning (MTL) enables joint learning over multiple tasks based on shared representations, but suffers from task interference issue during optimization.Existing works mainly focus on task balancing or probabilistic modeling but fail to address the issue since they struggle to learn sufficient representations for all target tasks.To address this, we propose a multi-task representation alignment (MTRA) framework to achieve task-specific alignment and self-alignment on the shared representations from a mutual information perspective.MTRA ensures that the learned representations contain task-relevant features while mitigating the negative effects of task-irrelevant features.First, we design a task-specific alignment objective to align the shared representations and task-specific representations with the expected targets of all tasks via information maximization.Besides, we design a self-alignment objective to eliminate task-irrelevant features via conditional information minimization.Experiments on two multi-task language benchmarks show that MTRA outperforms 13 representative MTL methods under the same settings, particularly under label-noisy and dataconstrained conditions.Further analysis shows that the learned shared representations exhibit sufficient task informativeness and superior alignment properties.
Dou Hu 0001, Lingwei Wei, Hongjiang Xiao, Songlin Hu 0001, Yuan Zhang 0013
ACL (1)5
2026 Toward Natural and Companionable Virtual Agents via Cross-Temporal Emotional Modeling
abstract
Recent advances in foundation models have enabled conversational agents that aim for sustained companionship rather than mere task completion. Yet most still remain unable to support natural, long-term companion-like interactions, resulting in experiences that feel episodic and inauthentic. We argue that current agents overlooked cross-temporal modeling of agents’ social behaviors and internal emotions: generated behaviors rarely influence an agent’s emotional state, and emotional states seldom shape subsequent behaviors. We present Cross-Temporal Emotion Modeling (CTEM), a framework that links long-term behavioral history to moment-to-moment emotional expression. CTEM establishes a closed loop where past experiences update an evolving emotional state; this state conditions immediate interactions; and user feedback continually revises both memory and emotional state, enabling reflection and anticipation. We instantiate CTEM as Auri, a companion agent on an instant-messaging platform, and report a 21-day in-the-wild study showing that CTEM shows improvements in perceived naturalness, coherence, and emotional harmony.
Feier Qin, Xiao Li 0030, Hanyao Wang, Yan Lu 0001, Yuan Zhang 0013
CHI8
2026 M2DE: A multi-stressor multi-dimensional dynamic evaluation framework for the trustworthiness of LLMs
Hongjiang Xiao, Xiuying Li, Ye Wang 0011, Liangfei Zhang, Yuan Zhang 0013
Pattern Recognit.7
2026 PsyMem: Fine-grained Psychological Alignment and Explicit Memory Control for Advanced Role-Playing LLMs
abstract
Abstract Existing LLM-based role-playing methods often rely on superficial textual descriptions or simplistic metrics, inadequately modeling both intrinsic and extrinsic character dimensions. Additionally, they typically simulate character memory with implicit model knowledge or basic retrieval augment generation without explicit memory alignment, compromising memory consistency. The two issues weaken reliability of role-playing LLMs in several applications, such as trustworthy social simulation. To address these limitations, we propose PsyMem, a novel framework integrating fine-grained psychological attributes and explicit memory control for role-playing. PsyMem supplements textual descriptions with 26 psychological indicators to detailed model character. Additionally, PsyMem implements memory alignment training, explicitly trains the model to align character’s response with memory, thereby enabling dynamic memory-controlled responding during inference. By training Qwen2.5-7B-Instruct on our specially designed dataset (including 5,414 characters and 38,962 dialogues extracted from novels), the resulting model, termed as PsyMem-Qwen, outperforms baseline models in role-playing, achieving the best performance in human-likeness and character fidelity.
Xilong Cheng, Yunxiao Qin, Yuting Tan 0002, Zhengnan Li, Ye Wang 0011, Hongjiang Xiao, Yuan Zhang 0013
Trans. Assoc. Comput. Linguistics7
2026 WiLD: Learning-Based Wireless Loss Diagnosis for Congestion Control With Ultra-Low Kernel Overhead
abstract
Current congestion control algorithms (CCAs) are inefficient in wireless networks due to the lack of distinction of congestion and wireless packet losses. In this work, we propose a simple yet effective learning-based wireless loss diagnosis (WiLD) solution for enhancing wireless congestion control. WiLD uses a neural network (NN) to accurately distinguish between wireless packet loss and congestion packet loss. To seamlessly cooperate with rule-based CCAs and make real-time decisions, we further implement WiLD in Linux kernel to avoid the frequent kernel-space communication. Specifically, we use a lightweight NN for inference and propose an integer quantization for WiLD deployment in various Linux versions. Real-world experiments and simulations demonstrate that WiLD can accurately differentiate the wireless and congestion packet loss with negligible CPU overhead (around 1% of WiLD vs. around 100% of learning-based algorithms such as Vivace and Aurora) and fast inference time (45% less compared to TensorFlow Lite). When combined with Cubic, WiLD-Cubic can achieve around 792%, 536%, 412%, 231%, 218%, 108%, 85% and 291% throughput improvement compared with BBRv2, Cubic, Westwood, Copa, Copa+, Vivace, Aurora and Indigo in the real network environment.
Jinyao Yan, Yuan Zhang 0013, Lingjun Pu
IEEE Trans. Netw. Serv. Manag.3
2026 Implicit Representation-based Volumetric Video Streaming for Photorealistic Full-scene Experience
abstract
The widespread integration of the Internet of Things with sensors like depth-of-field cameras, LiDAR scanners, and eye-tracking infrared sensors, in head-mounted devices, has ushered in a new era of immersive digital experiences. Full-scene volumetric video (VV), a key innovation in this integration, provides a deeply immersive experience by capturing the richness and detail of the 3D world. However, its massive data volume presents significant streaming challenges. While 3D tile-based viewport approaches have been proposed, they struggle to full-scene VV given the small video buffer limitation, high tile segmentation overhead, and lack of full-scene consideration. In this work, inspired by the advancements of implicit neural radiance field (NeRF), we present \({\mathsf{V}^{2}\mathsf{NeRF}}\) , a novel full-scene VV streaming system featured by layered representation. It harmonizes the NeRF with explicit point clouds to represent the static background and dynamic foreground, thereby avoiding large data transfers and achieving photorealistic content representation. To tackle the issues of intensive computation requirements and multiscale adaptation scheduling within \({\mathsf{V}^{2}\mathsf{NeRF}}\) system, we propose a lightweight non-visible background removal method and a two-stage decoupled architecture. In addition, an efficient buffer-aware simulated annealing algorithm is developed, alongside the utilization of a perceptually learned metric, to enhance user experience. We further discuss the concerns about practical development and deployment. Extensive prototype evaluations demonstrate \({\mathsf{V}^{2}\mathsf{NeRF}}\) ’s superior streaming and viewing performance on a wide variety of networks, viewing motions, and scenes. For instance, compared to state-of-the-art approaches, it achieves a 24% increment in perceptual quality, an 83% reduction in rebuffering time, and a 54% enhancement in user experience on average.
Jianxin Shi 0005, Miao Zhang 0003, Linfeng Shen, Jiangchuan Liu, Yuan Zhang 0013, Lingjun Pu, Jingdong Xu
ACM Trans. Multim. Comput. Commun. Appl.5
2025 Libra: Novel LLM Token Streaming via Region-Based Task Scheduling and Token Bundling
abstract
The LLM serving systems are increasingly growing in popularity, as they provide various capabilities ranging from realtime translation to AI-driven chatbots. Recently, significant effort has been made to optimize server-side metrics such as token generation throughput, while the optimization of token streaming is simply overlooked, resulting in excessive network traffic and poor network utilization. In this paper, we introduce user regions to relieve the network issues, since users from the same region (e.g., universities and business zones) are likely to share similar behaviors to access LLM serving systems (e.g., they are active in a period of time). In this context, we propose Libra, a proxy-cloud collaborative serving system, where the cloud generates and bundles the tokens in terms of user regions and region-based proxy extracts and repacks the received token bundle to their corresponding users. At its core, we design an online region-based task scheduling algorithm with a provable performance to optimize user QoE and system overhead over time. Our evaluations show that Libra outperforms the state-of-the-art LLM serving systems (without user regions), such as vLLM, VTC and Andes, by up to$2.1 \times$in the Time-Between-Tokens (TBT) metric and$78.4 \times$in the number of packets. In addition, it achieves a 32.6 % reduction in TBT compared to other alternative algorithms (with user regions).
Chengjin Zhou, Xinjing Yuan, Jianxin Shi 0005, Yuan Zhang 0013, Lingjun Pu
IWQoS5
2025 Bimodal Semantic-Driven 3D Immersive Telepresence System
abstract
3D immersive telepresence systems have dramatically transformed the way users communicate, yet existing point cloud and mesh-based approaches require large amounts of data transmission. Although recent 3D facial semantic-driven techniques can be used to reduce the network burden, they face the critical problem of the high computational cost of facial semantic extraction. To solve the problem, we design and implement an innovative bimodal semantic-driven real-time 3D telepresence system which leverages the low-cost audio-driven semantics for facial expression extraction and head movement semantics for 3D interaction. To improve the efficiency and accuracy of semantic information processing, we propose a speech separation module and optimize the data reception strategy. To optimize the quality of video content reconstruction, we employ frame interpolation and super-resolution techniques, and further propose an online resource scheduling algorithm to balance the rendering, interpolation, and super-resolution processes with limited terminal resources. Experimental results demonstrate that our system can achieve 1K rendering resolution and ~46FPS frame rate with ultra-low network bandwidth (375kbps, approximately 0.32% compared to the point cloud-based approach) and low latency (about 78% semantic feature extraction latency compared to the 3D facial semantic-driven approach) in the campus network.
Yuan Zhang 0013, Lingjun Pu, Tao Lin 0001, Jinyao Yan
NOSSDAV2
2025 3DGS-Enabled High-Fidelity Low-Cost Immersive Static 3D Video Streaming
abstract
3D Gaussian Splatting (3DGS), as the cutting-edge static three-dimensional (3D) content generation technology, revolutionizes the speed and fidelity of 3D model construction and provides immense potential for various applications, including e-commerce, 3D exhibitions, and virtual tourism. However, our pioneering analysis of firsthand user experiments uncovers a critical challenge: the unique user behavior patterns of static 3D scenarios render existing immersive video streaming solutions inadequate. To be concrete, the frequent switches between active and inactive states impair viewport prediction accuracy, while the fast glance and slow view pattern provides an opportunity for further quality of experience (QoE) improvement. To tackle these problems, this paper introduces innovative designs for static 3D video streaming. Specifically, we devise a viewport prediction and error correction mechanism on the client side to restore the user viewport with a low cost. Furthermore, we design a dynamic frame rate and bitrate control algorithm to improve user QoE under various network conditions. We implement the first 3DGS-enabled immersive static 3D video streaming system based on an edge-rendered architecture, ensuring efficient rendering and encoding on the edge server while providing broad accessibility for various client-side devices through a web browser. Extensive testing under real-world network conditions and with various kinds of devices demonstrates that the proposed approach exhibits robust and rapid adaptability to fluctuating network conditions, improving user QoE by over 20%, reducing interactive latency by 89%, and minimizing the stall duration by 26% compared to existing low-latency streaming solutions.
Rongji Liao, Yuan Zhang 0013, Wei Zhang 0324, Lingjun Pu, Yu Guan 0005, Yunpeng Jing, Tao Lin 0001, Jinyao Yan
IEEE J. Sel. Areas Commun.2
2025 MEAS: Multimodal Emotion Analysis System for Short Videos on Social Media Platforms
abstract
Short videos have surged in popularity on social media. The emotions expressed by short videos can trigger or even magnify the public sentiment. Hence, accurate computation of these emotions is vital for social affective computing. However, the multimodal emotion analysis of short videos on social platforms faces some challenges: the accuracy of the model is tested by the disunity of video resolution; the collection of large-scale social media data, the manual transcription and segmentation of audio content, and the precise labeling process require a lot of manpower. In this article, we have proposed an affective computing system MEAS for social short videos, which combines multiscale resolution adaptability and advanced RoBERTa model to optimize the preprocessing of high definition large size short videos and improve the contribution of text modality in emotion analysis. In addition, the system also adopts automatic audio segmentation and transcription technology to realize the efficient capture of speech forms in social short videos. Experimental results show that compared with the leading open source algorithm V2EM on the IEMOCAP dataset, the proposed method achieves a significant increase in weighted accuracy and F1 score of 4.17% and 7.29%, respectively. We constructed a novel dataset named “Bili-news” based on social platform news short videos, validating the effectiveness of the MEAS system. Through experimental verification, we also find a significant positive correlation between the emotions expressed in short videos and the social sentiments of the audience.
Qinglan Wei, Yaqi Zhou, Shenlian Xiang, Longhui Xiao, Yuan Zhang 0013
IEEE Trans. Comput. Soc. Syst.5
2024 Unifying Generation and Compression: Ultra-low bitrate Image Coding Via Multi-stage Transformer
abstract
Recent progress in generative compression technology has significantly improved the perceptual quality of compressed data. However, these advancements primarily focus on producing high-frequency details, often overlooking the ability of generative models to capture the prior distribution of image content, thus impeding further bitrate reduction in extreme compression scenarios (< 0.05 bpp). Motivated by the capabilities of predictive language models for lossless compression, this paper introduces a novel Unified Image Generation-Compression (UIGC) paradigm, merging the processes of generation and compression. A key feature of the UIGC framework is the adoption of vector-quantized (VQ) image models for tokenization, alongside a multi-stage transformer designed to exploit spatial contextual information for modeling the prior distribution. As such, the dual-purpose framework effectively utilizes the learned prior for entropy estimation and assists in the regeneration of lost tokens. Extensive experiments demonstrate the superiority of the proposed UIGC framework over existing codecs in perceptual quality and human perception, particularly in ultra-low bitrate scenarios (≤0.03 bpp), pioneering a new direction in generative compression.
Naifu Xue, Qi Mao 0002, Yuan Zhang 0013, Siwei Ma 0001
ICME4
2024 Focus, Distinguish, and Prompt: Unleashing CLIP for Efficient and Flexible Scene Text Retrieval
abstract
Scene text retrieval aims to find all images containing the query text from an image gallery. Current efforts tend to adopt an Optical Character Recognition (OCR) pipeline, which requires complicated text detection and/or recognition processes, resulting in inefficient and inflexible retrieval. Different from them, in this work we propose to explore the intrinsic potential of Contrastive Language-Image Pre-training (CLIP) for OCR-free scene text retrieval. Through empirical analysis, we observe that the main challenges of CLIP as a text retriever are: 1) limited text perceptual scale, and 2) entangled visual-semantic concepts. To this end, a novel model termed FDP (Focus, Distinguish, and Prompt) is developed. FDP first focuses on scene text via shifting the attention to the text area and probing the hidden text knowledge, and then divides the query text into content word and function word for processing, in which a semantic-aware prompting scheme and a distracted queries assistance module are utilized. Extensive experiments show that FDP significantly enhances the inference speed while achieving better or competitive retrieval accuracy compared to existing methods. Notably, on the IIIT-STR benchmark, FDP surpasses the state-of-the-art model by 4.37% with a 4 times faster speed. Furthermore, additional experiments under phrase-level and attribute-aware scene text retrieval settings validate FDP's particular advantages in handling diverse forms of query text. The source code will be available at https://github.com/Gyann-z/FDP.
Gangyan Zeng, Yuan Zhang 0013, Dongbao Yang, Peng Zhang 0044, Yiwen Gao 0001, Xugong Qin, Yu Zhou 0015
ACM Multimedia2
2024 OMMS: Multiple Control based Adaptive 360° Video Streaming
abstract
In the realm of 360° video streaming, how to deliver optimal viewing experience to users with minimal bandwidth cost has become an emerging challenge. Our research is driven by a comprehensive analysis of real-world user's head movement datasets of 360° video, revealing significant viewport shifts even during the playback of individual 360° video chunks. However, prevailing 360° video streaming algorithms fail to account for such viewport variations, thereby resulting in a substantial degradation of user experience. To this end, we propose OMMS, a multiple-control-based adaptive 360° video streaming algorithm. OMMS adopts an integrated approach by introducing multiple supplementary controls alongside the main control, enabling adaptation to changes in the user's viewport. Experimental results demonstrate that OMMS yields a significant improvement in video quality by 34%, while effectively reducing bandwidth resource consumption.
Ruisi Xu, Chenyu Liu 0001, Size Qian, Yuan Zhang 0013, Tao Lin 0001
MMSys5
2024 Towards Full-scene Volumetric Video Streaming via Spatially Layered Representation and NeRF Generation
abstract
Immersive full-scene volumetric video (VV) showcases the richness and detail of the 3D world, yet poses significant streaming challenges given its massive data volume. Existing 3D tile-based viewport approaches struggle to effectively adapt to full-scene VV owing to their small video buffer limitation, high tile segmentation overhead, and lack of full-scene consideration.
Jianxin Shi 0005, Miao Zhang 0003, Linfeng Shen, Jiangchuan Liu, Yuan Zhang 0013, Lingjun Pu, Jingdong Xu
NOSSDAV5
2024 To Distill or Not to Distill: Toward Fast, Accurate, and Communication-Efficient Federated Distillation Learning
abstract
Apart from the promising potential, federated learning (FL) faces challenges, such as high communication costs and client heterogeneity. Although numerous works have been proposed to address these issues, they lack a holistic perspective to balance all requirements. Moreover, these solutions have not fully utilized the underlying computation capability and network resources, resulting in suboptimal tradeoffs between communication efficiency and inference accuracy. To overcome these challenges, we propose FDL: a federated distillation (FD) learning framework that combines FD and FL to fully utilize computation and network resources. We theoretically prove the convergence bound of the proposed FDL framework. Furthermore, to minimize the training time while maintaining inference accuracy, we design HAD: a heterogeneity-aware FL/FD selection algorithm that determines the total communication rounds and selects the set of FL and FD nodes in each communication round. The optimality of HAD is also theoretically proved. The FDL framework and HAD algorithm together minimize the training time while satisfying the inference accuracy in a heterogeneous and dynamic environment. Extensive experiments on various learning algorithms and data sets show that the proposed FDL-HAD solution can obtain the optimal selection decision in overwhelmingly less selection time compared with the Gurobi solver and can reduce the overall training time by at least 44.8% compared with FL solutions with the same inference accuracy.
Yuan Zhang 0013, Lingjun Pu, Tao Lin 0001, Jinyao Yan
IEEE Internet Things J.1
2024 STOP: Joint send buffer and transmission control for user-perceived deadline guarantee via curriculum guided-deep reinforcement learning
Rongji Liao, Yuan Zhang 0013, Jinyao Yan, Narisu Tao
J. Netw. Comput. Appl.2
2024 DeCa360: Deadline-aware edge caching for two-tier 360° video streaming
Tao Lin 0001, Hao Yang 0057, Yuan Zhang 0013, Bo Jiang 0003, Jinyao Yan
J. Netw. Comput. Appl.4
2024 ${\sf NetDPI}$NetDPI: Efficient Deep Packet Inspection via Filtering-Plus-Verification in Programmable 5G Data Plane for Multi-Access Edge Computing
abstract
In this paper, we advocate${\sf NetDPI}$, a novel and efficient Deep Packet Inspection (DPI) solution built-in 5G Data Plane for multi-access edge computing, leveraging the unique forwarding while computing capability of emerging programmable switches. As the cornerstone, we propose${\sf FIVE}$, the firstFiltering-plus-Verification algorithm tailored to programmable switches to achieve efficient multiple pattern matching (i.e., the core of DPI). Briefly, the filtering phase introduces a multi-window parallel shift-or algorithm to rapidly screen out all the “suspicious” packet payloads. Meanwhile, the verification phase innovates a level-based state encoding scheme for the Aho–Corasick (AC) algorithm, which substantially increases the number of supported patterns and consequently figures out more “guilty” payloads. We implement the prototype of${\sf NetDPI}$in both software and hardware programmable switches (i.e., BMv2 and Barefoot Tofino2) and make them publicly available. Extensive evaluations indicate that${\sf NetDPI}$provides orders of magnitude improvement in throughput compared to the typical cloud-delivered DPI solutions, and besides${\sf FIVE}$greatly reduces the memory consumption compared to the alternative in-network exact match algorithms under a variety of system settings including different DPI pattern sets and malware-packet percentages.
Chengjin Zhou, Qiao Xiang, Lingjun Pu, Zheli Liu, Yuan Zhang 0013, Xinjing Yuan, Jingdong Xu
IEEE Trans. Mob. Comput.5
2023 Disentangled Feature Learning for Real-Time Neural Speech Coding
abstract
Recently end-to-end neural audio/speech coding has shown its great potential to outperform traditional signal analysis based audio codecs. This is mostly achieved by following the VQ-VAE paradigm where blind features are learned, vector-quantized and coded. In this paper, instead of blind end-to-end learning, we propose to learn disentangled features for real-time neural speech coding. Specifically, more global-like speaker identity and local content features are learned with disentanglement to represent speech. Such a compact feature decomposition not only achieves better coding efficiency by exploiting bit allocation among different features but also provides the flexibility to do audio editing in embedding space, such as voice conversion in real-time communications. Both subjective and objective results demonstrate its coding efficiency and we find that the learned disentangled features show comparable performance on any-to-any voice conversion with modern self-supervised speech representation learning models with far less parameters and low latency, showing the potential of our neural coding framework.
Xiulian Peng, Yuan Zhang 0013, Yan Lu 0001
ICASSP3
2023 Real-Time Speech Enhancement with Dynamic Attention Span
abstract
For real-time speech enhancement (SE) including noise suppression, dereverberation and acoustic echo cancellation, the time-variance of the audio signals becomes a severe challenge. The causality and memory usage limit that only the historical information can be used for the system to capture the time-variant characteristics. We propose to adaptively change the receptive field according to the input signal in deep neural network based SE model. Specifically, in an encoder-decoder framework, a dynamic attention span mechanism is introduced to all the attention modules for controlling the size of historical content used for processing the current frame. Experimental results verify that this dynamic mechanism can better track time-variant factors and capture speech-related characteristics, benefiting to both interference removing and speech quality retaining.
Xiulian Peng, Yuan Zhang 0013, Yan Lu 0001
ICASSP4
2023 Filling in the Blank: Rationale-Augmented Prompt Tuning for TextVQA
abstract
Recently, generative Text-based visual question answering (TextVQA) methods, which are often based on language models, have exhibited impressive results and drawn increasing attention. However, due to the inconsistencies in both input forms and optimization objectives, the power of pretrained language models is not fully explored, resulting in the need for large amounts of training data. In this work, we rethink the characteristics of the TextVQA task and find that scene text is indeed a special kind of language embedded in images. To this end, we propose a text-centered generative framework FITB (stands for Filling In The Blank), in which multimodal information is mainly represented in textual form and rationale-augmented prompting is involved. Specifically, an infilling-based prompt strategy is utilized to formulate TextVQA as a novel problem of filling in the blank with proper scene text according to the language context. Furthermore, aiming to prevent the model from language bias overfitting, we design a rough answer grounding module to provide visual rationales for promoting multimodal reasoning. Extensive experiments verify the superiority of FITB in both fully-supervised and zero-shot/few-shot settings. Notably, even with a saving of about 64M data, FITB surpasses the state-of-the-art method by 3.00% and 1.99% on TextVQA and ST-VQA datasets, respectively.
Gangyan Zeng, Yuan Zhang 0013, Yu Zhou 0015, Bo Fang 0003, Weiping Wang 0005
ACM Multimedia2
2023 Lambda-Domain Rate Control for Neural Image Compression
abstract
Rate control based on rate-distortion modeling is a classic problem in lossy image compression. Despite extensive research in neural image compression, its rate control remains understudied. In this paper, we introduce a variable rate neural image compression scheme that supports precise rate control with one-pass encoding. Our approach utilizes the Lagrangian multiplier method to transform rate control into an unconstrained optimization problem, mapping the target bitrate to λ for rate-distortion trade-off adjustment. We propose an improved exponential R-λ model and estimate the bitrates with a hybrid convolution-transformer network for model fitting. The encoder is controlled by λ, and a multi-layer modulation mechanism ensures variable rate ability. In our experiments, the proposed method outperforms the intra-frame coding of Versatile Video Coding (VVC). Meanwhile, the average rate control error is less than 5.1%, while maintaining almost identical rate-distortion performance and acceptable complexity.
Naifu Xue, Yuan Zhang 0013
MMAsia2
2023 DHP: A Joint Video Download and Dynamic Bitrate Adaptation Algorithm for Short Video Streaming
Wenhua Gao, Lanju Zhang, Hao Yang 0057, Yuan Zhang 0013, Jinyao Yan, Tao Lin 0001
MMM (2)4
2023 CLAPS: Curriculum Learning-Based Adaptive Bitrate and Preloading for Short Video Streaming
abstract
To provide high user QoE while maintaining low bandwidth waste, it is important to design adaptive bitrate and preloading algorithms for short video streaming. Current solutions either have relatively low performance, as observed in heuristic algorithms, or suffer the problem of poor generalization, as seen in deep reinforcement learning (DRL)-based algorithms. To address this issue, we propose CLAPS, a curriculum learning-based DRL model that enhances the generalization of the DRL model across a wide range of data, while ensuring high performance. CLAPS introduces a comprehensive metric that measures the curriculum difficulty by combing the performance gap between an existing heuristic algorithm and the DRL model with the prediction error of network bandwidth. Moreover, we design a training scheduler to control sampling proportion based on Markov transition probabilities to address the model forgetting problem. Extensive evaluations using real video datasets and network traces including 5G, 4G, and Wi-Fi demonstrate that CLAPS outperforms all the baseline algorithms. Specifically, CLAPS improves the overall performance by 10.04%-13.84% and the generalization by 22.52% -35.77% compared to the best-performing DRL baseline.
Fengzhou Sun, Hao Yang 0057, Tao Lin 0001, Yuan Zhang 0013, Zheng Chen 0019, Jinyao Yan
MMSP4
2023 Beyond OCR + VQA: Towards end-to-end reading and reasoning for robust and accurate textvqa
abstract
Text-based visual question answering (TextVQA), which answers a visual question by considering both visual contents and scene texts, has attracted increasing attention recently. Most existing methods employ an optical character recognition (OCR) module as a pre-processor to read texts, then combine it with a visual question answering (VQA) framework. However, inaccurate OCR results may lead to cumulative error propagation , and the correlation between text reading and text-based reasoning is not fully exploited. In this work, we integrate OCR into the flow of TextVQA, targeting the mutual reinforcement of OCR and VQA tasks. Specifically, a visually enhanced text embedding module is proposed to predict semantic features from the visual information of texts, by which texts can be reasonably understood even without accurate recognition. Further, two elaborate schemes are developed to leverage contextual information in VQA to modify OCR results. The first scheme is a reading modification module that adaptively selects the answer results according to the contexts. Second, we propose an efficient end-to-end text reading and reasoning network, where the downstream VQA signal contributes to the optimization of text reading. Extensive experiments show that our method outperforms existing alternatives in terms of accuracy and robustness, whether ground truth OCR annotations are used or not.
Gangyan Zeng, Yuan Zhang 0013, Yu Zhou 0015, Weiping Wang 0005, Xu-Cheng Yin
Pattern Recognit.2
2023 Latent-Domain Predictive Neural Speech Coding
abstract
Neural audio/speech coding has recently demonstrated its capability to deliver high quality at much lower bitrates than traditional methods. However, existing neural audio/speech codecs employ either acoustic features or learned blind features with a convolutional neural network for encoding, by which there are still temporal redundancies within encoded features. This article introduces latent-domain predictive coding into the VQ-VAE framework to fully remove such redundancies and proposes the TF-Codec for low-latency neural speech coding in an end-to-end manner. Specifically, the extracted features are encoded conditioned on a prediction from past quantized latent frames so that temporal correlations are further removed. Moreover, we introduce a learnable compression on the time-frequency input to adaptively adjust the attention paid to main frequencies and details at different bitrates. A differentiable vector quantization scheme based on distance-to-soft mapping and Gumbel-Softmax is proposed to better model the latent distributions with rate constraint. Subjective results on multilingual speech datasets show that, with low latency, the proposed TF-Codec at 1 kbps achieves significantly better quality than Opus at 9 kbps, and TF-Codec at 3 kbps outperforms both EVS at 9.6 kbps and Opus at 12 kbps. Numerous studies are conducted to demonstrate the effectiveness of these techniques.
Xiulian Peng, Huaying Xue, Yuan Zhang 0013, Yan Lu 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2023 QoE-Oriented Mobile Virtual Reality Game in Distributed Edge Networks
abstract
Mobile edge computing is a promising framework for mobile virtual reality (VR) game. Although there are several existing studies on the edge assisted mobile VR game system, they lack the consideration of provisioning services with satisfactory QoE to a large number of users. In this paper, we consider the problem of providing QoE-oriented edge assisted mobile VR game as a service to multiple users, with a comprehensive QoE concern of both visual and delay aspects. Due to the unique features of mobile VR game, the problem is formulated into a Mixed Integer Quadratically Constrained Quadratic Programming (MIQCQP) problem. We show that the problem is NP-hard with object placement decision and rendering level selection decision quadratically coupling together. To solve this problem, we propose the Alternating Directions Method of Multipliers (ADMM) algorithm which can iteratively decouple the quadratic terms and reform the problem into the efficiently solvable MIQCQP-1 (i.e., MIQCQP with one constraint) problem. Trace driven simulation shows that our algorithm fits the edge assisted mobile VR game scenario well with fast computation time (at least 4 orders of magnitude less computation time compared to Gurobi solver) and good performance (at least 18% of user visual QoE improvement compared to other mobile VR scheme).
Yuan Zhang 0013, Lingjun Pu, Tao Lin 0001, Jinyao Yan
IEEE Trans. Multim.1
2022 End-to-End Neural Speech Coding for Real-Time Communications
abstract
Deep-learning based methods have shown their advantages in audio coding over traditional ones but limited attention has been paid on real-time communications (RTC). This paper proposes the TFNet, an end-to-end neural speech codec with low latency for RTC. It takes an encoder-temporal filtering-decoder paradigm that has seldom been investigated in audio coding. An interleaved structure is proposed for temporal filtering to capture both short-term and long-term temporal dependencies. Furthermore, with end-to-end optimization, the TFNet is jointly optimized with speech enhancement and packet loss concealment, yielding a one-for-all network for three tasks. Both subjective and objective results demonstrate the efficiency of the proposed TFNet.
Xiulian Peng, Huaying Xue, Yuan Zhang 0013, Yan Lu 0001
ICASSP5
2022 Cross-Scale Vector Quantization for Scalable Neural Speech Coding
abstract
Bitrate scalability is a desirable feature for audio coding in real-time communications.Existing neural audio codecs usually enforce a specific bitrate during training, so different models need to be trained for each target bitrate, which increases the memory footprint at the sender and the receiver side and transcoding is often needed to support multiple receivers.In this paper, we introduce a cross-scale scalable vector quantization scheme (CSVQ), in which multi-scale features are encoded progressively with stepwise feature fusion and refinement.In this way, a coarse-level signal is reconstructed if only a portion of the bitstream is received, and progressively improves the quality as more bits are available.The proposed CSVQ scheme can be flexibly applied to any neural audio coding network with a mirrored auto-encoder structure to achieve bitrate scalability.Subjective results show that the proposed scheme outperforms the classical residual VQ (RVQ) with scalability.Moreover, the proposed CSVQ at 3 kbps outperforms Opus at 9 kbps and Lyra at 3kbps and it could provide a graceful quality boost with bitrate increase.
Xiulian Peng, Huaying Xue, Yuan Zhang 0013, Yan Lu 0001
INTERSPEECH4
2022 DAM: Deep Reinforcement Learning based Preload Algorithm with Action Masking for Short Video Streaming
abstract
Short video streaming has been increasingly popular in recent years. Due to its unique user behavior of watching and sliding, a critical technique issue is to design a preload algorithm deciding which video chunk to download next, bitrate selection and the pause time, in order to improve user experience while reducing bandwidth wastage. However, designing such a preload algorithm is non-trivial, especially taking into account conflicting goals of improving QoE and reducing bandwidth wastage. In this paper, we propose a deep reinforcement learning-based approach to simultaneously decide the aforementioned three decision variables via learning an optimal policy under a complex environment of varying network conditions and unpredictable user behavior. In particular, we incorporate domain knowledge into the decision procedure via action masking to make decisions more transparent, and accelerate the model training. Experimental results validate the proposed approach significantly outperforms baseline algorithms in terms of QoE metrics and bandwidth wastage.
Size Qian, Yuhong Xie, Zipeng Pan, Yuan Zhang 0013, Tao Lin 0001
ACM Multimedia4
2022 TextBlock: Towards Scene Text Spotting without Fine-grained Detection
abstract
Scene text spotting systems which integrate text detection and recognition modules have witnessed a lot of success in recent years. Existing works mostly follow the framework of word/character-level fine-grained detection and isolated-instance recognition, which overemphasize the role of detector and ignore the rich context information in recognition. After rethinking the conventional framework, and inspired by the glimpse-focus spotting pipeline of human beings, we ask:1) "can machine spot text without accurate detection just like human beings?", and if yes, 2) "is text block another alternative for scene text spotting other than word or character?". Based on these questions, we propose a new perspective of coarse-grained detection with multi-instance recognition for text spotting. Specifically, a pioneering network termed TextBlock is developed, and a heuristic text block generation method as well as a multi-instance block-level recognition module are proposed. In this way, the burden of detection is relieved, and the contextual semantic information is well explored for recognition. To train the block-level recognizer, a synthetic dataset including about 800K images is formed. As a by-product of attention, fine-grained detection can be recovered with the recognizer. Equipped with a detector without many bells and whistles (e.g., Faster R-CNN), TextBlock achieves competitive or even better performance compared with previous sophisticated text spotters on several public benchmarks. As a primary attempt, we expect this framework will have a potential impact on scene text spotting research in the future.
Yuan Zhang 0013, Yu Zhou 0015, Gangyan Zeng, Youhui Guo, Haiying Wu, Weiping Wang 0005
ACM Multimedia2
2022 Reinforcement Learning based Low Delay Rate Control for HEVC Region of Interest Coding
abstract
Rate control is one of the most critical technologies of real-time video coding. It aims to make bit rate allocation and quantization parameter(QP) decisions to minimize video distortion while reducing buffering delay. However, existing solutions suffer from the problem of extra encoding delay and lack joint optimization of the region of interest(ROI) quality and buffer latency. In this paper, we propose a deep reinforcement learning(RL) based low-delay rate control method without using the information of uncoded frames to avoid the extra delay. Particularly, our RL-based rate control algorithm outputs two types of policies. The first is to make frame-level QP decisions to stabilize buffer occupancy and optimize the overall quality, while the other is in charge of adjusting the QP offset between ROI and Non-ROI regions to enhance portrait quality. Extensive experiments verify that our proposed method improves ROI and overall quality while reducing the buffer occupancy variation, compared with baseline algorithms.
Naifu Xue, Yuan Zhang 0013, Tao Lin 0001
MMSP2
2022 Time-Variance Aware Dynamic Kernel Generation for Real-Time Acoustic Echo Cancellation
abstract
Time-variant factors including dynamic delay and varying echo path often occur in real-world acoustic echo cancellation (AEC) applications. Current end-to-end deep neural network (DNN) based methods usually model the time-variant components implicitly and can hardly handle the unpredictable time-variance in real-time AEC. To explicitly capture the time-variant components, we propose a dynamic kernel generation (DKG) module that can be introduced as a learnable plug-in to a DNN-based end-to-end pipeline. Specifically, the DKG module generates a convolutional kernel regarding to each input audio frame, so that the DNN model is able to dynamically adjust its weights according to the input signal during inference. Experimental results verify that DKG module improves the AEC performance of the model under time-variant scenarios, especially in the double-talk cases.
Xiulian Peng, Yuan Zhang 0013, Yan Lu 0001
IEEE Signal Process. Lett.4
2021 Interactive Speech and Noise Modeling for Speech Enhancement
abstract
Speech enhancement is challenging because of the diversity of background noise types. Most of the existing methods are focused on modelling the speech rather than the noise. In this paper, we propose a novel idea to model speech and noise simultaneously in a two-branch convolutional neural network, namely SN-Net. In SN-Net, the two branches predict speech and noise, respectively. Instead of information fusion only at the final output layer, interaction modules are introduced at several intermediate feature domains between the two branches to benefit each other. Such an interaction can leverage features learned from one branch to counteract the undesired part and restore the missing component of the other and thus enhance their discrimination capabilities. We also design a feature extraction module, namely residual-convolution-and-attention (RA), to capture the correlations along temporal and frequency dimensions for both the speech and the noises. Evaluations on public datasets show that the interaction module plays a key role in simultaneous modeling and the SN-Net outperforms the state-of-the-art by a large margin on various evaluation metrics. The proposed SN-Net also shows superior performance for speaker separation.
Xiulian Peng, Yuan Zhang 0013, Sriram Srinivasan 0003, Yan Lu 0001
AAAI3
2021 PIMNet: A Parallel, Iterative and Mimicking Network for Scene Text Recognition
abstract
Nowadays, scene text recognition has attracted more and more attention due to its various applications. Most state-of-the-art methods adopt an encoder-decoder framework with attention mechanism, which generates text autoregressively from left to right. Despite the convincing performance, the speed is limited because of the one-by-one decoding strategy. As opposed to autoregressive models, non-autoregressive models predict the results in parallel with a much shorter inference time, but the accuracy falls behind the autoregressive counterpart considerably. In this paper, we propose a Parallel, Iterative and Mimicking Network (PIMNet) to balance accuracy and efficiency. Specifically, PIMNet adopts a parallel attention mechanism to predict the text faster and an iterative generation mechanism to make the predictions more accurate. In each iteration, the context information is fully explored. To improve learning of the hidden layer, we exploit the mimicking learning in the training phase, where an additional autoregressive decoder is adopted and the parallel decoder mimics the autoregressive decoder with fitting outputs of the hidden layer. With the shared backbone between the two decoders, the proposed PIMNet can be trained end-to-end without pre-training. During inference, the branch of the autoregressive decoder is removed for a faster speed. Extensive experiments on public benchmarks demonstrate the effectiveness and efficiency of PIMNet. Our code is available in the supplementary material.
Yu Zhou 0015, Wei Wang 0315, Yuan Zhang 0013, Weiping Wang 0005
ACM Multimedia5
2021 Beyond OCR + VQA: Involving OCR into the Flow for Robust and Accurate TextVQA
abstract
Text-based visual question answering (TextVQA) requires analyzing both the visual contents and texts in an image to answer a question, which is more practical than general visual question answering (VQA). Existing efforts tend to regard optical character recognition (OCR) as a pre-processing and then combine it with a VQA framework. It makes the performance of multimodal reasoning and question answering highly depend on the accuracy of OCR. In this work, we address this issue with two perspectives. First, we take advantages of multimodal cues to complete the semantic information of texts. A visually enhanced text embedding is proposed to enable understanding of texts without accurately recognizing them. Second, we further leverage rich contextual information to modify the answer texts even if the OCR module does not correctly recognize them. In addition, the visual objects are endued with semantic representations to enable objects in the same semantic space as OCR tokens. Equipped with these techniques, the cumulative error propagation caused by poor OCR performance is effectively suppressed. Extensive experiments on TextVQA and ST-VQA datasets demonstrate that our approach achieves the state-of-the-art performance in terms of accuracy and robustness.
Gangyan Zeng, Yuan Zhang 0013, Yu Zhou 0015
ACM Multimedia2
2021 A Hybrid Receiver-side Congestion Control Scheme for Web Real-time Communication
abstract
Web real-time communication (WebRTC) employs congestion control to ensure the quality of experience (QoE). Different from congestion control schemes for TCP, WebRTC keeps a low-level playback buffer that considers excessively delayed packets as losses, which makes the congestion control for WebRTC more challenging. Existing heuristic schemes estimate the network conditions based on hand-crafted rules that may be suboptimal, leading to under-utilization or over-utilization of link capacity in many cases. On the other hand, the existing learning-based schemes train a model that acts in a large action space, which is hard to converge to a stable status and has low performance over unpredictable network conditions. In this paper, we propose a hybrid receiver-side congestion control (HRCC) framework, which combines a heuristic congestion control scheme with an RL-Agent that periodically generates a gain coefficient to tune the bandwidth estimated by the heuristic scheme. Extensive simulation experiments demonstrate that the HRCC's RL-Agent effectively tunes the bandwidth estimate of the heuristic scheme. The hybrid scheme achieves higher bandwidth utilization than the fully heuristic scheme with similar queuing delay and packet loss and outperforms the fully RL-based scheme on overall performance.
Yuan Zhang 0013, Size Qian, Zipeng Pan, Yuhong Xie
MMSys2
2021 A Cost-Efficient Framework for Scene Text Detection in the Wild
Gangyan Zeng, Yuan Zhang 0013, Yu Zhou 0015
PRICAI (1)2
2020 Dynamic Distributed Edge Resource Provisioning via Online Learning across Timescales
abstract
The strategic management of distributed resources of mobile edge computing networks often requires managing different system components over different timescales. In this paper, we formulate a nonlinear mixed-integer program to capture the online optimization of the edge network’s long-term cost, where we distribute workload more frequently on the fast timescale and provision resources less frequently on the slow timescale. We design a novel online learning framework consisting of three algorithms to make fast-timescale and slow-timescale fractional decisions, respectively, and round such decisions into integers. Our algorithms run in polynomial time in an online manner, jointly solving the original NP-hard problem that can contain arbitrary and unpredictable inputs. Via rigorous formal analysis, we prove a parameterized-constant competitive ratio as the performance guarantee for our approach. We conduct extensive evaluations with real-world data and confirm our approach’s superiority over existing practices and state-of-the-arts.
Wencong You, Lei Jiao 0002, Sourav Bhattacharya, Yuan Zhang 0013
SECON4
2020 Dynamic Component Placement and Request Scheduling for IoT Big Data Streaming
abstract
Internet-of-Things (IoT) big data streaming applications, such as video surveillance and automatic driving, tend to use mobile-edge computing (MEC) infrastructure to enhance their performance and augment their functionalities. Although extensive previous studies have worked on offloading requests to MEC servers, none of them has comprehensively and thoroughly considered the important features of IoT data streaming applications (i.e., component dependency and dynamic arrival) and the infrastructure provisioning (i.e., capacity constraint and colocation interference). In this article, we consider the offloading problem for dynamically arrived IoT data streaming requests on MEC servers in real time. We model it as a delay-sensitive multiuser multiresource online offloading problem respecting component dependency and capacity constraint. The problem is NP-hard with offloading decisions coupling together. To solve it, we decouple the problem into component placement problem and request scheduling problem and propose a two-stage DPGPD algorithm with polynomial time complexity. We show the first stage dynamic programming (DP) algorithm is the optimal solution and the second-stage greedy primal-dual (GPD) algorithm is asymptotic optimal. The simulation results show that our solution is effective yet efficient compared to benchmark solutions. (DP provides the optimal placement layout with 12× less decision time of Gurobi; and GPD provides the asymptotic optimal scheduling with 5× less average waiting time compared to least work left (LWL) in heavy workload.) We implement a dedicated prototype and exploit several representative big data streaming applications to evaluate it. Lab-scale experiment shows that our solution can provide over 3× less total completion time compared to local execution.
Yuan Zhang 0013, Jinyao Yan, Lingjun Pu
IEEE Internet Things J.1
2019 A Hybrid Control Scheme for Adaptive Live Streaming
abstract
The live streaming is more challenging than on-demand streaming, because the low latency is also a strong requirement in addition to the trade-off between video quality and jitters in playback. To balance several inherently conflicting performance metrics and improve the overall quality of experience (QoE), many adaptation schemes have been proposed. Bitrate adaptation is one of the major solution for video streaming under time-varying network conditions, which works even better combining with some latency control methods, such as adaptive playback rate control and frame dropping. However, it still remains a challenging problem to design an algorithm to combine these adaptation schemes together. To tackle this problem, we propose a hybrid control scheme for adaptive live streaming, namely HYSA, based on heuristic playback rate control, latency-constrained bitrate control and QoE-oriented adaptive frame dropping. The proposed scheme utilizes Kaufman's Adaptive Moving Average (KAMA) to predict segment bitrates for better rate decisions. Extensive simulations demonstrate that HYSA outperforms most of the existing adaptation schemes on overall QoE.
Huan Peng, Yuan Zhang 0013, Yongbei Yang, Jinyao Yan
ACM Multimedia2
2019 A neural network approach to GOP-level rate control of x265 using Lookahead
abstract
To optimize the perceived quality under a specific bitrate constraint, multi-pass encoding is usually performed with the rate control mode of the average bitrate (ABR) or the constant rate factor (CRF) to distribute bits as reasonably as possible in terms of perceived quality, leading to high computational complexity. In this paper, we propose to utilize the video information generated during the encoding to adaptively adjust the CRF setting at GOP level, ensuring the bits of frames in each GOP are allocated reasonably under the bitrate constraint with a single-pass encoding framework. In particular, due to the inherent relationship between CRF values and bitrates, we adopt a shallow neural network (NN) to map video content features to the CRF-bitrate model. The content-related features are collected from the lookahead module inside the x265 encoder, including encoding cost estimation, motion vector and so on. Further, a rate control method, called content adaptive rate factor (CARF), is proposed to adjust the CRF value of each GOP with the requirement of the target bitrate by using the predicted CRF- bitrate models of each GOP. The experimental results show that the proposed approach can make 84.5% testing data within 20% bitrate error (or better) and outperform the ABR mode in x265, leading to 5.23 % BD-rate reduction on average.
Boya Cheng, Yuan Zhang 0013
PCS2
2019 Dynamic Service Placement for Virtual Reality Group Gaming on Mobile Edge Cloudlets
abstract
To realize mobile virtual reality (VR) group gaming services which are currently hampered by the prohibitive bandwidth and the stringent delay requirements, we investigate the problem of provisioning such services using the emerging mobile edge cloudlet (MEC) networks with a distributed content rendering architecture. The underlying dynamic rendering-module placement problem requires to optimize the service’s operational cost and the users’ end-to-end performance, involving multiple intertwined conflicting system objectives that are discrete, nonconvex, and higher degree polynomial functions with coupled decisions and arbitrary user dynamics over time. We solve this online placement problem by leveraging model predictive control (MPC) and overcoming the aforementioned challenges over each prediction window. We explore the connection between the placement problem and the minimal$s$-$t$cut problem in graph theory and solve the former via solving a series of instances of the latter. We formally prove the performance guarantee of our approach. We also conduct extensive trace-driven evaluations and demonstrate the superior practical performance of our MPC-based approach compared to thede factopractices and the state-of-the-art alternatives.
Yuan Zhang 0013, Lei Jiao 0002, Jinyao Yan, Xiaojun Lin 0001
IEEE J. Sel. Areas Commun.1
2017 Perceptual optimized adaptive HTTP streaming
abstract
The paper presents a perceptual optimized adaptive HTTP streaming scheme to improve the quality of experience (QoE). In addition to barely controlling the bandwidth and buffer size in existing works, this paper integrates the video saliency based adaptation to improve the perceptual quality for end users. Resources (e.g., buffer) are well managed using saliency cues to ensure the smooth quality under a given network bandwidth. Algorithms are implemented on top of the open-source DASH platform - dash.js to demonstrate our superiority compared with the default throughput-based adaptation and well-known BOLA method. Our work can be a potential enhancement for current DASH standard to offer the smooth and perceptual optimized adaptive streaming in mobile networks.
Huaying Xue, Yuan Zhang 0013, Jinyao Yan
VCIP2
2014 Fine-grained multi-resource scheduling in cloud datacenters
abstract
Cloud datacenters typically require tenants to specify the resource demands for the virtual machines (VMs) they create using a set of pre-defined, fixed configurations, to ease the resource allocation problem. Unfortunately, this leads to low resource utilization of cloud datacenters as tenants are obligated to conservatively predict the maximum resource demand of their applications. We argue that instead of such a static VM resource allocation, a finer-grained dynamic resource allocation and scheduling can substantially improve the utilization of the datacenter resources by increasing the number of jobs accommodated and correspondingly, the cloud datacenter provider's revenue. The dynamic real-time scheduling of jobs can also ensure that the performance goals for the tenant VMs are achieved. Examining a typical publicly available cluster data center trace, we observe that a large number of jobs are short. Only a small proportion of jobs are long and which require substantial compute or memory resources. We propose an optimization based approach that exploits this division between the short and long jobs to dynamically allocate a cloud datacenter's resources to achieve significantly better utilization by increasing the number of jobs accommodated by the datacenter. We use a constraint programming solution to schedule the long jobs, and use simple heuristics to quickly, yet quite accurately schedule the short jobs. Using trace-driven simulations based on public traces collected on provider cluster we show that the overall revenue for the cloud provider can be improved by 30% over the traditional static VM resource allocation based on the coarse granularity specifications. We are able to increase the number of jobs accommodated using dynamic scheduling by 18%. We also compare the performance of our approach to multi-resource (CPU and memory) first-fit and best-fit algorithms and to the optimal offline solution, and demonstrate that our solution achieves within 76% of the offline optimal solution.
Yuan Zhang 0013, Xiaoming Fu 0001, K. K. Ramakrishnan
LANMAN1
2014 Fast mode decision for error resilient video coding
abstract
The error resilience and low-complexity video encoding are two major requirements of real-time visual communications on mobile devices. To address the two requirements simultaneously, this paper presents a fast mode decision algorithm for the error resilient video coding in packet loss environment. The proposed algorithm is a two-step method: early skip mode decision and early intra mode decision. Different from the existing methods for early skip mode decision, the proposed method takes the error-propagation distortion into account in estimating the coding cost. Considering the intra blocks are frequently used to terminate the error propagations, we also propose a method to fast estimate the intra block coding cost, so that the intra mode can be early determined. Overall, the proposed method can significantly reduce the encoding time while keeping the coding efficiency similar to the rate-distortion optimized mode decision method.
Yunong Wei, Yuan Zhang 0013, Jinyao Yan
MMSP2
2013 Classification based fast mode decision for stereo video coding
abstract
We propose a classification based fast mode decision scheme in stereo video coding. By treating mode decision as a classification problem, our scheme employs a decision tree classifier to separate out SKIP mode, which is the major and most computationally efficient mode in stereo video coding. Thus, we can pre-decide whether the current macroblock is coded as SKIP mode without going through the exhaustive mode decision process. Experimental results show that this scheme provides 30~75% of time saving over a wide range of quantization parameter values.
Yuan Zhang 0013, Pamela C. Cosman
ICIP2
2011 Fast mode decision for H.264 video coding in packet loss environment
abstract
In this paper, we propose a fast mode decision scheme for H.264 video coding to address the requirements of both low complexity and error resilience in realtime video communications. Traditional fast mode decision schemes are usually designed based on feature analysis on the source videos. However, the rate-distortion behaviors of the coding modes change when channel errors are involved. Therefore, the existing fast algorithms may not be applicable in error-resilient video coding. We first study the end-to-end rate-distortion behaviors of various coding modes, and then derive a hierarchical mode decision scheme. Different from the traditional fast algorithms that separate skip and inter modes at the beginning, we propose a fast estimation of the coding costs of skip and intra modes in a packet-loss environment, and then narrow the mode decision into one of the two paths: non-intra and non-skip. Testing shows the significant time savings of the proposed algorithm.
Yuan Zhang 0013, Pamela C. Cosman
ICIP1