Hsu-Feng Hsiao

dblp:58/3489 · DBLP profile ↗
← Back
30ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0003-3414-5622ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 5 first-author · 5 since 2021Systems, architecture and hardware · 12 · 1 first-author · 5 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 Saliency-guided video coding via recurrent learning and perceptual quality assessment
Tz-Cheng Chang, Hsu-Feng Hsiao
Signal Process. Image Commun.2
2025 DCM-VideoNet: A Densely-Connected Modulated Decoder Framework for Implicit Neural Video Compression
abstract
We introduce DCM-VideoNet, a novel implicit neural representation that enhances video reconstruction and compression through Densely-Connected Modulated Decoder (DCMD) blocks, enabling efficient feature fusion and robust representation learning. By incorporating Compact Inverted Bottleneck (CIB) structures, these decoders achieve better parameter efficiency without compromising expressive capability. Integrated modulation mechanisms further refine feature extraction, leading to improved reconstruction fidelity. To address the spectral bias commonly observed in neural networks, we propose a Spectral Bias Mitigation Loss (SBM Loss) with adaptive frequency weighting, ensuring a balanced evaluation of reconstruction quality during training. For video compression, our method employs a comprehensive strategy that includes layer-wise pruning, quantization-aware training, and arithmetic coding to realize efficient model compression. Extensive experiments demonstrate that DCM-VideoNet not only outperforms state-of-the-art implicit neural representation techniques but also achieves competitive performance relative to leading learning-based video compression methods.
Cih-Wei Wong, Hsu-Feng Hsiao
ICIP2
2025 BACF-Net: An Attention-Convolution Fusion Architecture for Learned Image Compression
abstract
In the era of big data, efficient image compression is essential for managing the rapid growth of visual content. While traditional codecs have advanced significantly, their dependence on handcrafted techniques limits their effectiveness. The emergence of deep learning has transformed image compression by enabling end-to-end optimization. However, existing learned methods typically utilize convolutional neural networks for local modeling or transformers for capturing long-range dependencies, indicating potential areas for enhancement. We introduce a novel hybrid hyperprior-based architecture that combines the advantages of residual CNNs and transformers to improve rate-distortion performance in learned image compression. Additionally, our proposed Bifurcated Attention-Convolution Fusion (BACF) block employs a parallel configuration of an enhanced residual CNN with split attention alongside mixed transformer variants for multi-axis attention and shifted window-based attention. This design allows the network to effectively process and integrate both local details and high-level semantic information. Extensive experiments on the Kodak, CLIC, and Tecnick datasets show that our proposed method achieves competitive rate-distortion performance.
Chen-Lin Chang, Hsu-Feng Hsiao
ISCAS2
2025 A Unified SpatioTemporal Network with Structural Pruning for Video Action Recognition
abstract
Video action recognition poses significant challenges in capturing and integrating the complex spatiotemporal patterns and motion dynamics necessary for robust understanding. Despite recent advancements, existing deep learning approaches often struggle to efficiently model these interactions over extended temporal ranges. To address this, we propose the Unified SpatioTemporal Network (USTN), a novel framework that fuses segment-level spatiotemporal features with long-range temporal difference information. By strategically employing sparse frame sampling, USTN constructs a rich, coarse-grained representation encapsulating both spatial structure and temporal evolution. Furthermore, we introduce a structural pruning technique to identify and remove redundant parameters, mitigating overfitting and enhancing computational efficiency without compromising performance. Extensive evaluations on the challenging UCF101 and HMDB51 benchmarks, using USTN instantiated with ResNet backbones, demonstrate the superiority of our approach.
Yang-Jie Chen, Hsu-Feng Hsiao
ISCAS3
2025 Spatiotemporally Modulated Dual-MLP Architecture for Neural Video Representation and Compression
abstract
This paper presents a novel modulated dual-MLP framework for efficient neural video representation and compression using implicit neural representations. Our approach overcomes limitations in existing methods by integrating a modulator network with sinusoidal activation functions within a dual-MLP architecture. Compression performance is further improved through advanced techniques such as layer-adaptive pruning, quantization, and context-adaptive binary arithmetic coding. Pruning is applied independently to each layer, and our pipeline retains all necessary metadata, enabling accurate recovery of compressed weights and model reconstruction for thorough evaluation. Additionally, our framework supports color space conversion; fine-tuning RGB-trained models with YUV420 data allows seamless transitions with minimal architectural changes, which is beneficial for digital broadcasting and streaming, where YUV formats are common. Extensive experiments on the challenging UVG dataset show that our method achieves up to 0.89 dB higher PSNR than leading INR approaches at 12.5 million parameters, while also delivering improved rate-distortion performance compared to other INR methods and popular codecs like AVC and HEVC.
Gong-Yao Wu, Cih-Wei Wong, Hsu-Feng Hsiao
ISCAS3
2025 Scene-Adaptive Neural Video Compression with Dynamic Resource Allocation
Yao-Wei Yang, Hsu-Feng Hsiao
PCS2
2023 Enhanced Pedestrian Trajectory Prediction via the Cross-Modal Feature Fusion Transformer
abstract
We address the challenge of predicting pedestrian trajectories in videos, a task inherently complex due to the diverse and intricate nature of human motion and interactions within their environment. The accurate anticipation of trajectories necessitates a holistic comprehension of the temporal evolution of past events in videos. Regrettably, existing methods often neglect the fusion of critical features, such as human behavior, motion, and interaction, thereby limiting their efficacy in tackling these challenges. To overcome these limitations, we propose the Cross-modal Feature Fusion Transformer, a novel approach for pedestrian trajectory prediction. Our model seamlessly integrates multimodal features, including human behavior, position, speed, and interaction with surroundings, to effectively encapsulate the temporal progression of observed frames. It consists of transformer-based cross-modal fusion encoder and decoder modules, adeptly melding the interactions between the multimodal features through a multi-head co-attentional mechanism. This enables the precise prediction of future trajectories. Additionally, we incorporate auxiliary self-supervised future prediction losses to learn the temporal evolution of past and future multimodal features. We evaluate our approach on ETH/UCY and ActEV/VIRAT datasets and demonstrate its superior performance compared to state-of-the-art methods.
Hsu-Feng Hsiao
VCIP2
2023 A Hybrid Convolutional and Transformer Network for Salient Object Detection
abstract
We present a novel hybrid architecture that seamlessly merges transformers and convolutional neural networks to enhance the performance of RGB-D salient object detection. Transformer-based models have recently demonstrated their potential in this field, owing to their unique ability to encode long-range information via the self-attention mechanism. This mechanism adeptly mirrors human visual perception by capturing long-distance dependencies and selectively focusing on the most relevant sections of the input image. In contrast, convolutional neural networks, with their robust generalization and trainability, have proven to be invaluable for a wide array of image processing tasks. By fusing the strengths of these two models, our proposed hybrid architecture outperforms the effectiveness of using either transformers or convolutional neural networks in isolation. Our architecture employs an encoder-decoder framework. Within this structure, the hybrid model functions as the feature encoder, while the decoder integrates a convolutional neural network with deep layer aggregation to adeptly merge features of varying resolutions derived from the transformer-based encoder. This strategic design choice exploits the computational modeling prowess of convolutional neural networks in tasks such as saliency prediction, while also benefiting from the long-range dependency modeling offered by the hybrid model. We also use a Siamese architecture with shared parameters in the encoder to concurrently learn salient features from RGB and depth data. By harnessing the complementary strengths of both models, the proposed hybrid architecture has demonstrated superior performance.
Bei-Sin Li, Hsu-Feng Hsiao
VCIP2
2022 A Scale-Reductive Pooling with Majority-Take-All for Salient Object Detection
abstract
With the rapid development of hardware and related technologies, salient object detection based on deep learning methods has become one of the popular research topics in computer vision applications. For the detection focused on the integrity of salient objects, edge accuracy of objects is one of the important indicators in the evaluation of visual saliency detection. However, in deep learning-based methods, complex networks and large amounts of data are usually required to achieve good boundary accuracy. To solve this issue, a scale-reductive pooling approach with clustering-based majority-take-all strategy is proposed in this paper. According to the experimental results, we show that the prediction results are improved with reasonable quantity of superpixels.
Chin-Han Shen, Yang-Jie Chen, Hsu-Feng Hsiao
ISCAS3
2021 Content Distribution Network for Streaming Using Multiple Galois Fields
abstract
In this paper, the architecture of random linear network coding based on the hybrid coding in multiple Galois field sizes is proposed. Random linear network coding is an efficient network coding approach that enables network to generate independently and randomly linear mapping between input and output symbols over finite field. With the proposed reduction method, coded symbols and coefficients in higher degree of Galois field can be converted to symbols and coefficients in GF(2) so that hybrid coding in multiple Galois field sizes can be made possible. Therefore, peers in the heterogeneous environments can all benefit from the proposed content distribution network for streaming to better utilize the network resource.
Tsung-Kang Hung, Sachin K. Kaushal, Hsu-Feng Hsiao
ISCAS3
2020 Supervoxel Segmentation using Spatio-Temporal Lazy Random Walks
abstract
Superpixel segmentation has been proved its effectiveness as a preprocessing step for many applications. Similarly, the over-segmentation with temporal consistency of a video can be useful for quite a few researches. In this paper, a novel supervoxel segmentation algorithm based on lazy random walk is proposed. In the proposed framework, a superpixel segmentation is first applied to each frame and the produced superpixels are called special-pixels. Then, we propose a generalized spatio-temporal adjacency matrix, which keeps the information of special-pixels with neighboring relationship, for the generation of supervoxels using lazy random walk approach. In addition, the center relocation and the splitting strategy of supervoxels are proposed to improve the quality of the supervoxels.
Yi-Xuan Zhan, Chin-Han Shen, Hsu-Feng Hsiao
ISCAS3
2019 Saliency Detection with Multi-Contextual Models and Spatially Coherent Loss Function
abstract
We have proposed a multi-contextual model architecture with color and depth information considered independently in this work. To utilize the feature maps of different levels better, short connection structures are used to integrate the knowledge from color and depth data separately. A novel loss function considering three criteria is proposed to improve the detection accuracy and spatial coherence of the detected results. The training process of the proposed network is divided into two stages, a pre-training phase and a refinement phase to increase the efficiency of the network.
Po-Sheng Huang, Chin-Han Shen, Hsu-Feng Hsiao
ISCAS3
2018 Expectation Model and Scheduling for Video Streaming
abstract
Resource scheduling is important to improve quality of experience for the services of multimedia streaming. In this paper, a scheduling algorithm of modulation and coding schemes for layered video streaming over wireless broadband networks is proposed. We suggest using the effective throughput, defined as the decompressible data rate, to formulate the utility of a streaming service, because of the stronger relationship between the video quality and the effective throughput. In the proposed method, we address the scheduling problem using a designed expectation model. The model is used to estimate the utility according to the expected contribution of coded blocks. Then, we develop a coarse to fine grained resource scheduling algorithm to solve the scheduling problem effectively.
Yu-Jui Chiang, Hsu-Feng Hsiao
ISCAS2
2017 Video streaming optimization using degradation estimation with unequal error protection
abstract
Videos compressed using modern compression techniques, such as HEVC, typically have the property of unequal importance. When a video with intra and inter coded frames is transmitted through a network, its different frames can suffer from quality degradation depending on their sizes and channel coding schemes. Moreover, errors in a reference frame can propagate forward and backward over several frames, while errors in a non-reference frame are localized within the same frame. In this paper, we design a system that dynamically determines the coding parameters of layer-aligned multipriority rateless codes depending on the video content and channel condition. For this purpose, we estimate the strength of error propagation and develop a model to estimate quality degradation of a transmitted video accurately. Through minimizing quality degradation, we are able to calculate optimal parameters for the system.
Philip Tovstogan, Hsu-Feng Hsiao
ISCAS2
2015 QoS-driven optimization for video streaming using layer-aligned multipriority rateless codes
abstract
Forward error correction techniques are commonly used for delay-sensitive video streaming applications. Streaming of compressed videos can usually withstand a certain level of packet loss, depending on the compression methods and error concealment techniques. In this paper, a QoS-driven optimization algorithm is proposed for the streaming of scalable videos using layer-aligned multipriority rateless codes over wideband wireless channels. Layer-aligned multipriority rateless codes provide controllable protection strengths for different layers of a scalable video. The coding is based on the N-cycle layer-aligned overlapping structure where each layer of scalable videos is protected by N windows. The objective of the proposed method in this paper is to decide a coding strategy by minimizing the total data rate for scalable video streaming, while fulfilling user's expectation of desired recovery rates of different video layers. The developed QoS-driven optimization algorithm is based on the prediction model of layer-aligned multipriority rateless codes. Its better performance is verified with simulations and real-world experiments over 4G wireless networks.
Lien-En Hung, Hsu-Feng Hsiao
ISCAS2
2014 QoE-driven performance analysis of cloud gaming services
abstract
With the popularity of cloud computing services and the endorsement from the video game industry, cloud gaming services have emerged promisingly. In a cloud gaming service, the contents of games can be delivered to the clients through either video streaming or file streaming. Due to the strict constraint on the end-to-end latency for real-time interaction in a game, there are still challenges in designing a successful cloud gaming system, which needs to deliver satisfying quality of experience to the customers. In this paper, the methodology for subjective and objective evaluation as well as the analysis of cloud gaming services was developed. The methodology is based on a nonintrusive approach, and therefore, it can be used on different kinds of cloud gaming systems. There are challenges in such objective measurements of important QoS factors, due to the fact that most of the commercial cloud gaming systems are proprietary and closed. In addition, satisfactory QoE is one of the crucial ingredients in the success of cloud gaming services. By combining subjective and objective evaluation results, cloud gaming system developers can infer possible results of QoE levels based on the measured QoS factors. It can also be used in an expert system for choosing the list of games that customers can appreciate at a given environment, as well as for deciding the upper bound of the number of users in a system.
Zi-Yi Wen, Hsu-Feng Hsiao
MMSP2
2014 Layer-Aligned Multipriority Rateless Codes for Layered Video Streaming
abstract
There exists a multitude of techniques, including automatic repeat request and error correction codes, to minimize data corruption when transmitting over error-prone networks. Streaming of multimedia data can usually withstand a certain level of data loss, yet have strict limitations on the latency tolerance. To enable acceptable reliability of transmission and low transmission latency, the channel coding approach is usually more appealing at the cost of additional bandwidth. In this paper, an \(N\) -cycle layer-aligned overlapping structure, which is good for layered data, is proposed. Accordingly, layer-aligned multipriority rateless codes were developed with favorable probabilities to control the protection strength for each layer of the streaming data. The major contribution of this paper is the analytical model developed to predict the failure decoding probabilities for each video layer and it is shown to achieve accurate estimation. A prediction model to estimate the expected decompressible video frames was developed for use with the developed codes for streaming scalable videos. By maximizing the number of expected decompressible video frames, the protection strength of the developed codes can then be determined. Simulation results show that the developed codes are good for streaming layered videos, which are difficult to deal with using traditional rateless codes, with or without unequal error protection.
Hsu-Feng Hsiao, Yong-Jhih Ciou
IEEE Trans. Circuits Syst. Video Technol.1
2013 Perceptual rate distortion optimization for block mode selection in hybrid video coding
abstract
Video compression technologies developed recently have taken advantage of a hybrid approach for better coding efficiency. For example, in the MPEG-4 Advanced Video Coding standard and also the new High Efficiency Video Coding, there are numerous possibilities of motion representation and residual coding options in transform domain. The coding parameters, including block modes, motion vectors, and indices of reference frames, are often correlated with each other in spatial and temporal directions. The selection of those coding parameters can affect the coding efficiency greatly. In many of the state-of-the-art solutions, the selection of coding options is based on the minimization of a Lagrangian cost function which is constructed of conventional distortion function and produced data rate. However, the conventional distortion function measured as either sum of absolute difference or mean square error does not often represent the observed deficiency in reconstructed videos. In this paper, a perceptual rate distortion optimization algorithm for block mode selection is developed, and the content-dependent Lagrangian multiplier is then derived. The just-noticeable distortion is used in the proposed method such that perceptual quality can be considered. The simulation results show that the rate-perceptual distortion performance is improved substantially without sacrificing the conventional rate-distortion performance.
Chen-Chou Huang, Hsu-Feng Hsiao
ISCAS2
2013 Balanced Parallel Scheduling for Video Encoding with Adaptive GOP Structure
abstract
Due to the nature of a dynamic group of picture (GOP) structure, parallel scheduling for video encoding becomes challenging. To address this, the balanced frame-level parallel scheduling algorithms are developed. The proposed approaches first determine the frame priority and then the thread priority assignment for scheduling. The concept of the algorithms lies in the analysis of coding complexity, temporal influence, and the required temporal burden to finish coding. To complete the scheduling with the dynamic GOP structure, a block-based abrupt and gradual scene change detection algorithm is also proposed to determine the GOP structure adaptively. The experiments show that the scheduling performance is close to the optimal. In addition, the concept of batch processing is incorporated so that the required buffer can be reduced.
Hsu-Feng Hsiao, Chen-Tsang Wu
IEEE Trans. Parallel Distributed Syst.1
2011 An error resilient multiple description video coder
abstract
The idea of multiple description video coding has been introduced to deal with the issues of bandwidth/path diversity and packet loss due to network congestion and/or error-prone channels which might cause serious quality degradation in video applications such as multimedia streaming and video conferencing services. In this paper, two approaches to description generation are proposed to produce multiple descriptions at higher coding efficiency. One of them is motivated by the multiple description scalar quantizer to reduce the distortion and the other is the coefficient partition in transform domain in order to balance the descriptions better. An estimation mechanism is further proposed to alleviate the drifting problem due to description fluctuation by synchronizing the reference frames at the encoder and the decoder as much as possible. The experiments show that the proposed methods offer substantial improvement at the event of description loss.
Yi-Jen Huang, Hsu-Feng Hsiao
MMSP2
2010 A depth refinement algorithm for multi-view video synthesis
abstract
With the recent progress of display, capture device, and coding technologies, multi-view video applications such as stereoscopic video, free viewpoint TV (FTV), and free viewpoint video (FVV) have been introduced to the world with growing interest. To achieve free navigation of such applications, depth information is required along with the video data. There have been many research activities in the area of depth estimation; however, it still poses us great challenge to estimate accurate depth map. In this paper, we propose a depth refinement algorithm for multi-view video synthesis. The proposed algorithm classifies the pixel-wise depth map into two categories, one is reliable and the other is unreliable, followed by the depth refinement algorithm for those pixels with unreliable depth values. Except for the depth refinement algorithm, we also propose a reliable weighted view interpolation algorithm. At last, the refined depth map is evaluated by the quality of the synthesized view.
Hsin-Chia Shih, Hsu-Feng Hsiao
ICASSP2
2010 TCP-friendly congestion control for the fair streaming of scalable video
Sheng-Shuen Wang, Hsu-Feng Hsiao
Comput. Commun.2
2008 TCP-friendly congestion control for layered video streaming using end-to-end bandwidth inference
abstract
More and more streaming protocols are developed for the multimedia applications. However, many streaming protocols only consider the network stability, but not the characteristics of streaming applications. In order to cooperate with H.264/MPEG-4 AVC scalable extension which can achieve fine granularity of scalability at bit level to the time-vary heterogeneous networks, we design a TCP-friendly congestion control algorithm based on the bandwidth estimation to smoothly change sending rate to avoid unnecessary oscillations so that the subscription decision of SVC layers can be made to better utilize the network resource. In case of the unavoidable network congestion, we unsubscribe scalable video layers according to the packet lost rate and the recently received throughput instead of only dropping one layer at a time to rapidly accommodate the streaming service to the channels and avoid persecuting the other flows at the same bottleneck. In addition, the probing packets for estimating the available bandwidth are encapsulated with RTP/RTCP. The simulations show that the proposed congestion control algorithm for real-time applications efficiently utilizes network bandwidth without hampering the performance of the existing TCP applications.
Sheng-Shuen Wang, Hsu-Feng Hsiao
MMSP2
2007 Dynamic FEC-Distortion Optimization for H.264 Scalable Video Streaming
abstract
Forward error correction codes have been shown to be a feasible solution either in application layer or in link layer to fulfill the need of quality of service for multimedia streaming over the fluctuant channels. In this paper, we propose FEC-distortion optimization algorithms to efficiently utilize the bandwidth for better video quality. The optimization criterions are based on the unequal error protection by taking account of the error drifting problems from both temporal motion compensation and inter-layer prediction of H.264/MPEG-4 AVC scalable video coding. Also, it can adapt to the content-dependent quality contribution of each video frame in a video layer. Lightweight error-concealment is also incorporated with the proposed algorithms for better H.264 SVC streaming. For some applications where either computation might be the bottleneck or the upper bound of non-decodable probability of each video layer is specified, alternative bandwidth allocation algorithm is provided with the trade-off of slight quality degradation.
Wei-Chung Wen, Hsu-Feng Hsiao, Jen-Yu Yu
MMSP2
2006 Fast End-to-End Available Bandwidth Estimation for Real-Time Multimedia Networking
abstract
Dynamic bandwidth estimation serves as an important basis for performance optimization of real-time distributed multimedia applications. The objective of this paper is to develop a bandwidth estimation algorithm for the fast fluctuated Internet. We analyze the relationship between the one way delay and the dispersion of packets train, and propose an available-bandwidth estimation algorithm which makes use of these two features without requiring administrative access to the intermediate routers along the network path. Instead of binary search or fixed-rate bandwidth adjustment of the probing data as described in literature, we use top-down approach to infer available bandwidth robustly and much more efficiently
Sheng-Shuen Wang, Hsu-Feng Hsiao
MMSP2
2005 A new multimedia packet loss classification algorithm for congestion control over wired/wireless channels
abstract
In a wireless network environment, common channel errors, due to multipath fading, shadowing and attenuation, may cause bit errors and packet loss quite different from the packet loss caused by network congestion. In congestion control, the packet loss information can serve as an index of network congestion for effective rate adjustment; therefore, wireless packet loss can mistakenly lead to dramatic performance degradation. The paper proposes a packet loss classification algorithm based on detecting the trend in relative one-way trip time (ROTT) when it falls in the ambiguous zone where the packet loss classification is not straightforward. We show that the proposed algorithm greatly benefits rate-based congestion control algorithms for multimedia over IP networks.
Hsu-Feng Hsiao, Aik Chindapol, James A. Ritcey, Yaw-Chung Chen, Jenq-Neng Hwang
ICASSP (2)1
2005 Adaptive FEC Scheme For Layered Multimedia Streaming over Wired/Wireless Channels
abstract
In wireless communication, noise and channel fluctuation often cause bit errors and subsequently the packet loss, which has different characteristic from the loss from network congestion. For most congestion control algorithms, the packet loss information serves as an index of network congestion and is used for effective rate adjustment; therefore wireless packet loss can mistakenly lead to dramatic performance degradation. We discuss the reasons leading to packet loss and propose a packet loss classification (PLC) algorithm that is based on trend detection of relative one-way trip time. With the assistance of PLC, not only can the wireless loss be separated from the congestion loss, the wireless packet error rate can be also estimated. With the combined information, the new adaptive congestion control algorithm is constructed so that a receiver can acquire an appropriate share of bandwidth. The transmitted data is protected by adaptive maximum distance separable erasure codes according to the wireless channel condition and end-to-end available bandwidth
Hsu-Feng Hsiao, Aik Chindapol, James A. Ritcey, Jenq-Neng Hwang
MMSP1
2004 A max-min fairness congestion control for streaming layered video
abstract
In a best-effort networking environment, efficient and fair congestion control is highly desired for every traffic flow, to share the bandwidth appropriately. This paper proposes a congestion control algorithm for UDP based layered video, whose bandwidth resolution in each layer has been predefined. This proposed congestion control mechanism is an extension of XCP, which is a newly proposed protocol believed to be superior to TCP, especially for high bandwidth-delay product networks. This paper also introduces reserved packet length so that the layered video traffic can share the bandwidth of a network better with the consideration of max-min fairness to other traffic.
Hsu-Feng Hsiao, Jenq-Neng Hwang
ICASSP (5)1
2003 Layered FGS video over active network with selective drop and adaptive rate control
abstract
Quality of service has been a great concern to video dissemination over the Internet, especially due to the heterogeneous networking environment. In contrast to the traditional passive networking, active networking by means of active routers/agents in a wide area network shows promise of better video service by offering packet control at finer resolution. This paper proposes an effective architecture based on active router and selective drop queue management as a solution to video unicast and multicast in a variable network environment.
Hsu-Feng Hsiao, Jenq-Neng Hwang
ICASSP (5)1
2000 Worst-Case Criterion for Content-Based Error-Resilient Video Coding
abstract
A video content based coding algorithm is presented to eliminate the error propagation effects caused by transmitting highly compressed video sequences over unreliable packet networks. We study the criteria to determine an appropriate threshold for it and conduct experiments in a bursty packet loss environment. Promising results are obtained in the preliminary simulations.
Wu-Hsiang Jonas Chen, Jenq-Neng Hwang, Hsu-Feng Hsiao
ICIP3