Yao Wang 0001

dblp:72/628-1 · DBLP profile ↗
← Back
210ranked-venue papers
25as first author
25since 2021 · last 2026
0000-0003-3199-3802ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 166 · 19 first-author · 14 since 2021Computer networks · 22 · 2 first-author · 5 since 2021Systems, architecture and hardware · 6 · 1 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-authorHuman-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Knowledge-Enhanced Intent-Driven Flow Scheduling for LEO Satellite Networks
abstract
Low Earth Orbit (LEO) satellite networks are characterized by dynamic network topologies and on-demand service requirements from Internet-of-Things (IoT) applications, which make efficient and intelligent flow scheduling challenging. Conventional schemes rely on static configurations or manual rules, thus making it difficult to capture and respond to diverse service demands. Moreover, they often fail to model task–resource relationships effectively, hindering the generation of real-time, executable scheduling policies. To address these challenges, we propose a knowledge-enhanced, intent-driven flow scheduling (KIFS) framework. Specifically, we design a unified pipeline that first translates user intents into precise Quality of Service (QoS) requirements. It incorporates a network state awareness module to estimate per-link bandwidth and utilization, and constructs a task–resource knowledge graph (KG) to enhance the Deep Q-Network (DQN) agent via state augmentation, action pruning, and reward shaping. Finally, the framework translates the resulting policies into standards-compliant SRv6 configurations for real-time deployment. In simulations, the proposed KIFS framework demonstrates superior performance compared to standard baselines in terms of flow success rate and QoS satisfaction.
Zhenzi Wang, Chungang Yang, Song Mao, Yao Wang 0001, Ying Ouyang, Zhu Han 0001
IEEE Internet Things J.4
2025 Bits-to-Photon: End-to-End Learned Scalable Point Cloud Compression for Direct Rendering
abstract
Point cloud is a promising 3D representation for volumetric streaming in emerging AR/VR applications. Despite recent advances in point cloud compression, decoding and rendering high-quality images from lossy compressed point clouds is still challenging in terms of quality and complexity, making it a major roadblock to achieve real-time 6-Degree-of-Freedom video streaming. In this paper, we address this problem by developing a point cloud compression scheme which generates a bit stream that can be directly decoded to renderable 3D Gaussians. The encoder and decoder are jointly optimized to consider both bit-rates and rendering quality. It significantly improves the rendering quality while substantially reducing decoding and rendering time, compared to existing point cloud compression methods. Furthermore, the proposed scheme generates a scalable bit stream, allowing multiple levels of details at different bit-rate ranges. Our method supports real-time color decoding and rendering of high quality point clouds, thus paving the way for interactive 3D streaming applications with free view points. The code is available at https://github.com/huzi96/bits2photon.
Yueyu Hu, Yao Wang 0001
ICIP3
2025 Multi-Satellite Collaboration Task Planning Based On Behavior Tree
abstract
As the number of satellites and satellite tasks continues to increase, multi-satellite collaboration task planning faces several challenges, including high complexity, high timeliness and limited scalability. This paper presents a method for task planning based on behavior tree, which leverages modularity and hierarchical structure of behavior tree to simplify planning process and enhance timeliness. We also propose an intelligent planning algorithm combining variable neighborhood search (VNS) algorithm with backtracking search algorithm (BSA) to reduce solution space and improve convergence speed. Experimental results show that this method improves satellite task planning timeliness by 16.7%, overcomes the flexibility and timeliness disadvantages of traditional methods in complex environments, and highlights the significant advantages and potential applications of behavior tree in multi-satellite collaboration field.
Mingji Wu, Ying Ouyang, Chungang Yang, Yao Wang 0001
IWCMC5
2025 Spatial Visibility and Temporal Dynamics: Rethinking Field of View Prediction in Adaptive Point Cloud Video Streaming
abstract
Field-of-View (FoV) adaptive streaming significantly reduces bandwidth requirement of immersive point cloud video (PCV) by only transmitting visible points inside a viewer's FoV. The traditional approaches often focus on trajectory-based 6 degree-of-freedom (6DoF) FoV predictions. The predicted FoV is then used to calculate point visibility. Such approaches do not explicitly consider video content's impact on viewer attention, and the conversion from FoV to point visibility is often error-prone and time-consuming. We reformulate the PCV FoV prediction problem from the cell visibility perspective, allowing for precise decision-making regarding the transmission of 3D data at the cell level based on the predicted visibility distribution. We develop a novel spatial visibility and object-aware graph model (CellSight) that leverages the historical 3D visibility data and incorporates spatial perception, occlusion between points, and neighboring cell correlation to predict the cell visibility in the future. We focus on multi-second ahead prediction to enable the use of long pre-fetching buffers in on-demand streaming, critical for enhancing the robustness to network bandwidth fluctuations. CellSight significantly improves the long-term cell visibility prediction, reducing the prediction Mean Squared Error (MSE) loss by up to 50% compared to the state-of-the-art models when predicting 2 to 5 seconds ahead, while maintaining real-time performance (more than 30fps) for point cloud videos with over 1 million points.
Chen Li 0043, Tongyu Zong, Yueyu Hu, Yao Wang 0001, Yong Liu 0013
MMSys4
2025 Knowledge-Enhanced Large Language Model for Intent Refinement Mechanism
abstract
Intent-Driven Networking (IDN) enables users to express high-level intents in natural language, which are then automatically refined into executable network configurations. Intent refinement plays a critical role in accurately refining human-declarative intents into network-level intents, and further refining into machine-readable policies. To address the limitations of existing intent refinement methods of lacking generalization and weak context awareness, this paper proposes a knowledge-enhanced Large Language Model (LLM) for intent refinement mechanism. The proposed approach achieves accurate and controllable refinement from natural language to structured network intents. Simulation results demonstrate the effectiveness and robustness of the proposed mechanism, particularly in complex and layered intent expression in flying ad hoc networking.
Chungang Yang, Tong Li 0019, Yulong Dai, Yao Wang 0001
VTC2025-Fall5
2025 Large Language Model-Empowered Intent-Driven Network Configuration Generator
abstract
With the advent of the sixth-generation (6G) era, the scale and complexity of communication networks have expanded dramatically, making conventional manual network management methods inefficient and error-prone. Intent-driven network (IDN) enables network operators to express high-level intents using natural language, which are then automatically translated into executable network configurations. However, the unstructured and ambiguous nature of intents poses challenges to achieving accurate intent-to-configuration translation. The emergence of a large language model (LLM) offers a promising solution to this problem. This paper proposes an LLM-empowered framework for generating IDN configurations, which integrates fine-tuning, retrieval-augmented generation, prompt engineering, and knowledge distillation techniques. We validate the effectiveness of the proposed framework through a network slicing use case implemented using open-source tools. Experimental results demonstrate that the proposed framework improves configuration generation time and accuracy by 24% and 25%, respectively, compared to the baseline schemes.
Chungang Yang, Yao Wang 0001, Rongqian Fan
VTC2025-Fall3
2025 Large Language Model-Enhanced Intent-Driven Management and Orchestration for 6G Networks
abstract
With emerging differentiated network services, intent-driven management and orchestration in the sixth-generation networks face challenges in on-demand and timely network configuration. Conventional intent-driven network approaches rely on structured templates for specific scenarios. This paper develops a large language model (LLM)-enhanced intent-driven management and orchestration framework that automatically refines user intent into abstract network policies. To improve the generality and accuracy of intent refinement, we design a novel intent decomposition mechanism based on a fine-tuned generic LLM, along with intent decomposition prompts. Moreover, we introduce an intent optimization method, leveraging a lightweight proximal policy optimization framework to model distributed energy-saving policies and select the optimal policy. Simulation results show that the proposed framework outperforms baseline schemes by 16% to 31% in terms of intent decomposition accuracy and intent optimization performance.
Yao Wang 0001, Chungang Yang, Rongqian Fan
VTC2025-Fall1
2025 Progressive Frame Patching for FoV-Based Point Cloud Video Streaming
abstract
Many XR applications require the delivery of volumetric video to users. Point Cloud has become a popular volumetric video format. A dense point cloud consumes much higher bandwidth than a 2D/360$^{\circ }$video frame. User Field of View (FoV) is more dynamic with 6-DoF movement than 3-DoF movement. To save bandwidth, FoV-adaptive streaming predicts a user's FoV and only downloads point cloud data falling in the predicted FoV. However, it is vulnerable to FoV prediction errors, which can be significant when a long buffer is utilized for smoothed streaming. In this work, we propose a multi-round progressive refinement framework for point cloud video streaming. Instead of sequentially downloading point cloud frames, our solution simultaneously downloads/patches multiple frames falling into a sliding time-window, leveraging the inherent scalability of octree-based point-cloud coding. The optimal rate allocation among all tiles of active frames are solved numerically using the heterogeneous tile rate-quality functions calibrated by the predicted user FoV. Multi-frame downloading/patching simultaneously takes advantage of the streaming smoothness resulting from long buffer and the FoV prediction accuracy at short buffer length. We evaluate our streaming solution using simulations driven by real point cloud videos, real bandwidth traces, and 6-DoF FoV traces of real users. Our solution is robust against the bandwidth/FoV prediction errors, and can deliver high and smooth view quality in the face of bandwidth variations and dynamic user and point cloud movements.
Tongyu Zong, Yixiang Mao, Chen Li 0043, Yong Liu 0013, Yao Wang 0001
IEEE Trans. Multim.5
2024 Standard Compatible Efficient Video Coding with Jointly Optimized Neural Wrappers
abstract
We present a standard-compatible video coding scheme with end-to-end optimized neural wrapper over standard video codecs that achieves significant rate-distortion (R-D) performance gains and is still efficient in decoding. We train a pair of pre- and post-processor using a differential JPEG proxy. The pre-processor applies a learned transform to the video and downsamples the video by a factor of 2. It generates a bottleneck video to be coded by a standard codec as a YUV sequence. The post-processor takes the decoded bottleneck video, does the inverse transform, and upsamples it to the original resolution. We follow the design in [1] , where we configure downsample using a layer of strided convolution. We optimize the post-processor for efficiency by replacing convolutions with kernel size larger than 1×1 to depth-wise convolutions [2] .
Yueyu Hu, Onur G. Guleryuz, Debargha Mukherjee, Yao Wang 0001
DCC5
2024 Standard Compliant Video Coding Using Low Complexity, Switchable Neural Wrappers
abstract
The proliferation of high resolution videos posts great storage and bandwidth pressure on cloud video services, driving the development of next-generation video codecs. Despite great progress made in neural video coding, existing approaches are still far from economical deployment considering the complexity and rate-distortion performance tradeoff. To clear the roadblocks for neural video coding, in this paper we propose a new framework featuring standard compatibility, high performance, and low decoding complexity. We employ a set of jointly optimized neural pre and post-processors, wrapping a standard video codec, to encode videos at different resolutions. The rate-distorion optimal downsampling ratio is signaled to the decoder at the per-sequence level for each target rate. We design a low complexity neural post-processor architecture that can handle different upsampling ratios. The change of resolution exploits the spatial redundancy in high-resolution videos, while the neural wrapper further achieves rate-distortion performance improvement through end-to-end optimization with a codec proxy. Our light-weight post-processor architecture has a complexity of 516 MACs / pixel, and achieves 9.3% BD-Rate reduction over VVC on the UVG dataset, and $6.4 \%$ on AOM CTC Class A1. Our approach has the potential to further advance the performance of the latest video coding standards using neural processing with minimal added complexity.
Yueyu Hu, Onur G. Guleryuz, Debargha Mukherjee, Yao Wang 0001
ICIP5
2024 5G Edge Vision: Wearable Assistive Technology for People with Blindness and Low Vision
abstract
In an increasingly visual world, people with blindness and low vision (pBLV) face substantial challenges in navigating their surroundings and interpreting visual information. From our previous work, VIS4ION is a smart wearable that helps pBLV in their daily challenges. It enables multiple microservices based on artificial intelligence (AI), such as visual scene processing, navigation, and vision-language inference. These microservices require powerful computational resources and, in some cases, stringent inference times, hence the need to offload computation to edge servers. This paper introduces a novel video streaming platform that improves the capabilities of VIS4ION by providing real-time support of the microservices at the network edge. When video is offloaded wirelessly to the edge, the time-varying nature of the wireless network requires adaptation strategies for a seamless video service. We demonstrate the performance of our adaptive real-time video streaming platform through experimentation with an open-source 5G deployment based on open air interface (OAI). The experiments demonstrate the ability to provide microservices robustly in time-varying network conditions.
Tommy Azzino, Marco Mezzavilla, Sundeep Rangan, Yao Wang 0001, John-Ross Rizzo
WCNC4
2024 Intent-Driven Closed-Loop Control and Management Framework for 6G Open RAN
abstract
Future mobile networks should provide on-demand services for various industries and applications with the stringent guarantees of Quality of Experience (QoE), which highly challenge the flexibility of network management. However, the diverse requirements of QoE and the management of heterogeneous networks create significant pressure toward communication service providers (CSPs). In the sixth-generation mobile networks, the CSPs should guarantee resilient performance for the communication service consumers with less human involvement. In this work, we turn to Intent-driven network and on-demand slice management, and to decrease the complexity and cost in full life cycle slice management, we first present an intent-driven closedloop (CL) control and management framework that automates the deployment of network slices and manages resources intelligently based on the extended CL architecture. And then, we explore and exploit the deep reinforcement learning algorithm to address the problem of resource allocation, which is formulated as a Markov decision process. Finally, we demonstrate the feasibility of the proposed framework by deploying the open radio access network (RAN) infrastructure in the OpenAirInterface platform and realizing the CL control and management with a near real-time RAN intelligent controller. The emulation results demonstrate the effectiveness of slicing performance, measured in terms of delay and rate.
Chungang Yang, Ru Dong, Yao Wang 0001, Alagan Anpalagan, Qiang Ni, Mohsen Guizani
IEEE Internet Things J.4
2024 Split Computing With Scalable Feature Compression for Visual Analytics on the Edge
abstract
Running deep visual analytics models for real-time applications is challenging for mobile devices. Offloading the computation to edge server can mitigate computation bottleneck at the mobile device, but may decrease the analytics performance due to the necessity of compressing the image data. We consider a “split computing” system to offload a part of the deep learning model's computation and introduce a novel learned feature compression approach with lightweight computation. We demonstrate the effectiveness of the split computing pipeline in performing computation offloading for the problems of object detection and image classification. Compared to compressing the raw images at the mobile, and running the analytics model on the decompressed images at the server, the proposed feature-compression approach can achieve significantly higher analytics performance at the same bit rate, while reducing the complexity at the mobile. We further propose a scalable feature compression approach, which facilitates adaptation to network bandwidth dynamics, while having comparable performance to the non-scalable approach.
Zhongzheng Yuan, Samyak Rawlekar, Siddharth Garg, Elza Erkip, Yao Wang 0001
IEEE Trans. Multim.5
2023 Understanding the Impact of Image Quality and Distance of Objects to Object Detection Performance
abstract
Object detection is a fundamental task for autonomous driving, which aim to identify and localize objects within an image. Deep learning has made great strides for object detection, with popular models including Faster R-CNN, YOLO, and SSD. The detection accuracy and computational cost of object detection depend on the spatial resolution of an image, which may be constrained by both the camera and storage considerations. Furthermore, original images are often compressed and uploaded to a remote server for object detection. Compression is often achieved by reducing either spatial or amplitude resolution or, at times, both, both of which have well-known effects on performance. Detection accuracy also depends on the distance of the object of interest from the camera. Our work examines the impact of spatial and amplitude resolution, as well as object distance, on object detection accuracy and computational cost. As existing models are optimized for uncompressed (or lightly compressed) images over a narrow range of spatial resolution, we develop a resolution-adaptive variant of YOLOv5 (RA-YOLO), which varies the number of scales in the feature pyramid and detection head based on the spatial resolution of the input image. To train and evaluate this new method, we created a dataset of images with diverse spatial and amplitude resolutions by combining images from the TJU and Eurocity datasets and generating different resolutions by applying spatial resizing and compression. We first show that RA-YOLO achieves a good trade-off between detection accuracy and inference time over a large range of spatial resolutions. We then evaluate the impact of spatial and amplitude resolutions on object detection accuracy using the proposed RA-YOLO model. We demonstrate that the optimal spatial resolution that leads to the highest detection accuracy depends on the ‘tolerated’ image size (constrained by the available bandwidth or storage). We further assess the impact of the distance of an object to the camera on the detection accuracy and show that higher spatial resolution enables a greater detection range. These results provide important guidelines for choosing the image spatial resolution and compression settings predicated on available bandwidth, storage, desired inference time, and/or desired detection range, in practical applications.
Haoyang Pei, Yixuan Lyu, Zhongzheng Yuan, John-Ross Rizzo, Yao Wang 0001, Yi Fang 0006
IROS6
2023 Resilience In-Band Control Path Routing in Blockchain-Based Multi-Domain SDN
abstract
As Software Defined Networking (SDN) continues to evolve, the demand for enhanced reliability and security within network infrastructures has surged. To address these pressing needs, we put forth a novel, blockchain-based SDN multi-domain network security architecture designed to bolster the safety measures of distributed systems. Furthermore, we’ve developed an innovative primary backup path optimization algorithm, which capitalizes on maximum disjoint in in-band mode to elevate the resilience of communication services. When a network encounters a failure, seamlessly transitioning the control path is crucial for maintaining uninterrupted, dependable operations and fortifying network resilience. In order to validate our approach, we conducted a series of simulation experiments, leveraging Pica8 and ONOS controllers to design the system architecture. Remarkably, the numerical outcomes revealed that our pioneering primary and backup control path algorithms significantly outperform the conventional shortest path algorithm in the face of network failure. By enhancing throughput and reducing packet loss rates, our proposal ultimately augments communication reliability.
Xingpeng Lei, Yanbo Song, Yao Wang 0001, Chungang Yang
IWCMC6
2023 Privacy-Aware Laser Wireless Power Transfer for Aerial Multi-Access Edge Computing: A Colonel Blotto Game Approach
abstract
This article studies the integration of laser-beamed wireless power transfer (WPT) into high-altitude platform (HAP)-aided multiaccess edge computing (MEC) systems for the HAP-connected aerial user equipments (AUEs). By discretizing the 3-D coverage space of the HAP, we present a multitier tile grid-based spatial structure to provide the aerial locations in the form of tile grids to AUEs for laser charging. We identify a new privacy vulnerability caused by the openness during the WPT signaling transfer in the presence of a terrestrial adversary, which is able to launch the attacks by distributing the false tile grids to the AUEs. To address this vulnerability and enhance the location privacy of AUEs, we then propose a Colonel Blotto (CB) game framework to formulate the competitive tile grid allocation problem for the HAP and the adversary. The attack-defense interaction between the adversary and the HAP as a defender in their tile grid allocations to the AUEs is formulated as a CB game, which models the competition of two players for limited resources over multiple battlefields. Moreover, we derive the mixed-strategy Nash equilibria of the game for both symmetric and asymmetric tile grids between two players. Simulation results show that the proposed framework significantly outperforms the design baselines with a given privacy protection level in terms of system-wide expected total utilities.
Long Zhang 0003, Yao Wang 0001, Minghui Min, Chao Guo 0002, Vishal Sharma 0001, Zhu Han 0001
IEEE Internet Things J.2
2023 Live 360 Degree Video Delivery Based on User Collaboration in a Streaming Flock
abstract
Streaming of live 360-degree video allows users to follow a live event from any view point and has already been deployed on some commercial platforms. However, the current systems can only stream the video at relatively low-quality because the entire 360-degree video is delivered to the users under limited bandwidth. Streaming video falling into user field of view (FoV) can improve bandwidth efficiency of 360-degree video delivery. In this paper, we propose to use the idea of “flocking” to simultaneously improve the accuracy of user FoV prediction and video delivery efficiency for live 360-degree video streaming. By assigning variable playback latencies to users in a streaming session based on their network conditions, a “streaming flock” is formed and led by “strong” users with low playback latencies in the front of the flock. We propose a long short-term memory (LSTM) based collaborative FoV prediction scheme where the FoV traces of users in the front of the flock are utilized to predict the FoV of users behind them. Given a predicted FoV, we develop an optimal rate allocation strategy to maximize the perceptual quality. By conducting experiments using real-world user FoV traces and LTE/5 G network bandwidth traces, we evaluate the gains of the proposed strategies over several benchmarks. Our experimental results demonstrate that the proposed streaming system can increase the overall quality dramatically by about 10 dB compared with heuristic FoV prediction strategy. In addition, the network-aware flocking formation can further reduce the video freeze without influencing video quality.
Liyang Sun, Yixiang Mao, Tongyu Zong, Yong Liu 0013, Yao Wang 0001
IEEE Trans. Multim.5
2022 Learning Neural Volumetric Field for Point Cloud Geometry Compression
abstract
Due to the diverse sparsity, high dimensionality, and large temporal variation of dynamic point clouds, it remains a challenge to design an efficient point cloud compression method. We propose to code the geometry of a given point cloud by learning a neural volumetric field. Instead of representing the entire point cloud using a single overfit network, we divide the entire space into small cubes and represent each non-empty cube by a neural network and an input latent code. The network is shared among all the cubes in a single frame or multiple frames, to exploit the spatial and temporal redundancy. The neural field representation of the point cloud includes the network parameters and all the latent codes, which are generated by using back-propagation over the network parameters and its input. By considering the entropy of the network parameters and the latent codes as well as the distortion between the original and reconstructed cubes in the loss function, we derive a rate-distortion (R-D) optimal representation. Experimental results show that the proposed coding scheme achieves superior R-D performances compared to the octree-based G-PCC, especially when applied to multiple frames of a point cloud video. The code is available at https://github.com/huzi96/NVFPCC/.
Yueyu Hu, Yao Wang 0001
PCS2
2022 End-to-End Neural Video Coding Using a Compound Spatiotemporal Representation
abstract
Recent years have witnessed rapid advances in learnt video coding. Most algorithms have solely relied on the vector-based motion representation and resampling (e.g., optical flow based bilinear sampling) for exploiting the inter frame redundancy. In spite of the great success of adaptive kernel-based resampling (e.g., adaptive convolutions and deformable convolutions) in video prediction for uncompressed videos, integrating such approaches with rate-distortion optimization for inter frame coding has been less successful. Recognizing that each resampling solution offers unique advantages in regions with different motion and texture characteristics, we propose a hybrid motion compensation (HMC) method that adaptively combines the predictions generated by these two approaches. Specifically, we generate a compound spatiotemporal representation (CSTR) through a recurrent information aggregation (RIA) module using information from the current and multiple past frames. We further design a one-to-many decoder pipeline to generate multiple predictions from the CSTR, including vector-based resampling, adaptive kernel-based resampling, compensation mode selection maps and texture enhancements, and combines them adaptively to achieve more accurate inter prediction. Experiments show that our proposed inter coding system can provide better motion-compensated prediction and is more robust to occlusions and complex motions. Together with jointly trained intra coder and residual coder, the overall learnt hybrid coder yields the state-of-the-art coding efficiency in low-delay scenario, compared to the traditional H.264/AVC and H.265/HEVC, as well as recently published learning-based methods, in terms of both PSNR and MS-SSIM metrics.
Ming Lu 0003, Zhiqi Chen 0001, Xun Cao, Zhan Ma 0001, Yao Wang 0001
IEEE Trans. Circuits Syst. Video Technol.6
2021 Motion Adaptive Pose Estimation from Compressed Videos
abstract
Human pose estimation from videos has many real-world applications. Existing methods focus on applying models with a uniform computation profile on fully decoded frames, ignoring the freely-available motion signals and motion-compensation residuals from the compressed stream. A novel model, called Motion Adaptive Pose Net is proposed to exploit the compressed streams to efficiently decode pose sequences from videos. The model incorporates a Motion Compensated ConvLSTM to propagate the spatially aligned features, along with an adaptive gate to dynamically determine if the computationally expensive features should be extracted from fully decoded frames to compensate the motion-warped features, solely based on the residual errors. Leveraging the informative yet readily available signals from compressed streams, we propagate the latent features through our Motion Adaptive Pose Net efficiently Our model outperforms the state-of-the-art models in pose-estimation accuracy on two widely used datasets with only around half of the computation complexity.
Zhipeng Fan 0001, Jun Liu 0036, Yao Wang 0001
ICCV3
2021 Tightrope walking in low-latency live streaming: optimal joint adaptation of video rate and playback speed
abstract
It is highly challenging to simultaneously achieve high-rate and low-latency in live video streaming. Chunk-based streaming and playback speed adaptation are two promising new trends to achieve high user Quality-of-Experience (QoE). To thoroughly understand their potentials, we develop a detailed chunk-level dynamic model that characterizes how video rate and playback speed jointly control the evolution of a live streaming session. Leveraging on the model, we first study the optimal joint video rate-playback speed adaptation as a non-linear optimal control problem. We further develop model-free joint adaptation strategies using deep reinforcement learning. Through extensive experiments, we demonstrate that our proposed joint adaptation algorithms significantly outperform rate-only adaptation algorithms and the recently proposed low-latency video streaming algorithms that separately adapt video rate and playback speed without joint optimization. In a wide-range of network conditions, the model-based and model-free algorithms can achieve close-to-optimal trade-offs tailored for users with different QoE preferences.
Liyang Sun, Tongyu Zong, Siquan Wang, Yong Liu 0013, Yao Wang 0001
MMSys5
2021 Block-based Learned Image Coding with Convolutional Autoencoder and Intra-Prediction Aided Entropy Coding
abstract
Recent works on learned image coding using autoencoder models have achieved promising results in rate-distortion performance. Typically, an autoencoder is used to transform an image into a latent tensor, which is then quantized and entropy coded. Based on a work by Ballé et al., we adapted the autoencoder with a hyperprior model to code images in a block-based approach. When the autoencoder model is directly applied to code small image blocks, spatial redundancy in the larger image cannot be fully utilized, resulting in a decrease in ratedistortion performance. We propose a method to utilize border information in the entropy coding of latent and hyper-latent tensors, which has achieved promising results. We show that using intra-prediction to help entropy coding is more effective than applying a convolutional autoencoder with hyper priors to intra-prediction residual blocks.
Zhongzheng Yuan, Debargha Mukherjee, Balu Adsumilli, Yao Wang 0001
PCS5
2021 Neural Video Coding Using Multiscale Motion Compensation and Spatiotemporal Context Model
abstract
Over the past two decades, traditional block-based video coding has made remarkable progress and spawned a series of well-known standards such as MPEG-4, H.264/AVC and H.265/HEVC. On the other hand, deep neural networks (DNNs) have shown their powerful capacity for visual content understanding, feature extraction and compact representation. Some previous works have explored the learnt video coding algorithms in an end-to-end manner, which show the great potential compared with traditional methods. In this paper, we propose an end-to-end deep neural video coding framework (NVC), which uses variational autoencoders (VAEs) with joint spatial and temporal prior aggregation (PA) to exploit the correlations in intra-frame pixels, inter-frame motions and inter-frame compensation residuals, respectively. Novel features of NVC include: 1) To estimate and compensate motion over a large range of magnitudes, we propose an unsupervised multiscale motion compensation network (MS-MCN) together with a pyramid decoder in the VAE for coding motion features that generates multiscale flow fields, 2) we design a novel adaptive spatiotemporal context model for efficient entropy coding for motion information, 3) we adopt nonlocal attention modules (NLAM) at the bottlenecks of the VAEs for implicit adaptive feature extraction and activation, leveraging its high transformation capacity and unequal weighting with joint global and local information, and 4) we introduce multi-module optimization and a multi-frame training strategy to minimize the temporal error propagation among P-frames. NVC is evaluated for the low-delay causal settings and compared with H.265/HEVC, H.264/AVC and the other learnt video compression methods following the common test conditions, demonstrating consistent gains across all popular test sequences for both PSNR and MS-SSIM distortion metrics.
Ming Lu 0003, Zhan Ma 0001, Zhihuang Xie, Xun Cao, Yao Wang 0001
IEEE Trans. Circuits Syst. Video Technol.7
2021 End-to-End Learnt Image Compression via Non-Local Attention Optimization and Improved Context Modeling
abstract
This article proposes an end-to-end learnt lossy image compression approach, which is built on top of the deep nerual network (DNN)-based variational auto-encoder (VAE) structure with Non-Local Attention optimization and Improved Context modeling (NLAIC). Our NLAIC 1) embeds non-local network operations as non-linear transforms in both main and hyper coders for deriving respective latent features and hyperpriors by exploiting both local and global correlations, 2) applies attention mechanism to generate implicit masks that are used to weigh the features for adaptive bit allocation, and 3) implements the improved conditional entropy modeling of latent features using joint 3D convolutional neural network (CNN)-based autoregressive contexts and hyperpriors. Towards the practical application, additional enhancements are also introduced to speed up the computational processing (e.g., parallel 3D CNN-based context prediction), decrease the memory consumption (e.g., sparse non-local processing) and reduce the implementation complexity (e.g., a unified model for variable rates without re-training). The proposed model outperforms existing learnt and conventional (e.g., BPG, JPEG2000, JPEG) image compression methods, on both Kodak and Tecnick datasets with the state-of-the-art compression efficiency, for both PSNR and MS-SSIM quality measurements. We have made all materials publicly accessible at https://njuvision.github.io/NIC for reproducible research.
Tong Chen 0004, Zhan Ma 0001, Qiu Shen, Xun Cao, Yao Wang 0001
IEEE Trans. Image Process.6
2021 Towards Optimal Low-Latency Live Video Streaming
abstract
Low-latency is a critical user Quality-of-Experience (QoE) metric for live video streaming. It poses significant challenges for streaming over the Internet. In this paper, we explore the design space of low-latency live streaming by developing dynamic models and optimal adaptation strategies to establish QoE upper bounds as a function of the allowable end-to-end latency. We further develop practical live streaming algorithms within the iterative Linear Quadratic Regulator (iLQR) based Model Predictive Control and Deep Reinforcement Learning frameworks, namely MPC-Live and DRL-Live, to maximize user live streaming QoE by adapting the video bitrate while maintaining low end-to-end video latency in dynamic network environment. Through extensive experiments driven by real network traces, we demonstrate that our live streaming algorithms can achieve close-to-optimal performance within the latency range of two to five seconds.
Liyang Sun, Tongyu Zong, Siquan Wang, Yong Liu 0013, Yao Wang 0001
IEEE/ACM Trans. Netw.5
2020 Convolutional Neural Network-Based Coefficients Prediction for HEVC Intra-Predicted Residues
abstract
We propose a convolutional neural network-based coefficients prediction (CNNCP) method for intra-predicted residues in the High Efficiency Video Coding (HEVC) standard. In HEVC, discrete cosine transform (DCT) or discrete sine transform (DST) is adopted to convert the intra-predicted residues in the spatial domain into coefficients in the frequency domain. Each coefficient is scalar quantized and entropy coded into the bitstream. As DCT or DST is non-optimal linear transform, there still exist linear and non-linear correlations among different coefficients after the transform. In addition, there exist coefficients' correlations between current block and neighboring blocks, as these correlations cannot be completely exploited in the intra prediction. We thus propose to perform coefficients prediction to further reduce the redundancy among coefficients. The coefficients prediction is achieved using trained convolutional neural networks (CNNs), as CNNs can build complex relationship between input and output by training with a lot of data. In addition, a flag that signals whether to perform coefficients prediction or not at the coding unit level is transmitted to decoder. The proposed CNNCP method is implemented upon the HEVC reference software. Experimental results show that the proposed method achieves on average 1.8%, 4.1%, and 4.5% BD-rate reduction ratios in Y, U, V, respectively, compared with the HEVC baseline in all-intra configuration. In particular, the average BD-rate reduction ratios for 4K test sequences are 2.9%, 6.5%, and 6.6%.
Changyue Ma, Dong Liu 0002, Li Li 0040, Yao Wang 0001, Feng Wu 0001
DCC4
2020 Adaptive Computationally Efficient Network for Monocular 3D Hand Pose Estimation
Zhipeng Fan 0001, Jun Liu 0036, Yao Wang 0001
ECCV (4)3
2020 Low-latency FoV-adaptive Coding and Streaming for Interactive 360° Video Streaming
abstract
Virtual Reality (VR) and Augmented Reality (AR) technologies have become popular in recent years. Encoding and transmitting the omni-directional or $360^\circ $ video is critical and challenging for those applications. The $360^\circ $ video requires much higher bandwidth than the traditional planar video. A premium quality $360^\circ $ video with 120 frames per second (fps) and 24K resolution can easily consume bandwidth in the range of Gigabits-per-second~\cite1. On the other hand, at any given time, a user only watches a small portion of the $360^\circ$ scope within her Field-of-View (FoV). An effective way to reduce the bandwidth requirement of $360^\circ $ video is through FoV-adaptive streaming, which codes and delivers the predicted FoV region at higher quality, and discards or codes at lower quality the remaining regions. Such strategy has been quite extensively studied for video-on-demand \citefov_adapt_2,fov_adapt_3,1,tile_based_3,qian2016optimizing and live video streaming applications\citelive_1,live_2,live_3, sun2020flocking. Interactive applications, such as conferencing, gaming, and remote collaboration, can also benefit from $360^\circ $ video by creating an immersive environment for participants to interact with each other citeinteractive_gamming \citevr_conferencing \citelee2015outatime. However, realtime coding and streaming of $360^\circ $ video with extremely low latency, required for interactive applications, has not been sufficiently addressed. This work focuses on developing low-latency and FoV-adaptive coding and streaming strategies for interactive $360^\circ$ video streaming. We assume the sender and the receiver are connected by a network path with dynamically varying throughput without short-latency guarantee. The sender is either the video source, or a proxy server relaying the source video. The receiver is either the end user device that directly renders the video, or a local edge server that renders the video and transmit to the end user \citeHou2017.
Yixiang Mao, Liyang Sun, Yong Liu 0013, Yao Wang 0001
ACM Multimedia4
2020 Flocking-based live streaming of 360-degree video
abstract
Streaming of live 360-degree video allows users to follow a live event from any view point and has already been deployed on some commercial platforms. However, the current systems can only stream the video at relatively low-quality because the entire 360-degree video is delivered to the users under limited bandwidth. In this paper, we propose to use the idea of "flocking" to improve the performance of both prediction of field of view (FoV) and caching on the edge servers for live 360-degree video streaming. By assigning variable playback latencies to all the users in a streaming session, a "streaming flock" is formed and led by low latency users in the front of the flock. We propose a collaborative FoV prediction scheme where the actual FoV information of users in the front of the flock are utilized to predict of users behind them. We further propose a network condition aware flocking strategy to reduce the video freeze and increase the chance for collaborative FoV prediction on all users. Flocking also facilitates caching as video tiles downloaded by the front users can be cached by an edge server to serve the users at the back of the flock, thereby reducing the traffic in the core network. We propose a latency-FoV based caching strategy and investigate the potential gain of applying transcoding on the edge server. We conduct experiments using real-world user FoV traces and WiGig network bandwidth traces to evaluate the gains of the proposed strategies over benchmarks. Our experimental results demonstrate that the proposed streaming system can roughly double the effective video rate, which is the video rate inside a user's actual FoV, compared to the prediction only based on the user's own past FoV trajectory, while reducing video freeze. Furthermore, edge caching can reduce the traffic in the core network by about 80%, which can be increased to 90% with transcoding on edge server.
Liyang Sun, Yixiang Mao, Tongyu Zong, Yong Liu 0013, Yao Wang 0001
MMSys5
2020 An adaptive cache management approach in ICN with pre-filter queues
Dapeng Man, Yao Wang 0001, Wu Yang 0001, Xiaojiang Du, Mohsen Guizani
Comput. Commun.3
2019 Optimal Strategies for Live Video Streaming in the Low-latency Regime
abstract
Low-latency is a critical user Quality-of-Experience (QoE) metric for live video streaming. It poses significant challenges for streaming over the Internet. In this paper, we explore the design space of low-latency live video streaming by developing dynamic models and optimal control strategies. We further develop practical live video streaming algorithms within the Model Predictive Control (MPC) framework, namely MPC-Live, to maximize user QoE by adapting the video bitrate while maintaining low end-to-end video latency in dynamic network environment. Through extensive experiments driven by real network traces, we demonstrate that our live video streaming algorithms can improve the performance dramatically within latency range of two to five seconds.
Liyang Sun, Tongyu Zong, Yong Liu 0013, Yao Wang 0001, Haihong Zhu
ICNP4
2019 An HEVC-Compliant Fast Screen Content Transcoding Framework Based on Mode Mapping
abstract
This paper presents a novel fast transcoding framework to efficiently bridge the state-of-art high efficiency video coding (HEVC) standard and its screen content coding (SCC) extension to support the bitstream compatibility over legacy HEVC devices. By exploiting the side information from the SCC bitstream, fast mode and partition decisions are made to accurately translate the novel SCC modes to conventional HEVC modes based on statistical mode mapping techniques. Compared with the full-decoding-full-encoding (FDFE) solution, the proposed framework achieves on average 51% and 82% complexity reductions with 0.57% Bjøntegaard-delta rate (BD-Rate) loss and 9.74% BD-Rate gain under all-intra (AI) and low-delay (LD) configurations, respectively. Compared with the direct transcoding reusing intra mode and inter motion, the proposed mode mapping framework introduces additional 23% and 6% complexity reductions for AI and LD encoding configurations with 0.43% BD-Rate loss and 1.10% BD-Rate saving, respectively. The proposed solution is extended to support the single-input-multiple-output screen content adaptive streaming at the edge clouds, where an SCC bitstream coded in high quality is transcoded into multiple HEVC bitstreams in reduced qualities. Our proposed solution achieves on average 49% and 76% complexity reductions with 0.78% BD-Rate loss and 7.40% BD-Rate gain under AI and LD configurations, respectively.
Fanyi Duanmu, Zhan Ma 0001, Meng Xu 0011, Yao Wang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2019 An ADMM Approach to Masked Signal Decomposition Using Subspace Representation
abstract
Signal decomposition is a classical problem in signal processing, which aims to separate an observed signal into two or more components, each with its own property. Usually, each component is described by its own subspace or dictionary. Extensive research has been done for the case where the components are additive, but in real-world applications, the components are often non-additive. For example, an image may consist of a foreground object overlaid on a background, where each pixel either belongs to the foreground or the background. In such a situation, to separate signal components, we need to find a binary mask which shows the location of each component. Therefore, it requires solving a binary optimization problem. Since most of the binary optimization problems are intractable, we relax this problem to the approximated continuous problem and solve it by alternating optimization technique. We show the application of the proposed algorithm for three applications: separation of text from a background in images, separation of moving objects from a background undergoing global camera motion in videos, and separation of sinusoidal and spike components in 1-D signals. We demonstrate in each case that considering the non-additive nature of the problem can lead to a significant improvement.
Shervin Minaee, Yao Wang 0001
IEEE Trans. Image Process.2
2019 MTBI Identification From Diffusion MR Images Using Bag of Adversarial Visual Features
abstract
In this paper, we propose bag of adversarial features (BAFs) for identifying mild traumatic brain injury (MTBI) patients from their diffusion magnetic resonance images (MRIs) (obtained within one month of injury) by incorporating unsupervised feature learning techniques. MTBI is a growing public health problem with an estimated incidence of over 1.7 million people annually in USA. Diagnosis is based on clinical history and symptoms, and accurate, concrete measures of injury are lacking. Unlike most of the previous works, which use hand-crafted features extracted from different parts of brain for MTBI classification, we employ feature learning algorithms to learn more discriminative representation for this task. A major challenge in this field thus far is the relatively small number of subjects available for training. This makes it difficult to use an end-to-end convolutional neural network to directly classify a subject from MRIs. To overcome this challenge, we first apply an adversarial auto-encoder (with convolutional structure) to learn patch-level features, from overlapping image patches extracted from different brain regions. We then aggregate these features through a bag-of-words approach. We perform an extensive experimental study on a dataset of 227 subjects (including 109 MTBI patients, and 118 age and sex-matched healthy controls) and compare the bag-of-deep-features with several previous approaches. Our experimental results show that the BAF significantly outperforms earlier works relying on the mean values of MR metrics in selected brain regions.
Shervin Minaee, Yao Wang 0001, Alp Aygar, Sohae Chung, Xiuyuan Wang 0001, Yvonne W. Lui, Els Fieremans, Steven Flanagan 0001, Joseph Rath
IEEE Trans. Medical Imaging2
2018 Multispectral Image Intrinsic Decomposition via Subspace Constraint
abstract
Multispectral images contain many clues of surface characteristics of the objects, thus can be used in many computer vision tasks, e.g., recolorization and segmentation. However, due to the complex geometry structure of natural scenes, the spectra curves of the same surface can look very different under different illuminations and from different angles. In this paper, a new Multispectral Image Intrinsic Decomposition model (MIID) is presented to decompose the shading and reflectance from a single multispectral image. We extend the Retinex model, which is proposed for RGB image intrinsic decomposition, for multispectral domain. Based on this, a subspace constraint is introduced to both the shading and reflectance spectral space to reduce the ill-posedness of the problem and make the problem solvable. A dataset of 22 scenes is given with the ground truth of shadings and reflectance to facilitate objective evaluations. The experiments demonstrate the effectiveness of the proposed method.
Weixin Zhu, Linsen Chen, Yao Wang 0001, Tao Yue 0003, Xun Cao
CVPR5
2018 Hybrid Cubemap Projection Format for 360-Degree Video Coding
abstract
360-degree video has become popular in recent years with the advances in virtual reality (VR) and augmented reality (AR) technologies and has been rapidly commercialized. To provide viewers with an immersive experience, 360-degree video requires higher resolution and much higher bandwidth compared with conventional 2D video. In a typical 360-degree video compression and delivery framework, the stitched input 360-degree videos, represented in a native projection format, e.g., equirectangular (ERP), are converted into another projection format, e.g., cubemap (CMP), octahedron (OHP), etc. and frame packed before being fed into existing video codecs. The intermediate projection format is important and would potentially improve the representation efficiency and coding performance. Among all the projection solutions, CMP is very popular and has been widely used in the computer graphics community. The intrinsic rectilinear properties of the CMP format are advantageous for the translational motion model in the modern codec architecture. However, in the CMP representation, the samples on the sphere are not evenly distributed within the faces, resulting in a higher density near the face boundaries and a lower density near the face center. Such non-uniform sampling scheme penalizes the video representation efficiency and degrades the coding performance. Adjusted cubemap projection (ACP) was proposed to address such non-uniform sampling by introducing transform functions to improve the sampling uniformity. However, the transform function parameters in ACP are fixed regardless of the content inside each cube face. In this paper, a generalized hybrid cubemap projection (HCP) is proposed to improve the 360-degree video coding efficiency beyond ACP. HCP is defined by a pair of forward transform and inverse transform functions with a pair of horizontal and vertical transform parameters per cube face. The encoder can choose the optimal sampling for each face by adjusting the parameters in the horizontal and vertical directions based on the 360-degree video content characteristics inside each cube face. In order to maintain the boundary continuities between two neighboring faces, in a 3x2 packing layout, vertical parameter constraints are imposed such that faces in each face-row have the same vertical parameters. The HCP parameters are chosen to minimize the end-to-end weighted conversion error and determined using iterative search between the horizontal and the vertical directions. Significant changes in HCP parameter values can cause drastic change in sampling distribution, and may affect the inter-picture coding efficiency. Therefore, an efficient HCP parameter estimation algorithm is proposed to achieve a better trade-off between the temporal sampling adaptation and the inter-picture prediction efficiency by reducing the temporal variation of HCP parameters. The proposed HCP parameter search algorithm reduces the computational complexity by 5x compared to the exhaustive search method. The HCP parameters are selected by the encoder using the first picture of each Intra Random-Access Point (IRAP) and signalled once per IRAP. In SPS, projection format, frame packing parameters including number of faces in horizontal and vertical directions and each face's position and orientation are signalled. In PPS, the horizontal and vertical HCP parameters in 6-bit precision are encapsulated. The proposed HCP solution is implemented upon JEM-6.0 and 360Lib-3.0 software. Simulation results are reported using the test conditions specified in the JVET Call-for-Evidence (CfE) document. Compared with the CMP and ACP formats, the proposed HCP format demonstrates average 3.0 dB (up to 3.6 dB) and 0.2dB (up to 0.4 dB) End-to-End WS-PSNR improvement for the luma (Y) component, respectively, and average luma (Y) BD-rate reductions of 11.5% (up to 23.0%) and 0.5% (up to 1.0%), respectively.
Fanyi Duanmu, Yuwen He, Xiaoyu Xiu, Philippe Hanhart, Yan Ye 0003, Yao Wang 0001
DCC6
2018 A Subjective Study of Viewer Navigation Behaviors When Watching 360-Degree Videos on Computers
abstract
Virtual reality (VR) applications have become popular recently and rapidly commercialized. The behaviors of users watching 360-degree omni-directional videos have not been fully investigated. In this paper, a dataset of view trajectories for users watching 360-degree videos over computer (desktop/laptop) environment is presented. The dataset includes view center trajectory data collected from viewers watching twelve 360-degree videos over the computer, using mouse to navigate and explore the environment. The selected videos cover a variety of contents, leading to different navigation patterns and behaviors. Based on the dataset, we demonstrate that the viewers share similar viewing patterns over certain 360-degree video category. We also compare the view motion patterns and statistics with prior datasets captured using head-mounted display (HMD). The dataset has been made available online, to facilitate the studies on 360-degree video view prediction, content saliency analysis and VR streaming, etc.
Fanyi Duanmu, Yixiang Mao, Sumanth Srinivasan, Yao Wang 0001
ICME5
2018 Multi-path multi-tier 360-degree video streaming in 5G networks
abstract
360° video streaming is a key component of the emerging Virtual Reality (VR) and Augmented Reality (AR) applications. In 360° video streaming, a user may freely navigate through the captured 360° video scene by changing her desired Field-of-View. High-throughput and low-delay data transfers enabled by 5G wireless networks can potentially facilitate untethered 360° video streaming experience. Meanwhile, the high volatility of 5G wireless links present unprecedented challenges for smooth 360° video streaming. In this paper, novel multi-path multi-tier 360° video streaming solutions are developed to simultaneously address the dynamics in both network bandwidth and user viewing direction. We systematically investigate various design trade-offs on streaming quality and robustness. Through simulations driven by real 5G network bandwidth traces and user viewing direction traces, we demonstrate that the proposed 360° video streaming solutions can achieve a high-level of Quality-of-Experience (QoE) in the challenging 5G wireless network environment.
Liyang Sun, Fanyi Duanmu, Yong Liu 0013, Yao Wang 0001, Yinghua Ye, David Dai
MMSys4
2018 Content-Adaptive 360-Degree Video Coding Using Hybrid Cubemap Projection
abstract
In this paper, a novel hybrid cubemap projection (HCP) is proposed to improve the 360-degree video coding efficiency. HCP allows adaptive sampling adjustments in the horizontal and vertical directions within each cube face. HCP parameters of each cube face can be adjusted based on the input 360-degree video content characteristics for a better sampling efficiency. The HCP parameters can be updated periodically to adapt to temporal content variation. An efficient HCP parameter estimation algorithm is proposed to reduce the computational complexity of parameter estimation. Experimental results demonstrate that HCP format achieves on average luma (Y) BD-rate reduction of 11.51%, 8.0%, and 0.54% compared to equirectangular projection format, cubemap projection format, and adjusted cubemap projection format, respectively, in terms of end-to-end WS-PSNR.
Yuwen He, Xiaoyu Xiu, Philippe Hanhart, Yan Ye 0003, Fanyi Duanmu, Yao Wang 0001
PCS6
2018 A Novel Video Coding Framework Using a Self-Adaptive Dictionary
abstract
In this paper, we propose to use a self-adaptive redundant dictionary, consisting of all possible inter and intra prediction candidates, to directly represent the frame blocks in a video sequence. The self-adaptive dictionary generalizes the conventional predictive coding approach by allowing adaptive linear combinations of prediction candidates, which is solved by an rate-distortion aware L0-norm minimization problem using orthogonal least squares (OLS). To overcome the inefficiency in quantizing and coding coefficients corresponding to correlated chosen atoms, we orthonormalize the chosen atoms recursively as part of OLS process. We further propose a two-stage video coding framework, in which a second stage codes the residual from the chosen atoms using a modified discrete cosine transform (DCT) dictionary that is adaptively orthonormalized with respect to the subspace spanned by the first stage atoms. To determine the transition from the first stage to the second stage, we propose a rate-distortion (RD) aware adaptive switching algorithm. The proposed framework is further extended to accommodate variable block sizes (16×16, 8×8, and 4×4), and the partition mode is derived by a fast partition mode decision algorithm. A contextadaptive binary arithmetic entropy coder is designed to code the symbols of the proposed coding framework. The proposed coder shows competitive, and in some cases better RD performance, compared with the HEVC video coding standard for P-frames.
Yuanyi Xue, Yao Wang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2018 Perceptual Quality Maximization for Video Calls With Packet Losses by Optimizing FEC, Frame Rate, and Quantization
abstract
We consider video calls affected by bursty packet losses, where frame-level forward error correction (FEC) is employed due to delay constraints, and damaged frames and others predicted from them are discarded. Here, a high frame rate (FR) at low bitrates leads to large quantization step sizes (QS), small frames, and suboptimal FEC, whereas a low FR at high bitrates reduces the perceptual quality. To mitigate frame losses and freezing, hierarchical-P (hierP) temporal layering can be used with lower coding efficiency than IPPP coding. We study the received quality maximization for hierP and IPPP, by jointly optimizing the encoding FR, QS, and the FEC redundancy, under sending bitrate constraints. Building upon Q-STAR perceptual quality and R-STAR bitrate models, which depend on QS, and the decoded and encoding FR, respectively, we cast the problem as a combinatorial optimization problem. We solve for the encoding FR and the bitrate using exhaustive search, hill-climbing, and a greedy FEC distribution algorithm to determine the FEC redundancies. We show that, for random losses, the FEC bitrate ratio is an affine function of the packet-loss rate; low encoding FR is preferred at wider sending bitrate ranges with higher packet loss; layer protection is more even at higher bitrates; and IPPP, while achieving higher Q-STAR scores, is prone to freezing. For bursty losses, we show that layer redundancies are higher, rising with the mean burst length, reaching 80%; and hierP achieves higher Q-STAR scores than IPPP for longer bursts, and a smaller mean and variance of decoded frame distances.
Eymen Kurdoglu, Yong Liu 0013, Yao Wang 0001
IEEE Trans. Multim.3
2017 An object based graph representation for video comparison
abstract
This paper develops a novel object based graph model for semantic video comparison. The model describes a video with detected objects as nodes, and relationship between the objects as edges in a graph. We investigated several spatial and temporal features as the graph node attributes, and different ways to describe the spatial-temporal relationship between objects as the edge attributes. To tackle the problem of erratic camera motion on the detected object, a global motion estimation and correction approach is proposed to reveal the true object trajectory. We further propose to evaluate the similarity between two videos by establishing the object correspondence between two object graphs through graph matching. The model is verified on a challenging user generated video dataset. Experiments show that our method outperforms other video representation frameworks in matching videos with the same semantic content. The proposed object graph provides a compact and robust semantic descriptor for a video, which can be used for applications such as video retrieval, clustering and summarization. The graph representation is also flexible to incorporate other features as node and edge attributes.
Yuanyi Xue, Yao Wang 0001
ICIP3
2017 View direction and bandwidth adaptive 360 degree video streaming using a two-tier system
abstract
360 degree video compression and delivery is one of the key components of virtual reality (VR) applications. In such applications, the users may freely control and navigate the captured 3D environment from any viewing direction. Given that only a small portion of the entire video is watched at any time, fetching the entire 360 degree raw video is therefore unnecessary and bandwidth-consuming. In this work, a novel two-tier 360 degree video streaming scheme is proposed to accommodate the dynamics in both network bandwidth and viewing direction. Based on the real-trace driven simulations, we demonstrate that the proposed framework can significantly outperform conventional 360 video streaming schemes.
Fanyi Duanmu, Eymen Kurdoglu, Yong Liu 0013, Yao Wang 0001
ISCAS4
2017 Palmprint recognition using deep scattering network
abstract
Palmprint recognition has drawn a lot of attentions during recent years. Different features and algorithms have been proposed for palmprint recognition in the past such as Gabor-based features, wavelet features, and histogram of oriented lines. In this paper, a powerful image representation, so called deep scattering network, is used for recognition. Scattering network is a convolutional network where its architecture and filters are predefined wavelet transforms. Scattering transform is designed such that the features in its first layer are similar to SIFT descriptors and the higher layers' features capture higher frequency content of the signal which are lost in SIFT. After extraction of scattering features, their dimensionality is reduced by applying principal component analysis (PCA). By doing so, a great amount of computation complexity can be reduced. At the end, the recognition is performed using two different classifiers, multi-class SVM and minimum-distance classifier. The proposed scheme has been tested on a well-known palmprint database and achieved accuracy rates of 99.4% and 99.9% using minimum distance classifier and SVM respectively, outperforming previous algorithms on this dataset.
Shervin Minaee, Yao Wang 0001
ISCAS2
2017 Subspace learning in the presence of sparse structured outliers and noise
abstract
Subspace learning is an important problem, which has many applications in image and video processing. It can be used to find a low-dimensional representation of signals and images. But in many applications, the desired signal is heavily distorted by outliers and noise, which negatively affect the learned subspace. In this work, we present a novel algorithm for learning a subspace for signal representation, in the presence of structured outliers and noise. The proposed algorithm tries to jointly detect the outliers and learn the subspace for images. We present an alternating optimization algorithm for solving this problem, which iterates between learning the subspace and finding the outliers. This algorithm has been trained on a large number of image patches, and the learned subspace is used for image segmentation, and is shown to achieve better segmentation results than prior methods, including least absolute deviation fitting, k-means clustering based segmentation in DjVu, and shape primitive extraction and coding algorithm.
Shervin Minaee, Yao Wang 0001
ISCAS2
2017 HEVC-compliant screen content transcoding based on mode mapping and fast termination
abstract
In this paper, a novel screen content (SC) transcoding framework is presented to efficiently bridge High Efficiency Video Coding (HEVC) and its Screen Content Coding (SCC) extension and support the bitstream backward-compatibility over legacy HEVC devices. Based on the side information analysis of SCC bit-stream, fast mode and partition heuristics are designed to accurately determine the mapping between novel SCC modes and conventional HEVC modes. Compared with the trivial SCC-HEVC transcoding solution, the proposed framework achieves a 42% re-encoding complexity reduction over standard JCT-VC screen content sequences with only 0.52% negligible BD-Rate loss under All-Intra (AI) configuration.
Fanyi Duanmu, Meng Xu 0011, Yao Wang 0001, Zhan Ma 0001
VCIP3
2016 Robust vehicle tracking for urban traffic videos at intersections
abstract
We develop a robust, unsupervised vehicle tracking system for videos of very congested road intersections in urban environments. Raw tracklets from the standard Kanade-Lucas-Tomasi tracking algorithm are treated as sample points and grouped to form different vehicle candidates. Each tracklet is described by multiple features including position, velocity, and a foreground score derived from robust PCA background subtraction. By considering each tracklet as a node in a graph, we build the adjacency matrix for the graph based on the feature similarity between the tracklets and group these tracklets using spectral embedding and Dirichelet Process Gaussian Mixture Models. The proposed system yields excellent performance for traffic videos captured in urban environments and highways.
A. Chiang, Gregory Dobler, Yao Wang 0001, Kun Xie 0002, Kaan Özbay, Masoud Ghandehari
AVSS4
2016 A novel screen content fast transcoding framework based on statistical study and machine learning
abstract
In this paper, a novel screen content transcoding framework is presented to efficiently bridge the state-of-art High Efficiency Video Coding (HEVC) standard and its incoming screen content coding (SCC) extension currently pending finalization. The proposed scheme is implemented as an Intra-coding “pre-processing” module on top of official SCC test model software (SCM). Both Coding Unit (CU) statistical features (such as CU color quantity, CU pixel variance, CU edge directionality distribution, etc.) and decoded video side information (such as CU partitions, modes, residual, etc.) are jointly analyzed. Accordingly, fast CU mode decisions and CU partitions bypass / termination heuristics are designed. Compared with SCM-4.0 official release, the proposed fast transcoding scheme can achieve an average of 48% re-encoding complexity reduction over JCT-VC screen content testing sequences with less than 2.14% marginal BD-Rate increase under SCC common testing conditions for All-Intra (AI) configuration.
Fanyi Duanmu, Zhan Ma 0001, Meng Xu 0011, Yao Wang 0001
ICIP5
2016 Screen content image segmentation using sparse decomposition and total variation minimization
abstract
Sparse decomposition has been widely used for different applications, such as source separation, image classification, image denoising and more. This paper presents a new algorithm for segmentation of an image into background and foreground text and graphics using sparse decomposition and total variation minimization. The proposed method is designed based on the assumption that the background part of the image is smoothly varying and can be represented by a linear combination of a few smoothly varying basis functions, while the foreground text and graphics can be modeled with a sparse component overlaid on the smooth background. The background and foreground are separated using a sparse decomposition framework regularized with a few suitable regularization terms which promotes the sparsity and connectivity of foreground pixels. This algorithm has been tested on a dataset of images extracted from HEVC standard test sequences for screen content coding, and is shown to have superior performance over some prior methods, including least absolute deviation fitting, k-means clustering based segmentation in DjVu and shape primitive extraction and coding (SPEC) algorithm.
Shervin Minaee, Yao Wang 0001
ICIP2
2016 Real-time bandwidth prediction and rate adaptation for video calls over cellular networks
abstract
We study interactive video calls between two users, where at least one of the users is connected over a cellular network. It is known that cellular links present highly-varying network bandwidth and packet delays. If the sending rate of the video call exceeds the available bandwidth, the video frames may be excessively delayed, destroying the interactivity of the video call. In this paper, we present Rebera, a cross-layer design of proactive congestion control, video encoding and rate adaptation, to maximize the video transmission rate while keeping the one-way frame delays sufficiently low. Rebera actively measures the available bandwidth in real-time by employing the video frames as packet trains. Using an online linear adaptive filter, Rebera makes a history-based prediction of the future capacity, and determines a bit budget for the video rate adaptation. Rebera uses the hierarchical-P video encoding structure to provide error resilience and to ease rate adaptation, while maintaining low encoding complexity and delay. Furthermore, Rebera decides in real time whether to send or discard an encoded frame, according to the budget, thereby preventing self-congestion and minimizing the packet delays. Our experiments with real cellular link traces demonstrate Rebera can, on average, deliver higher bandwidth utilization and shorter packet delays than Apple's FaceTime.
Eymen Kurdoglu, Yong Liu 0013, Yao Wang 0001, Yongfang Shi, Chenchen Gu, Jing Lyu
MMSys3
2016 Depth recovery via decomposition of polynomial and piece-wise constant signals
abstract
This paper proposes a novel decomposition model for high-quality depth recovery (DMDR) from low quality depth measurement accompanied by high-resolution RGB image. We observe that depth patches extracted from the depth map containing smooth regions separated by curves, can be decomposed simultaneously by a low-order polynomial surface and a piece-wise constant signal. In our model, the polynomial surface component is regularized by least-square polynomial smoothing, while the piece-wise constant component is constrained by total variation filtering. The model is effectively solved by the alternating direction method under the augmented Lagrangian multiplier (ALM-ADM) algorithm. Experimental results show that our method is able to handle various types of depth degradation under the designed signal decomposition model, and produces high-quality depth recovery results.
Xinchen Ye, Jing-Yu Yang 0002, Chunping Hou, Yao Wang 0001
VCIP5
2016 Nested Graph Cut for Automatic Segmentation of High-Frequency Ultrasound Images of the Mouse Embryo
abstract
We propose a fully automatic segmentation method called nested graph cut to segment images (2D or 3D) that contain multiple objects with a nested structure. Compared to other graph-cut-based methods developed for multiple regions, our method can work well for nested objects without requiring manual selection of initial seeds, even if different objects have similar intensity distributions and some object boundaries are missing. Promising results were obtained for separating the brain ventricles, the head, and the uterus region in the mouse-embryo head images obtained using high-frequency ultrasound imaging. The proposed method achieved mean Dice similarity coefficients of 0.87 ±0.04 and 0.89 ±0.06 for segmenting BVs and the head, respectively, compared to manual segmentation results by experts on 40 3D images over five gestation stages.
Jen-Wei Kuo, Jonathan Mamou, Orlando Aristizábal, Jeffrey A. Ketterling, Yao Wang 0001
IEEE Trans. Medical Imaging6
2016 Dealing With User Heterogeneity in P2P Multi-Party Video Conferencing: Layered Distribution Versus Partitioned Simulcast
abstract
We consider peer-to-peer multi-party video conferencing (P2P-MPVC), where users with different uplink -downlink capacities send their videos using multicast trees. One way to deal with user bandwidth heterogeneity is employing layered video coding, generating multiple layers with different rates, whereas an alternative is partitioning the receivers of each source and disseminating a different non-layered video version within each group. In this paper, we aim to maximize the received video quality for both systems under uplink-downlink capacity constraints, while constraining the number of hops the packets traverse to two. We first show any multicast tree is equivalent to a collection of 1-hop and 2-hop trees, under user uplink-downlink capacity constraints. This reveals that the packet overlay hop count can be limited to two without sacrificing the achievable rate performance. Assuming a fine granularity scalable stream that can be truncated at any rate, we propose an algorithm that solves for the number of video layers, layer rates, and distribution trees for the layered system. For the partitioned simulcast system, we develop an algorithm to determine the receiver partitions along with the video rate and the distribution trees for each group. Through numerical comparison, we show that the partitioned simulcast system achieves the same average receiving quality as the ideal layered system without any coding overhead for the four-user systems simulated, and better quality than the layered system when the layered coding overhead is only 20%. The two systems perform similarly for the six-user case if the layered coding overhead is 10%.
Eymen Kurdoglu, Yong Liu 0013, Yao Wang 0001
IEEE Trans. Multim.3
2015 Fast CU partition decision using machine learning for screen content compression
abstract
Screen Content Coding (SCC) extension is currently being developed by Joint Collaborative Team on Video Coding (JCT-VC), as the final extension for the latest High-Efficiency Video Coding (HEVC) standard. It employs some new coding tools and algorithms (including palette coding mode, intra block copy mode, adaptive color transform, adaptive motion compensation precision, etc.), and outperforms HEVC by over 40% bitrate reduction on typical screen contents. However, enormous computational complexity is introduced on encoder primarily due to heavy optimization processing, especially rate distortion optimization (RDO) for Coding Unit (CU) partition decision and mode selection. This paper proposes a novel machine learning based approach for fast CU partition decision using features that describe CU statistics and sub-CU homogeneity. The proposed scheme is implemented as a "preprocessing" module on top of the Screen Content Coding reference software (SCM-3.0). Compared with SCM-3.0, experimental results show that our scheme can achieve 36.8% complexity reduction on average with only 3.0% BD-rate increase over 11 JCT-VC testing sequences when encoded using "All Intra" (AI) configuration.
Fanyi Duanmu, Zhan Ma 0001, Yao Wang 0001
ICIP3
2015 Assessing the visual effect of non-periodic temporal variation of quantization stepsize in compressed video
abstract
This work investigates the impact of non-periodic quantization step size variation on perceptual video quality. We composed non-periodic temporal variation patterns from periodic variation patterns. We present subjective test results and examine how does the overall quality under non-periodic variation relate to the qualities of individual periodic components, and how does it relate to the instantaneous qualities at different time points. We propose two quality models based on our observations and data analysis. Such quality assessment and modeling are essential for designing video adaptation strategy when delivering video over dynamically changing wireless links.
Zhili Guo, Yao Wang 0001
ICIP2
2015 Screen content image segmentation using least absolute deviation fitting
abstract
We propose an algorithm for separating the foreground (mainly text and line graphics) from the smoothly varying background in screen content images. The proposed method is designed based on the assumption that the background part of the image is smoothly varying and can be represented by a linear combination of a few smoothly varying basis functions, while the foreground text and graphics create sharp discontinuity and cannot be modeled by this smooth representation. The algorithm separates the background and foreground using a least absolute deviation method to fit the smooth model to the image pixels. This algorithm has been tested on several images from HEVC standard test sequences for screen content coding, and is shown to have superior performance over other popular methods, such as k-means clustering based segmentation in DjVu and shape primitive extraction and coding (SPEC) algorithm. Such background/foreground segmentation are important pre-processing steps for text extraction and separate coding of background and foreground for compression of screen content images.
Shervin Minaee, Yao Wang 0001
ICIP2
2015 A two-stage video coding framework with both self-adaptive redundant dictionary and adaptively orthonormalized DCT basis
abstract
In this work, we propose a two-stage video coding framework, as an extension of our previous one-stage framework in [1]. The two-stage frameworks consists two different dictionaries. Specifically, the first stage directly finds the sparse representation of a block with a self-adaptive dictionary consisting of all possible inter-prediction candidates by solving an L0-norm minimization problem using orthogonal least squares (OLS), and the second stage codes the residual using altered DCT dictionary orthonormalized to the subspace spanned by the first stage atoms. The transition of the first stage and the second stage is adaptively determined based on the estimated residual reduction per bit. We further propose a complete context adaptive entropy coder to efficiently code the locations and the coefficients of chosen first stage atoms. Simulation results show that the proposed coder significantly improves the RD performance over our previous one-stage coder. More importantly, the two-stage coder, using a fixed block size and inter-prediction only, outperforms the H.264 coder (x264) and is competitive with the HEVC reference coder (HM) over a large rate range.
Yuanyi Xue, Yao Wang 0001
ICIP3
2015 Foreground-Background Separation From Video Clips via Motion-Assisted Matrix Restoration
abstract
Separation of video clips into foreground and background components is a useful and important technique, making recognition, classification, and scene analysis more efficient. In this paper, we propose a motion-assisted matrix restoration (MAMR) model for foreground-background separation in video clips. In the proposed MAMR model, the backgrounds across frames are modeled by a low-rank matrix, while the foreground objects are modeled by a sparse matrix. To facilitate efficient foreground-background separation, a dense motion field is estimated for each frame, and mapped into a weighting matrix which indicates the likelihood that each pixel belongs to the background. Anchor frames are selected in the dense motion estimation to overcome the difficulty of detecting slowly moving objects and camouflages. In addition, we extend our model to a robust MAMR model against noise for practical applications. Evaluations on challenging datasets demonstrate that our method outperforms many other state-of-the-art methods, and is versatile for a wide range of surveillance videos.
Xinchen Ye, Jing-Yu Yang 0002, Kun Li 0001, Chunping Hou, Yao Wang 0001
IEEE Trans. Circuits Syst. Video Technol.6
2015 One-Pass Mode and Motion Decision for Multilayer Quality Scalable Video Coding
abstract
This paper presents a novel low-complexity motion estimation and mode decision algorithm for encoding multiple quality layers following the H.264/scalable video coding standard, considering both coarse grain scalability (CGS) and medium grain scalability (MGS). The proposed algorithm conducts motion estimation and mode decision only at the base layer (BL) and enforces the higher layers to inherit the motion and mode decisions of the BL. In order for the decision made at the BL to be nearly optimal for all layers, we use the highest layer reconstructed frame as the reference frame for motion estimation and set the Lagrangian multipliers according to the quantization parameter of the current and higher layers. We also propose a simple early skip/direct decision to further boost the encoding speed. Mode decision and motion estimation is conducted at a higher layer only if the layer below it uses the skip/direct mode for a block. Significant complexity reduction can be achieved because the mode and motion estimation is performed at most once for each macroblock. Because the mode and motion information only needs to be transmitted once, we also achieve a slightly better rate-distortion (R-D) performance for typical videos. Experiments have shown more than 2× (up to 5×) speedup for a three-layer encoder against the conventional R-D optimized reference software JSVM on both CIF and HD sequences, and for both CGS and MGS, with the tradeoff of the coding efficiency measured by the Bjontegaard delta rate.
Meng Xu 0011, Zhan Ma 0001, Yao Wang 0001
IEEE Trans. Image Process.3
2015 Wireless Video Multicast With Cooperative and Incremental Transmission of Parity Packets
abstract
This paper introduces a novel and efficient approach for user cooperation in wireless video multicast using randomized distributed space time codes (R-DSTC), in which the sender first transmits the source packets, and the sender and receivers that have received all source packets then generate and send the parity packets simultaneously using R-DSTC. As more parity packets are delivered, more receivers can recover all source packets and join the parity packet transmission. Four variations of the proposed systems are considered. The first one requires complete channel information between the sender and all receivers and between all receivers to derive the optimal transmission rates for sending source and parity packets, and employs receiver feedback to determine when to terminate parity transmission. The other three suboptimal systems do not require full channel information and/or receiver feedback, and hence are more feasible in practice. All four versions can support significantly higher video rates and correspondingly higher quality of decoded video, than prior approaches in the literature, which require full channel information but not feedback.
Zhili Guo, Yao Wang 0001, Elza Erkip, Shivendra S. Panwar
IEEE Trans. Multim.2
2015 A Novel No-Reference Video Quality Metric for Evaluating Temporal Jerkiness due to Frame Freezing
abstract
In this work, we propose a novel no-reference (NR) video quality metric that evaluates the impact of frame freezing due to either packet loss or late arrival. Our metric uses a trained neural network acting on features that are chosen to capture the impact of frame freezing on the perceived quality. The considered features include the number of freezes, freeze duration statistics, inter-freeze distance statistics, frame difference before and after the freeze, normal frame difference, and the ratio of them. We use the neural network to find the mapping between features and subjective test scores. We optimize the network structure and the feature selection through a cross-validation procedure, using training samples extracted from both VQEG and LIVE video databases. The resulting feature set and network structure yields accurate quality prediction for both the training data containing 54 test videos and a separate testing dataset including 14 videos, with Pearson correlation coefficients greater than 0.9 and 0.8 for the training set and the testing set, respectively. Our proposed metric has low complexity and could be utilized in a system with real-time processing constraint.
Yuanyi Xue, Beril Erkin, Yao Wang 0001
IEEE Trans. Multim.3
2014 Robust shape-constrained active contour for whole heart segmentation in 3-D CT images for radiotherapy planning
abstract
Automatic segmentation of the whole heart in computed tomography(CT) image is crucial for efficient treatment planning of thoracic radiotherapy. In this paper, we propose a fully automatic method for whole heart segmentation of thoracic CT images. A robust active shape model (Robust ASM) is proposed using shape models developed from training data to reduce outliers due to similar intensity of neighboring organs. A novel shape constrained active contour model is presented to further improve the segmentation result. A mean point-to-surface error of 2.37mm was measured based on 38 images. The averaged Dice index is 0.90.
Yao Wang 0001, Gabor Jozsef
ICIP2
2014 Video coding using a self-adaptive redundant dictionary consisting of spatial and temporal prediction candidates
abstract
All standard video coders are based on the prediction plus transform representation of an image block, which predicts the current block using various intra- and inter-prediction modes and then represents the prediction error using a fixed orthonormal transform. We propose to directly represent a mean-removed block using a redundant dictionary consisting of all possible inter-prediction candidates with integer motion vectors (mean-removed). In general the dictionary may also contain some intra-prediction candidates and some pre-designed fixed dictionary atoms. However, simulation results reported in this papers are obtained by using the inter-prediction candidates only. We determine the coefficients by minimizing the L0 norm of the coefficients subject to a constraint on the sparse approximation error. We show that using such a self-adaptive dictionary can lead to a very sparse representation, with significantly fewer non-zero coefficients than using the DCT transform on the prediction error. We further propose a modified orthogonal matching pursuit (OMP) algorithm which othonormalizes each new chosen atom with respect to all previously chosen and orthonormalized atoms. Each image block is represented by the quantized coefficients corresponding to the othonormalized atoms, to overcome the inefficiency associated with using non-orthonormal atoms. Each image block is represented by its mean, which is predictively coded, the indices of the chosen atoms, and the quantized coefficients. Each variable is coded based on its unconditional distribution. Simulation results show that the proposed coder can achieve significant gain over the H.264 coder (implemented using x264) and achieve similar performance comparing to the HEVC reference encoder (HM).
Yuanyi Xue, Yao Wang 0001
ICME2
2014 Q-STAR: A Perceptual Video Quality Model Considering Impact of Spatial, Temporal, and Amplitude Resolutions
abstract
In this paper, we investigate the impact of spatial, temporal, and amplitude resolution on the perceptual quality of a compressed video. Subjective quality tests were carried out on a mobile device and a total of 189 processed video sequences with 10 source sequences included in the test. Subjective data reveal that the impact of spatial resolution (SR), temporal resolution (TR), and quantization stepsize (QS) can each be captured by a function with a single content-dependent parameter, which indicates the decay rate of the quality with each resolution factor. The joint impact of SR, TR, and QS can be accurately modeled by the product of these three functions with only three parameters. The impact of SR and QS on the quality are independent of that of TR, but there are significant interactions between SR and QS. Furthermore, the model parameters can be predicted accurately from a few content features derived from the original video. The proposed model correlates well with the subjective ratings with a Pearson correlation coefficient of 0.985 when the model parameters are predicted from content features. The quality model is further validated on six other subjective rating data sets with very high accuracy and outperforms several well-known quality models.
Yen-Fu Ou, Yuanyi Xue, Yao Wang 0001
IEEE Trans. Image Process.3
2014 Color-Guided Depth Recovery From RGB-D Data Using an Adaptive Autoregressive Model
abstract
This paper proposes an adaptive color-guided autoregressive (AR) model for high quality depth recovery from low quality measurements captured by depth cameras. We observe and verify that the AR model tightly fits depth maps of generic scenes. The depth recovery task is formulated into a minimization of AR prediction errors subject to measurement consistency. The AR predictor for each pixel is constructed according to both the local correlation in the initial depth map and the nonlocal similarity in the accompanied high quality color image. We analyze the stability of our method from a linear system point of view, and design a parameter adaptation scheme to achieve stable and accurate depth recovery. Quantitative and qualitative evaluation compared with ten state-of-the-art schemes show the effectiveness and superiority of our method. Being able to handle various types of depth degradations, the proposed method is versatile for mainstream depth sensors, time-of-flight camera, and Kinect, as demonstrated by experiments on real systems.
Jing-Yu Yang 0002, Xinchen Ye, Kun Li 0001, Chunping Hou, Yao Wang 0001
IEEE Trans. Image Process.5
2013 Proxy-Based Multi-Stream Scalable Video Adaptation Over Wireless Networks Using Subjective Quality and Rate Models
abstract
Despite growing maturity in broadband mobile networks, wireless video streaming remains a challenging task, especially in highly dynamic environments. Rapidly changing wireless link qualities, highly variable round trip delays, and unpredictable traffic contention patterns often hamper the performance of conventional end-to-end rate adaptation techniques such as TCP-friendly rate control (TFRC). Furthermore, existing approaches tend to treat all flows leaving the network edge equally, without accounting for heterogeneity in the underlying wireless link qualities or the different rate utilities of the video streams. In this paper, we present a proxy-based solution for adapting the scalable video streams at the edge of a wireless network, which can respond quickly to highly dynamic wireless links. Our design adopts the recently standardized scalable video coding (SVC) technique for lightweight rate adaptation at the edge. Leveraging previously developed rate and quality models of scalable video with both temporal and amplitude scalability, we derive the rate-quality model that relates the maximum quality under a given rate by choosing the optimal frame rate and quantization stepsize. The proxy iteratively allocates rates of different video streams to maximize a weighted sum of video qualities associated with different streams, based on the periodically observed link throughputs and the sending buffer status. The temporal and amplitude layers included in each video are determined to optimize the quality while satisfying the rate assignment. Simulation studies show that our scheme consistently outperforms TFRC in terms of agility to track link qualities and overall subjective quality of all streams. In addition, the proposed scheme supports differential services for different streams, and competes fairly with TCP flows.
Yao Wang 0001, Jiang Zhu 0002, Flavio Bonomi
IEEE Trans. Multim.3
2013 Modeling and Analysis of Skype Video Calls: Rate Control and Video Quality
abstract
Video-conferencing has recently gained its momentum and is widely adopted by end-consumers. But there have been very few studies on the network impacts of video calls and the user Quality-of-Experience (QoE) under different network conditions. In this paper, we study the rate control and video quality of Skype video call, and analyze the network impacts in large-scale networks. We first measure the behaviors of Skype video call on a controlled network testbed. By varying packet loss rate, propagation delay and available network bandwidth, we observe how Skype adjusts its sending rate, FEC redundancy, video rate and frame rate. It is found that Skype is robust against mild packet losses and propagation delays, and can efficiently utilize the available network bandwidth. We also find that it employs an overly aggressive FEC protection strategy. Based on the measurement results, we develop rate control model, FEC model, and video quality model for Skype video calls. Extrapolating from the models, we conduct numerical analysis to study the network impacts. We demonstrate that user back-offs upon quality degradation serve as an effective user-level rate control scheme. We also show that Skype video calls are indeed TCP-friendly and respond to congestion quickly when the network is overloaded. Through a case study of a 4G wireless network, we demonstrate that the proposed models can be used in user-QoE-aware network provisioning.
Xinggong Zhang, Yang Xu 0011, Yong Liu 0013, Zongming Guo, Yao Wang 0001
IEEE Trans. Multim.6
2012 QoE-based multi-stream scalable video adaptation over wireless networks with proxy
abstract
In this paper, we present a proxy-based solution for adapting the scalable video streams at the edge of a wireless network, which can respond quickly to highly dynamic wireless links. Our design adopts the scalable video coding (SVC) technique for lightweight rate adaptation at the edge. We derive a QoE model, i.e., rate-quality tradeoff model, that relates the maximum subjective quality under a given rate by choosing the optimal frame rate and quantization stepsize. The proxy iteratively allocates rates of different video streams to maximize a weighted sum of video qualities associated with different streams, based on the periodically observed link throughputs and the sending buffer status. Simulation studies show that our scheme consistently outperforms TFRC in terms of agility to track link qualities and overall quality of all streams. In addition, the proposed scheme supports differential services for different streams, and competes fairly with TCP flows.
Yao Wang 0001, Jiang Zhu 0002, Flavio Bonomi
ICC3
2012 Two-way wireless video communication using Randomized cooperation, Network Coding and packet level FEC
abstract
Two-way real-time video communication in wireless networks requires high bandwidth, low delay and error resiliency. This paper addresses these demands by proposing a system with the integration of Network Coding (NC), user cooperation using Randomized Distributed Space-time Coding (R-DSTC) and packet level Forward Error Correction (FEC) under a one-way delay constraint. Simulation results show that the proposed scheme significantly outperforms both conventional direct transmission as well as R-DSTC based two-way cooperative transmission, and is most effective when the distance between the users is large.
Xiaozhong Xu, Özgü Alay, Elza Erkip, Yao Wang 0001, Shivendra S. Panwar
ICC4
2012 Optimization of spatial, temporal and amplitude resolution for rate-constrained video coding and scalable video adaptation
abstract
This paper considers how to choose the frame size, frame rate, and quantization stepsize to optimize the perceptual quality for a given rate constraint. The proposed solution leverages previously developed quality and rate models that explicitly consider the impact of spatial, temporal, and amplitude resolution (STAR) on the quality and rate. Using these models we further propose algorithms for ordering the STAR layers to form a rate-quality optimized stream, which can greatly facilitate scalable video adaptation.
Zhan Ma 0001, Yao Wang 0001
ICIP3
2012 Perceptual quality of video with quantization variation: A subjective study and analytical modeling
abstract
This work investigates the impact of temporal variation of quantization stepsize (QS) on perceptual video quality. Among many dimensions of QS variation, as a first step we focus on videos in which two QS's, alternate over fixed intervals. We present subjective test results, and analyze the influence of several factors (including the QS difference, QS ratio, changing intervals, and video content). According the observation and data analysis, we propose analytical models that relate the perceived quality with the two QS's. Such quality assessment and modeling are essential in making video adaptation decisions when delivering video over dynamically changing wireless links.
Yen-Fu Ou, Huiqi Zeng, Yao Wang 0001
ICIP3
2012 Profiling Skype video calls: Rate control and video quality
abstract
Video telephony has recently gained its momentum and is widely adopted by end-consumers. But there have been very few studies on the network impacts of video calls and the user Quality-of-Experience (QoE) under different network conditions. In this paper, we study the rate control and video quality of Skype video calls. We first measure the behaviors of Skype video calls on a controlled network testbed. By varying packet loss rate, propagation delay and bandwidth, we observe how Skype adjusts its rates, FEC redundancy and video quality. We find that Skype is robust against mild packet losses and propagation delays, and can efficiently utilize the available network bandwidth. We also find that Skype employs an overly aggressive FEC protection strategy. Based on the measurement results, we develop rate control model, FEC model, and video quality model for Skype. Extrapolating from the models, we conduct numerical analysis to study the network impacts of Skype. We demonstrate that user back-offs upon quality degradation serve as an effective user-level rate control scheme. We also show that Skype video calls are indeed TCP-friendly and respond to congestion quickly when the network is overloaded.
Xinggong Zhang, Yang Xu 0011, Yong Liu 0013, Zongming Guo, Yao Wang 0001
INFOCOM6
2012 Modeling of Rate and Perceptual Quality of Compressed Video as Functions of Frame Rate and Quantization Stepsize and Its Applications
abstract
This paper first investigates the impact of frame rate and quantization on the bit rate and perceptual quality of compressed video. We propose a rate model and a quality model, both in terms of the quantization stepsize and frame rate. Both models are expressed as the product of separate functions of quantization stepsize and frame rate. The proposed models are analytically tractable, each requiring only a few content-dependent parameters. The rate model is validated over videos coded using both scalable and nonscalable encoders, under a variety of encoder settings. The quality model is validated only for a scalable video, although it is expected to be applicable to a single-layer video as well. We further investigate how to predict the model parameters using the content features extracted from original videos. Results show accurate bit rate and quality prediction (average Pearson correlation >;0.99) can be achieved with model parameters predicted using three features. Finally, we apply rate and quality models for rate-constrained scalable bitstream adaptation and frame rate adaptive rate control. Simulations show that our model-based solutions produce better video quality compared with conventional video adaptation and rate control.
Zhan Ma 0001, Meng Xu 0011, Yen-Fu Ou, Yao Wang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2011 Modeling of rate and perceptual quality of video and its application to frame rate adaptive rate control
abstract
In a prior work, we have developed both rate and perceptual quality models for temporal and amplitude (i.e., SNR) scalable video produced by the H.264/SVC encoder. In this paper, we validate from experimental data that the functional form of the rate model is applicable to H.264/AVC encoded video, which has the same temporal scalability but no SNR scalability, but the model parameter values differ. We further investigate how to predict both rate and quality model parameters using content features computed from the original video. Experimental data show that with proper feature combination, we can estimate the model parameters very accurately, and the estimated bit rate and quality using the predicted model parameters match with the measured bit rate and quality with high Pearson correlation (PC) and small root mean square error (RMSE). We have implemented a simple pre-processor in the H.264/AVC encoder to guide the frame rate adaptive rate control. Results show that our model-based frame rate adaptive rate control outperforms the default rate control algorithm with better quality.
Zhan Ma 0001, Meng Xu 0011, Kyeong Yang, Yao Wang 0001
ICIP4
2011 Perceptual Quality Assessment of Video Considering Both Frame Rate and Quantization Artifacts
abstract
In this paper, we explore the impact of frame rate and quantization on perceptual quality of a video. We propose to use the product of a spatial quality factor that assesses the quality of decoded frames without considering the frame rate effect and a temporal correction factor, which reduces the quality assigned by the first factor according to the actual frame rate. We find that the temporal correction factor follows closely an inverted falling exponential function, whereas the quantization effect on the coded frames can be captured accurately by a sigmoid function of the peak signal-to-noise ratio. The proposed model is analytically simple, with each function requiring only a single content-dependent parameter. The proposed overall metric has been validated using both our subjective test scores as well as those reported by others. For all seven data sets examined, our model yields high Pearson correlation (higher than 0.9) with measured mean opinion score (MOS). We further investigate how to predict parameters of our proposed model using content features derived from the original videos. Using predicted parameters from content features, our model still fits with measured MOS with high correlation.
Yen-Fu Ou, Zhan Ma 0001, Yao Wang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2011 Cooperative Layered Video Multicast Using Randomized Distributed Space Time Codes
abstract
With the increased popularity of mobile multimedia services, efficient and robust video multicast strategies are of critical importance. Cooperative communications has been shown to improve the robustness and the data rates for point-to-point transmission. In this paper, a two-hop cooperative transmission scheme for multicast in infrastructure-based networks is used, where multiple relays forward the data simultaneously using randomized distributed space time codes (RDSTC). This randomized cooperative transmission is further integrated with layered video coding and packet level forward error correction (FEC) to enable efficient and robust video multicast. Three different schemes are proposed to find the system operating parameters based on the availability of the channel information at the source station: RDSTC with full channel information, RDSTC with limited channel information, and RDSTC with node count. The performance of these three schemes are compared with rate adaptive direct transmission and conventional multicast that does not use rate adaptation. The results show that while rate-adaptive direct transmission provides better video quality than conventional multicast, all three proposed randomized cooperative schemes outperform both strategies significantly as long as the network has enough nodes. Furthermore, the performance gap between RDSTC with full channel information and RDSTC with limited channel information or node count is relatively small, indicating the robustness of the proposed cooperative multicast system using RDSTC.
Özgü Alay, Pei Liu 0001, Yao Wang 0001, Elza Erkip, Shivendra S. Panwar
IEEE Trans. Multim.3
2011 On Complexity Modeling of H.264/AVC Video Decoding and Its Application for Energy Efficient Decoding
abstract
This paper proposes a new complexity model for H.264/AVC video decoding. The model is derived by decomposing the entire decoder into several decoding modules (DM), and identifying the fundamental operation unit (termed complexity unit or CU) in each DM. The complexity of each DM is modeled by the product of the average complexity of one CU and the number of CUs required. The model is shown to be highly accurate for software video decoding both on Intel Pentium mobile 1.6-GHz and ARM Cortex A8 600-MHz processors, over a variety of video contents at different spatial and temporal resolutions and bit rates. We further show how to use this model to predict the required clock frequency and hence perform dynamic voltage and frequency scaling (DVFS) for energy efficient video decoding. We evaluate achievable power savings on both the Intel and ARM platforms, by using analytical power models for these two platforms as well as real experiments with the ARM-based TI OMAP35x EVM board. Our study shows that for the Intel platform where the dynamic power dominates, a power saving factor of 3.7 is possible. For the ARM processor where the static leakage power is not negligible, a saving factor of 2.22 is still achievable.
Zhan Ma 0001, Yao Wang 0001
IEEE Trans. Multim.3
2010 Error resilient video multicast using Randomized Distributed Space Time Codes
abstract
In this paper we study a two-hop cooperative transmission scheme where multiple relays forward the data simultaneously using Randomized Distributed Space Time Codes (R-DSTC). We propose to integrate this randomized cooperative transmission with layered video coding and packet level Forward Error Correction (FEC) to enable error resilient video multicast. Data rates in both hops as well as the FEC rate are adopted to maximize the video quality. Our results show that while rate-adaptive direct transmission provides better video quality than conventional multicast, randomized cooperative scheme outperforms both strategies significantly.
Özgü Alay, Pei Liu 0001, Yao Wang 0001, Elza Erkip, Shivendra S. Panwar
ICASSP3
2010 On compressed sensing in parallel MRI of cardiac perfusion using temporal wavelet and TV regularization
abstract
Imaging of cardiac perfusion with MR is a challenging area of research especially due to the motion of the heart and limited time of data acquisition. Compressed sensing is a popular signal estimation method recently adopted by researchers in MRI which can improve the spatial and/or temporal resolution of the acquired images by reducing the number of necessary samples for image reconstruction. This paper focuses on performance of temporal regularization with total variation and wavelets in compressed sensing. The impact of the choice of regularization parameters on the image quality and the temporal variation of intensity in region of interests (ROIs) are discussed. It is found that selecting the regularization parameter so as to optimize the quality of the reconstructed image sequence as a whole, leads to erroneous reconstruction of certain regions due to over regularization.
Cagdas Bilen, Ivan W. Selesnick, Yao Wang 0001, Ricardo Otazo, Leon Axel, Daniel K. Sodickson
ICASSP3
2010 Perceptual quality of video with frame rate variation: A subjective study
abstract
This work investigates the impact of periodic frame rate variation on perceptual video quality. Among many dimensions of frame rate variation, as a first step we focus on videos in which two frame rates alternate over fixed intervals. We present subjective test results, and analyze the influence of several factors (including the average frame rate, the frame rate deviation, and the video content) on the perceptual quality.
Yen-Fu Ou, Yao Wang 0001
ICASSP3
2010 P2P Trading in Social Networks: The Value of Staying Connected
abstract
The success of future P2P applications ultimately depends on whether users will contribute their bandwidth, CPU and storage resources to a larger community. In this paper, we propose a new incentive paradigm, Networked Asynchronous Bilateral Trading (NABT), which can be applied to a broad range of P2P applications. In NABT, peers belong to an underlying social network, and each pair of friends keeps track of a credit balance between them. When user Alice provides a service (a file, storage space, computation and so on) to her friend Bob, she charges Bob credits. Thus, in NABT, there is no global currency; instead, there are only credit balances maintained between pairs of friends. NABT allows peers to supply each other asynchronously and further allows peers to trade with remote peers through intermediaries. We theoretically show that NABT is perfectly efficient with balanced demands and supports "networked tit-for-tat". The efficiency of NABT with unbalanced demands is determined by the min-cut of credit limits of the underlying social network. Using simulations driven by MySpace traces, we demonstrate that a simple two-hop NABT design can have high trading efficiency, provide service differentiation, exploit trading intermediaries, and discourage free-riders.
Zhengye Liu, Yong Liu 0013, Keith W. Ross, Yao Wang 0001, Markus Mobius
INFOCOM5
2010 Enhanced parity packet transmission for Video multicast using R-DSTC
abstract
In this paper, a cooperative multicast scheme that uses Randomized Distributed Space Time Codes (R-DSTC), along with packet level Forward Error Correction (FEC), is studied. For the source packets, two-hop transmission is considered, where a packet is transmitted first by the access point (AP), and then forwarded using R-DSTC by the nodes that receive the packet. On the other hand, parity packets are generated by the nodes that receive all the source packets correctly and are transmitted using R-DSTC. The optimum transmission rates for source and parity packets, as well as the number of parity packets required, are determined such that the video quality at all nodes is maximized. It is shown that this scheme can support a higher video rate than a previously developed R-DSTC based scheme where both source and parity packets go through a two-hop transmission, as well as non-cooperative direct transmission.
Özgü Alay, Zhili Guo, Yao Wang 0001, Elza Erkip, Shivendra S. Panwar
PIMRC3
2010 Layer bargaining: multicast layered video over wireless networks
abstract
Wireless video multicast efficiently streams video to multiple receivers. When designing a wireless video multicast system, the system must be able to handle (i) packet losses induced by the underlying wireless channel and (ii) receiver heterogeneity in channel condition in a multicast group. We propose Layer Bargaining, a wireless video multicast design in infrastructure-based wireless networks that simultaneously addresses both of the above problems. To combat packet losses and improve received video quality, we propose a layered hybrid ARQ scheme that provides unequal protection to layered video by exploring light-weight feedback. To deal with receiver heterogeneity, we propose a framework based on Nash bargaining game for operating point selection in a multicast group. Our simulation results show that the layered hybrid ARQ scheme significantly outperforms the conventional hybrid ARQ scheme with single layer video and the layered FEC scheme. The results also show that the game-based operating point selection provides a high overall system performance while facilitating fairness among receivers. We carefully examine the overhead and the computational complexity of our proposed schemes, through both theoretical analysis and OPNET simulation/experiment, and show that Layer Bargaining is practically feasible under representative settings.
Zhengye Liu, Pei Liu 0001, Hang Liu 0003, Yao Wang 0001
IEEE J. Sel. Areas Commun.5
2010 Dynamic Rate and FEC Adaptation for Video Multicast in Multi-rate Wireless Networks
Özgü Alay, Thanasis Korakis, Yao Wang 0001, Shivendra S. Panwar
Mob. Networks Appl.3
2010 Layered Wireless Video Multicast Using Relays
abstract
Wireless video multicast enables delivery of popular events to many mobile users in a bandwidth efficient manner. However, providing good and stable video quality to a large number of users with varying channel conditions remains elusive. In this paper, an integration of layered video coding, packet level forward error correction, and two-hop relaying is proposed to enable efficient and robust video multicast in infrastructure-based wireless networks. First, transmission with conventional omni-directional antennas is considered where relays have to transmit in non-overlapping time slots in order to avoid collision. In order to improve system efficiency, we next investigate a system in which relays transmit simultaneously using directional antennas. In both systems, we consider a non-layered configuration, where the relays forward all received video packets and all users receive the same video quality, as well as a layered setup, where the relays forward only the base-layer video. For each system setup, we consider optimization of the relay placement, user partition, transmission rates of each hop, and time scheduling between source and relay transmissions. Our analysis shows that the non-layered system can provide better video quality to all users than the conventional direct transmission system, and the layered system enables some users to enjoy significantly better quality, while guaranteeing other users the same or better quality than direct transmission. The directional relay system can provide substantial improvements over the omni-directional relay system. To support our results, a prototype is implemented using open source drivers and socket programming, and the system performance is validated with real-world experiments.
Özgü Alay, Thanasis Korakis, Yao Wang 0001, Elza Erkip, Shivendra S. Panwar
IEEE Trans. Circuits Syst. Video Technol.3
2009 An Experimental Study of Packet Loss and Forward Error Correction in Video Multicast over IEEE 802.11b Network
abstract
Video multicast over wireless local area networks (WLANs) faces many challenges due to varying channel conditions and limited bandwidth. A promising solution to this problem is the use of packet level forward error correction (FEC) mechanisms. However, the adjustment of the FEC rate is not a trivial issue due to the dynamic wireless environment. This decision becomes more complicated if we consider the multi-rate capability of the existing wireless LAN technology that adjusts the transmission rates based on the channel conditions and the coverage range. In order to explore the above issues we conducted an experimental study of the packet loss behavior of the IEEE 802.11b protocol. In our experiments we considered different transmission rates under the broadcast mode in indoor and outdoor environments. We further explored the effectiveness of packet level FEC for video multicast over wireless networks with multi-rate capability. In order to evaluate the system quantitatively, we implemented a prototype using open source drivers and socket programming. Based on the experimental results, we provide guidelines on how to efficiently use FEC for wireless video multicast in order to improve the overall system performance. We show that the Packet Error Rate (PER) increases exponentially with distance and using a higher transmission rate together with stronger FEC is more efficient than using a lower transmission rate with weaker FEC for video multicast.
Özgü Alay, Thanasis Korakis, Yao Wang 0001, Shivendra S. Panwar
CCNC3
2009 Is Physical Layer Error Correction Sufficient for Video Multicast over IEEE 802.11g Networks?
abstract
Wireless video multicast enables delivery of popular events to many mobile users in a bandwidth efficient manner. However, providing good and stable video quality to a large number of users with varying channel conditions remains elusive. A promising solution to this problem is the use of packet level (FEC) mechanisms. However, the adjustment of the FEC rate is not a trivial issue due to the dynamic wireless environment. This decision becomes more complicated if we consider the multi-rate capability of the existing wireless LAN technology that adjusts the transmission rates based on the channel conditions and the coverage range. In this paper, we explore the dynamics of Forward Error Correction (FEC) schemes in multi-rate wireless local area networks. We study the fundamental behavior of a 802.11g network which already has embedded error correction in physical layer, under unicast and broadcast modes in a real outdoor environment. We then explore the effectiveness of packet level FEC over wireless networks with multi-rate capability. In order to evaluate the system quantitatively, we implemented a prototype using open source drivers, and ran experiments. Based on the experimental results, we provide guidelines on how to efficiently use FEC for wireless multicast services in order to improve the overall system performance. We argue that even there is a physical layer error correction, using a higher transmission rate together with stronger FEC is more efficient than using a lower transmission rate with weaker FEC for multicast.
Özgü Alay, Thanasis Korakis, Yao Wang 0001, Shivendra S. Panwar
CCNC3
2009 Implementing a cooperative MAC protocol for wireless video multicast
abstract
Wireless video multicast enables delivery of popular events to many wireless users in a bandwidth efficient manner. However, providing good and stable video quality to a large number of users with varying channel conditions remains elusive. In our previous work, we integrated layered video coding with cooperative communication to enable efficient and robust video multicast in infrastructure-based wireless networks. Through simulation and analysis, we showed that cooperative multicast improves the multicast system performance and the coverage area. In this work, we integrate the proposed system with packet level forward error correction (FEC) and evaluate the viability of the system in a realistic environment. We implement the system at the MAC layer and report the experimental results in a medium size (i.e., 8 stations) testbed. The experimental results confirm that the new cooperative MAC protocol for multicast, delivers superior performance.
Özgü Alay, Thanasis Korakis, Yao Wang 0001, Shivendra S. Panwar
WCNC4
2009 Image and Video Denoising Using Adaptive Dual-Tree Discrete Wavelet Packets
abstract
We investigate image and video denoising using adaptive dual-tree discrete wavelet packets (ADDWP), which is extended from the dual-tree discrete wavelet transform (DDWT). With ADDWP, DDWT subbands are further decomposed into wavelet packets with anisotropic decomposition, so that the resulting wavelets have elongated support regions and more orientations than DDWT wavelets. To determine the decomposition structure, we develop a greedy basis selection algorithm for ADDWP, which has significantly lower computational complexity than a previously developed optimal basis selection algorithm, with only slight performance loss. For denoising the ADDWP coefficients, a statistical model is used to exploit the dependency between the real and imaginary parts of the coefficients. The proposed denoising scheme gives better performance than several state-of-the-art DDWT-based schemes for images with rich directional features. Moreover, our scheme shows promising results without using motion estimation in video denoising. The visual quality of images and videos denoised by the proposed scheme is also superior.
Jing-Yu Yang 0002, Yao Wang 0001, Wenli Xu, Qionghai Dai
IEEE Trans. Circuits Syst. Video Technol.2
2009 LayerP2P: Using Layered Video Chunks in P2P Live Streaming
abstract
Although there are several successful commercial deployments of live P2P streaming systems, the current designs 1) lack incentives for users to contribute bandwidth resources, 2) lack adaptation to aggregate bandwidth availability, and 3) exhibit poor video quality when bandwidth availability falls below bandwidth supply. In this paper, we propose, prototype, deploy, and validateLayerP2P, a P2P live streaming system that addresses all three of these problems. LayerP2P combines layered video, mesh P2P distribution, and a tit-for-tat-like algorithm, in a manner such that a peer contributing more upload bandwidth receives more layers and consequently better video quality. We implement LayerP2P (including seeds, clients, trackers, and layered codecs), deploy the prototype in PlanetLab, and perform extensive experiments. We also examine a wide range of scenarios using trace-driven simulations. The results show that LayerP2P has high efficiency, provides differentiated service, adapts to bandwidth deficient scenarios, and provides protection against free-riders.
Zhengye Liu, Yanming Shen, Keith W. Ross, Shivendra S. Panwar, Yao Wang 0001
IEEE Trans. Multim.5
2008 Layered wireless video multicast using omni-directional relays
abstract
Wireless video multicast enables delivery of popular events to many wireless users in a bandwidth efficient manner. However, providing good and stable video quality to a large number of users with varying channel conditions remains elusive. We propose to integrate layered video coding with cooperative communication to enable efficient and robust video multicast in infrastructure-based wireless networks. We determine the user partition and transmission time scheduling that can optimize a multicast performance criterion.
Özgü Alay, Thanasis Korakis, Yao Wang 0001, Elza Erkip, Shivendra S. Panwar
ICASSP3
2008 Complexity modeling of scalable video decoding
abstract
This paper addresses the computational complexity of scalable video decoding using emerging scalable extension of H.-264/AVC (SVC) standard compliant decoder. Scalable functionalities provided by SVC standard encompass temporal, spatial, quality enhancements and their combinations. The complexity model for decoding a bit stream with only temporal, spatial, or quality scalability are developed first. We then extend to a more general model for decoding a bit stream with arbitrarily combined scalability. Comparison with the number of clock cycles used in SVC decoding on a PC shows that the proposed model is very accurate.
Zhan Ma 0001, Yao Wang 0001
ICASSP2
2008 Layered wireless video multicast using directional relays
abstract
In this paper, we explore the use of directional antennas in relay transmission to improve the performance of video multicast with omni-directional relays in infrastructure- based wireless networks. We describe the system setup with directional relays and determine the user partition along with transmission time scheduling that can optimize a multicast performance criterion. We demonstrate that directional relays significantly improve the multicast system performance compared to omni-directional relays. Furthermore, it also provides larger coverage area.
Özgü Alay, Thanasis Korakis, Yao Wang 0001, Shivendra S. Panwar
ICIP3
2008 Saliency based objective quality assessment of decoded video affected by packet losses
abstract
In this work, we propose a novel saliency-based objective quality assessment metric, for assessing the perceptual quality of decoded video sequences affected by packet loses. The proposed method weights the error at each pixel by the visual saliency of the pixel. Different weighting methods are explored and compared. Our test results show that the predicted scores by the proposed metrics correlate very well with mean subjective scores, significantly better than the mean square error (MSE), mean absolute difference (MAD) or structure similarity (SSIM).
Dan Yang 0001, Yao Wang 0001
ICIP4
2008 Complexity modeling of H.264 entropy decoding
abstract
This paper models the complexity consumption of entropy decoding in a H.264 decoder, which includes many modes such as fix-length code (FLC), variable-length code (Exp-Golomb code in particular), content-adaptive variable length code (CAVLC) and content-adaptive binary arithmetic code (CABAC). By analyzing the computations involved in different entropy decoding modes, we derive a complexity model that relates the total complexity with the total bit rate and slice mode (I or P or B). The model is shown to be very accurate when compared with the measured complexity in a PC running the H.264 decoder.
Zhan Ma 0001, Zhongbo Zhang, Yao Wang 0001
ICIP3
2008 Modeling the impact of frame rate on perceptual quality of video
abstract
This study aims to understand how perceived quality of a video varies as the frame rate changes. We find that an inverted falling exponential function can accurately reflect the trend observed from subjective testing. We further explore the relationship between the falling rate and the video content features, such as frame difference, motion, and contrast.
Yen-Fu Ou, Zhan Ma 0001, Yao Wang 0001
ICIP5
2008 2-D anisotropic dual-tree complex wavelet packets and its application to image denoising
abstract
In this paper, we extend the 2-D dual-tree complex wavelet transform (DTCWT) to an adaptive anisotropic dual-tree complex wavelet packets (ADTCWP). The DTCWT subbands are iteratively decomposed into anisotropic complex wavelet packets, generating anisotropic wavelets and increasing the number of wavelets orientations without introducing extra redundancy. Then a basis selection procedure is applied so that the selected complex wavelet packets are well adapted to image characteristics. The effectiveness of ADTCWP is examined in image denoising with a bivariate statistical model. The ADTCWP-based denoising scheme shows better denoising results than several DTCWT-based methods for image with rich directional features. The denoised images recovered by the proposed scheme are visually more appealing.
Jing-Yu Yang 0002, Wenli Xu, Yao Wang 0001, Qionghai Dai
ICIP3
2008 Substream Trading: Towards an open P2P live streaming system
abstract
We consider the design of an open P2P live-video streaming system. When designing a live video system that is both open and P2P, the system must include mechanisms that incentivize peers to contribute upload capacity. We advocate an incentive principle for live P2P streaming: a peer’s video quality is commensurate with its upload rate. We propose Substream Trading, a new P2P streaming design which not only enables differentiated video quality commensurate with a peer’s upload contribution but can also accommodate different video coding schemes, including single-layer coding, layered coding, and multiple description coding. Extensive trace-driven simulations show that substream trading has high efficiency, provides differentiated service, low start-up latency, synergies among peers with different Internet access rates, and protection against free-riders.
Zhengye Liu, Yanming Shen, Keith W. Ross, Shivendra S. Panwar, Yao Wang 0001
ICNP5
2008 Image Coding Using Dual-Tree Discrete Wavelet Transform
abstract
In this paper, we explore the application of 2-D dual-tree discrete wavelet transform (DDWT), which is a directional and redundant transform, for image coding. Three methods for sparsifying DDWT coefficients, i.e., matching pursuit, basis pursuit, and noise shaping, are compared. We found that noise shaping achieves the best nonlinear approximation efficiency with the lowest computational complexity. The interscale, intersubband, and intrasubband dependency among the DDWT coefficients are analyzed. Three subband coding methods, i.e., SPIHT, EBCOT, and TCE, are evaluated for coding DDWT coefficients. Experimental results show that TCE has the best performance. In spite of the redundancy of the transform, our DDWT _ TCE scheme outperforms JPEG2000 up to 0.70 dB at low bit rates and is comparable to JPEG2000 at high bit rates. The DDWT _TCE scheme also outperforms two other image coders that are based on directional filter banks. To further improve coding efficiency, we extend the DDWT to an anisotropic dual-tree discrete wavelet packets (ADDWP), which incorporates adaptive and anisotropic decomposition into DDWT. The ADDWP subbands are coded with TCE coder. Experimental results show that ADDWP _ TCE provides up to 1.47 dB improvement over the DDWT _TCE scheme, outperforming JPEG2000 up to 2.00 dB. Reconstructed images of our coding schemes are visually more appealing compared with DWT-based coding schemes thanks to the directionality of wavelets.
Jing-Yu Yang 0002, Yao Wang 0001, Wenli Xu, Qionghai Dai
IEEE Trans. Image Process.2
2007 Subjective Quality Evaluation of Decoded Video in the Presence of Packet Losses
abstract
This paper investigates the perceptual quality of decoded video bitstreams after packet losses. We focus on low-resolution and low-bit rate video coded by the H.264/AVC encoder, and packet loss patterns likely in 3G wireless networks. We examine the impact of several factors on the perceptual quality, including the error length (the error propagation duration after a loss), the loss severity (measured by the drop in PSNR due to a loss), and loss location (forgiveness effect). We further propose an objective quality measure based on the findings of our study.
Yao Wang 0001, Jill M. Boyce
ICASSP (1)2
2007 Image Coding using 2-D Anisotropic Dual-Tree Discrete Wavelet Transform
abstract
We propose an image coding scheme using 2-D anisotropic dual-tree discrete wavelet transform (DDWT). First, we extend 2-D DDWT to anisotropic decomposition, and obtain more directional subbands. Second, an iterative projection-based noise shaping algorithm is employed to further sparsify anisotropic DDWT coefficients. At last, the resulting coefficients are rearranged to preserve zero-tree relationship so that they can be efficiently coded with SPIHT. Experimental results show that our proposed scheme outperforms JPEG2000 and SPIHT at low bit rates despite the redundancy of DDWT.
Jing-Yu Yang 0002, Jizheng Xu, Feng Wu 0001, Qionghai Dai, Yao Wang 0001
ICIP (3)5
2007 P2P Video Live Streaming with MDC: Providing Incentives for Redistribution
abstract
In this paper, we consider applying multiple description coding in mesh-pull P2P live streaming networks to provide incentives for redistribution. In our system, a video is encoded into multiple descriptions with each description having equal importance. We consider a heterogeneous system with peers having different uplink bandwidths. We design a distributed protocol in which a peer contributing more uplink bandwidth receives more descriptions and consequently better video quality. Previous approaches consider single-layer video, where each peer receives the same video quality no matter how much bandwidth it contributes to the system. The simulation results show that our approach can provide differentiated video quality commensurate with a peer's contribution to other peers.
Zhengye Liu, Yanming Shen, Shivendra S. Panwar, Keith W. Ross, Yao Wang 0001
ICME5
2007 Video Coding using 3-D Anisotropic Dual-Tree Wavelet Transform
abstract
This paper investigates the use of the anisotropic 3-D dual-tree discrete wavelet transform (DDWT) for video coding. The 3-D DDWT is an attractive video representation because it isolates image patterns with different spatial orientations and motion directions and speeds in separate subbands. Our previous codecs using the 3-D isotropic DDWT provides better performance than the 3-D SPIHT codec on the 3-D DWT. In this paper, we explore the use of anisotropic DDWT (ADDWT) for video coding. The ADDWT extends the superiority of the normal DDWT with more directional subbands without adding to the redundancy. The proposed codec applies SPIHT to each of the ADDWT trees. This codec provides significantly better performance than the 3-D SPIHT codec using the standard DWT and the DDWT both objectively and subjectively. None of these video codecs requires motion compensation.
Jing-Yu Yang 0002, Beibei Wang 0005, Yao Wang 0001, Wenli Xu
ICME3
2007 Image Compression using 2D Dual-tree Discrete Wavelet Transform (DDWT)
abstract
In this paper, we investigate image compression using 2D dual-tree discrete wavelet transform (DDWT), which is an overcomplete transform with direction-selective basis functions. To further sparsify DDWT coefficients, an iterative projection-based noise shaping method is employed. We analyze the statistics of DDWT coefficients as well as the inter-scale, inter-subband, and intra-subband dependency among the DDWT coefficients. We further evaluate the application of SPIHT and EBCOT for coding DDWT coefficients. Experimental results show that SPIHT is more effective than EBCOT for DDWT, and the DDWT-SPIHT coder outperforms JPEG2000 at low bit rates and is comparable to JPEG2000 at high bit rates.
Jing-Yu Yang 0002, Wenli Xu, Qionghai Dai, Yao Wang 0001
ISCAS4
2007 A comparative study of image compression based on directional wavelets
abstract
Discrete wavelet transform is an effective tool to generate scalable stream, but it cannot efficiently represent edges which are not aligned in horizontal or vertical directions, while natural images often contain rich edges and textures of this kind. Hence, recently, intensive research has been focused particularly on the directional wavelets which can effectively represent directional attributes of images. Specifically, there are two categories of directional wavelets: redundant wavelets (RW) and adaptive directional wavelets (ADW). One representative redundant wavelet is the dual-tree discrete wavelet transform (DDWT), while adaptive directional wavelets can be further categorized into two types: with or without side information. In this paper, we briefly introduce directional wavelets and compare their directional bases and image compression performances.
Kun Li 0001, Wenli Xu, Qionghai Dai, Yao Wang 0001
VCIP4
2007 Total Power Minimization for Multiuser Video Communications Over CDMA Networks
abstract
In this work, we consider a CDMA cell with multiple terminals transmitting video signals. We adapt the system parameters to minimize the sum of compression powers and transmitter powers of all users while guaranteeing the received video quality at each terminal. The adjustable parameters at user i include the transmitter power Pt,i, the video coding bit rate Rs,i, and video encoder parameters that control the complexity and hence power consumption of the video coder (referred simply as complexity betai). Instead of determining Pt,idirectly, we first determine the desired signal to interference-noise ratio (SINR) gammai. Based on the optimal gammaiand Rs,i, we then determine Pt,i. Our analysis shows that the product of Rs,iand gammaiis an important quantity. Given the complexity betai(i.e., given the compression power) and quality constraint, in order to reduce the transmission power, one should choose Rs,iand gammaito minimize their product. When only the total transmission power is concerned, the optimal operating points can be determined at individual users separately: each user should run the encoder to minimize the product of Rs,iand gammai. When the objective is to minimize the sum of compression and transmission powers of all users, the optimal solution can be found in two steps. The first step searches the optimal Rs,iand gammaithat minimize Rs,itimesgammaifor each video category and each possible betaiwhile satisfying the quality constraint at user i. The second step searches the optimal {betai}i=1,...,Nfor all users jointly, that minimizes the sum of transmission and compression powers of all users. The first step can be completed offline in advance, only the second step needs to be computed in real time based on channel conditions of the users. Our results indicate that for the same class of video users, the one who is closer to the base station compresses at a lower complexity. Simulation results show that significant power savings are obtained by our adaptive algorithms over nonadaptive approaches, where {Rs,i,betai,gammai} are fixed regardless the channel conditions
Xiaoan Lu, Yao Wang 0001, Elza Erkip, David J. Goodman
IEEE Trans. Circuits Syst. Video Technol.2
2007 Major Cast Detection in Video Using Both Speaker and Face Information
abstract
Major casts, for example, the anchor persons or reporters in news broadcast programs and the principle characters in movies, play an important role in video, and their occurrences provide meaningful indices for organizing and presenting video content. This paper describes a new approach for automatically generating a list of major casts in a video sequence based on multiple modalities, specifically, speaker information in audio track and face information in video track. The core algorithm is composed of three steps. First, speaker boundaries are detected and speaker segments are clustered in audio stream. Second, face appearances are tracked and face tracks are clustered in video stream. Finally, correspondences between speakers and faces are determined based on their temporal co-occurrence. A list of major casts is constructed and ranked in an order that reflects each cast's importance, which is determined by the accumulative temporal and spatial presence of the cast. The proposed algorithm has been integrated in a major cast based video browsing system, which presents the face icon and marks the speech locations in time stream for each detected major cast. The system provides a semantically meaningful summary of the video content, which helps the user to effectively digest the theme of the video
Zhu Liu 0001, Yao Wang 0001
IEEE Trans. Multim.2
2006 Cooperative Source and Channel Coding for Wireless Video Transmission
abstract
Past work on cooperative communications has indicated substantial improvement in channel reliability through cooperative transmission strategies. To exploit this benefit for video transmission, we propose to jointly allocate bits among video coding, channel coding and cooperation to optimize the decoded video quality. Recognizing that not all source bits are equal, we further propose to protect the more important bits through user cooperation. We simulate and compare four modes of video transmission that differ in their error protection strategy (equal vs. layered, with vs. without cooperation). Our simulation uses the H.263+ video codec and the RCPC channel code over quasi-static Rayleigh fading channels. We show that cooperation can provide significant improvement over no-cooperation when the average channel SNR is in the low to medium range, and that layered cooperation can extend this benefit to the entire range of channel quality.
Hoi Yin Shutoy, Yao Wang 0001, Elza Erkip
ICIP2
2006 A Two Pass H.264-Based Matching Pursuit Video Coder
abstract
This paper proposes a matching pursuit video codec built on top of the H.264 framework. H.264/MPEG4 AVC is currently the most efficient video coding standard and it applies block-based motion compensation and DCT-like transform as all other standards. In order to further improve the coding efficiency and enable bit stream scalability, we explore the use of redundant frame-based Gabor transforms for coding the prediction error. The motion compensation and mode decision parts are compatible to H.264. The overlapped block motion compensation (OBMC) method can be optionally invoked to smooth the predicted images. Because H.264/AVC incorporates both inter and intra prediction, and allows many different block sizes, it is more challenging to apply matching pursuit on H.264/AVC than on prior standards. We investigate how to apply MP to prediction error images containing both inter- and intra-predicted blocks. We also investigate how to adapt the quantizer based on the H.264 QP. The preliminary simulation results show that the proposed codec outperforms H.264/AVC at the low to medium bit rate range, both objectively and subjectively, while generating a scalable bit stream.
Beibei Wang 0005, Yao Wang 0001, Peng Yin 0002
ICIP2
2006 Cross Layer Adaptation for H.264 Video Multicasting Over Wireless Lan
abstract
This paper describes cross-layer optimization strategies and simulation results for H.264 videomulticast over wireless LAN. The proposed scheme takes into account the varying channe conditions of multiple users, and dynamically allocates available bandwidth between source coding and channel coding. In particular, source coding parameters (intra update and quantization) and application-layer FEC code rate are chosen jointly to optimize a multicast performance criterion, based on feedbacks from all multicast receivers. Two performance criteria for video multicast are investigated and compared.
Zhengye Liu, Hang Liu 0003, Yao Wang 0001
ICME3
2006 On the Design of Prefetching Strategies in a Peer-Driven Video on-Demand System
abstract
In this paper, we examine the prefetching strategies in a peer-driven video on-demand system. In our design, each video is encoded into multiple low bit-rate substreams and copies of the substreams are distributed to the participating peers. When a peer streams in a substream of rate r, it instead streams at rate rcirc, where r>rcirc. In this manner, if one of the peer's suppliers disconnects, the client peer can tap the reservoir of prefetched bits while searching for a replacement server, thereby avoiding any glitches or reduced visual quality. We examine how to assign prefetching rates to each of substreams as a function of their importance. Our studies show that appropriate prefetching strategies can bring significant performance improvements for both multiple description and layered videos
Yanming Shen, Zhengye Liu, Shivendra S. Panwar, Keith W. Ross, Yao Wang 0001
ICME5
2006 Modeling of transmission-loss-induced distortion in decoded video
abstract
This paper analyzes the distortion in decoded video caused by random packet losses in the underlying transmission network. A recursion model is derived that relates the average channel-induced distortion in successive P-frames. The model is applicable to all video encoders using the block-based motion-compensated prediction framework (including the H.261/263/264 and MPEG1/2/4 video coding standards) and allows for any motion-compensated temporal concealment method at the decoder. The model explicitly considers the interpolation operation invoked for motion-compensated temporal prediction and concealment with sub-pel motion vectors. The model also takes into account the two new features of the H.264/AVC standard, namely intraprediction and inloop deblocking filtering. A comparison with simulation data shows that the model is very accurate over a large range of packet loss rates and encoder intrablock rates. The model is further adapted to characterize the channel distortion in subsequent received frames after a single lost frame. This allows one to easily evaluate the impact of a single frame loss.
Yao Wang 0001, Jill M. Boyce
IEEE Trans. Circuits Syst. Video Technol.1
2005 Minimize the total power consumption for multiuser video transmission over CDMA wireless network: a two-step approach
abstract
We consider a CDMA cell with multiple terminals transmitting video signals. We minimize the sum of signal processing and transmitter power while the received quality at each terminal is guaranteed. The system parameters to be adjusted include video coding bit rate, video compression complexity and transmitter power. Instead of full search in the space of {bit rate, complexity, transmitter power} for all users, we design a two-step fast algorithm to reduce the computation burden in the base station. In our algorithm, the search in the base station is over the space of complexity only. Our results indicate that, for the same class of video users, the one who is closest to the base station compresses at least complexity. This is used to further reduce the computation required by our algorithm.
Xiaoan Lu, Yao Wang 0001, Elza Erkip, David J. Goodman
ICASSP (3)2
2005 Video Coding Using 3-D Dual-Tree Discrete Wavelet Transforms
abstract
The paper explores the use of a recently introduced 3D dual-tree discrete wavelet transform (DDWT) for video coding. The 3D DDWT is an attractive video representation because it isolates motion along different directions in separate subbands. However, it is an overcomplete transform with 8:1 or 4:1 redundancy. Based on the effectiveness of the iterative projection-based noise shaping scheme proposed by Kingsbury in reducing the number of coefficients, and our prior investigation about the correlation between subbands at the same spatial/temporal location, both in the significance map and in actual coefficient values, a new video coding scheme using 3D DDWT is proposed. The proposed video codec does not require motion compensation and provides better performance than the 3D SPIHT codec, both objectively and subjectively, despite the fact that the raw number of coefficients resulting from the 3D DDWT is much more than that of the conventional 3D DWT. The proposed coder allows full scalability in spatial, temporal and quality dimensions.
Beibei Wang 0005, Yao Wang 0001, Ivan W. Selesnick, Anthony Vetro
ICASSP (2)2
2005 Layered cooperative source and channel coding
abstract
Cooperative techniques form a new wireless communication paradigm in which terminals help each other in relaying information to combat the random fading and to provide diversity in radio channels. Past work has focused on improving channel reliability through cooperation. We propose to jointly allocate bits among source coding, channel coding and cooperation to minimize the expected source distortion. Recognizing that not all source bits are equal, we further propose to protect the more important bits through user cooperation. To evaluate the gain of layered cooperation, we simulate four modes of communications that differ in their error protection strategy (equal vs. layered, with vs. without cooperation) with a practical channel coder, and show that, for i.i.d. Gaussian sources, layered cooperation can achieve significant performance gains over non-layered/non-cooperative communication. We also carry out an information theoretic analysis illustrating fundamental benefits of layered cooperation.
Deniz Gündüz, Elza Erkip, Yao Wang 0001
ICC4
2005 Streaming layered encoded video using peers
abstract
Peer-to-peer video streaming has emerged as an important means to transport stored video. The peers are less costly and more scalable than an infrastructure-based video streaming network which deploys a dedicated set of servers to store and distribute videos to clients. In this paper, we investigate streaming layered encoded video using peers. Each video is encoded into hierarchical layers which are stored on different peers. The system serves a client request by streaming multiple layers of the requested video from separate peers. The system provides unequal error protection for different layers by varying the number of copies stored for each layer according to its importance. We evaluate the performance of our proposed system with different copy number allocation schemes through extensive simulations. Finally, we compare the performance of layered coding with multiple description coding.
Yanming Shen, Zhengye Liu, Shivendra S. Panwar, Keith W. Ross, Yao Wang 0001
ICME5
2005 Modelling of distortion caused by packet losses in video transport
abstract
This paper analyzes transmission-error induced distortion in decoded video. A recursion model is derived that relates the distortion in successive P-frames. The model takes into account of non-integer motion vectors used for motion-compensated temporal prediction and concealment, unconstrained intra prediction, and in-loop deblocking filtering. Experimental data show that the model is quite accurate over a large range of packet loss rates and encoder intra rates.
Yao Wang 0001, Jill M. Boyce, Xiaoan Lu
ICME1
2005 Multiple Description Coding for Video Delivery
abstract
Multiple description coding (MDC) is an effective means to combat bursty packet losses in the Internet and wireless networks. MDC is especially promising for video applications where retransmission is unacceptable or infeasible. When combined with multiple path transport (MPT), MDC enables traffic dispersion and hence reduces network congestion. This work describes principles in designing MD video coders employing temporal prediction and presents several predictor structures that differ in their tradeoffs between mismatch-induced distortion and coding efficiency. The paper also discusses example video communication systems integrating MDC and MPT.
Yao Wang 0001, Amy R. Reibman, Shunan Lin
Proc. IEEE1
2005 Joint scene classification and segmentation based on hidden Markov model
abstract
Scene classification and segmentation are fundamental steps for efficient accessing, retrieving and browsing large amount of video data. We have developed a scene classification scheme using a Hidden Markov Model (HMM)-based classifier. By utilizing the temporal behaviors of different scene classes, HMM classifier can effectively classify presegmented clips into one of the predefined scene classes. In this paper, we describe three approaches for joint classification and segmentation based on HMM, which search for the most likely class transition path by using the dynamic programming technique. All these approaches utilize audio and visual information simultaneously. The first two approaches search optimal scene class transition based on the likelihood values computed for short video segment belonging to a particular class but with different search constrains. The third approach searches the optimal path in a super HMM by concatenating HMM's for different scene classes.
Jincheng Huang 0001, Zhu Liu 0001, Yao Wang 0001
IEEE Trans. Multim.3
2004 Complexity-bounded power control in video transmission over a CDMA wireless network
abstract
In this work, we consider a CDMA cell with multiple terminals transmitting video signals. The concept of a utility function is used to maximize the number of received picture frames with adequate quality per Joule of energy. For a reconstructed signal at the video decoder, the quality is controlled by the encoded bit rate, compression complexity as well as received signal-to-interference-noise ratio (SINR). In this work, video quality is measured in peak signal-to-noise ratio (PSNR) rather than SINR. We find that for a given compression complexity, maximum utility is achieved when the product of bit rate and required SINR is minimized. This maximum usually occurs at maximum video coding complexity. We also investigate the capacity in terms of the number of users that can be supported simultaneously for this system, and how the total utility varies with the number of users in this system. This can be used as an admission policy by the central station.
Xiaoan Lu, David J. Goodman, Yao Wang 0001, Elza Erkip
GLOBECOM3
2004 Optimal coding rate and power allocation for the streaming of scalably encoded video over a wireless link
abstract
Scalably encoded information results in files which can be truncated at an arbitrary point and decoded, as supported by the JPEG-2000 (image) and MPEG-4 (video) standards. This work introduces a tractable, yet flexible analytical model for resource management involving scalably encoded video. Each segment of video of a predetermined length yields a file that can be truncated and decoded independently of other segments. The problem is set up as a joint optimization of transmission power, and coding rate (where to truncate?). The analysis reveals that any one of these variables uniquely determines the other. The terminal should truncate the file at the point that maximizes quality per unit of power employed.
Virgilio Rodriguez, David J. Goodman, Yao Wang 0001
ICASSP (5)3
2004 Fine-granular-scalability video streaming over wireless LANs using cross layer error control
abstract
Real-time streaming of audiovisual content over wireless LANs (WLANs) is emerging as an important technology area in multimedia communications. Due to the error-prone, time-varying characteristics of wireless channels, there is a need for strong protection of video bitstreams. To cope with the variation of channel conditions, we introduce a novel cross-layer protection strategy that combines adaptive application-layer forward error correction (FEC) and physical-layer modulation with fine-granular-scalability (FGS) coding to improve the robustness of wireless transmission. Unlike data streams, different parts of a video stream have different priorities and hence merit the use of unequal error protection. Our schemes can dynamically select the optimal combination of FEC and modulation based on channel conditions and video data content. Experimental results show the stability of the scheme over a variety of channel conditions.
Mihaela van der Schaar, Santhana Krishnamachari, Sunghyun Choi 0001, Yao Wang 0001
ICASSP (5)5
2004 Power optimization of source encoding and radio transmission in multiuser CDMA systems
abstract
We investigate the power consumed by two mobile terminals transmitting compressed source signals to a base station in a CDMA cellular system. The aim is to minimize total power consumption at the terminals by simultaneously adjusting the complexity of source compression and the transmitter power. We find that in general the complexity of source compression should increase with increasing distance between a terminal and the base station. However, the exact configuration of compression and transmitter power depends on the interaction of the two terminals. The optimum operating points of the two terminals are contained in a pair of nonlinear equations. We use the example of a transform coder processing signals from a Gauss-Markov source to explore the advantages of an adaptive system over a system with fixed compression. The analysis can be applied to a variety of source coders and be extended to a system with more than two users per cell.
Xiaoan Lu, Yao Wang 0001, Elza Erkip, David J. Goodman
ICC2
2004 An investigation of 3D dual-tree wavelet transform for video coding
abstract
This paper examines the properties of a recently introduced 3D dual-tree discrete wavelet transform (DDWT) for video coding. The 3D DDWT is an attractive video representation because it isolates motion along different directions in separate subbands. However, it is an overcomplete transform with 8:1 redundancy. We examine the effectiveness of the iterative projection-based noise shaping scheme proposed by Kingsbury et al., (2002) on reducing the number of coefficients. We also investigate the correlation between subbands at the same spatial/temporal location, both in the significance map and in actual coefficient values.
Beibei Wang 0005, Yao Wang 0001, Ivan W. Selesnick, Anthony Vetro
ICIP2
2004 A peer-to-peer video-on-deniand system using multiple description coding and server diversity
Yao Wang 0001, Shivendra S. Panwar, Keith W. Ross
ICIP2
2004 Joint PHY and MAC layer power optimization for video transmission over wireless LAN
abstract
In this paper, we consider the problem of transmitting compressed video over wirless LAN under transmission power constraints, where retransmission is adopted as the error control scheme. For different retry limits, transmitter energy per bit is adjusted to keep a constant video quality at the receiver, which costs different transmission power. We propose an algorithm to minimize transmission power by carefully choosing the maximum number of retransmission times and transmission energy level based on the quality and delay requirement. We also examine the effect of distance on the choice of the optimal points.
Xiaoan Lu, Yingwei Chen, Yao Wang 0001
VCIP3
2003 Rate-distortion analysis of the multiple description motion compensation video coding scheme
abstract
Multiple description motion compensation (MDMC) is a multiple description video coding scheme that has shown good error resilience performance. MDMC enables one to vary coding parameters according to the desired trade-off between coding efficiency and error resilience. To fully utilize this advantage, one needs to establish a set of models, relating the rate, encoder distortion, and the end-to-end distortion after transmission, with the encoder parameters and channel parameters. Using these models, one can find the optimal coding parameters for given channel parameters and rate (or distortion) constraints. In this paper, we formulate and validate the rate and encoder distortion models.
Shunan Lin, Anthony Vetro, Yao Wang 0001
ICASSP (3)3
2003 Rate-distortion analysis of the multiple description motion compensation video coding scheme
abstract
Multiple description motion compensation (MDMC) is a multiple description video coding scheme that has shown good error resilience performance. MDMC enables one to vary coding parameters according to the desired trade-off between coding efficiency and error resilience. To fully utilize this advantage, one needs to establish a set of models, relating the rate, encoder distortion, and the end-to-end distortion after transmission, with the encoder parameters and channel parameters. Using these models, one can find the optimal coding parameters for given channel parameters and rate (or distortion) constraints. In this paper, we formulate and validate the rate and encoder distortion models.
Shunan Lin, Anthony Vetro, Yao Wang 0001
ICME3
2003 Adaptive error control for fine-granular-scalability video coding over IEEE 802.11 wireless LANs
abstract
Robust streaming of video over the IEEE 802.11 wireless LANs (WLANs) poses many challenges, including coping with bandwidth variations and data losses. To address these challenges, we evaluate and compare various error control strategies in this paper, namely, medium access control (MAC) layer forward error correction (FEC), MAC retransmission, and application-layer FEC, combined with FGS scalable video coding to achieve reliable communication over the IEEE 802.11 WLANs. Moreover, we assess the performance of fine-grained loss protection (FGLP) for different multipath channel conditions. Furthermore, the performance of cross-layer strategies such as application-layer FEC with MAC retransmissions is also determined.
Mihaela van der Schaar, Santhana Krishnamachari, Sunghyun Choi 0001, Yao Wang 0001
ICME5
2003 Power efficient multimedia communication over wireless channels
abstract
In this work, we introduce an approach for minimizing the total power consumption of a mobile transmitter due to source compression, channel coding and transmission subject to a fixed end-to-end source distortion. We illustrate our approach both on an abstract class of sources and channels and on a realistic H.263 video transmission system through a wireless channel. Performance under different channel environments and implementation schemes are investigated. Our numerical analysis shows that optimized settings can reduce the total power consumption by a significant factor and prolong battery life considerably compared with fixed parameter settings.
Xiaoan Lu, Elza Erkip, Yao Wang 0001, David J. Goodman
IEEE J. Sel. Areas Commun.3
2003 Video transport over ad hoc networks: multistream coding with multipath transport
abstract
Enabling video transport over ad hoc networks is more challenging than over other wireless networks. The wireless links in an ad hoc network are highly error prone and can go down frequently because of node mobility, interference, channel fading, and the lack of infrastructure. However, the mesh topology of ad hoc networks implies that it is possible to establish multiple paths between a source and a destination. Indeed, multipath transport provides an extra degree of freedom in designing error resilient video coding and transport schemes. In this paper, we propose to combine multistream coding with multipath transport, to show that, in addition to traditional error control techniques, path diversity provides an effective means to combat transmission error in ad hoc networks. The schemes that we have examined are: 1) feedback based reference picture selection; 2) layered coding with selective automatic repeat request; and 3) multiple description motion compensation coding. All these techniques are based on the motion compensated prediction technique found in modern video coding standards. We studied the performance of these three schemes via extensive simulations using both Markov channel models and OPNET Modeler. To further validate the viability and performance advantages of these schemes, we implemented an ad hoc multiple path video streaming testbed using notebook computers and IEEE 802.11b cards. The results show that great improvement in video quality can be achieved over the standard schemes with limited additional cost. Each of these three video coding/transport techniques is best suited for a particular environment, depending on the availability of a feedback channel, the end-to-end delay constraint, and the error characteristics of the paths.
Shiwen Mao, Shunan Lin, Shivendra S. Panwar, Yao Wang 0001, Emre Celebi
IEEE J. Sel. Areas Commun.4
2003 Bit allocation for MPEG-4 video coding with spatio-temporal tradeoffs
abstract
This paper describes rate-control algorithms that consider the tradeoff between coded quality and temporal rate. We target improved coding efficiency for both frame-based and object-based video coding. We propose models that estimate the rate-distortion characteristics for coded frames and objects, as well as skipped frames and objects. Based on the proposed models, we propose three types of rate-control algorithms. The first is for frame-based coding, in which the distortion of coded frames is balanced with the distortion incurred by frame skipping. The second algorithm applies to object-based coding, where the temporal rate of all objects is constrained to be the same, but the bit allocation is performed at the object level. The third algorithm also targets object-based coding, but in contrast to the second algorithm, the temporal rates of each object may vary. The algorithm also takes into account the composition problem, which may cause holes in the reconstructed frame when objects are encoded at different temporal rates. We propose a solution to this problem that is based on first detecting changes in the shape boundaries over time at the encoder, then employing a hole detection and recovery algorithm at the decoder. Overall, the proposed algorithms are able to achieve the target bit rate, effectively code frames and objects with different temporal rates, and maintain a stable buffer level.
Anthony Vetro, Yao Wang 0001, Yo-Sung Ho
IEEE Trans. Circuits Syst. Video Technol.3
2003 Rate-distortion modeling for multiscale binary shape coding based on Markov random fields
abstract
The purpose of this paper it to explore the relationship between the rate-distortion characteristics of multiscale binary shape and Markov random field (MRF) parameters. For coding, it is important that the input parameters that will be used to define this relationship be able to distinguish between the same shape at different scales, as well as different shapes at the same scale. We consider an MRF model, referred to as the Chien model, which accounts for high-order spatial interactions among pixels. We propose to use the statistical moments of the Chien model as input to a neural network to accurately predict the rate and distortion of the binary shape when coded at various scales.
Anthony Vetro, Yao Wang 0001, Huifang Sun
IEEE Trans. Image Process.2
2002 A probabilistic approach for rate-distortion modeling of multiscale binary shape
abstract
The purpose of this paper it to explore the relationship between the rate-distortion (R-D) characteristics of multi scale binary shape and Markov Random Field (MRF) parameters. In our experiments, we consider two prior models. The first MRF model takes into account pair-wise interaction between pels, and for the binary case, is typically referred to as the auto-logistic model; the second MRF model accounts for higher order spatial interactions and is referred to as the Chien model. Experimental results indicate that the auto-logistic model is not sufficient to characterize the R-D characteristics of multi scale binary shape data. However, higher order models, such as the Chien model, do seem feasible. We propose to use the statistical moments of the Chien model as input to a neural network to accurately predict the rate and distortion of the binary shape when coded at various scales.
Anthony Vetro, Yao Wang 0001, Huifang Sun
ICASSP2
2002 MPEG-4 video object-based rate allocation with variable temporal rates
abstract
This paper describes a bit allocation algorithm to achieve a constant bit rate when coding multiple video objects (MVOs), while improving the rate-distortion (R-D) performance over the reference method for MPEG-4 object-based rate control. In object-based coding, bit allocation is performed at the object level and temporal rates of different objects may vary. In this paper, we deal with these two issues. We pay particular attention to maintenance of buffer occupancy levels and propose a new method for spatiotemporal trade-offs for object-based coding. The proposed algorithm is able to successfully achieve the target bit rate, effectively code arbitrarily-shaped MVOs with different temporal rates, and maintain a stable buffer level.
Anthony Vetro, Yao Wang 0001, Yo-Sung Ho
ICIP (3)3
2002 Analysis and improvement of multiple description motion compensation video coding for lossy packet networks
abstract
Multiple description motion compensation (MDMC) is a multiple description coding scheme which has shown good error resilience performance. This paper investigates the behavior of an MDMC video codec in a packet loss environment. A model for the decoder distortion induced by transmission errors is developed and their accuracy is validated by simulation studies. The model shows the necessity of coding mismatch signals in an MDMC codec. Also, a new mismatch signal coding method and the corresponding decoding strategy are presented which improves upon the original MDMC scheme in a packet lossy environment.
Shunan Lin, Yao Wang 0001
ICIP (2)2
2002 Error resilience property of multihypothesis motion-compensated prediction
abstract
The multihypothesis motion-compensated prediction (MHMCP) approach has shown significant gain in terms of coding efficiency both in theory and practice. However, the fact that it also enhances the error resilience of a codec has not been fully realized. This paper analyzes the error resilience gain of MHMCP. Models for the decoder distortion induced by transmission errors and the rate-distortion function of the encoder in a codec using MHMCP are presented and their accuracy is validated by simulation studies. Using these models, we show how to design the multihypothesis predictor to jointly consider the coding gain and the error resilience gain for given channel error characteristics.
Shunan Lin, Yao Wang 0001
ICIP (3)2
2002 Power efficient H.263 video transmission over wireless channels
abstract
We introduce an approach for adaptive minimization of the total power consumption of wireless video communications subject to a given level of quality of service. Our approach exploits tradeoffs between the power consumption of the H.263 encoder, the Reed-Solomon channel encoder and the transmitter. Simulation results show that source and channel coding parameters and transmit energy per bit should vary based on channel conditions. Optimized settings can reduce the total power consumption by a significant factor compared to fixed parameter settings which do not match with the channel conditions.
Xiaoan Lu, Yao Wang 0001, Elza Erkip
ICIP (1)2
2002 Wireless video transport using path diversity: multiple description vs layered coding
abstract
Typical video applications may need a higher bandwidth and/or higher reliability connection than that provided by a single link in current or emerging wireless networks. We propose to employ path diversity to provide higher bandwidth and more robust end-to-end connections than that affordable by a single path. Under this transport environment, two viable strategies for video coding are multiple description coding (MDC) and layered coding (LC). MDC is more effective when the underlying application has a very stringent delay constraint and the round trip time on each path is relatively long. LC can be a good alternative when limited retransmission of the base layer is acceptable and when it is feasible to apply unequal error protection over different paths. The paper describes the general issues involved in integrating MDC/LC with multiple path transport, and compares the performances of MDC and LC, under different path conditions.
Yao Wang 0001, Shivendra S. Panwar, Shunan Lin, Shiwen Mao
ICIP (1)1
2002 Constrained video object segmentation by color masks and MPEG-7 descriptors
abstract
We present an automatic and computationally conservative boundary extraction method using available priori information in a framework consists of change detection mask, region growing, and trajectory motion. Instead of segmenting a entire video frame, only the regions belong to a target object specified by a set of rules are detected. One example of such rules is skin color features of a human body part. The processing domain is limited to the pixels that satisfy color or geometric rules. These rules are represented as a detection mask and implemented as a look-up table. As a result, significant computational reduction and real-time performance are achieved. The framework utilizes color consistency within a centroid-linkage growing technique to grow initial regions. Region seeds are selected among the pixels in the detection mask. The similarity thresholds are adapted from the MPEG-7 dominant color descriptors. The segmentation results of a frame diffused to the next frame and region statistics such as trajectory, percentage of the changed pixels, etc., are registered to determine the moving regions. A computational load comparison of the constrained region growing and regular region growing shows significant reduction in the complexity.
Fatih Porikli, Yao Wang 0001
ICME (1)2
2002 Object-based rate allocation with spatio-temporal trade-offs
Anthony Vetro, Yao Wang 0001, Yo-Sung Ho
VCIP3
2002 Lapped orthogonal transforms designed for error-resilient image coding
abstract
This paper describes a new design method for lapped orthogonal transforms (LOTs) that can provide a desired tradeoff between coding efficiency and resilience to transmission errors. Traditionally, LOT bases have been designed solely to maximize coding efficiency. When certain coefficients are lost due to transmission channel impairments, the reconstructed image is often unsatisfactory. Previously, we have developed a maximally smooth recovery method for image reconstruction from incomplete LOT coefficients (see Chung and Wang, ibid., vol.9, p.895-908, 1999). The reconstruction quality depends on the LOT basis used. We describe a new LOT-basis design method, which maximizes the weighted average of a coding gain and a reconstruction gain, with the latter being defined according to the maximally smooth recovery method. A coder using the designed basis with a high weighting factor toward the reconstruction gain can achieve significantly better reconstruction quality than a LOT basis that is designed to optimize the coding efficiency only. The newly designed bases are evaluated by their redundancy-rate-distortion performance. Simulation results show that the new bases are more efficient than the bases designed previously by S.S. Hemami (see ibid., vol.6, p.168-81, 1996), in that the new bases require fewer redundancy bits to achieve the same reconstruction quality under the same channel error pattern.
Doo-Man Chung, Yao Wang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2002 Supporting image and video applications in a multihop radio environment using path diversity and multiple description coding
abstract
This paper examines the effectiveness of combining multiple description coding (MDC) and multiple path transport (MPT) for video and image transmission in a multihop mobile radio network. The video and image information is encoded nonhierarchically into multiple descriptions with the following objectives. The received picture quality should be acceptable, even if only one description is received and every additional received description contributes to enhanced picture quality. Typical applications will need a higher bandwidth/higher reliability connection than that provided by a single link in current mobile networks. To support these applications, a mobile node may need to set up and use multiple paths to the desired destination, either simply because of the lack of raw bandwidth on a single channel or because of its poor error characteristics, which reduce its effective throughput. The principal reason for considering such an architecture is to provide high bandwidth and more robust end-to-end connections. We describe a protocol architecture that addresses this need and, with the help of simulations, we demonstrate the feasibility of this system and compare the performance of the MDC-MPT scheme to a system using layered coding and asymmetrical paths for the base and enhancement layers.
Nitin Gogate, Doo-Man Chung, Shivendra S. Panwar, Yao Wang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2002 Multiple-description video coding using motion-compensated temporal prediction
abstract
We propose multiple description (MD) video coders which use motion-compensated predictions. Our MD video coders utilize MD transform coding and three separate prediction paths at the encoder to mimic the three possible scenarios at the decoder: both descriptions received or either of the single descriptions received. We provide three different algorithms to control the mismatch between the prediction loops at the encoder and decoder. We present simulation results comparing the three approaches to two standards-based approaches to MD video coding. We show that when the main prediction loop at the encoder uses a two-channel reconstruction, it is important to have side prediction loops and transmit some redundancy information to control mismatch. We also examine the performance of our MD video coder with partial mismatch control in the presence of random packet loss, and demonstrate a significant improvement compared to more traditional approaches.
Amy R. Reibman, Hamid Jafarkhani, Yao Wang 0001, Michael T. Orchard, Rohit Puri
IEEE Trans. Circuits Syst. Video Technol.3
2002 Error-resilient video coding using multiple description motion compensation
abstract
A new approach for multiple description video coding is proposed. It employs a second-order predictor for motion-compensation, which predicts a current frame from two previously coded frames. The coder generates two descriptions, containing the coded even and odd frames, respectively. When only a single description (say, that containing even frames) is received, the decoder can only use previous even frames for prediction. The mismatch between the predicted frames at the encoder and decoder is explicitly coded to avoid error propagation in the ideal MD channels, where a description is either received intact or lost completely. By using the second-order predictor and coding the mismatch signal, one can also suppress error propagation in packet lossy networks where packets in either description can be lost. The predictor and the mismatch signal quantizer can be varied to achieve a wide range of tradeoffs between coding efficiency and error resilience.
Yao Wang 0001, Shunan Lin
IEEE Trans. Circuits Syst. Video Technol.1
2001 Major cast detection in video using both audio and visual information
abstract
Major casts, for example, the anchor persons or reporters in news broadcast programs and principle characters in movies play an important role in video, and their occurrences provide good indices for organizing and presenting video content. This paper describes a new approach for automatically generating the list of major casts in a video sequence based on multiple modalities, specifically, both speaker and face information. A list of major casts is created and ordered by the accumulative temporal and spatial presence of corresponding casts. Preliminary simulation results show that the detected major casts are meaningful and the proposed approach is promising.
Zhu Liu 0001, Yao Wang 0001
ICASSP2
2001 An unsupervised multi-resolution object extraction algorithm using video-cube
abstract
We propose a fast video object segmentation method that detects object boundaries accurately, and does not require any user assistance. Video streams are considered as 3D data, called video-cubes, to take advantage of 3D signal processing techniques. After a video sequence is filtered, marker nodes are selected from the color gradient. A volume around each marker is grown by using color/texture distance criteria. Then volumes that have similar characteristics are merged. Self-descriptors for each volume, mutual descriptors for each pair of volumes are computed. These descriptors capture motion and spatial information of volumes. In the clustering stage, volumes are classified into objects in a fine-to-coarse hierarchy. While applying and relaxing descriptor based adaptive, similarity scores are estimated for each possible pair-wise combination of volumes. The pair that gives the maximum score is clustered iteratively. Finally, an object-based multi-resolution representation tree is assembled.
Fatih Porikli, Yao Wang 0001
ICIP (2)2
2001 Multiple description video using rate-distortion splitting
abstract
We consider a simple multiple description (MD) video coder, that uses redundancy-rate-distortion criteria to split a one-layer stream generated by a standard video coder into two correlated streams. Our simulation results demonstrate that this MD coder has much better performance for large redundancies than our previous MDTC video coder, although it cannot perform as well at low redundancies. This MD video coder is very simple to implement and is compatible with H.263 to the extent that each description can be decoded by a standard H.263 decoder. This MD coder was used in a previous study on the transport of MD and layered video over an EGPRS wireless network, where the fact that it creates two streams with very balanced rates was a strong advantage.
Amy R. Reibman, Hamid Jafarkhani, Yao Wang 0001, Michael T. Orchard
ICIP (1)3
2001 Rate-distortion optimized video coding considering frameskip
abstract
The general problem of optimized video encoding has received a great deal of attention in recent years. This paper focuses on the optimization of video coding with frameskip. We propose models that estimate the distortion for coded frames as well as non-coded frames. Using these models in conjunction with well-know models that estimate the rate allows us to formulate a rate control problem that trades-off spatial and temporal quality. Simulation results indicate moderate improvements for low motion test sequences.
Anthony Vetro, Huifang Sun, Yao Wang 0001
ICIP (3)3
2001 Evaluation Of Different Descriptors For Identifying Similar Video Shots
abstract
In this paper, three techniques for video sequence retrieval which use statistical measures of color patterns in shots are pro- posed and compared. The first technique is based on a correla- tion of the MPEG7 Dominant Color Descriptor (DC) [6] as single characteristic feature of a shot. The second approach is to model color pattern distribution in a shot with a codebook, obtained by VQ (Vector Quantization) of the frame blocks composing the shot. The last one models the color pattern distribution using a GMM (Gaussian Mixture Model). Such descriptors are used to establish correspondence between non consecutive camera records through an appropriately designed similarity measure. As such, a new dis- tance measure is used in the comparison between the shot descrip- tors, by extending the metric proposed in [9]. A comparison is made of the dissimilarity performance associated with each of the three proposed descriptors, demonstrating the superior results ob- tainable with the VQ based approach.
Nicola Adami, Riccardo Leonardi, Yao Wang 0001
ICME3
2001 A Reference Picture Selection Scheme For Video Transmission Over Ad-Hoc Networks Using Multiple Paths
abstract
Enabling video transmission over ad-hoc networks is more challenging than over conventional mobile networks because a connection path in an ad-hoc network is highly error-prone and the path can go down frequently. On the other hand, it is possible to establish multiple paths between a source and a destination, which provides an extra degree of freedom in coding algorithm design. This paper presents a feedback-based reference picture selection scheme for video transmission over ad-hoc networks. Encoded video streams are transmitted over multiple paths and the reference frames for motion compensated prediction are selected according to the feedback information about the paths' condition. Simulations under the two paths scenario have shown significant improvement over two standard techniques, layered coding and video redundancy coding, which do not use feedback. A novel statistical model for the ad-hoc multi-path environment is also proposed and used in our simulation of transmission loss.
Shunan Lin, Shiwen Mao, Yao Wang 0001, Shivendra S. Panwar
ICME3
2001 Estimating Distortion Of Coded And Non-Coded Frames For Frameskip-Optimized Video Coding
abstract
This paper focuses on the problem of estimating the distortion for coded and non-coded frames in a video coder that employs variable frameskip. The distortion for coded frames is given by classic rate-distortion models, however the distortion for non-coded frames has not been considered. Based on the optical flow equation, we formulate a method for estimating the distortion of the non-coded frames.
Anthony Vetro, Yao Wang 0001, Huifang Sun
ICME2
2001 Reliable transmission of video over ad-hoc networks using automatic repeat request and multipath transport
abstract
The increase in the bandwidth of the wireless channels and the computing power of the mobile devices makes it possible to offer video service for wireless networks in the near future. In an ad-hoc network, strong error protection is required because of the lack of a fixed infrastructure. On the other hand, the mesh structure of an ad-hoc network implies that there may be multiple paths existing between a source and destination, which can be used to enhance video transmissions. We propose a simple but robust scheme for reliable transmission of video in bandwidth limited ad-hoc networks. In our scheme, a video stream is layer coded. The base layer (BL) packets and the enhancement layer (EL) packets are transmitted separately on two disjoint paths using multipath transport (MPT). BL packets are protected by automatic repeat request (ARQ), and a lost BL packet is retransmitted through the path where EL packets are transmitted. An EL packet has lower priority than a retransmitted BL packet and may be dropped at the sender when congestion occurs. Simulation results show that this scheme can guarantee a graceful video quality in adverse channel conditions. It is effective for video transmission over the high loss environment found in ad-hoc networks.
Shiwen Mao, Shunan Lin, Shivendra S. Panwar, Yao Wang 0001
VTC Fall4
2001 Object-based transcoding for adaptable video content delivery
abstract
This paper introduces a new framework for video content delivery that is based on the transcoding of multiple video objects. Generally speaking, transcoding can be defined as the manipulation or conversion of data into another more desirable format. We consider manipulations of object-based video content, and more specifically, from one set of bit streams to another. Given the object-based framework, we present a set of new algorithms that are responsible for manipulating the original set of video bit streams. Depending on the particular strategy that is adopted, the transcoder attempts to satisfy network conditions or user requirements in various ways. One of the main contributions of this paper is to discuss the degrees of freedom within an object-based transcoder and demonstrate the flexibility that it has in adapting the content. Two approaches are considered: a dynamic programming approach and an approach that is based on available meta-data. Simulations with these two approaches provide insight regarding the bit allocation among objects and illustrates the tradeoffs that can be made in adapting the content. When certain meta-data about the content is available, we show that bit allocation can be significantly improved, key objects can be identified, and varying the temporal resolution of objects can be considered.
Anthony Vetro, Huifang Sun, Yao Wang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2001 Multiple description coding using pairwise correlating transforms
abstract
The objective of multiple description coding (MDC) is to encode a source into multiple bitstreams supporting multiple quality levels of decoding. In this paper, we only consider the two-description case, where the requirement is that a high-quality reconstruction should be decodable from the two bitstreams together, while lower, but still acceptable, quality reconstructions should be decodable from either of the two individual bitstreams. This paper describes techniques for meeting MDC objectives in the framework of standard transform-based image coding through the design of pairwise correlating transforms. The correlation introduced by the transform helps to reduce the distortion when only a single description is received, but it also increases the bit rate beyond that prescribed by the rate-distortion function of the source. We analyze the relation between the redundancy (i.e., the extra bit rate) and the single description distortion using this transform-based framework. We also describe an image coder that incorporates the pairwise transform and show its redundancy-rate-distortion performance for real images.
Yao Wang 0001, Michael T. Orchard, Vinay A. Vaishampayan, Amy R. Reibman
IEEE Trans. Image Process.1
2000 Lapped Orthogonal Transform Designed for Error Resilient Image Coding
abstract
This paper describes a new design method for lapped orthogonal transforms (LOTs) that can provide a desired trade-off between coding efficiency and robustness to transmission errors. Traditionally, the LOT bases were designed to maximize the coding efficiency solely. When certain coefficients are lost due to impairments in the transmission channel, these bases often provide unsatisfactory reconstruction quality. In our previous work, we have developed a new reconstruction method, the maximally smooth recovery (MSR) method, which can achieve better image reconstruction quality from incomplete LOT coefficients than simple interpolation methods. We present a new LOT basis design method, which maximizes a weighted average of a coding gain and a reconstruction gain, with the latter being defined according to the MSR method. We show that these new bases can achieve significantly better redundancy-rate-distortion (RRD) performance than the set of basis designed by Hemami (1996).
Doo-Man Chung, Yao Wang 0001
ICIP2
2000 Face Detection and Tracking in Video Using Dynamic Programming
abstract
Face detection and tracking are important in video content analysis since the most important objects in most video are human beings. This paper proposes a new approach for combined face detection and tracking in video. The face detection algorithm is a fast template matching procedure using iterative dynamic programming (DP). Although the face detection algorithm is designed for frontal face, the same mechanism can also be applied to track non-frontal faces with online adapted face models. Due to the essence of template matching, the algorithm is capable of comparing the similarity among different faces, which makes it suitable for tracking the same face that occur at disjointed temporal locations in video. While the proposed face detection method provides comparable accuracy as the neural network based approach, it is much faster.
Zhu Liu 0001, Yao Wang 0001
ICIP2
2000 Transmission of Multiple Description and Layered Video over an EGPRS Wireless Network
abstract
We investigate the ability of multiple descriptions (MD) and layered coding to improve the quality of video transmitted over EGPRS networks. One-layer video sent over a single channel on such a network has a fairly sharp quality transition, depending on a user's location. Either the video can be transmitted reliably (if the video rate is less than or equal to what the channel can sustain), or it is subjected to many lost packets. In this system, MD and layered video may offer two ways to improve the video quality beyond that of the one-layer video. First, each sub-stream can be sent on a separate channel, essentially doubling the assigned bandwidth and increasing the video quality. Second, MD and layered video are more error resilient than one-layer video, potentially improving the video quality seen by users in poor locations. We find that for the system scenarios considered, one and two-layer coding outperform MD coding, depending upon the number of wireless channels used for the video transport.
Amy R. Reibman, Yao Wang 0001, Xiaoxin Qiu, Zhimei Jiang, Kapil K. Chawla
ICIP2
2000 Object-based transcoding for scalable quality of service
abstract
In this paper, we focus on the methods for delivering object-based video data. More specifically, me exploit the fact that a finer level of scalability can be achieved when the video frame has been decomposed into objects and coded using MPEG-4. A new framework is proposed.
Anthony Vetro, Huifang Sun, Yao Wang 0001
ISCAS3
2000 Guest editorial
King Ngi Ngan, Michael G. Strintzis, Masayuki Tanimoto, Yao Wang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2000 Multiview video sequence analysis, compression, and virtual viewpoint synthesis
abstract
This paper considers the problem of structure and motion estimation in multiview teleconferencing-type sequences and its application for video-sequence compression and intermediate-view generation. First, we introduce a new approach for structure estimation from a stereo pair acquired by two parallel cameras. It is based on a 2-D mesh representation of both views of the imaged scene and a parametrization of the structure information by the disparity between corresponding nodes in the image pair. Next, we describe a novel image alignment approach which can convert images captured using nonparallel cameras to coplanar-like images. This approach greatly eases the computational burden incurred by the nonparallel camera geometry, where one must consider both horizontal and vertical disparities. Finally, we present a coder for multiview sequences, which exploits the proposed alignment and structure estimation algorithm. By extracting the foreground objects and estimating the disparity field between a selected view and a reference view the coder can compress the image pair very efficiently. In the meantime, by using the coded structure information, the decoder can generate virtual viewpoints between decoded views, which can be very helpful for telepresence applications.
Ru-Shang Wang, Yao Wang 0001
IEEE Trans. Circuits Syst. Video Technol.2
1999 Performance of multiple description coders on a real channel
abstract
We explore the ability of multiple description (MD) source coders to achieve good performance on channels other than ideal MD channels. We examine both the overall system design and compare the performance of a system with MD source coder to that of a more traditional system using a layered source coder. For the memoryless channels we consider, MD source coding cannot achieve acceptable performance for a memoryless Gaussian source without appropriate channel coding. Also, in memoryless channels, a system with MD source coding outperforms a layered source coding system only in very poor channels. The introduction of memory in the channel degrades the performance of both systems equally. Using interleaving to reduce the impact of memory in the channel has more influence on performance than the choice of source coder.
Amy R. Reibman, Hamid Jafarkhani, Michael T. Orchard, Yao Wang 0001
ICASSP4
1999 Multiple Description Coding for Video Using Motion Compensated Prediction
abstract
We propose multiple description (MD) video coders which use motion compensated predictions. Our MD video coders utilize MD transform coding and three separate prediction paths at the encoder, to mimic the three possible scenarios at the decoder: both descriptions received or either of the single descriptions received. We provide three different algorithms to control the mismatch between the prediction loops at the encoder and decoder. The results show that when the main prediction loop is the central loop, it is important to have side prediction loops and transmit some redundancy information to control mismatch.
Amy R. Reibman, Hamid Jafarkhani, Yao Wang 0001, Michael T. Orchard, Rohit Puri
ICIP (3)3
1999 Rate-Distortion Modeling of Binary Shape Using State Partitioning
abstract
In this paper, the rate-distortion (R-D) characteristics of binary shapes are modeled. Specifically we are interested in predicting the rate and distortion that is produced by the shape coding techniques that have been adopted into the MPEG-4 standard. The shape coding algorithm is a context-based arithmetic encoder and operates on a per block basis. Currently, there is no efficient way of estimating the rate and distortion at various levels of resolution. Consequently, we propose a model that is based on a set of parameters that can be easily extracted from the binary blocks. The parameters represent states that arise from the possible binary patterns that can occur in a small neighborhood around the current pixel. Symmetry is exploited to keep the number of states to a minimum. It is shown that the proposed model is computationally efficient and provides accurate estimates of the R-D characteristics of a binary shape.
Anthony Vetro, Huifang Sun, Yao Wang 0001, Onur G. Guleryuz
ICIP (2)3
1999 Coding of Motion Compensation Residuals Using Edge Information
abstract
In most current video coders, a block is first predicted from its best matching block in a previous frame, and the prediction error is then coded using discrete transform coding (DCT). Because of the inadequacy of the block-wise translational motion model, edges in the predicted block are often shifted from their true positions, leading to errors that are clustered around edges in the predicted block. DCT is inefficient for coding such errors. Independent searching of block motion vectors also lead to discontinuities of edges across block boundaries. Existing coders ignore such correlation between error location and edge discontinuity. We describe a coder that corrects edge-misalignment before applying DCT coding. The correlation between edge-discontinuity and edge-misalignment is exploited in the coding of the misalignment parameters.
Yao Wang 0001, Michael T. Orchard
ICIP (1)1
1999 Integration of multimodal features for video scene classification based on HMM
abstract
Along with the advances in multimedia and Internet technology, a huge amount of data, including digital video and audio, are generated daily. Tools for the efficient indexing and retrieval of such data are indispensable. With multi-modal information present in the data, effective integration is necessary and is still a challenging problem. In this paper, we present four different methods for integrating audio and visual information for video classification based on a hidden Markov model (HMM): direct concatenation, product HMM, two-stage HMM, and integration by neural network. Our results have shown significant improvements over using a single modality.
Jincheng Huang 0001, Zhu Liu 0001, Yao Wang 0001, Edward K. Wong
MMSP3
1999 Multiple description image coding using signal decomposition and reconstruction based on lapped orthogonal transforms
abstract
This paper considers the use of multiple description coding (MDC) for image transmission in communication systems where long burst errors and sometimes complete channel failures are inevitable. A general framework for MDC is proposed, which uses nonhierarchical signal decomposition at the encoder and image reconstruction at the decoder. A realization of this framework using lapped orthogonal transforms (LOTs) is developed. In the encoder, the bitstream generated by a conventional LOT-based image coder is decomposed so that each description consists of a subsampled set of the coded LOT coefficient blocks. In the decoder, instead of using the inverse LOT directly, a novel image reconstruction technique is employed, which makes use of the constraints between adjacent LOT coefficient blocks and the smoothness property of common image signals. To guarantee a satisfactory reconstruction quality, the transform should introduce a desired amount of correlation among adjacent LOT coefficient blocks. The tradeoff between coding efficiency and reconstruction quality obtainable by using different LOT bases is investigated.
Doo-Man Chung, Yao Wang 0001
IEEE Trans. Circuits Syst. Video Technol.2
1999 MPEG-4 rate control for multiple video objects
abstract
This paper describes an algorithm which can achieve a constant bit rate when coding multiple video objects. The implementation is a nontrivial extension of the MPEG-4 rate control algorithm for single video objects which employs a quadratic rate quantizer model. The algorithm is organized into two stages: a pre- and a post-encoding stage. In the pre-encoding stage, an initial target estimate is made for each object. Based on the buffer fullness, the total target is adjusted and then distributed proportional to the relative size, motion, and variance of each object. Based on the new individual targets and rate-quantizer relation for texture, appropriate quantization parameters are calculated. After each object is encoded, the model parameters for each object are updated, and if necessary, frames are skipped to ensure that the buffer does not overflow. A preframeskip control is exercised to avoid buffer overflow when the motion and shape information occupies a significant portion of the bit budget. The rate control algorithm switches between two operation modes so that the coder can reduce the spatial coding accuracy for an improved temporal resolution. A shape-coding control mechanism is also proposed, which provides a tradeoff between texture and shape coding accuracy. Overall, the algorithm is able to successfully achieve the target bit rate, effectively code arbitrarily shaped objects, and maintain a stable buffer level. These techniques have been adopted by the MPEG committee in July 1997 as part of the video verification model (VM8).
Anthony Vetro, Huifang Sun, Yao Wang 0001
IEEE Trans. Circuits Syst. Video Technol.3
1999 Compression of color facial images using feature correction two-stage vector quantization
abstract
A feature correction two-stage vector quantization (FC2VQ) algorithm was previously developed to compress gray-scale photo identification (ID) pictures. This algorithm is extended to color images in this work. Three options are compared, which apply the FC2VQ algorithm in RGB, YCbCr, and Karhunen-Loeve transform (KLT) color spaces, respectively. The RGB-FC2VQ algorithm is found to yield better image quality than KLT-FC2VQ or YCbCr-FC2VQ at similar bit rates. With the RGB-FC2VQ algorithm, a 128 x 128 24-b color ID image (49,152 bytes) can be compressed down to about 500 bytes with satisfactory quality. When the codeword indices are further compressed losslessly using a first order Huffman coder, this size is further reduced to about 450 bytes.
Jincheng Huang 0001, Yao Wang 0001
IEEE Trans. Image Process.2
1999 Regularized total least squares approach for nonconvolutional linear inverse problems
abstract
In this correspondence, a solution is developed for the regularized total least squares (RTLS) estimate in linear inverse problems where the linear operator is nonconvolutional. Our approach is based on a Rayleigh quotient (RQ) formulation of the TLS problem, and we accomplish regularization by modifying the RQ function to enforce a smooth solution. A conjugate gradient algorithm is used to minimize the modified RQ function. As an example, the proposed approach has been applied to the perturbation equation encountered in optical tomography. Simulation results show that this method provides more stable and accurate solutions than the regularized least squares and a previously reported total least squares approach, also based on the RQ formulation.
Wenwu Zhu 0001, Yao Wang 0001, Nikolas P. Galatsanos, Jun Zhang 0006
IEEE Trans. Image Process.2
1998 Multiple Description Image Coding based on Lapped Orthogonal Transforms
abstract
This paper considers the use of multiple description coding (MDC) for image transmission in communication systems where long burst errors and sometimes complete channel failures are inevitable. A general framework for MDC is proposed, which uses non-hierarchical signal decomposition at the encoder and image reconstruction at the decoder. A relaxation of this framework using lapped orthogonal transforms (LOTs) is developed. A novel image reconstruction technique is employed, which makes use of the constraints between adjacent LOT coefficient blocks and the smoothness property of common image signals. The tradeoff between coding efficiency and reconstruction quality obtainable by using different LOT bases is investigated.
Doo-Man Chung, Yao Wang 0001
ICIP (1)2
1998 Integration of Audio and Visual Information for Content-based Video Segmentation
Jincheng Huang 0001, Zhu Liu 0001, Yao Wang 0001
ICIP (3)3
1998 Optimal Pairwise Correlating Transforms for Multiple Description Coding
abstract
Multiple description coding (MDC) addresses the problem of encoding a source into two (or more) bitstreams such that a high-quality reconstruction is decodable from the two bitstreams together, while a lower, but still acceptable, quality reconstruction is decodable if either of the two bitstreams is lost. Recent research has proposed using transforms to introduce a controlled amount of correlation between the two bitstreams in order to achieve MDC objectives. This paper considers several optimality issues related to such transform based MDC methods. Redundancy rate-distortion (RRD) performance of a general class of transforms is derived and used to identify the optimal transform for achieving any given amount of redundancy. Then, the paper introduces a more general transform-based MDC framework incorporating both the transform mode of redundancy and a second mode of redundancy. The optimal allocation of redundancy among these two modes is analyzed.
Yao Wang 0001, Michael T. Orchard, Amy R. Reibman
ICIP (1)1
1998 Classification TV programs based on audio information using hidden Markov model
abstract
This paper describes a technique for classifying TV broadcast video using a hidden Markov model (HMM). Here we consider the problem of discriminating five types of TV programs, namely commercials, basketball games, football games, news reports, and weather forecasts. Eight frame-based audio features are used to characterize the low-level audio properties, and fourteen clip-based audio features are extracted based on these frame-based features to characterize the high-level audio properties. For each type of these five TV programs, we build an ergodic HMM using the clip-based features as observation vectors. The maximum likelihood method is then used for classifying testing data using the trained models.
Zhu Liu 0001, Jincheng Huang 0001, Yao Wang 0001
MMSP3
1998 Stereo sequence analysis, compression, and virtual viewpoint synthesis
abstract
This paper considers the problem of structure and motion estimation in stereoscopic tele-conferencing type sequences and its application for stereo sequence compression and for intermediate view generation. Generally, this type of sequence consists of one or several foreground objects (head and shoulders) and a more or less static background. By extracting the foreground objects and investigating the relationship between right and left images, we can compress the image-pair by making use of the redundancy of the left and right image-pair. In the meantime, by using the estimated structure information, we ran generate virtual viewpoints between the left and right image-pair, which can be very helpful for tele-presence applications.
Ru-Shang Wang, Yao Wang 0001
MMSP2
1998 Error control and concealment for video communication: a review
abstract
The problem of error control and concealment in video communication is becoming increasingly important because of the growing interest in video delivery over unreliable channels such as wireless networks and the Internet. This paper reviews the techniques that have been developed for error control and concealment. These techniques are described in three categories according to the roles that the encoder and decoder play in the underlying approaches. Forward error concealment includes methods that add redundancy at the source end to enhance error resilience of the coded bit streams. Error concealment by postprocessing refers to operations at the decoder to recover the damaged areas based on characteristics of image and video signals. Last, interactive error concealment covers techniques that are dependent on a dialogue between the source and destination. Both current research activities and practice in international standards are covered.
Yao Wang 0001, Qin-Fan Zhu
Proc. IEEE1
1998 End-to-end modeling and simulation of MPEG-2 transport streams over ATM networks with jitter
abstract
The operation of MPEG-2 systems is modeled and simulated when an MPEG-2 transport stream is delivered through an ATM network with jitter. End-to-end packet-based analysis is performed for delivery of MPEG-2 transport streams over ATM networks. A novel approach to analyzing the decoder buffer behaviour in the presence of network jitter is presented. The probability density function of the interarrival time of the ATM adaptation layer 5 (AAL5) protocol data unit (PDU) is derived from an MPEG-2 video source model and an ATM network jitter model. Based on a real-time decoding requirement of the MPEG-2 transport stream (TS) system target decoder (T-STD), the decoder buffer behaviour is simulated. In this simulation, the packets' arrivals follow the derived probability density function of the AAL5 PDU interarrival time. The modeling and simulation results show the interactions among packet loss ratio, decoder buffer size, and network jitter level. We found that jitter affects decoder buffer size and packet loss ratio in a significant way.
Wenwu Zhu 0001, Y. Thomas Hou 0001, Yao Wang 0001, Ya-Qin Zhang
IEEE Trans. Circuits Syst. Video Technol.3
1998 Second-order derivative-based smoothness measure for error concealment in DCT-based codecs
abstract
We study the recovery of lost or erroneous transform coefficients in image and video communication systems employing discrete cosine transform (DCT)-based codecs. Previously, we have developed a technique that exploits the smoothness property of image signals and recovers the damaged blocks by maximizing a smoothness measure. There, the first-order derivative was used as the smoothness measure, which can lead to the blurring of sharp edges. In order to alleviate this problem, we propose to use second-order derivatives as the smoothness measure. Our simulation results show that a weighted combination of the quadratic variation and the Laplacian operator can significantly reduce the blurring across the edges while enforcing smoothness along the edges.
Wenwu Zhu 0001, Yao Wang 0001, Qin-Fan Zhu
IEEE Trans. Circuits Syst. Video Technol.2
1997 Adaptive stripe based patch matching for depth estimation
abstract
A novel stereo matching technique for depth estimation in stereoscopic image pairs is presented. The input image pair is preprocessed in the intensity domain and together with edge maps an adaptive mesh in which individual elements approximate linearly modeled regions are obtained. Then, an iterative stripe based, quadrilateral patch matching technique is employed to estimate the depth map from the image pair in a hierarchical manner. Finally, the resultant map is postprocessed to smooth the depth map at the patch borders. The quality of the test results demonstrates the effectiveness of the technique.
Fatih Porikli, Yao Wang 0001, Cassandra T. Swain
ICASSP2
1997 Check image compression: a comparison of JPEG, wavelet and layered coding method
abstract
An emerging trend in the banking industry is to digitize checks for storage and transmission. An immediate requirement for efficient storage and transmission is check image compression. General purpose compression algorithms such as JPEG and wavelet-based methods produce annoying ringing and blocking artifacts at high compression ratios. A layered approach to check image compression is proposed, based on which a check image is represented in several layers. The first layer describes the foreground map; the second layer specifies the gray levels of foreground pixels; the third layer is a lossy representation of the background image; and the fourth layer describes the error between the original and the reconstructed image based on the first three layers. The layered coding approach produces images of better quality than traditional JPEG and wavelet coding methods, especially in the foreground, i.e., the text and graphics. In addition, this approach allows progressive retrieval or transmission of different image layers.
Jincheng Huang 0001, Yao Wang 0001, Edward K. Wong
ICIP (3)2
1997 Redundancy Rate-Distortion Analysis Of Multiple Description Coding Using Pairwise Correlating Transforms
abstract
The objective of multiple description coding (MDC) is to encode a source into two (or more) bitstreams supporting two quality levels of decoding. A high-quality reconstruction should be decodable from the two bitstreams together, while lower, but still acceptable, quality reconstructions should be decodable from either of the two individual bitstreams. This paper describes techniques for meeting MDC objectives in the framework of standard transform-based image coding through the design of pairwise transforms.
Yao Wang 0001, Michael T. Orchard, Amy R. Reibman, Vinay A. Vaishampayan
ICIP (1)1
1997 Regularized Total Least Squares Reconstruction for Optical Tomographic Imaging Using Conjugate Gradient Method
abstract
A regularized total least square (RTLS) approach to solve a linear perturbation equation encountered in optical tomography is developed based on the Rayleigh quotient formulation. To compute efficiently the solution, the Rayleigh quotient form of the RTLS filter (RQF-RTLS) is used and a conjugate gradient algorithm is implemented. Simulation results show that the RQF-RTLS method obtains more stable and accurate solutions than the regularized least squares (RLS) approach which does not account for the errors in the operator.
Wenwu Zhu 0001, Yao Wang 0001, Nikolas P. Galatsanos, Jun Zhang 0006
ICIP (1)2
1997 Audio feature extraction and analysis for scene classification
abstract
Analysis and classification of the scene content of a video sequence are very important for content-based indexing and retrieval of multimedia databases. We report our research on using the associated audio information for video scene classification. We describe several audio features that have been found effective in distinguishing audio characteristics of different scene classes. Based on these features, a neural net classifier was quite successful in separating audio clips from different TV programs.
Zhu Liu 0001, Jincheng Huang 0001, Yao Wang 0001, Tsuhan Chen
MMSP3
1997 Multiple description image coding for noisy channels by pairing transform coefficients
abstract
Multiple description coding (MDC) is a way of trading off coding gain with robustness to channel errors. This paper presents a new method for MDC using the framework of transform coding. Instead of using the Karhunen-Loeve transform (KLT) that decorrelates all the coefficients, we choose the transform bases so that the coefficients are correlated pair-wise. This is accomplished by rotating every two basis vectors in the KLT. Each pair of correlated coefficients are then split between two descriptions. Only 45/spl deg/ rotation is considered which leads to two balanced streams. In the actual implementation, the DCT is employed in place of the KLT and the rotation of transform bases is accomplished by rotating the DCT coefficients. Experimental results show that this method can lead to satisfactory image reconstruction from any one description with a relatively small (20% for "lena") overhead over a standard JPEG coder.
Yao Wang 0001, Michael T. Orchard, Amy R. Reibman
MMSP1
1997 Facial feature extraction and tracking in video sequences
abstract
In this paper, we describe a new approach for facial feature extraction, The features considered include eyes and mouth. A novel nonlinear filter is proposed for valley detection, which forms the low-level processing stage. A combination of temporal difference, color, image intensity and geometrical constraint is used to determine facial features in an initial frame. A temporal matching technique is used to track these features in the following frames. We also present a simple profile scanning method which can accurately locate special feature points such as corners in the eyes and mouth.
Ru-Shang Wang, Yao Wang 0001
MMSP2
1997 Modeling and simulation of MPEG-2 video transport over ATM networks considering the jitter effect
abstract
In this paper, the operation of MPEG-2 systems is modeled and simulated when an MPEG-2 transport stream is delivered through a ATM network with jitter. A novel approach to analyzing the decoder buffer behavior in the presence of network jitter is presented. The probability density function of the interarrival time of the ATM adaptation layer 5 (AAL5) Protocol Data Unit (PDU) is derived from a MPEG-2 video source model and an ATM network jitter model. Based on a real-time decoding requirement of the MPEG-2 transport stream (TS) system target decoder (T-STD), the decoder buffer behavior is simulated. The modeling; and simulation results show that jitter affects decoder buffer size and packet loss ratio in a significant way.
Wenwu Zhu 0001, Y. Thomas Hou 0001, Yao Wang 0001, Ya-Qin Zhang
MMSP3
1997 Segmented Adaptive DPCM for Lossy Compression of Multispectral MR Images
abstract
This paper reports a multispectral segmented differential pulse coded modulation (MSDPCM) method for well registered multispectral magnetic resonance (MR) images. Given a set of multispectral MR images, the MSDPCM method first segments it into statistically distinct regions which by and large correspond to different tissue classes. It then finds a suitable linear prediction model (LPM) for each class. The LPMs used here are the well known causalautoregressive(AR) andautoregressive moving average(ARMA) models. Finally, the MSDPCM method quantizes the prediction error in each class using a vector quantizer. The original image set is described by the segmentation map, the model parameters for each class, and the quantized prediction errors. The MSDPCM method can produce very high compression gains, because the specification of the segmentation map and model parameters requires significantly fewer bits than that for the original intensity values. The MSDPCM method using the backward adaptive ARMA model has been applied to head MR images with three spectral bands (one T1 weighted and two T2 weighted, 256 × 256 × 12 bits/image). In an informal validation, the compressed images have been evaluated against the originals by three neuroradiologists. Images compressed by an average factor of more than 23 have been regarded asacceptablefor clinical film reading.
Jian-Hong Hu, Yao Wang 0001, Patrick T. Cahill
J. Vis. Commun. Image Represent.2
1997 Multispectral code excited linear prediction coding and its application in magnetic resonance images
abstract
This paper reports a multispectral code excited linear prediction (MCELP) method for the compression of multispectral images. Different linear prediction models and adaptation schemes have been compared. The method that uses a forward adaptive autoregressive (AR) model has been proven to achieve a good compromise between performance, complexity, and robustness. This approach is referred to as the MFCELP method. Given a set of multispectral images, the linear predictive coefficients are updated over nonoverlapping three-dimensional (3-D) macroblocks. Each macroblock is further divided into several 3-D micro-blocks, and the best excitation signal for each microblock is determined through an analysis-by-synthesis procedure. The MFCELP method has been applied to multispectral magnetic resonance (MR) images. To satisfy the high quality requirement for medical images, the error between the original image set and the synthesized one is further specified using a vector quantizer. This method has been applied to images from 26 clinical MR neuro studies (20 slices/study, three spectral bands/slice, 256x256 pixels/band, 12 b/pixel). The MFCELP method provides a significant visual improvement over the discrete cosine transform (DCT) based Joint Photographers Expert Group (JPEG) method, the wavelet transform based embedded zero-tree wavelet (EZW) coding method, and the vector tree (VT) coding method, as well as the multispectral segmented autoregressive moving average (MSARMA) method we developed previously.
Jian-Hong Hu, Yao Wang 0001, Patrick T. Cahill
IEEE Trans. Image Process.2
1997 A Wavelet-Based Multiresolution Regularized Least Squares Reconstruction Approach for Optical Tomography
abstract
In this paper, we present a wavelet-based multigrid approach to solve the perturbation equation encountered in optical tomography. With this scheme, the unknown image, the data, as well as the weight matrix are all represented by wavelet expansions, thus yielding a multiresolution representation of the original perturbation equation in the wavelet domain. This transformed equation is then solved using a multigrid scheme, by which an increasing portion of wavelet coefficients of the unknown image are solved in successive approximations. One can also quickly identify regions of interest (ROI's) from a coarse level reconstruction and restrict the reconstruction in the following fine resolutions to those regions. At each resolution level a regularized least squares solution is obtained using the conjugate gradient descent method. This approach has been applied to continuous wave data calculated based on the diffusion approximation of several two-dimensional (2-D) test media. Compared to a previously reported one grid algorithm, the multigrid method requires substantially shorter computation time under the same reconstruction quality criterion.
Wenwu Zhu 0001, Yao Wang 0001, Yining Deng, Yuqi Yao, Randall L. Barbour
IEEE Trans. Medical Imaging2
1996 Use of two-dimensional deformable mesh structures for video coding .I. The synthesis problem: mesh-based function approximation and mapping
abstract
This paper explores the use of a deformable mesh (also known as the control grid) structure for motion analysis and synthesis in an image sequence. We focus on the synthesis problem, i.e., how to interpolate an image function given nodal positions and values and how to predict a present image frame from a reference one given nodal displacements between the two images. For this purpose, we review the fundamental theory and numerical techniques that have been developed in the finite element method for function approximation and mapping using a mesh structure. Specifically, we focus on (i) the use of shape functions for node-based function interpolation and mapping; and (ii) the use of regular master elements to simplify numerical calculations involved in dealing with irregular mesh structures. In addition to a general introduction that is applicable to an arbitrary mesh structure, we also present specific results for triangular and quadrilateral mesh structures, which are the most useful two-dimensional (2-D) meshes. Finally, we describe how to apply the above results for motion compensated frame prediction and interpolation. It is shown that the concepts of shape functions and master elements are crucial for developing computationally efficient algorithms for both the analysis and synthesis problems.
Yao Wang 0001, Ouseb Lee
IEEE Trans. Circuits Syst. Video Technol.1
1996 Use of two-dimensional deformable mesh structures for video coding. II. The analysis problem and a region-based coder employing an active mesh representation
abstract
For pt.I see ibid., vol.6, no.6, p.636-46 (1996). This paper explores the use of the deformable mesh structure for motion/shape analysis and synthesis in an image sequence. We present algorithms for the analysis problem, including scene-adaptive mesh generation and node tracking over successive frames. We also describe a region-based video coder that integrates the analysis and synthesis algorithms presented. The coder describes each region by an ensemble of connected quadrilateral elements embedded in a mesh structure. For each region, its shape and texture are described by the nodal positions and image functions of the elements in this region in an initial frame, while its motion (including shape deformation) is characterized by the nodal trajectories in the following frames, which are in turn specified by a few motion parameters. This coder has been applied to a typical common intermediate format (CIF) resolution, head-and-shoulder type sequence. The visual quality is significantly better than the H.263-TMN4 algorithm at about 50 kb/s (for the luminance component only, 30 Hz).
Yao Wang 0001, Ouseb Lee, Anthony Vetro
IEEE Trans. Circuits Syst. Video Technol.1
1995 Speech-assisted lip synchronization in audio-visual communications
abstract
We utilize speech information to improve the quality of audio-visual communications such as video telephony and videoconferencing. We show that the marriage of speech analysis and image processing can solve problems related to lip synchronization. We present a technique called speech-assisted frame-rate conversion, and apply it to coding of talking head video. Demonstration sequences are presented. Extensions and other applications are outlined.
Tsuhan Chen, Hans Peter Graf, Barry G. Haskell, Eric Petajan, Yao Wang 0001, Homer H. Chen, Wu Chou
ICIP5
1995 A new frame interpolation scheme for talking head sequences
abstract
A video codec typically-skips frames to satisfy the bit rate constraint. This results in a jerky motion and a loss of lip synchronization in sequences of talking persons. We propose a frame interpolation technique to solve these problems. The main components of this technique include: foreground/background segmentation, mesh-based frame interpolation, and image analysis/synthesis for mouth movements. This technique generates smooth head motion and renders lip motion that is synchronized with the voice of the person.
Tsuhan Chen, Yao Wang 0001, Hans Peter Graf, Cassandra T. Swain
ICIP2
1995 Region segmentation based on active mesh representation of motion: comparison of parallel and sequential approaches
abstract
An active mesh representation of motion has been developed by the authors. It describes the motion in an image sequence by the trajectories of a set of nodes that form a mesh (also known as a control grid). In this paper we investigate how to cluster the nodes, or equivalently the mesh elements, so that the nodes in the same group have similar global motions. We present and compare two approaches. The parallel approach first finds a set of global motion parameters for each element and then partitions these parameters using the K-means clustering algorithm; The sequential (or layered) approach recursively extracts the nodes with the dominant motion from all the nodes, using a modified robust regression algorithm.
Yao Wang 0001, Xia-Ming Hsieh, Jian-Hong Hu, Ouseb Lee
ICIP1
1995 Motion-Compensated Prediction Using Nodal-Based Deformable Block Matching
abstract
One well-known deficiency of the block matching algorithm (BMA) is that it cannot handle nontranslational motion within a block. This paper proposes a nodal-displacement-based deformation model. It assumes that a selected number of control nodes in a block can move freely, and that the displacement of any interior point can be interpolated from nodal displacements. This model includes the translational model as a special case with a single node and can characterize increasingly more complex deformation by the use of more nodes. Of particular importance is the four-node-based rectangle to quadrangle mapping, which covers essentially all possible 2D deformations caused by 3D motions over a reasonably small block. A deformable block matching algorithm (DBMA) is developed for the estimation of nodal displacements. It is a gradient-based search algorithm and each iteration requires the solution of a simple linear equation. A hybrid three-mode method is further proposed, which chooses among the BMA and DBMA, both without error correction, and the BMA with error correction using a DCT coding method (BMA-DCT). The DBMA and the three-mode method have been simulated for the case of the four-node-based mapping. In comparison with the BMA-DCT, the reduction of prediction errors by the DBMA outweighs the increase in required bits for specifying motion. At similar bit rates, the DBMA has yielded better image quality than the BMA-DCT. On the other hand, the three-mode method can significantly reduce the bit rate while maintaining a similar quality, at the expense of a slight increase in the computational cost.
Ouseb Lee, Yao Wang 0001
J. Vis. Commun. Image Represent.2
1994 Segmentation Based Linear Predictive Coding of Mulitspectral Images
abstract
This paper presents a segmentation based linear predictive coding (SLPC) method for multispectral images. Given a set of multispectral images, the SLPC method first segments it into statistically distinct regions. It then finds a suitable linear prediction model for each region. Finally, it quantizes the prediction error in each class using a vector quantizer. The original image set is described by the segmentation map, the model parameters for each class, and the quantized prediction errors. The SLPC method can produce very high compression gains, because the specification of the segmentation map and model parameters requires significantly fewer bits than that for the original intensity values. This method has been applied to magnetic resonance head images with three spectral bands (one T1 weighted and two T2 weighted, 256/spl times/256/spl times/12 bits/image). Images compressed by a factor of more than 22 have been regarded as indistinguishable from the originals, by several radiologists.>
Jian-Hong Hu, Yao Wang 0001, Patrick T. Cahill
ICIP (3)2
1994 Active mesh-a feature seeking and tracking image sequence representation scheme
abstract
This paper introduces a representation scheme for image sequences using nonuniform samples embedded in a deformable mesh structure. It describes a sequence by nodal positions and colors in a starting frame, followed by nodal displacements in the following frames. The nodal points in the mesh are more densely distributed in regions containing interesting features such as edges and corners; and are dynamically updated to follow the same features in successive frames. They are determined automatically by maximizing feature (e.g., gradient) magnitudes at nodal points, while minimizing interpolation errors within individual elements, and matching errors between corresponding elements. In order to avoid the mesh elements becoming overly deformed, a penalty term is also incorporated, which measures the irregularity of the mesh structure. The notions of shape functions and master elements commonly used in the finite element method have been applied to simplify the numerical calculation of the energy functions and their gradients. The proposed representation is motivated by the active contour or snake model proposed by Kass, Witkin, and Terzopoulos (1988). The current representation retains the salient merit of the original model as a feature tracker based on local and collective information, while facilitating more accurate image interpolation and prediction. Our computer simulations have shown that the proposed scheme can successfully track facial feature movements in head-and-shoulder type of sequences, and more generally, interframe changes that can be modeled as elastic deformation. The treatment for the starting frame also constitutes an efficient representation of arbitrary still images.
Yao Wang 0001, Ouseb Lee
IEEE Trans. Image Process.1
1993 Active mesh: a video representation scheme for feature seeking and tracking
abstract
This paper introduces a representation scheme for images and video sequences using nonuniform samples embedded in a mesh structure. It describes a video sequence by the nodal positions and colors in a starting frame, followed by the nodal displacements in the following frames. The nodal points are more densely distributed in regions containing interesting features such as edges and corners, and are dynamically updated to follow the same features in successive frames. They are determined automatically by maximizing feature (e.g., gradient) magnitudes at nodal points, while minimizing interpolation errors within individual elements, and matching errors between corresponding elements. In order to avoid the mesh elements becoming overly deformed, a penalty term is also incorporated which measures the irregularity of the mesh structure. The notions of shape functions and master elements commonly used in the finite element method have been employed to simplify the numerical calculation of the energy functions and their gradients. The proposed representation is motivated by the active contour or snake model proposed by Kass, Witkin, and Terzopoulos. The current representation retains the salient merit of the original model as a feature tracker based on local and collective information, while facilitating more accurate image interpolation and prediction.
Yao Wang 0001, Ouseb Lee
VCIP1
1993 Semiadaptive vector quantization and its application in medical image compression
abstract
In this paper, we introduce a semi-adaptive vector quantization (SAVQ) method, which is a combination of the traditional VQ scheme using a fixed code book and the locally adaptive VQ (LAVQ) method which dynamically constructs a code book according to the input data stream. The code book in SAVQ consists of two parts: a fixed part that is designed based on certain training signals as in VQ, and an adaptive part that it updated based on the input vectors to be compressed. The proposed method is more effective than VQ and LAVQ for semi-stationary signals that have patterns common over different images as well as features specific to a particular image. Such is the case with medical images, which have similar tissue characteristics over different images, as well as with local variations that are patient and pathology dependent. The SAVQ as well as VQ and LAVQ methods have been applied to multispectral magnetic resonance brain images. The SAVQ has achieved higher compression ratios than the VQ and LAVQ methods over a wide range of reproduction quality, with more significant improvement in the mid to high quality range. Furthermore, under the same quality criterion, SAVQ requires a much smaller code book than VQ, making the former less time and memory demanding. Readings by neuroradiologists have suggested that images produced by SAVQ at compression ratios up to 40 (for MRI data with 3 or 4 images/set, 256 X 256 pixels/image, and 16 bits/pixel) are acceptable for primary reading.
Jian-Hong Hu, Yao Wang 0001, Patrick T. Cahill
VCIP2
1993 Image Representation Using Block Pattern Models and Its Image Processing Applications
abstract
An image representation scheme using a set of block pattern models (BPMs) consisting of three categories (constant, oriented, and irregular) is introduced. Algorithms for model classification, model parameter estimation, and image reconstruction from model parameters are presented, and these provide the necessary vehicles for applying the proposed representation scheme to various image processing tasks. The applications of the proposed models in image coding, image zooming, and image smoothing are described. Satisfactory coded images have been obtained at bit rates between 0.5 approximately 0.6 b.p.p. (bits per pixel) with a high-rate realization and between 0.3 approximately 0.5 b.p.p. with a low-rate realization. The high-rate realization has a simple structure suitable for real-time implementation. The methods for image zooming and smoothing are similar, where both adapt the processing for each pixel according to the model of its neighborhood. By using directional filters in oriented regions, edges and lines are rendered sharper in a smoother manner than with conventional linear filtering approaches, which leads to significant improvement in perceived image quality.>
Yao Wang 0001, Sanjit K. Mitra
IEEE Trans. Pattern Anal. Mach. Intell.1
1993 Maximally smooth image recovery in transform coding
abstract
The authors consider the reconstruction of images from partial coefficients in block transform coders and its application to packet loss recovery in image transmission over asynchronous transfer mode (ATM) networks. The proposed algorithm uses the smoothness property of common image signals and produces a maximally smooth image among all those with the same coefficients and boundary conditions. It recovers each damaged block by minimizing the intersample variation within the block and across the block boundary. The optimal solution is achievable through two linear transformations, where the transform matrices depend on the loss pattern and can be calculated in advance. The reconstruction of contiguously damaged blocks is accomplished iteratively using the previous solution as the boundary conditions in each new step. This technique is applicable to any unitary block-transform and is effective for recovering the DC and low-frequency coefficients. When applied to still image coders using the discrete cosine transform (DCT), high quality images are reconstructed in the absence of many DC and low-frequency coefficients over spatially adjacent blocks. When the damaged blocks are isolated by block interleaving, satisfactory results have been obtained even when all the coefficients are missing.>
Yao Wang 0001, Qin-Fan Zhu, Leonard Shaw
IEEE Trans. Commun.1
1993 Coding and cell-loss recovery in DCT-based packet video
abstract
The applications of discrete cosine transform (DCT)-based image- and video-coding methods in the asynchronous transfer mode (ATM) environment are considered. Coding and reconstruction mechanisms are jointly designed to achieve a good compromise among compression gain, system complexity, processing delay, error-concealment capability, and reconstruction quality. The Joint Photographic Experts Group (JPEG) and Motion Picture Experts Group (MPEG) algorithms for image and video compression are modified to incorporate block interleaving in the spatial domain and DCT coefficient segmentation in the frequency domain to conceal the errors due to packet loss. A new algorithm is developed that recovers the damaged regions by adaptive interpolation in the spatial, temporal, and frequency domains. The weights used for spatial and temporal interpolations are varied according to the motion content and loss patterns of the damaged regions. When combined with proper layered transmission, the proposed coding and reconstruction methods can handle very high packet-loss rates at only a slight cost in compression gain, system complexity, and processing delay.>
Qin-Fan Zhu, Yao Wang 0001, Leonard Shaw
IEEE Trans. Circuits Syst. Video Technol.2
1992 Vector Run-length Coding of Bilevel Images
abstract
Run-length coding (RC) is a simple and yet quite effective technique for bi-level image coding. A problem with the conventional RC which describes an image by alternating runs of white and black pixels is that it only exploits the redundancy within the same scan line. The modified relative address run-length coding (MRC) used in Group III facsimile transmission is more efficient by making use of the correlation between adjacent lines. The paper presents a vector run-length coding (VRC) technique which exploits the spatial redundancy more thoroughly by representing images with vector or black patterns and vector run-lengths. Depending on the coding method for the block patterns, various algorithms have been developed, including single run-length VRC (SVRC), double run-length VRC (DVRC), and block VRC (BVRC). The conventional RC is a special case of BVRC with block size of 1*1. The proposed methods have been applied to the CCITT standard test documents and the best result has been obtained with the BVRC method. With a block dimension of 4*4, it has yielded compression gains higher than the MRC with k=4 by 15.5% and 22.7%, when using a single and multiple run-length codebooks, respectively.>
Yao Wang 0001, J. M. Wu
Data Compression Conference1
1992 Image Reconstruction for Hybrid Video Coding Systems
abstract
Presents a new technique for image reconstruction from partially received information for hybrid video coding systems using DCT and motion compensated prediction and interpolation. The technique makes use of the smoothness property of typical video signals by requiring the reconstructed samples be smoothly connected with their adjacent samples, both spatially and temporally. This is fulfilled by minimizing the differences between neighboring pixels in the current as well as adjacent frames. The optimal solution is obtained through three linear transformations. This approach can yield more satisfactory results than the existing algorithms, especially for images with large motions or scene changes.>
Qin-Fan Zhu, Yao Wang 0001, Leonard Shaw
Data Compression Conference2
1991 Edge detection based on orientation distribution of gradient images
abstract
An edge detection scheme which exploits both the magnitude and orientation information of a gradient image is presented. A pixel with a large gradient is considered as an edge element only if the samples in its neighborhood have a unique orientation (straight edge) or a few strong directions (mixed edge). By making use of the orientation information, the proposed scheme can effectively distinguish between the edge points defining object boundaries and those constituting texture patterns. It is also insensitive to noise since the detection is based on the orientation assumed by the majority of the pixels in a neighborhood instead of the individual orientation or intensity variation.>
Yao Wang 0001, Sanjit K. Mitra
ICASSP1
1991 Motion/pattern adaptive interpolation of interlaced video sequences
abstract
A hybrid motion/pattern adaptive scheme for the conversion of interlaced video sequences to the progressive format is presented. It normally operates in the temporal interpolation model and switches to spatial interpolation when motion is detected. The filter for spatial interpolation is adapted according to the local image pattern. Directional filters are designed for oriented features based on their representation by oriented polynomials. Although the proposed scheme has only been applied to the interpolation of interlaced signals, its principle is applicable to other problems in video signal resolution enhancement.>
Yao Wang 0001, Sanjit K. Mitra
ICASSP1
1991 Image reconstruction from partial subband images and its application in packet video transmission
abstract
This paper addresses the problem of image reconstruction in a subband coding system when certain parts of one or several down-sampled sub-images are missing. By requiring that the sub-images produced from the reconstructed image be similar to those interpolated from the received sub-images, the loss recovery problem has been formulated as a quadratic optimization problem. Two reconstruction algorithms have been developed: a relaxational algorithm that achieves the optimal solution and a fast algorithm that leads to a sub-optimal solution. The interpolation scheme for the sub-images has been derived by characterizing each small image region by a texture or edge model. The proposed algorithms can be applied to any subband system and can accomodate various loss patterns. For the algorithm to work well, the analysis filters should have substantial overlap in their passbands such that the sub-images before down-sampling are correlated. Very good results have been obtained with some short kernel filter banks. The reconstructed image is satisfactory even when many parts of the low-low image is missing. It becomes unacceptable only if the lost regions contain certain periodic line structures which can cause Moiré patterns in the down-sampled sub-images. The results of our investigation suggest that signal loss problem such as packet loss can be combatted by using subband systems with overlapping filters. Although the coding efficiency is reduced compared to the conventional subband system using non-overlapping filters, the more disastrous signal loss can be prevented.
Yao Wang 0001, V. Ramamoorthy
Signal Process. Image Commun.1
1989 The recognition of shapes in binary images using a gradient classifier
abstract
The authors consider a prototype-based binary image classifier that makes comparisons based on blurred representations of the images. The blurring induces a metric on the space of all images that varies continuously under continuous deformation of the image plane. This blurred representation is suitable for direct implementation of a nearest-neighbor classifier. However, it is still desirable to have a representation which is invariant under certain spatial deformations, such as rotation, translation, and scaling of the image plane. A representation which is invariant under these transformation is produced by transforming an input to a local minimum of its distance from each prototype simultaneously. These minima are found by performing a gradient descent on an appropriate error surface over the transformation parameters. The error functional is the L/sub 2/-norm of the difference between the blurred prototype and the blurred input. The resulting classifier makes more efficient use of prototypes than does the nearest-neighbor classifier.>
Robert D. Brandt, Yao Wang 0001, Alan J. Laub, Sanjit K. Mitra
IEEE Trans. Syst. Man Cybern.2
1988 Edge-preserving image coding based on local modeling of images
abstract
A coding scheme which encodes each contiguous block of an image based on its local structure is developed. Three different types of models are constructed to characterize the image patterns occurring most frequently in a small block. The average bit rate is reduced by making use of the special constraint among the pixels in an oriented model, where the intensity variation is along only one direction. Judging from the image quality around edges, the reconstructed images with a bit rate less than 0.68 bit/pixel are perceptually superior to those obtained by the block and truncation coding at a bit rate of 1.64 bits/pixel. The complexity of this algorithm increases linearly with the size of the block and the number of orientations considered, which is much lower than that of the standard vector quantization techniques.>
Yao Wang 0001, Sanjit K. Mitra
ICASSP1
1988 A fast algorithm for the Fourier transform over finite fields and its VLSI implementation
abstract
The Fourier transform over finite fields is mainly required in the encoding and decoding of Reed-Solomon and BCH codes. An algorithm for computing the Fourier transform over any finite field GF(p/sup m/) is introduced. It requires only O(n(log n)/sup 2//4) additions and the same number of multiplications for an n-point transform and allows in some fields a further reduction of the number of multiplications to O(n log n). Because of its highly regular structure, this algorithm can be easily implementation by VLSI technology.>
Yao Wang 0001, Xuelong Zhu
IEEE J. Sel. Areas Commun.1