Zongming Guo

dblp:02/894 · DBLP profile ↗
← Back
208ranked-venue papers
0as first author
33since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 155 · 22 since 2021Artificial intelligence and machine learning · 19 · 3 since 2021Systems, architecture and hardware · 19Computer networks · 18 · 8 since 2021Databases, data management, data science and information retrieval · 8Security and privacy · 5 · 2 since 2021Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HybridPrompt: Bridging Generative Priors and Traditional Codecs for Mobile Streaming
abstract
In Video on Demand (VoD) scenarios, traditional codecs are the industry standard due to their high decoding efficiency. However, they suffer from severe quality degradation under low bandwidth conditions. While emerging generative neural codecs offer significantly higher perceptual quality, their reliance on heavy frame-by-frame generation makes real-time playback on mobile devices impractical. We ask: is it possible to combine the blazing-fast speed of traditional standards with the superior visual fidelity of neural approaches? We present HybridPrompt, the first generative-based video system capable of achieving real-time 1080p decoding at over 150 FPS on a commercial smartphone. Specifically, we employ a hybrid architecture that encodes Keyframes using a generative model while relying on traditional codecs for the remaining frames. A major challenge is that the two paradigms have conflicting objectives: the "hallucinated" details from generative models often misalign with the rigid prediction mechanisms of traditional codecs, causing bitrate inefficiency. To address this, we demonstrate that the traditional decoding process is differentiable, enabling an end-to-end optimization loop. This allows us to use subsequent frames as additional supervision, forcing the generative model to synthesize keyframes that are not only perceptually high-fidelity but also mathematically optimal references for the traditional codec. By integrating a two-stage generation strategy, our system outperforms pure neural baselines by orders of magnitude in speed while achieving an average LPIPS gain of 8% over traditional codecs at 200kbps.
Jiangkai Wu, Haoyang Wang 0015, Peiheng Wang, Zongming Guo, Xinggong Zhang
NOSSDAV5
2026 Syntax-Driven Multi-Realism Image Compression With Consistency Guided Diffusion Model
abstract
Given the challenge of balancing high fidelity with perceptual quality, multi-realism image compression is developed to adapt flexibly to varying requirements. It allows images with different levels of realism to be decoded from the same bit stream. Diffusion models are known for generating images with high perceptual quality. However, their inherent process of adding noise and denoising is often difficult to control and will bring more distortion. This limits their direct application in image compression, especially in multi-realism image compression which requires precise control to adapt to different requirements. To address this issue, we propose aConsistency Guided Diffusion Modelas a post-processing network for multi-realism image compression, aiming to control the addition of detail representations, thereby adjusting the trade-off between subjective quality and fidelity. In detail, our proposed novel method is crafted to introduce an additional consistency guided feature branch into the diffusion model to constrain the deviation caused by randomness in the diffusion process to ensure fidelity. Furthermore, a syntax-driven feature fusion module is constructed to guide the information adaptive fusion of two branches with an input extra ultra-low stream, which contains the context information and trade-off control information. In addition, we design a warm-up based training strategy and adopt a continuous online optimization method to improve coding efficiency and trade-off control precision. Extensive experiments validate the superiority of our method over existing compression techniques, as well as the effectiveness of each component.
Haowei Kuang, Wenhan Yang, Zongming Guo, Jiaying Liu 0001
IEEE Trans. Circuits Syst. Video Technol.3
2025 PromptMobile: Efficient Promptus for Low Bandwidth Mobile Video Streaming
abstract
Traditional video compression algorithms exhibit significant quality degradation at extremely low bitrates. Promptus emerges as a new paradigm for video streaming, substantially cutting down the bandwidth essential for video streaming. However, Promptus is computationally intensive and can not run in real-time on mobile devices. This paper presents PromptMobile, an efficient acceleration framework tailored for on-device Promptus. Specifically, we propose (1) a two-stage efficient generation framework to reduce computational cost by 8.1x, (2) a fine-grained inter-frame caching to reduce redundant computations by 16.6%, (3) system-level optimizations to further enhance efficiency. The evaluations demonstrate that compared with the original Promptus, PromptMobile achieves a 13.6x increase in image generation speed. Compared with other streaming methods, PromptMobile achives an average LPIPS improvement of 0.016 (compared with H.265), reducing 60% of severely distorted frames (compared to VQGAN).
Jiangkai Wu, Haoyang Wang 0015, Peiheng Wang, Xinggong Zhang, Zongming Guo
APNet6
2025 Modeling Virtual Reality Traffic with Head Movement in Remote Rendering
abstract
The proliferation of virtual reality (VR) content, particularly in resource-intensive applications, has been met by remote rendering to overcome local hardware limitations. Along with numerous advantages, remote rendering VR brings about a new traffic type that features huge throughput and burstiness generally, the understanding and modeling of which is critical for performing VR networking optimization to guarantee the Quality of Experience (QoE) of VR traffic transmission, including synthetic traffic generation and Network Slicing orchestrators. However, existing VR traffic modeling studies are limited in that they do not consider the impact of user interactions on VR traffic. In contrast, we carry out extensive traffic measurements in this paper, and discover that head movements actively affect the VR frame sizes generated. We analyze traffic features and further model the relationship between angular velocities and frame sizes quantitatively. A linear regressor is modeled to predict the frame size by considering history frame sizes and angular velocities jointly. We use Air Light VR (ALVR) to stream VR content in various scenarios, construct the datasets, and validate our model on top of them. The result shows that the our model is capable of reducing the 95% square prediction error by 18-30 compared to the state-of-the-art model. To the best of our knowledge, this is the first investigation into the intricate relationship between remote rendering VR traffic and head movement. Our dataset and results will be publicly available and reproducible.
Yihang Zhang 0007, Zhidong Jia, Li Jiang 0021, Qingyang Li 0010, Xinggong Zhang, Zongming Guo
ICC6
2025 Cross-Granularity Online Optimization with Masked Compensated Information for Learned Image Compression
Haowei Kuang, Wenhan Yang, Zongming Guo, Jiaying Liu 0001
ICCV3
2025 Sync5D: Novel View Synthesis from a Single Image with 5D Consistency
Junlin Hao, Yunpeng Tan, Jiangkai Wu, Peiheng Wang, Xinggong Zhang, Zongming Guo
PRCV (10)7
2024 SRFC: Scalable Radiance Fields Streaming with Planar Codec
abstract
Volumetric videos afford comprehensive and immer-sive viewing experiences with six degrees of freedom (6DoF) for navigation, allowing users to move freely within a three-dimensional space. Radiance fields (RF) is emerging to reproduce photorealistic 3D scenes, which achieves lighting consistency and realistic transfer between the real and virtual worlds. In this paper, we investigate how to deliver photorealistic volumetric video under the network bandwidth constraints and low-power device. We design a novel Scalable Radiance Fields Video Streaming, SRFC, to enable streaming RF video with planar codec. An Orthogonal Layered Depth Image (OLDI) mapping is introduced to map 3D to view-dependent 2D plane images, in order to reduce bitrates and decoding complexity by the commercial 2D codec. Moreover, to address the issue of deviation of the user viewpoint, we propose a dual-tiered structure in radiance fields and optimizes user's perceptual quality by adaptive bitrate and viewpoint decision. The evaluations demonstrate that SRFC is capable of reducing bandwidth requirements by 4 × and increasing decoding frame rate by nearly 3 × while maintaining satisfactory user perception of video quality.
Quanlu Jia, Haodan Zhang, Haoyang Wang 0015, Jiangkai Wu, Xinggong Zhang, Zongming Guo
ICC7
2024 Consistency Guided Diffusion Model with Neural Syntax for Perceptual Image Compression
abstract
Diffusion models show impressive performances in image generation with excellent perceptual quality. However, its tendency to introduce additional distortion prevents its direct application in image compression. To address the issue, this paper introduces a Consistency Guided Diffusion Model (CGDM) tailored for perceptual image compression, which integrates an end-to-end image compression model with a diffusion-based post-processing network, aiming to learn richer detail representations with less fidelity loss. In detail, the compression and post-processing networks are cascaded and a branch of consistency guided features is added to constrain the deviation in the diffusion process for better reconstruction quality. Furthermore, a Syntax driven Feature Fusion (SFF) module is constructed to take an extra ultra-low bitstream from the encoding end as input, guiding the adaptive fusion of information from the two branches. In addition, we design a globally uniform boundary control strategy with overlapped patches and adopt a continuous online optimization mode to improve both coding efficiency and global consistency. Extensive experiments validate the superiority of our method to existing perceptual compression techniques. Our project is publicly available at: https://ellisonkuang.github.io/CGDM.github.io/.
Haowei Kuang, Yiyang Ma, Wenhan Yang, Zongming Guo, Jiaying Liu 0001
ACM Multimedia4
2024 Cross-Task Knowledge Transfer for Semi-supervised Joint 3D Grounding and Captioning
abstract
3D visual grounding is a fundamental yet important task in multimedia understanding, which aims to locate a specific object in a complicated 3D scene semantically according to a text description. However, this task requires a large number of annotations of labeled text-object pairs for training, so the scarcity of annotated data has been a key obstacle in this task. To this end, this paper makes the first attempt to introduce and address a new semi-supervised setting, where only a few text-object labels are provided during training. Considering most scene data has no annotation, we explore a new solution for unlabeled 3D grounding by additionally training and transferring knowledge from a correlated task, i.e., 3D captioning. Our main insight is that 3D grounding and captioning are complementary and can be iteratively trained with unlabeled data to provide object and text contexts for each other with pseudo-label learning. Specifically, we propose a novel 3D Cross-Task Teacher-Student Framework (3D-CTTSF) for joint 3D grounding and captioning in the semi-supervised setting, where each branch contains parallel grounding and captioning modules. We first pre-train the two modules of the teacher branch with limited labeled data for warm-up. Then, we train the student branch to mimic the ability of the teacher model and iteratively update both branches with the unlabeled data. In particular, we transfer the learned knowledge between the grounding and captioning modules across two branches to generate and refine the pseudo-labels of unlabeled data for providing reliable supervision. To further improve the quality of the pseudo-labels, we design a cross-task pseudo-label generation scheme, filtering low-quality pseudo-labels at the detection, captioning, and grounding levels, respectively. Experimental results on various datasets show competitive performances in both tasks compared to previous fully- and weakly-supervised methods, demonstrating the proposed 3D-CTTSF can serve as an effective solution to overcome the data scarcity issue.
Daizong Liu, Zongming Guo, Wei Hu 0003
ACM Multimedia3
2024 RTCC: Enable End-to-end Sub-RTT Congestion Control for Next-generation Network
abstract
The advancement of next-generation networks such as 5G/6G and satellite systems has significantly increased available network bandwidth, while also exacerbating network burstiness. This surge presents a formidable challenge for congestion control (CC), a pivotal mechanism for achieving high bandwidth utilization and low latency by adjusting congestion windows or modifying sending rates. Traditional end-to-end CC algorithms fall short of optimality due to their reliance on congestion signals in acknowledgment packets, which introduce a delay of one round-trip time (RTT). In this paper, to mitigate end-to-end delayed feedback, we introduce a novel Real-Time Congestion Control (RTCC) algorithm that integrates machine learning with conventional model-based CC. RTCC employs a Multi-Layer Perceptron (MLP) to model network conditions and predict current congestion signals accurately. A tailored network model then utilizes these predictions to manage packet accumulation in the network bottleneck. Moreover, an online model-updating mechanism is proposed to adapt to diverse network environments. We integrate RTCC into QUIC and conduct comprehensive experiments in both emulated test-beds and real-world settings, including WiFi/4G/5G and cross-continent networks. The results demonstrate RTCC's efficacy, with up to a 32% increase in average throughput, a reduction in RTT by up to 21% compared to BBR V2, and the maintenance of fair bandwidth allocation.
Yihang Zhang 0007, Zhidong Jia, Qingyang Li 0010, Xinggong Zhang, Zongming Guo
SECON5
2024 Inferring Video Streaming Quality of Real-Time Communication Inside Network
abstract
Real-time video streaming is getting indispensable in people’s daily life, and poses heavy loads and stringent performance requirements on the network. For Internet Service Providers (ISPs), ensuring high-quality real-time video communication is a widely concerned issue. However, inferring the quality of real-time video streaming based on passively-collected network traffic is a great challenge due to limited information in the User Datagram Protocol (UDP) header and the encryption of the application-level protocol. In this paper, we propose IReaV-T to Infer Real-time Video streaming quality with a generalized Transformer, which understands the intrinsic state of the network and predicts the future real-time video quality. By applying novel embedding methods, IReaV-T could make full use of observed traffic features and distinguish different real-time video applications. Extensive comparative experiments demonstrate the effectiveness of IReaV-T, showing that IReaV-T could predict future real-time video quality with mean squared Video Multimethod Assessment Fusion (VMAF) score error less than 6.
Yihang Zhang 0007, Sheng Cheng 0002, Zongming Guo, Xinggong Zhang
IEEE Trans. Circuits Syst. Video Technol.3
2024 Toward Real-World Super Resolution With Adaptive Self-Similarity Mining
abstract
Despite efforts to construct super-resolution (SR) training datasets with a wide range of degradation scenarios, existing supervised methods based on these datasets still struggle to consistently offer promising results due to the diversity of real-world degradation scenarios and the inherent complexity of model learning. Our work explores a new route: integrating the sample-adaptive property learned through image intrinsic self-similarity and the universal knowledge acquired from large-scale data. We achieve this by uniting internal learning and external learning by an unrolled optimization process. With the merits of both, the tuned fully-supervised SR models can be augmented to broadly handle the real-world degradation in a plug-and-play style. Furthermore, to promote the efficiency of combining internal/external learning, we apply an attention-based weight-updating method to guide the mining of self-similarity, and various data augmentations are adopted while applying the exponential moving average strategy. We conduct extensive experiments on real-world degraded images and our approach outperforms other methods in both qualitative and quantitative comparisons. Our project is available at: https://github.com/ZahraFan/AdaSSR/.
Zejia Fan, Wenhan Yang, Zongming Guo, Jiaying Liu 0001
IEEE Trans. Image Process.3
2023 Collaborative Spatial-Temporal Distillation for Efficient Video Deraining
abstract
In this paper, we propose a novel knowledge distillation framework to improve the efficiency of deep networks for video deraining. The knowledge is transferred from a large-scale powerful teacher network to a compact efficient student network via the proposed collaborative spatial-temporal distillation framework. The framework is equipped with three collaboration schemes of different granularities that make use of spatial-temporal redundancy in a complementary way for better distillation performance. First, the spatial alignment module applies distillation constraints at different spatial scales to achieve better scale invariance in transferred knowledge. Second, the temporal alignment module traces both temporal status between teacher and student separately and collaboratively, to comprehensively utilize inter-frame information. Third, these two alignment modules interact through a spatial-temporal adaptor, where spatial-temporal knowledge is transferred in a unified framework. Extensive experiments demonstrate the superiority of our distillation framework as well as the effectiveness of each module. Our code is available at: https://github.com/HuYuzhang/Knowledge-Distillation.
Yuzhang Hu, Minghao Liu 0019, Wenhan Yang, Jiaying Liu 0001, Zongming Guo
ICME5
2023 An Improved Reversible Database Watermarking Method based on Histogram Shifting
abstract
Database watermarking is typically employed to address the issues of data theft, illegal replication, and copyright infringement that may arise during the sharing of databases. Unfortunately, the existing methods often cause permanent distortion to the original data, and it is challenging to strike a balance between the watermark embedding capacity and data distortion. Therefore, this paper proposes a reversible database watermarking method based on histogram shifting, rhombus prediction, and double embedding with high capacity and low distortion, called RPDE-HSW. By utilizing the rhombus prediction, we respectively constructed two prediction error histograms in each subgroup and expanded the watermark capacity through the adoption of double-layer embedding and single-bin embedding 2 bits. A scrambling algorithm is used to make the attribute value distribution more discretized, resulting in a sparse distribution of the database histogram. Subsequently, we optimized the selection rules for the watermark embedding carrier, effectively eliminating the redundant distortion caused by histogram shifting. Experimental results demonstrate that the proposed method achieves smaller data distortion and higher watermark embedding capacity, outperforming some other state-of-the-art works, and does not affect the classification results and data mining.
Cheng Li 0045, Xinhui Han, Wenfa Qi, Zongming Guo
IH&MMSec4
2023 Rebuffering but not Suffering: Exploring Continuous-Time Quantitative QoE by User's Exiting Behaviors
Sheng Cheng 0002, Xinggong Zhang, Zongming Guo
INFOCOM4
2023 QUTY: Towards Better Understanding and Optimization of Short Video Quality
abstract
Short video applications such as TikTok and Instagram have attracted tremendous attention recently. However, it is very limited for industry and academia to understand the user's Quality of Experience (QoE) on short video, let alone how to improve the QoE in short video streaming.
Haodan Zhang, Yixuan Ban, Zongming Guo, Zhimin Xu 0001, Yue Wang 0032, Xinggong Zhang
MMSys3
2023 ZGaming: Zero-Latency 3D Cloud Gaming by Image Prediction
abstract
In cloud gaming, interactive latency is one of the most important factors in users' experience. Although the interactive latency can be reduced through typical network infrastructures like edge caching and congestion control, the interactive latency of current cloud-gaming platforms is still far from users' satisfaction.
Jiangkai Wu, Yu Guan 0005, Qi Mao 0002, Yong Cui 0001, Zongming Guo, Xinggong Zhang
SIGCOMM5
2023 Reversible data hiding based on prediction-error value ordering and multiple-embedding
Wenfa Qi, Tong Zhang 0024, Xiaolong Li 0001, Bin Ma 0003, Zongming Guo
Signal Process.5
2023 RAM360: Robust Adaptive Multi-Layer 360$^\circ$ Video Streaming With Lyapunov Optimization
abstract
Viewport-adaptive streaming approaches are emerging as the most promising way to deliver high-quality 360 videos over mobile networks. However, the viewport prediction is only reliable within a short prediction window, i.e., a short playback buffer, which conicts with maintaining a long buffer to avoid playback rebuffering. To deal with this problem, we present RAM360, a Robust Adaptive Multi-layer 360 video streaming system, to ensure high viewport quality and low stall ratio concurrently. We make three technical contributions. First, we design a QoE-driven robust multi-layer streaming framework, where each chunk is encoded by multiple independent layers with different quality levels. The client dynamically decides which chunk and layer to be downloaded by their QoE contributions. Thus, the base-layer could be prefetched to avoid the risk of stalling while the viewport quality is improved by downloading enhancement layer. Second, we establish a novel QoE model to represent the quality of whole playback session, not that of chunk. It aims to maximize the overall QoE of playback session. Third, we introduce the Lyapunov optimization theory to solve the QoE optimization problem, which is an online algorithm with near-optimality solution. We demonstrate that RAM360 can significantly outperform the existing schemes regarding viewport quality, stall ratio, and QoE through extensive experiments with public datasets.
Haodan Zhang, Yixuan Ban, Zongming Guo, Xinggong Zhang
IEEE Trans. Multim.3
2023 Deep Inter Prediction with Error-Corrected Auto-Regressive Network for Video Coding
abstract
Modern codecs remove temporal redundancy of a video via inter prediction, i.e., searching previously coded frames for similar blocks and storing motion vectors to save bit-rates. However, existing codecs adopt block-level motion estimation, where a block is regressed by reference blocks linearly and is doomed to fail to deal with non-linear motions. In this article, we generate virtual reference frames (VRFs) with previously reconstructed frames via deep networks to offer an additional candidate, which is not constrained to linear motion structure and further significantly improves coding efficiency. More specifically, we propose a novel deep Auto-Regressive Moving-Average (ARMA) model, Error-Corrected Auto-Regressive Network (ECAR-Net), equipped with the powers of the conventional statistic ARMA models and deep networks jointly for reference frame prediction. Similar to conventional ARMA models, the ECAR-Net consists of two stages: Auto-Regression (AR) stage and Error-Correction (EC) stage, where the first part predicts the signal at the current time-step based on previously reconstructed frames, while the second one compensates for the output of the AR stage to obtain finer details. Different from the statistic AR models only focusing on short-term temporal dependency, the AR model of our ECAR-Net is further injected with the long-term dynamics mechanism, where long temporal information is utilized to help predict motions more accurately. Furthermore, ECAR-Net works in a configuration-adaptive way, i.e., using different dynamics and error definitions for the Low Delay B and Random Access configurations, which helps improve the adaptivity and generality in diverse coding scenarios. With the well-designed network, our method surpasses HEVC on average 5.0% and 6.6% BD-rate saving for the luma component under the Low Delay B and Random Access configurations and also obtains on average 1.54% BD-rate saving over VVC. Furthermore, ECAR-Net works in a configuration-adaptive way, i.e., using different dynamics and error definitions for the Low Delay B and Random Access configurations, which helps improve the adaptivity and generality in diverse coding scenarios.
Yuzhang Hu, Wenhan Yang, Jiaying Liu 0001, Zongming Guo
ACM Trans. Multim. Comput. Commun. Appl.4
2022 Self-Learned Video Super-Resolution with Augmented Spatial and Temporal Context
abstract
Video super-resolution methods typically rely on paired training data, in which the low-resolution frames are usually synthetically generated under predetermined degradation conditions (e.g., Bicubic downsampling). However, in real applications, it is labor-consuming and expensive to obtain this kind of training data, which limits the practical performance of these methods. To address the issue and get rid of the synthetic paired data, in this paper, we make exploration in utilizing the internal self-similarity redundancy within the video to build a Self-Learned Video Super-Resolution (SLVSR) method, which only needs to be trained on the input testing video itself. We employ a series of data augmentation strategies to make full use of the spatial and temporal context of the target video clips. The idea is applied to two branches of mainstream SR methods: frame fusion and frame recurrence methods. Since the former takes advantage of the short-term temporal consistency and the latter of the long-term one, our method can satisfy different practical situations. The experimental results show the superiority of our proposed method, especially in addressing the video super-resolution problems in real applications.
Zejia Fan, Jiaying Liu 0001, Wenhan Yang, Zongming Guo
ICASSP5
2022 Rain-Prior Injected Knowledge Distillation for Single Image Deraining
abstract
This paper makes efforts in improving the efficiency of deep networks for single image deraining with a newly proposed knowledge distillation framework. Specifically, we propose a rain-prior injected distillation scheme to transfer the knowledge from a large-scale teacher network to a more compact student network. Previous works directly calculate the distillation loss between the features extracted from the student and teacher networks. Differently, our distillation scheme adaptively removes the noisy background patterns by calculating the distillation loss based on the residual feature, which is inferred from the features extracted from the rain and ground truth images. This residual operation makes the student network focus on transferring only the knowledge on the rain streaks instead of the background, which facilitates more effective distillation results. Furthermore, our method can be applied to reduce both the network size and the deraining recurrence stage, which makes it a plug-and-play module that can be integrated into diverse existing deraining methods. Experimental results prove the efficiency of our method to build an efficient deraining network and the superiority over existing distillation methods.
Yuzhang Hu, Wenhan Yang, Jiaying Liu 0001, Zongming Guo
ICIP4
2022 DIG: A Data-Driven Impact-Based Grouping Method for Video Rebuffering Optimization
Shengbin Meng, Chunyu Qiao, Yue Wang 0032, Zongming Guo
MMM (2)5
2022 Research on Reversible Visible Watermarking Algorithms Based on Vectorization Compression Method
abstract
Abstract In current research on reversible visible watermarking algorithm, the original visible watermark image plays an important auxiliary role, and some algorithms also entirely depend on it to restore host image without any distortion. Therefore, in order to realize semi-blind reversible visible watermarking algorithm, the conventional reversible watermarking algorithm is used to embed compressed visible watermark image data into non-visible-watermarked region of host image. However, the amount of compressed image data obtained by conventional image compression algorithm is relatively large. Therefore, a method based on vectorization compression for the visible watermark image is proposed in this paper. Firstly, it performs edge detection on visible watermark image to obtain a discrete points set $\Gamma $ of vector contour curve. Then, the discrete points in $\Gamma $ are simplified by improved Douglas–Peucker algorithm, after that it obtains compressed vector contour data of visible watermark image. In addition, a reversible visible watermarking algorithm based on convolutional relief and image alpha fusion is proposed, which realizes reversible embedding of visible watermark image and lossless restoration of host image. The experimental results show that the proposed vectorization compression method has more advantages than traditional image compression algorithms, which greatly reduces the storage space of visible watermark image with high fidelity. Additionally, the embedded watermarking image has translucent 3D relief effect, and the fusion of host image and visible watermark image becomes more natural and harmonious.
Wenfa Qi, Sirui Guo, Yuxin Liu 0005, Xiang Wang 0009, Zongming Guo
Comput. J.5
2022 STC: FoV Tracking Enabled High-Quality 16K VR Video Streaming on Mobile Platforms
abstract
The ultra-high-definition 16K Virtual Reality (VR) video is coming to ages with more ”real” virtual experience and less cybersickness. However, the huge bitrate and decoding overhead would overwhelm today’s network and mobile hardware. The widely-known Field-of-View (FoV) adaptation streaming method still has severe bitrate wastes and decoding overhead as it delivers FoV areas with grid-like static tiles. Inspired by this, we present a novel ShiftTile-traCking (STC) streaming scheme, which crops and delivers tiles by tracking FoV movement. It is equivalent to deliver an FoV planar video instead of VR videos. This would save huge bit-rate and reduce decoding complexity. We mainly entail three contributions. 1) To reduce projection distortions, a novel FoV-centric sphere projection is proposed, which projects VR videos with the center of users’ FoVs. 2) To cover diverse FoV movement trajectories with a limited number of tiles, we propose an optimal tiling algorithm by trajectory clustering. 3) To be resilient to FoV prediction errors, we propose an accuracy-sensitive streaming algorithm, which scales FoV areas by the prediction accuracy. The evaluation shows that under the same network conditions, STC improves up to 1.3dB V-PSNR, reduces up to 13.2% buffering ratio, and achieves 60% faster decoding speed (61.5 frames per second) compared with state-of-the-art solutions.
Chengyuan Zheng, Jinyu Yin, Fangzhen Wei, Yu Guan 0005, Zongming Guo, Xinggong Zhang
IEEE Trans. Circuits Syst. Video Technol.5
2022 Learning to Recognize Human Actions From Noisy Skeleton Data Via Noise Adaptation
abstract
Recent studies have made great progress on skeleton-based action recognition. However, most of them are developed with relatively clean skeletons without the presence of intensive noise. We argue that the models learned from relatively clean data are not well generalizable to handle noisy skeletons commonly appeared in the real world. In this paper, we address the challenge of recognizing human actions from noisy skeletons, which is seldom explored by previous methods. Beyond exploring the new problem, we further take a new perspective to address it, \textit{i.e.}, noise adaptation, which gets rid of explicit skeleton noise modeling and reliance on skeleton ground truths. Specifically, we develop regression-based and generation-based adaptation models according to whether pairs of noisy skeletons are available. The regression-based model aims to learn noise-suppressed intrinsic feature representations by mapping pairs of noisy skeletons into a noise-robust space. When only unpaired skeletons are accessible, the generation-based model aims to adapt the features from noisy skeletons to a low-noise space by adversarial learning. To verify our proposed model and facilitate research on noisy skeletons, we collect a new dataset Noisy Skeleton Dataset (NSD), the skeletons of which are with much noise and more similar to daily-life data than previous datasets. Extensive experiments are conducted on the NSD, VV-RGBD and N-UCLA datasets, and results consistently show the outstanding performance of our proposed model.
Sijie Song, Jiaying Liu 0001, Lilang Lin, Zongming Guo
IEEE Trans. Multim.4
2021 Co-Grounding Networks With Semantic Attention for Referring Expression Comprehension in Videos
abstract
In this paper, we address the problem of referring expression comprehension in videos, which is challenging due to complex expression and scene dynamics. Unlike previous methods which solve the problem in multiple stages (i.e., tracking, proposal-based matching), we tackle the problem from a novel perspective, co-grounding, with an elegant one-stage framework. We enhance the single-frame grounding accuracy by semantic attention learning and improve the cross-frame grounding consistency with co-grounding feature learning. Semantic attention learning explicitly parses referring cues in different attributes to reduce the ambiguity in the complex expression. Co-grounding feature learning boosts visual feature representations by integrating temporal correlation to reduce the ambiguity caused by scene dynamics. Experiment results demonstrate the superiority of our framework on the video grounding datasets VID and LiOTB in generating accurate and stable results across frames. Our model is also applicable to referring expression comprehension in images, illustrated by the improved performance on the RefCOCO dataset. Our project is available at https://sijiesong.github.io/co-grounding.
Sijie Song, Xudong Lin 0003, Jiaying Liu 0001, Zongming Guo, Shih-Fu Chang
CVPR4
2021 LightFEC: Network Adaptive FEC with a Lightweight Deep-Learning Approach
abstract
Nowadays, the interest of real-time video streaming reaches a peak. To deal with the problem of packet loss and optimize users' Quality of Experience (QoE), Forward error correction (FEC) has been studied and applied extensively. The performance of FEC depends on whether the future loss pattern is precisely predicted, while the previous researches have not provided a robust packet loss prediction method. In this work, we propose LightFEC to make accurate and fast prediction of packet loss pattern. By applying long short-term memory (LSTM) networks, clustering algorithms and model compression methods, LightFEC is able to accurately predict packet loss in various network conditions without consuming too much time. According to the results of well-designed experiments, we find out that LightFEC outperforms other schemes on prediction accuracy, which improves the packet recovery ratio while keeping the redundancy ratio at a low level.
Sheng Cheng 0002, Xinggong Zhang, Zongming Guo
ACM Multimedia4
2021 Unpaired Person Image Generation With Semantic Parsing Transformation
abstract
In this paper, we tackle the problem of pose-guided person image generation with unpaired data, which is a challenging problem due to non-rigid spatial deformation. Instead of learning a fixed mapping directly between human bodies as previous methods, we propose a new pathway to decompose a single fixed mapping into two subtasks, namely, semantic parsing transformation and appearance generation. First, to simplify the learning for non-rigid deformation, a semantic generative network is developed to transform semantic parsing maps between different poses. Second, guided by semantic parsing maps, we render the foreground and background image, respectively. A foreground generative network learns to synthesize semantic-aware textures, and another background generative network learns to predict missing background regions caused by pose changes. Third, we enable pseudo-label training with unpaired data, and demonstrate that end-to-end training of the overall network further refines the semantic map prediction and final results accordingly. Moreover, our method is generalizable to other person image generation tasks defined on semantic maps, e.g., clothing texture transfer, controlled image manipulation, and virtual try-on. Experimental results on DeepFashion and Market-1501 datasets demonstrate the superiority of our method, especially in keeping better body shapes and clothing attributes, as well as rendering structure-coherent backgrounds.
Sijie Song, Wei Zhang 0031, Jiaying Liu 0001, Zongming Guo, Tao Mei 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2021 An Adaptive Visible Watermark Embedding Method based on Region Selection
abstract
Aiming at the problem that the robustness, visibility, and transparency of the existing visible watermarking technologies are difficult to achieve a balance, this paper proposes an adaptive embedding method for visible watermarking. Firstly, the salient region of the host image is detected based on superpixel detection. Secondly, the flat region with relatively low complexity is selected as the embedding region in the nonsalient region of the host image. Then, the watermarking strength is adaptively calculated by considering the gray distribution and image texture complexity of the embedding region. Finally, the visible watermark image is adaptively embedded into the host image with slight adjustment by just noticeable difference (JND) coefficient. The experimental results show that our proposed method improves the robustness of visible watermarking technology and greatly reduces the risk of malicious removal of visible watermark image. Meanwhile, a good balance between the visibility and transparency of the visible watermark image is achieved, which has the advantages of high security and ideal visual effect.
Wenfa Qi, Yuxin Liu 0005, Sirui Guo, Xiang Wang 0009, Zongming Guo
Secur. Commun. Networks5
2021 PrefCache: Edge Cache Admission With User Preference Learning for Video Content Distribution
abstract
With the deployment of video streaming in 4G/5G mobile network, Content Delivery Networks (CDN) are extending to the network edge to provide end-users better Quality of Experience (QoE). However, small cache size and irregular request patterns make it a great challenge for edge caching in video content distribution. Most of the existing cache policies are item-wise, they admit each video object separately, which performs poorly on the network edge due to irregular request patterns. We observe that compared with single video objects, users' preferences for video topics are much more constant, thus are easier to be predicted. So we propose PrefCache, a novel cache admission policy based on preference learning, for video content edge caching. PrefCache enables an edge cache to learn users' preferences for videos in real-time. Once receiving a video object, PrefCache decides whether to admit it to the cache by whether it is under users' preference. We make three contributions in this work. (1) First, we design an information collector, which can proactively collect the preference-related information without any modification of clients and video providers. (2) Second, we propose a tree-structure model to learn and compress users' preferences. (3) Third, to decide which videos should be admitted to the cache in real-time, an explore-and-exploit method is applied. We carried out extensive experiments with 24 hours of trace data from a large commercial video content provider. The experimental results demonstrate that PrefCache can improve hit ratio up to 12%, and save 92% memory / 98% CPU overhead, compared to the state-of-the-art cache policies.
Yu Guan 0005, Xinggong Zhang, Zongming Guo
IEEE Trans. Circuits Syst. Video Technol.3
2021 Predictive Generalized Graph Fourier Transform for Attribute Compression of Dynamic Point Clouds
abstract
As 3D scanning devices and depth sensors advance, dynamic point clouds have attracted increasing attention as a format for 3D objects in motion, with applications in various fields such as immersive telepresence, navigation for autonomous driving and gaming. Nevertheless, the tremendous amount of data in dynamic point clouds significantly burden transmission and storage. To this end, we propose a complete compression framework for attributes of 3D dynamic point clouds, focusing on optimal inter-coding. Firstly, we derive the optimal inter-prediction and predictive transform coding assuming the Gaussian Markov Random Field model with respect to a spatio-temporal graph underlying the attributes of dynamic point clouds. The optimal predictive transform proves to be the Generalized Graph Fourier Transform in terms of spatio-temporal decorrelation. Secondly, we propose refined motion estimation via efficient registration prior to inter-prediction, which searches the temporal correspondence between adjacent frames of irregular point clouds. Finally, we present a complete framework based on the optimal inter-coding and our previously proposed intra-coding, where we determine the optimal coding mode from rate-distortion optimization with the proposed offline-trained λ-Q model. Experimental results show that we achieve around 17% bit rate reduction on average over competitive dynamic point cloud compression methods.
Yiqun Xu, Wei Hu 0003, Shanshe Wang, Xinfeng Zhang 0001, Shiqi Wang 0001, Siwei Ma 0001, Zongming Guo, Wen Gao 0001
IEEE Trans. Circuits Syst. Video Technol.7
2021 Controllable Sketch-to-Image Translation for Robust Face Synthesis
abstract
In this paper, we propose a novel controllable sketch-to-image translation framework that allows users to interactively and robustly synthesize and edit face images with hand-drawn sketches. Inspired by the coarse-to-fine painting process of human artists, we propose a novel dilation-based sketch refinement method to refine sketches at varied coarse levels without the need for real sketch training data. We further investigate multi-level refinement that enables users to flexibly define how "reliable" the input sketch should be considered for the final output through a refinement level control parameter, which helps balance between the realism of the output and its structural consistency with the input sketch. It is realized by leveraging scale-aware style transfer to model and adjust the style features of sketches at different coarse levels. Moreover, advanced user controllability in terms of the editing region control, facial attribute editing, and spatially non-uniform refinement is further explored for fine-grained and semantic editing. We demonstrate the effectiveness of the proposed method in terms of visual quality and user controllability through extensive experiments including qualitative and quantitative comparison with state-of-the-art methods, ablation studies and various applications.
Shuai Yang 0001, Zhangyang Wang, Jiaying Liu 0001, Zongming Guo
IEEE Trans. Image Process.4
2020 Deep Plastic Surgery: Robust and Controllable Image Editing with Human-Drawn Sketches
Shuai Yang 0001, Zhangyang Wang, Jiaying Liu 0001, Zongming Guo
ECCV (15)4
2020 STC: Enabling 16K VR streaming on mobile platforms with FoV tracking
abstract
16K VR videos are coming to ages. But it could overwhelm mobile hardware for its huge bandwidth consumption and decoding complexity. To enable 16K VR video streaming over mobile platforms, we present a novel ShiftTile-Tracking (STC) streaming system, which crops and transmits video by tracking the Field-of-View (FoV) movement of users. The video chunk is split into ShiftTiles with frame granularity, which always covers FoV areas along the FoV movement trajectory. This transforms a 360-degree VR video into a traditional planar video, which leads to huge bandwidth saving and faster decoding speed. In the system design, we mainly entail two contributions. 1) To accommodate various FoV movement trajectories with a limited number of ShiftTiles, we propose an optimal tiling algorithm by FoV trajectory clustering. 2) To be resilient to the FoV prediction errors, we propose an accuracy-sensitive streaming algorithm, which expands the FoV area if the FoV prediction errors are high. The evaluation shows that under the same real-world 4G network conditions, the proposed STC improves 0. 9dBV-PSNR, reduces 12.4% buffering ratio, and achieves 45% faster decoding speed (64 frames per second) on average compared with the state-of-the-art solutions. This enables 16K VR video streaming on current mobile platforms.
Chengyuan Zheng, Jinyu Yin, Yu Guan 0005, Xinggong Zhang, Zongming Guo
GLOBECOM5
2020 MA360: Multi-Agent Deep Reinforcement Learning Based Live 360-Degree Video Streaming on Edge
abstract
The mobile edge caching has made video service providers deliver live 360-degree videos worldwide. However, these services still suffer from the huge network traffic on the core network due to the spherical nature and the diverse requests generated from large user populations. It is challenging to optimize the Quality of Experience (QoE) and the bandwidth consumption simultaneously under the significant number of users as well as dynamic network and playback status. In this paper, we propose a Multi-Agent deep reinforcement learning based 360-degree video streaming system, named MA360, to tackle this multi-user live 360-degree video streaming problem in the context of the edge cache network. Specifically, MA360 employs the Mean Field Actor-Critic (MFAC) algorithm to make clients collaboratively and distributively request tiles aiming at maximizing the overall QoE while minimizing the total bandwidth consumption. Experiments over real-world datasets show that MA360 can improve the QoE while significantly reducing the bandwidth consumption compared with several state-of-the-art edge-assisted 360-degree video streaming strategies.
Yixuan Ban, Yuanxing Zhang, Haodan Zhang, Xinggong Zhang, Zongming Guo
ICME5
2020 3d Dynamic Point Cloud Inpainting Via Temporal Consistency On Graphs
abstract
With the development of 3D laser scanning techniques and depth sensors, 3D dynamic point clouds have attracted increasing attention as a representation of 3D objects in motion, enabling various applications such as 3D immersive tele-presence, gaming and navigation. However, dynamic point clouds usually exhibit holes of missing data, mainly due to the fast motion, the limitation of acquisition and complicated structure. Leveraging on graph signal processing tools, we represent irregular point clouds on graphs and propose a novel inpainting method exploiting both intra-frame self-similarity and inter-frame consistency in 3D dynamic point clouds. Specifically, for each missing region in every frame of the point cloud sequence, we search for its self-similar regions in the current frame and corresponding ones in adjacent frames as references. Then we formulate dynamic point cloud inpainting as an optimization problem based on the two types of references, which is regularized by a graph-signal smoothness prior. Experimental results show the proposed approach outperforms three competing methods significantly, both in objective and subjective quality.
Zeqing Fu, Wei Hu 0003, Zongming Guo
ICME3
2020 Exploring Structure-Adaptive Graph Learning for Robust Semi-Supervised Classification
abstract
Graph Convolutional Neural Networks (GCNNs) are generalizations of CNNs to graph-structured data, in which convolution is guided by the graph topology. In many cases where graphs are unavailable, existing methods manually construct graphs or learn task-driven adaptive graphs. In this paper, we propose Graph Learning Neural Networks (GLNNs), which exploit the optimization of graphs (the adjacency matrix in particular) from both data and tasks. Leveraging on spectral graph theory, we propose the objective of graph learning from a sparsity constraint, properties of a valid adjacency matrix as well as a graph Laplacian regularizer via maximum a posteriori estimation. The optimization objective is then integrated into the loss function of the GCNN, which adapts the graph topology to not only labels of a specific task but also the input data. Experimental results show that our proposed GLNN significantly outperforms state-of-the-art approaches over widely adopted social network datasets and citation network datasets for semi-supervised classification.
Xiang Gao 0014, Wei Hu 0003, Zongming Guo
ICME3
2020 Exploring Hypergraph Representation On Face Anti-Spoofing Beyond 2d Attacks
abstract
Face anti-spoofing plays a crucial role in protecting face recognition systems from various attacks. Previous model-based and deep learning approaches achieve satisfactory performance for 2D face spoofs, but remain limited for more advanced 3D attacks such as vivid masks. In this paper, we address 3D face anti-spoofing via the proposed Hypergraph Convolutional Neural Networks (HGCNN). Firstly, we construct a computation-efficient and posture-invariant face representation with only a few key points on hypergraphs. The hypergraph representation is then fed into the designed HGCNN with hypergraph convolution for feature extraction, while the depth auxiliary is also exploited for 3D mask anti-spoofing. Further, we build a 3D face attack database with color, depth and infrared light information to validate the proposed paradigm and overcome the deficiency of 3D face anti-spoofing data. Experiments show that our method achieves the state-of-the-art performance over widely used 3D databases as well as the proposed one under various tests.
Gusi Te, Wei Hu 0003, Zongming Guo
ICME3
2020 DeepRS: Deep-Learning Based Network-Adaptive FEC for Real-Time Video Communications
abstract
As real-time multimedia streaming thriving, Forward Error Correction (FEC) methods have been studied and applied extensively these years. Most of researchers paid their attention to the coding algorithms, attempted to balance the trade off between recovery ratio and delay with fewer redundance. However, when packet loss pattern changes dynamically, the redundance waste is too serious to be ignored. In this work, we propose a novel algorithm which adjusts the redundance ratio of FEC encoder according to the prediction of packet loss. Receivers are additionally required to feedback observed packet loss pattern. Streaming sender collects the feedbacked packet loss pattern and predicts the number of packet loss in the incoming short period. As for implementation, we adopt long short-term memory (LSTM) network as our deep learning algorithm, and exquisitely embed it in our adaptive FEC system. With the extensive experiments, our proposed scheme outperforms other FEC methods greatly both in the simulations and evaluations on traces observed from the real world.
Sheng Cheng 0002, Xinggong Zhang, Zongming Guo
ISCAS4
2020 Prediction-Error Value Ordering for High-Fidelity Reversible Data Hiding
Tong Zhang 0024, Xiaolong Li 0001, Wenfa Qi, Zongming Guo
MMM (1)4
2020 APL: Adaptive Preloading of Short Video with Lyapunov Optimization
abstract
Short video applications, like TikTok, have attracted many users across the world. It can feed short videos based on users' preferences and allow users to slide the boring content anywhere and anytime. To reduce the loading time and keep playback smoothness, most of the short video apps will preload the recommended short videos in advance. However, these apps preload short videos in fixed size and fixed order, which can lead to huge playback stall and huge bandwidth waste. To deal with these problems, we present an Adaptive Preloading mechanism for short videos based on Lyapunov Optimization, also called APL, to achieve near-optimal playback experience, i.e., maximizing playback smoothness and minimizing bandwidth waste considering users' sliding behaviors. Specifically, we make three technical contributions: (1) We design a novel short video streaming framework which can dynamically preload the recommended short videos before the current video is downloaded completely. (2) We formulate the preloading problem into a playback experience optimization problem to maximize the playback smoothness and minimize the bandwidth waste. (3) We transform the playback experience optimization problem during the whole viewing process into a single-step greedy algorithm based on the Lyapunov optimization theory to make the online decisions during playback. Through extensive experiments based on the real datasets that generously provided by TikTok, we demonstrate that APL can reduce the stall ratio by 81%/12% and bandwidth waste by 11%/31% compared with no-preloading/fixed-preloading mechanism.
Haodan Zhang, Yixuan Ban, Xinggong Zhang, Zongming Guo, Zhimin Xu 0001, Shengbin Meng, Yue Wang 0032
VCIP4
2020 Joint Rain Detection and Removal from a Single Image with Contextualized Deep Networks
abstract
Rain streaks, particularly in heavy rain, not only degrade visibility but also make many computer vision algorithms fail to function properly. In this paper, we address this visibility problem by focusing on single-image rain removal, even in the presence of dense rain streaks and rain-streak accumulation, which is visually similar to mist or fog. To achieve this, we introduce a new rain model and a deep learning architecture. Our rain model incorporates a binary rain map indicating rain-streak regions, and accommodates various shapes, directions, and sizes of overlapping rain streaks, as well as rain accumulation, to model heavy rain. Based on this model, we construct a multi-task deep network, which jointly learns three targets: the binary rain-streak map, rain streak layers, and clean background, which is our ultimate output. To generate features that can be invariant to rain steaks, we introduce a contextual dilated network, which is able to exploit regional contextual information. To handle various shapes and directions of overlapping rain streaks, our strategy is to utilize a recurrent process that progressively removes rain streaks. Our binary map provides a constraint and thus additional information to train our network. Extensive evaluation on real images, particularly in heavy rain, shows the effectiveness of our model and architecture.
Wenhan Yang, Robby T. Tan, Jiashi Feng, Zongming Guo, Shuicheng Yan, Jiaying Liu 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2020 Optimal Reversible Data Hiding Scheme Based on Multiple Histograms Modification
abstract
Recently, a method based on multiple histograms modification (MHM) is proposed for reversible data hiding (RDH), in which a sequence of prediction-error histograms are generated and two expansion bins are selected in each histogram for expansion embedding. However, although efficient, it only chooses a single pair of expansion bins which limits the embedding capacity. On the other hand, the exhaustive expansion-bin-selection procedure in MHM takes huge computation time, so that it cannot be extended for high capacity RDH. In order to overcome the aforementioned drawbacks, an optimal RDH scheme based on MHM for high capacity embedding is proposed in this paper. First, to improve the embedding capacity, instead of a single pair of expansion bins, multiple pairs of expansion bins are utilized for each histogram, and the multiple-expansion-bin-selection for optimal embedding is formulated as an optimization problem. Then, unlike the exhaustive searching way used in MHM, a computationally efficient algorithm is proposed to solve the optimization problem, so that the optimal expansion bins can be adaptively determined to optimize the embedding performance. By the proposed approach, high embedding capacity can be achieved with good marked image quality, and the experimental results show that it is better than the original MHM and some other state-of-the-art methods.
Wenfa Qi, Xiaolong Li 0001, Tong Zhang 0024, Zongming Guo
IEEE Trans. Circuits Syst. Video Technol.4
2020 Location-Based PVO and Adaptive Pairwise Modification for Efficient Reversible Data Hiding
abstract
Pixel-value-ordering (PVO) is an efficient technique of reversible data hiding (RDH). By PVO, the maximum and minimum in each cover image block are first predicted and then modified to embed data. Actually, many PVO-based methods are essentially based on high-dimensional histogram modification. For these methods, a two-dimensional (2D) prediction-error histogram (PEH) is first generated and then modified based on a 2D mapping. However, these methods have two drawbacks. On one hand, the generated 2D PEH is irregular so that it is difficult to design suitable histogram modification strategy. On the other hand, the employed 2D mapping is empirically designed, and thus the embedding performance is far from optimal. Based on these considerations, a new PVO-based RDH scheme is proposed in this paper. By considering both pixel value orders and pixel locations, a new predictor is proposed so that the generated 2D PEH is regular in shape and suitable for reversible embedding. Moreover, instead of manually designing 2D mappings, to optimize the embedding performance, a self-learning mechanism is proposed to adaptively select the 2D mapping according to the image content. With the new predictor and the self-learning mechanism for 2D mapping selection, the proposed method works well with a good marked image quality, e.g., the PSNR of the image Lena is as high as 61.53 dB for an embedding capacity of 10 000 bits. Besides, compared with some state-of-the-art RDH methods, the superiority of the proposed method is experimentally verified.
Tong Zhang 0024, Xiaolong Li 0001, Wenfa Qi, Zongming Guo
IEEE Trans. Inf. Forensics Secur.4
2020 Modality Compensation Network: Cross-Modal Adaptation for Action Recognition
abstract
With the prevalence of RGB-D cameras, multimodal video data have become more available for human action recognition. One main challenge for this task lies in how to effectively leverage their complementary information. In this work, we propose a Modality Compensation Network (MCN) to explore the relationships of different modalities, and boost the representations for human action recognition. We regard RGB/ optical flow videos as source modalities, skeletons as auxiliary modality. Our goal is to extract more discriminative features from source modalities, with the help of auxiliary modality. Built on deep Convolutional Neural Networks (CNN) and Long Short Term Memory (LSTM) networks, our model bridges data from source and auxiliary modalities by a modality adaptation block to achieve adaptive representation learning, that the network learns to compensate for the loss of skeletons at test time and even at training time. We explore multiple adaptation schemes to narrow the distance between source and auxiliary modal distributions from different levels, according to the alignment of source and auxiliary data in training. In addition, skeletons are only required in the training phase. Our model is able to improve the recognition performance with source data when testing. Experimental results reveal that MCN outperforms stateof- the-art approaches on four widely-used action recognition benchmarks.
Sijie Song, Jiaying Liu 0001, Yanghao Li, Zongming Guo
IEEE Trans. Image Process.4
2020 Statistical Learning Based Congestion Control for Real-Time Video Communication
abstract
The existing congestion control is hard to simultaneously achieve low latency, high throughput, good adaptability and fair bandwidth allocation, mainly because of the hardwired control strategy and egocentric convergence objective. To address these issues, we propose an end-to-end statistical learning based congestion control, named Iris. By exploring the underlying principles of self-inflicted delay, we find that RTT variation is linearly related to the difference between sending rate and receiving rate, which inspires us to control video bit rate using a statistical-learning congestion control model. The key idea of Iris is to force all flows to converge to the same queue load and adjust bit rate by the model. All flows keep a small and fixed number of packets queuing in the network, thus the fair bandwidth allocation and low latency are both achieved. Besides, the adjustment step size of sending rate is updated by online learning, to better adapt to dynamically changing networks. We carried out extensive experiments to evaluate the performance of Iris, with the implementations over transport layer and application layer respectively. The testing environment includes emulated network, real-world Internet and commercial cellular networks. Compared against Transmission Control Protocol (TCP) flavors and state-of-the-art protocols, Iris is able to achieve high bandwidth utilization, low latency and good fairness concurrently. Especially for HyperText Transfer Protocol (HTTP) video streaming service, Iris is able to increase the video bitrate up to 25% and Peak Signal to Noise Ratio (PSNR) up to 1 dB.
Tongyu Dai, Xinggong Zhang, Yihang Zhang 0007, Zongming Guo
IEEE Trans. Multim.4
2019 TET-GAN: Text Effects Transfer via Stylization and Destylization
abstract
Text effects transfer technology automatically makes the text dramatically more impressive. However, previous style transfer methods either study the model for general style, which cannot handle the highly-structured text effects along the glyph, or require manual design of subtle matching criteria for text effects. In this paper, we focus on the use of the powerful representation abilities of deep neural features for text effects transfer. For this purpose, we propose a novel Texture Effects Transfer GAN (TET-GAN), which consists of a stylization subnetwork and a destylization subnetwork. The key idea is to train our network to accomplish both the objective of style transfer and style removal, so that it can learn to disentangle and recombine the content and style features of text effects images. To support the training of our network, we propose a new text effects dataset with as much as 64 professionally designed styles on 837 characters. We show that the disentangled feature representations enable us to transfer or remove all these styles on arbitrary glyphs using one network. Furthermore, the flexible network design empowers TET-GAN to efficiently extend to a new text style via oneshot learning where only one example is required. We demonstrate the superiority of the proposed method in generating high-quality stylized text over the state-of-the-art methods.
Shuai Yang 0001, Jiaying Liu 0001, Wenjing Wang 0001, Zongming Guo
AAAI4
2019 Typography With Decor: Intelligent Text Style Transfer
abstract
Text effects transfer can dramatically make the text visually pleasing. In this paper, we present a novel framework to stylize the text with exquisite decor, which is ignored by the previous text stylization methods. Decorative elements pose a challenge to spontaneously handle basal text effects and decor, which are two different styles. To address this issue, our key idea is to learn to separate, transfer and recombine the decors and the basal text effect. A novel text effect transfer network is proposed to infer the styled version of the target text. The stylized text is finally embellished with decor where the placement of the decor is carefully determined by a novel structure-aware strategy. Furthermore, we propose a domain adaptation strategy for decor detection and a one-shot training strategy for text effects transfer, which greatly enhance the robustness of our network to new styles. We base our experiments on our collected topography dataset including 59,000 professionally styled text and demonstrate the superiority of our method over other state-of-the-art style transfer methods.
Wenjing Wang 0001, Jiaying Liu 0001, Shuai Yang 0001, Zongming Guo
CVPR4
2019 Optimal Viewport-Adaptive 360-Degree Video Streaming Against Random Head Movement
abstract
Recently, a significant interest in 360-degree virtual reality (VR) video has been formed. However, a key problem is how to design a robust adaptive streaming approach and implement a practical system. The traditional streaming methods which are not sensitive to user's viewport could cause huge bandwidth budget with low video quality, while the viewport-adaptive schemes may be not accurate enough especially under random head movement. In this paper, we have designed an optimal viewport-adaptive 360-degree video streaming scheme, which is to maximize Quality of Experience (QoE) by predicting user's viewport with a probabilistic model, prefetching video segments into the buffer and replacing some unbefitting segments. In this way, continuous and smooth playback, high bandwidth utilization, low viewport prediction error and high peak signal-to-noise ratio in the viewport (V-PSNR) can be obtained. In order to deal with user's head movement, we have reduced prediction error rate by employing a probabilistic viewport prediction model, and a replacement strategy has been applied to update the downloaded segments when the user's viewport suddenly changes. To reduce smoothness loss, the segments' oscillation during playback has been considered. We also developed a prototype system with our method. The well-designed experiments provided numerous results which proved the better performance of our scheme.
Zhimin Xu 0001, Xinggong Zhang, Zongming Guo
ICC4
2019 UtilCache: Effectively and Practicably Reducing Link Cost in Information-Centric Network
abstract
Minimizing total link cost in Information-Centric Network (ICN) by optimizing content placement is challenging in both effectiveness and practicality. To attain better performance, upstream link cost caused by a cache miss should be considered in addition to content popularity. To make it more practicable, a content placement strategy is supposed to be distributed, adaptive, with low coordination overhead as well as low computational complexity. In this paper, we present such a content placement strategy, UtilCache, that is both effective and practicable. UtilCache is compatible with any cache replacement policy. When the cache replacement policy tends to maintain popular contents, UtilCache attains low link cost. In terms of practicality, UtilCache introduces little coordination overhead because of piggybacked collaborative messages, and its computational complexity depends mainly on content replacement policy, which means it can be O(1) when working with LRU. Evaluations prove the effectiveness of UtilCache, as it saves nearly 40% link cost more than current ICN design.
Lemei Huang, Yu Guan 0005, Xinggong Zhang, Zongming Guo
ICC4
2019 Generating Diverse and Descriptive Image Captions Using Visual Paraphrases
abstract
Recently there has been significant progress in image captioning with the help of deep learning. However, captions generated by current state-of-the-art models are still far from satisfactory, despite high scores in terms of conventional metrics such as BLEU and CIDEr. Human-written captions are diverse, informative and precise, but machine-generated captions seem to be simple, vague and dull. In this paper, aimed at improving diversity and descriptiveness characteristics of generated image captions, we propose a model utilizing visual paraphrases (different sentences describing the same image) in captioning datasets. We explore different strategies to select useful visual paraphrase pairs for training by designing a variety of scoring functions. Our model consists of two decoding stages, where a preliminary caption is generated in the first stage and then paraphrased into a more diverse and descriptive caption in the second stage. Extensive experiments are conducted on the benchmark MS COCO dataset, with automatic evaluation and human evaluation results verifying the effectiveness of our model.
Jiajun Tang 0001, Xiaojun Wan 0001, Zongming Guo
ICCV4
2019 Controllable Artistic Text Style Transfer via Shape-Matching GAN
abstract
Artistic text style transfer is the task of migrating the style from a source image to the target text to create artistic typography. Recent style transfer methods have considered texture control to enhance usability. However, controlling the stylistic degree in terms of shape deformation remains an important open challenge. In this paper, we present the first text style transfer network that allows for real-time control of the crucial stylistic degree of the glyph through an adjustable parameter. Our key contribution is a novel bidirectional shape matching framework to establish an effective glyph-style mapping at various deformation levels without paired ground truth. Based on this idea, we propose a scale-controllable module to empower a single network to continuously characterize the multi-scale shape features of the style image and transfer these features to the target text. The proposed method demonstrates its superiority over previous state-of-the-arts in generating diverse, controllable and high-quality stylized text.
Shuai Yang 0001, Zhangyang Wang, Ning Xu 0007, Jiaying Liu 0001, Zongming Guo
ICCV6
2019 Point Cloud Attribute Inpainting in Graph Spectral Domain
abstract
With the prevalence of depth sensors and 3D scanning devices, point clouds have attracted increasing attention as a format for 3D object representation, with applications in various fields such as tele-presence, navigation for autonomous driving and heritage reconstruction. However, point clouds usually exhibit holes of missing data, mainly due to the limitation of acquisition techniques and complicated structure. Hence, we propose an efficient inpainting method for the attribute (e.g., color) of point clouds, exploiting non-local self-similarity in graph spectral domain. Specifically, we represent irregular point clouds naturally on graphs, and split a point cloud into fixed-sized cubes as the processing unit. We then globally search for the most similar cubes to the target cube with holes inside, and compute the graph Fourier transform (GFT) basis from the similar cubes, which will be leveraged for the GFT representation of the target patch. We then formulate attribute inpainting as a sparse coding problem, imposing sparsity on the GFT representation of the attribute for hole filling. Experimental results demonstrate the superiority of our method.
Ju He, Zeqing Fu, Wei Hu 0003, Zongming Guo
ICIP4
2019 Feature Preserving and Uniformity-Controllable Point Cloud Simplification on Graph
abstract
With the development of 3D sensing technologies, point clouds have attracted increasing attention in a variety of applications for 3D object representation, such as autonomous driving, 3D immersive tele-presence and heritage reconstruction. However, it is challenging to process large-scale point clouds in terms of both computation time and storage due to the tremendous amounts of data. Hence, we propose a point cloud simplification algorithm, aiming to strike a balance between preserving sharp features and keeping uniform density during resampling. In particular, leveraging on graph spectral processing, we represent irregular point clouds naturally on graphs, and propose concise formulations of feature preservation and density uniformity based on graph filters. The problem of point cloud simplification is finally formulated as a trade-off between the two factors and efficiently solved by our proposed algorithm. Experimental results demonstrate the superiority of our method, as well as its efficient application in point cloud registration.
Junkun Qi, Wei Hu 0003, Zongming Guo
ICME3
2019 Optimized Skeleton-based Action Recognition via Sparsified Graph Regression
abstract
With the prevalence of accessible depth sensors, dynamic human body skeletons have attracted much attention as a robust modality for action recognition. Previous methods model skeletons based on RNN or CNN, which has limited expressive power for irregular skeleton joints. While graph convolutional networks (GCN) have been proposed to address irregular graph-structured data, the fundamental graph construction remains challenging. In this paper, we represent skeletons naturally on graphs, and propose a graph regression based GCN (GR-GCN) for skeleton-based action recognition, aiming to capture the spatio-temporal variation in the data. As the graph representation is crucial to graph convolution, we first propose graph regression to statistically learn the underlying graph from multiple observations. In particular, we provide spatio-temporal modeling of skeletons and pose an optimization problem on the graph structure over consecutive frames, which enforces the sparsity of the underlying graph for efficient representation. The optimized graph not only connects each joint to its neighboring joints in the same frame strongly or weakly, but also links with relevant joints in the previous and subsequent frames. We then feed the optimized graph into the GCN along with the coordinates of the skeleton sequence for feature learning, where we deploy high-order and fast Chebyshev approximation of spectral graph convolution. Further, we provide analysis of the variation characterization by the Chebyshev approximation. Experimental results validate the effectiveness of the proposed graph regression and show that the proposed GR-GCN achieves the state-of-the-art performance on the widely used NTU RGB+D, UT-Kinect and SYSU 3D datasets.
Xiang Gao 0014, Wei Hu 0003, Jiaxiang Tang, Jiaying Liu 0001, Zongming Guo
ACM Multimedia5
2019 CACA: Learning-based Content-aware Cache Admission for Video Content in Edge Caching
abstract
In the last decades, network caches (Content Distribution Network, CDN) have been widely deployed in video delivery system. As cache has been pushed to network edge as far as possible, small cache size and irregular request pattern make it a great challenge for edge cache to catch popular video contents. Although we can apply cache admission policies to block cold contents out, however, all current admission policies are still based on request pattern (content size, frequency), which perform poorly in edge cache. This paper proposes a novel feature-based cache admission policy, Content-feature Aware Cache Admission(CACA). It admits video objects to cache by video features, not by request pattern anymore. The intuition behind that is, for a group of users, their preferred contents may change at any time, but their preferred content features would maintain for a while. Popularity of video features (such as topic, author), is much more predicable than that of single video object. To mine critical features from huge feature space, this paper proposes a tree-structure reinforcement learning algorithm. Critical features are learned from a feature-partition tree which is spanned and pruned by history popularity. Then, an Exploration-and-Exploitation method is used to select the Top-K critical features. Video contents with these features will be admitted to cache. We carried out extensive experiments with 24-hours data traces from a commercial video content provider. The experimental results demonstrate that the proposed CACA is able to improve hit ratio up to 15%, reduce back-to-origin up to 20% and save 95% memory, compared with state-of-art cache admission policies.
Yu Guan 0005, Xinggong Zhang, Zongming Guo
ACM Multimedia3
2019 Learning Diachronic Word Embeddings with Iterative Stable Information Alignment
Zefeng Lin, Xiaojun Wan 0001, Zongming Guo
NLPCC (1)3
2019 PKU Paraphrase Bank: A Sentence-Level Paraphrase Corpus for Chinese
Weiwei Sun 0007, Xiaojun Wan 0001, Zongming Guo
NLPCC (1)4
2019 Pano: optimizing 360° video streaming with a better understanding of quality perception
abstract
Streaming 360° videos requires more bandwidth than non-360° videos. This is because current solutions assume that users perceive the quality of 360° videos in the same way they perceive the quality of non-360° videos. This means the bandwidth demand must be proportional to the size of the user's field of view. However, we found several quality-determining factors unique to 360° videos, which can help reduce the bandwidth demand. They include the moving speed of a user's viewpoint (center of the user's field of view), the recent change of video luminance, and the difference in depth-of-fields of visual objects around the viewpoint.
Yu Guan 0005, Chengyuan Zheng, Xinggong Zhang, Zongming Guo, Junchen Jiang
SIGCOMM4
2019 Reduced-reference quality assessment of image super-resolution by energy change and texture variation
Yuming Fang 0001, Jiaying Liu 0001, Yabin Zhang 0002, Weisi Lin, Zongming Guo
J. Vis. Commun. Image Represent.5
2019 Improved reversible visible image watermarking based on HVS and ROI-selection
Wenfa Qi, Guangyuan Yang, Tong Zhang 0024, Zongming Guo
Multim. Tools Appl.4
2019 QoE-Driven Adaptive K-Push for HTTP/2 Live Streaming
abstract
Dynamic adaptive streaming (DAS) over HTTP has been widely deployed over the Internet. However, due to the pull-based nature of HTTP/1.1, there exists intolerable streaming latency and high request overhead in the current DAS systems. With dynamic k-push, HTTP/2 live streaming promises to achieve low live latency with less overhead and small segment duration. In this paper, we propose a quality of experience (QoE) driven adaptive k-push mechanism (QK-Push) for HTTP/2 live streaming. The client just sends one request to set push length ($K$ ) and bitrate (v) parameters and the server would push back $K$ segments in a batch. To determine k-push parameters, a probabilistic buffer model is first designed to avoid buffer underflow/overflow. Also, three QoE objective functions are designed to ensure the high streaming quality (bitrate), playback continuity, and smoothness. QK-Push casts this multi-objective optimization problem as a Pareto optimal problem. To solve it, a Nash bargaining solution is designed to balance the needs for video quality, bitrate smoothness, and request overhead. Finally, the segments in each push cycle are selected by solving the Nash problem with a discrete space Lagrangian method. We implement an HTTP/2 live streaming prototype system, with the QK-Push algorithm over modified dash.js and media presentation description. To evaluate the performances, the extensive live streaming experiments are carried out over a controllable network test bed and real Internet trace. The results demonstrate that the proposed QK-Push algorithm is able to improve the average bitrate up to 13%, reduce the bitrate oscillations up to 81%, decrease the startup delay up to 58%, and increase the estimate the mean opinion score up to 12% compared to the current HTTP/1.1 system.
Zhimin Xu 0001, Xinggong Zhang, Zongming Guo
IEEE Trans. Circuits Syst. Video Technol.3
2019 Reference-Guided Deep Super-Resolution via Manifold Localized External Compensation
abstract
The rapid development of social network and online multimedia technology makes it possible to address traditional image and video enhancement problems, with the aid of online similar reference data. In this paper, we tackle the problem of super-resolution (SR) in this way, specifically aiming to handle the “one-to-many” problem between the image patches of low resolution (LR) and high resolution (HR). We propose a manifold localized deep external compensation (MALDEC) network to additionally utilize reference images, i. e., retrieved similar images in cloud database and reference HR frame in a video, to provide an accurate localization and mapping to the HR manifold, and compensate the lost high-frequency details. The proposed network employs a three-step recovery: 1) internal structure inference, which uses the LR image itself and the internally inferred high frequency information to preserve main structure of the HR image; 2) manifold localization, which localizes the HR manifold and constructs the correspondence between the internal inferred image and the external images; and 3) external compensation, which introduces the external references of retrieved similar patches based on manifold localization information to reconstruct the high-frequency details. The learnable components of MALDEC, internal structure inference, and external compensation, are trained jointly to make a good tradeoff between these two terms for an optimal SR result. Finally, the proposed method is examined under three tasks: cloud-based image SR, multi-pose face reconstruction, and reference frame-guided video SR. Extensive experiments demonstrate the superiority of our method than the state-of-the-art SR methods in both objective and subjective evaluations, and our method offers new state-of-the-art performance.
Wenhan Yang, Sifeng Xia, Jiaying Liu 0001, Zongming Guo
IEEE Trans. Circuits Syst. Video Technol.4
2019 TFDASH: A Fairness, Stability, and Efficiency Aware Rate Control Approach for Multiple Clients Over DASH
abstract
Dynamic adaptive streaming over HTTP (DASH) has recently been widely deployed in the Internet and adopted in the industry. It, however, does not impose any adaptation logic for selecting the quality of video segments requested by clients and suffers from lackluster performance with respect to a number of desirable properties: efficiency, stability, and fairness when multiple players compete for a bottleneck link. In this paper, we propose a throughput-friendly DASH rate control scheme for video streaming with multiple clients over DASH to well balance the tradeoffs among efficiency, stability, and fairness. The core idea behind guaranteeing fairness and high efficiency (bandwidth utilization) is to avoid OFF periods during the downloading process for all clients, i.e., the bandwidth is in perfect-subscription or over-subscription with bandwidth utilization approach to 100%. We also propose a dual-threshold buffer model to solve the instability problem caused by the above idea. As a result, by integrating these novel components, we also propose a probability-driven rate adaption logic taking into account several key factors that most influence visual quality, including buffer occupancy, video playback quality, video bit-rate switching frequency and amplitude, to guarantee high-quality video streaming. Our experiments evidently demonstrate the superior performance of the proposed method.
Chao Zhou 0003, Chia-Wen Lin, Xinggong Zhang, Zongming Guo
IEEE Trans. Circuits Syst. Video Technol.4
2019 Local Frequency Interpretation and Non-Local Self-Similarity on Graph for Point Cloud Inpainting
abstract
As 3D scanning devices and depth sensors mature, point clouds have attracted increasing attention as a format for 3D object representation, with applications in various fields such as tele-presence, navigation and heritage reconstruction. However, point clouds usually exhibit holes of missing data, mainly due to the limitation of acquisition techniques and complicated structure. Further, point clouds are defined on irregular non- Euclidean domains, which is challenging to address especially with conventional signal processing tools. Hence, leveraging on recent advances in graph signal processing, we propose an efficient point cloud inpainting method, exploiting both the local smoothness and the non-local self-similarity in point clouds. Specifically, we first propose a frequency interpretation in graph nodal domain, based on which we derive the smoothing and denoising properties of a graph-signal smoothness prior in order to describe the local smoothness of point clouds. Secondly, we explore the characteristics of non-local self-similarity, by globally searching for the most similar area to the missing region. The similarity metric between two areas is defined based on the direct component and the anisotropic graph total variation of normals in each area. Finally, we formulate the hole-filling step as an optimization problem based on the selected most similar area and regularized by the graph-signal smoothness prior. Besides, we propose voxelization and automatic hole detection methods for the point cloud prior to inpainting. Experimental results show that the proposed approach outperforms four competing methods significantly, both in objective and subjective quality.
Wei Hu 0003, Zeqing Fu, Zongming Guo
IEEE Trans. Image Process.3
2019 D3R-Net: Dynamic Routing Residue Recurrent Network for Video Rain Removal
abstract
In this paper, we address the problem of video rain removal by considering rain occlusion regions, i.e., very low light transmittance for rain streaks. Different from additive rain streaks, in such occlusion regions, the details of backgrounds are completely lost. Therefore, we propose a hybrid rain model to depict both rain streaks and occlusions. Integrating the hybrid model and useful motion segmentation context information, we present a Dynamic Routing Residue Recurrent Network (D3R-Net). D3R-Net first extracts the spatial features by a residual network. Then, the spatial features are aggregated by recurrent units along the temporal axis. In the temporal fusion, the context information is embedded into the network in a "dynamic routing" way. A heap of recurrent units takes responsibility for handling the temporal fusion in given contexts, e.g., rain or non-rain regions. In the certain forward and backward processes, one of these recurrent units is mainly activated. Then, a context selection gate is employed to detect the context and select one of these temporally fused features generated by these recurrent units as the final fused feature. Finally, this last feature plays a role of "residual feature." It is combined with the spatial feature and then used to reconstruct the negative rain streaks. In such a D3R-Net, we incorporate motion segmentation, which denotes whether a pixel belongs to fast moving edges or not, and rain type indicator, indicating whether a pixel belongs to rain streaks, rain occlusions, and non-rain regions, as the context variables. Extensive experiments on a series of synthetic and real videos with rain streaks verify not only the superiority of the proposed method over state of the art but also the effectiveness of our network design and its each component.
Jiaying Liu 0001, Wenhan Yang, Shuai Yang 0001, Zongming Guo
IEEE Trans. Image Process.4
2019 Context-Aware Text-Based Binary Image Stylization and Synthesis
abstract
In this work, we present a new framework for the stylization of text-based binary images. First, our method stylizes the stroke-based geometric shape like text, symbols and icons in the target binary image based on an input style image. Second, the composition of the stylized geometric shape and a background image is explored. To accomplish the task, we propose legibilitypreserving structure and texture transfer algorithms, which progressively narrow the visual differences between the binary image and the style image. The stylization is then followed by a contextaware layout design algorithm, where cues for both seamlessness and aesthetics are employed to determine the optimal layout of the shape in the background. Given the layout, the binary image is seamlessly embedded into the background by texture synthesis under a context-aware boundary constraint. According to the contents of binary images, our method can be applied to many fields.We show that the proposed method is capable of addressing the unsupervised text stylization problem and is superior to stateof- the-art style transfer methods in automatic artistic typography creation. Besides, extensive experiments on various tasks, such as visual-textual presentation synthesis, icon/symbol rendering and structure-guided image inpainting, demonstrate the effectiveness of the proposed method.
Shuai Yang 0001, Jiaying Liu 0001, Wenhan Yang, Zongming Guo
IEEE Trans. Image Process.4
2019 Scale-Free Single Image Deraining Via Visibility-Enhanced Recurrent Wavelet Learning
abstract
In this paper, we address a rain removal problem from a single image, even in the presence of large rain streaks and rain streak accumulation (where individual streaks cannot be seen, and thus visually similar to mist or fog). For rain streak removal, the mismatch problem between different streak sizes in training and testing phases leads to a poor performance, especially when there are large streaks. To mitigate this problem, we embed a hierarchical representation of wavelet transform into a recurrent rain removal process: 1) rain removal on the low-frequency component; 2) recurrent detail recovery on highfrequency components under the guidance of the recovered lowfrequency component. Benefiting from the recurrent multi-scale modeling of wavelet transform-like design, the proposed network trained on streaks with one size can adapt to those with larger sizes, which significantly favors real rain streak removal. The dilated residual dense network is used as the basic model of the recurrent recovery process. The network includes multiple paths with different receptive fields, thus can make full use of multi-scale redundancy and utilize context information in large regions. Furthermore, to handle heavy rain cases where rain streak accumulation is presented, we construct a detail appearing rain accumulation removal to not only improve the visibility but also enhance the details in dark regions. The evaluation on both synthetic and real images, particularly on those containing large rain streaks and heavy accumulation, shows the effectiveness of our novel models, which significantly outperforms the state-ofthe- art methods.
Wenhan Yang, Jiaying Liu 0001, Shuai Yang 0001, Zongming Guo
IEEE Trans. Image Process.4
2018 Erase or Fill? Deep Joint Recurrent Rain Removal and Reconstruction in Videos
abstract
In this paper, we address the problem of video rain removal by constructing deep recurrent convolutional networks. We visit the rain removal case by considering rain occlusion regions, i.e. the light transmittance of rain streaks is low. Different from additive rain streaks, in such rain occlusion regions, the details of background images are completely lost. Therefore, we propose a hybrid rain model to depict both rain streaks and occlusions. With the wealth of temporal redundancy, we build a Joint Recurrent Rain Removal and Reconstruction Network (J4R-Net) that seamlessly integrates rain degradation classification, spatial texture appearances based rain removal and temporal coherence based background details reconstruction. The rain degradation classification provides a binary map that reveals whether a location is degraded by linear additive streaks or occlusions. With this side information, the gate of the recurrent unit learns to make a trade-off between rain streak removal and background details reconstruction. Extensive experiments on a series of synthetic and real videos with rain streaks verify the superiority of the proposed method over previous state-of-the-art methods.
Jiaying Liu 0001, Wenhan Yang, Shuai Yang 0001, Zongming Guo
CVPR4
2018 Soft Decoding of Light Field Images Using Pocs and Fast Graph Spectrayl Filters
abstract
Light field data captured by a lenslet-based image sensor is typically demosaicked, aligned and rearranged into a series of sub-aperture (viewpoint) images, before a disparity-compensated coding scheme is employed for compression. In this paper, we focus on the problem of soft decoding of block-based compressed sub-aperture images at the decoder: given quantization bin indices of DCT coefficients of non-overlapping code blocks, we select appropriate coefficient values that are low-pass filtered using graph spectral filters and view-consistent across sub-aperture images via projection on convex sets (POCS). Specifically, after an initial pixel estimate, we low-pass filter each pixel block using accelerated graph filters based on the Lanczos method. We then map filtered pixels to a neighborhood of sub-aperture images based on estimated disparity to enforce indexed quantization bin constraints of multiple images. Experimental results show that our algorithm achieves PSNR gain of 2.34dB over JPEG hard decoding.
Shuai Yang 0001, Gene Cheung, Jiaying Liu 0001, Zongming Guo
ICASSP4
2018 Name-Based Routing with On-Path Name Lookup in Information-Centric Network
abstract
Name-based routing is one of the core ideas in Information-centric network (ICN). In name-based routing, there is a tradeoff between the cost of name announcement and name lookup. Some ICN architectures introduce an efficient way of name lookup but pay high price in name announcement, others cut off most information exchange in name announcement yet introduce heavy burden in name lookup. In order to solve this problem and balance the cost of name announcement and lookup, we propose Name-based routing with On-Path Name Lookup (OPNL). OPNL looks up name prefixes on the path to name's guaranteed destination. It accomplishes distributed name lookup with lighter burden while maintaining little information exchange in name announcement. Results of simulation experiments show that OPNL makes a tradeoff between the cost of name announcement and lookup to have better scalability, eliminates storage overhead and communication overhead compared with prior works and attains even better performance.
Yu Guan 0005, Lemei Huang, Xinggong Zhang, Zongming Guo
ICC4
2018 Point Cloud Inpainting on Graphs from Non-Local Self-Similarity
abstract
As 3D scanning devices and depth sensors advance, point clouds have attracted increasing attention as a format for 3D object representation, with applications in various fields such as tele-presence, navigation and heritage reconstruction. However, point clouds usually exhibit holes of missing data, mainly due to the limitation of acquisition techniques and complicated structure. Hence, we propose an efficient point cloud inpainting method, leveraging on graph signal processing and based on the observation of non-local self-similarity in point clouds. Specifically, we split a point cloud into fixed-size cubes as the processing unit, and globally search for the most similar cube to the target cube with holes inside. The similarity metric between two cubes is defined based on the direct component and the proposed anisotropic graph total variation of normals in each cube. We then formulate the hole-filling step as an optimization problem, based on the selected most similar cube and regularized by a graph-signal smoothness prior. Experimental results show that the proposed approach outperforms three competing methods significantly, both in objective and subjective quality.
Zeqing Fu, Wei Hu 0003, Zongming Guo
ICIP3
2018 Restoration of Unevenly Illuminated Images
abstract
In this paper, we tackle the problem of restoring unevenly illuminated images. Generally, there exist three kinds of exposure conditions in these images: under-, normal-, and over-exposures. Thus, a three-component generalized Gaussian mixture model (3GGMM) is used to fit the histogram of the illuminance image, and probabilistically characterize the three exposure states. Based on the 3GGMM, separate optimal tone mapping functions are designed to enhance under- and overexposed regions by maximizing expected contrast of these regions. The output illumination can be obtained by fusing the restoration results in different exposure states. Experimental results validate the effectiveness of the proposed image restoration approach.
Mading Li, Jiaying Liu 0001, Zongming Guo
ICIP4
2018 CUB360: Exploiting Cross-Users Behaviors for Viewport Prediction in 360 Video Adaptive Streaming
abstract
To ensure 360-degree video's continuous playback and reduce the bandwidth waste, predicting user's future fixation is indispensable. However, existing methods concentrate either on user's motion information or content information. None of them consider users watching behaviors' inconsistency which embodies user's attention distribution more explicitly. So in this paper, we exploit Cross-Users Behaviors for viewport prediction in 360-degree video adaptive streaming, namely CUB360, trying to concurrently consider user's personalized information and cross-users behaviors information to predict future viewport. Besides, we use a QoE-driven framework to optimize existing video streaming approaches and propose a general algorithm aiming at solving the NP problem at a low complexity. Extensive experimental results over real datasets demonstrate that compared with traditional adaptive streaming method, our proposal can significantly boost the prediction accuracy by 20.2% absolutely and 48.1 % relatively. Besides, the mean quality can get 30.28% gain while quality variance can be reduced by 29.89%.
Yixuan Ban, Lan Xie, Zhimin Xu 0001, Xinggong Zhang, Zongming Guo, Yue Wang 0032
ICME5
2018 Learning-based Congestion Control for Internet Video Communication over Wireless Networks
abstract
With the deployment of real-time video applications and wireless networks, the real-time congestion control becomes a hot topic. Most existing congestion control algorithms are not designed for low-latency real-time flows, or perform poorly in the face of highly variable channel capacities. In this paper, we proposed a novel Learning-based Congestion Control (LCC) for real-time video communication over wireless networks. The key idea of LCC is employing Kernel Density Estimation for one-way delay and sending rate to capture the underlying information about channel state. Then LCC bases on the estimated probability density and Bayesian theorem to quickly adapt sending rate to the changing channel. We implemented LCC in WebRTC framework and extensive experiments were carried out. Compared with the native WebRTC congestion control (GCC), experimental results show that LCC achieves higher channel utilization, even more than 4.2× throughput in lossy links. LCC is also much better at adapting to the variable channel than GCC. Besides, LCC performs well in delay constraint and intra-protocol fairness.
Tongyu Dai, Xinggong Zhang, Zongming Guo
ISCAS3
2018 Dual Recovery Network with Online Compensation for Image Super-Resolution
abstract
Image super-resolution (SR) methods essentially lead to a loss of some high-frequency (HF) information when predicting high-resolution (HR) images from low-resolution (LR) images without using external references. To address this issue, we additionally utilize online retrieved data to facilitate image SR in a unified deep framework. A novel dual high-frequency recovery network (DHN) is proposed to predict an HR image with three parts: an LR image, an internal inferred HF (IHF) map (HF missing part inferred solely from the LR image) and an external extracted HF (EHF) map. In particular, we infer the HF information based on both the LR image and similar HR references which are retrieved online. For the EHF map, we align the references with affine transformation and then in the aligned references, part of HF signals are extracted by the proposed DHN to compensate for the HF loss. Extensive experimental results demonstrate that our DHN achieves notably better performance than state-of-the-art SR methods.
Sifeng Xia, Wenhan Yang, Jiaying Liu 0001, Zongming Guo
ISCAS4
2018 Probabilistic Viewport Adaptive Streaming for 360-degree Videos
abstract
Recently, there has been a significant interest towards 360-degree virtual reality (VR) video. However, it is a big challenge for them to stream over Internet for huge bit-rates. In this paper, we have designed a novel viewport adaptive streaming scheme for 360-degree videos with probabilistic viewport prediction and optimal segments prefetching by Dynamic Adaptive Streaming over HTTP (DASH). In this way, continuous and smooth video playback, low viewport prediction error and high PSNR are obtained. To avoid head-movement prediction error, a probabilistic viewport prediction model is proposed, which leverages the probability distribution of user's orientation. Further, an optimal segments prefetching method is implemented. Finally, we also implement our method in a real system. The numerous experiment results have demonstrated that the proposed method has achieved significant performance gains compared with the existing methods. Our related work also win the Runner-up in ICME 2017 DASH-IF Grand Challenge: Dynamic Adaptive Streaming over HTTP.
Zhimin Xu 0001, Xinggong Zhang, Kai Zhang 0007, Zongming Guo
ISCAS4
2018 Gradient Based Interpolation for Intra Angular Prediction in HEVC
abstract
In the High Efficiency Video Coding (HEVC), the predicted pixels generated by the intra angular prediction are the same along the prediction direction. In this paper, an improved gradient based interpolation for intra angular prediction is proposed. The gradient is generated by both row and column reference samples according to the prediction direction and changed dynamically for each pixel. It is appended to the original intra prediction process to improve the performance of interpolation. This method describes the features of directional gradient changes, which improves the performance of intra interpolation. The method achieves up to 1.90% BD-Rate reduction and ignorable decoding time increase under intra main configuration based on HM 16.7.
Yushan Zheng, Jun Sun 0012, Zongming Guo
ISCAS4
2018 Images2Poem: Generating Chinese Poetry from Image Streams
abstract
Natural language generation from visual inputs has attracted extensive research attention recently. Generating poetry from visual content is an interesting but very challenging task. We propose and address the new multimedia task of generating classical Chinese poetry from image streams. In this paper, we propose an Images2Poem model with a selection mechanism and an adaptive self-attention mechanism for the problem. The model first selects representative images to summarize the image stream. During decoding, it adaptively pays attention to the information from either source-side image stream or target-side previously generated characters. It jointly summarizes the images and generates relevant, high-quality poetry from image streams. Experimental results demonstrate the effectiveness of the proposed approach. Our model outperforms baselines in different human evaluation metrics.
Xiaojun Wan 0001, Zongming Guo
ACM Multimedia3
2018 RGCNN: Regularized Graph CNN for Point Cloud Segmentation
abstract
Point cloud, an efficient 3D object representation, has become popular with the development of depth sensing and 3D laser scanning techniques. It has attracted attention in various applications such as 3D tele-presence, navigation for unmanned vehicles and heritage reconstruction. The understanding of point clouds, such as point cloud segmentation, is crucial in exploiting the informative value of point clouds for such applications. Due to the irregularity of the data format, previous deep learning works often convert point clouds to regular 3D voxel grids or collections of images before feeding them into neural networks, which leads to voluminous data and quantization artifacts. In this paper, we instead propose a regularized graph convolutional neural network (RGCNN) that directly consumes point clouds. Leveraging on spectral graph theory, we treat features of points in a point cloud as signals on graph, and define the convolution over graph by Chebyshev polynomial approximation. In particular, we update the graph Laplacian matrix that describes the connectivity of features in each layer according to the corresponding learned features, which adaptively captures the structure of dynamic graphs. Further, we deploy a graph-signal smoothness prior in the loss function, thus regularizing the learning process. Experimental results on the ShapeNet part dataset show that the proposed approach significantly reduces the computational complexity while achieving competitive performance with the state of the art. Also, experiments show RGCNN is much more robust to both noise and point cloud density in comparison with other methods. We further apply RGCNN to point cloud classification and achieve competitive results on ModelNet40 dataset.
Gusi Te, Wei Hu 0003, Amin Zheng, Zongming Guo
ACM Multimedia4
2018 CLS: A Cross-user Learning based System for Improving QoE in 360-degree Video Adaptive Streaming
abstract
Viewport adaptive streaming is emerging as a promising way to deliver high quality 360-degree video. It is still a critical issue to predict user's viewpoint and deliver partial video within the viewport. Current widely-used motion-based or content-saliency methods have low precision, especially for long-term prediction. In this paper, benefiting from data-driven learning, we propose a Cross-user Learning based System (CLS) to improve the precision of viewport prediction. Since users have similar region-of-interest (ROI) when watching a same video, it is possible to exploit cross-users' ROI behavior to predict viewport. We use a machine learning algorithm to group users according to historical fixations, and predict the viewing probability by the class. Additionally, we present a QoE-driven rate allocation to minimize the expected streaming distortion under bandwidth constraint, and give a Multiple-Choice Knapsack solution. Experiments demonstrate that CLS provides 2dB quality improvement than full-image streaming and 1.5 dB quality improvement than linear regression (LR) method. On average, the precision of viewpoint prediction improve 15% compared with LR.
Lan Xie, Xinggong Zhang, Zongming Guo
ACM Multimedia3
2018 Context-Aware Unsupervised Text Stylization
abstract
In this work, we present a novel algorithm to stylize the text without supervision, which provides a flexible and convenient way to invoke fantastic text expressions. Rather than employing the fixed pair of target text and source style images, our unsupervised framework establishes an implicit mapping for them by using an abstract imagery of the style image as bridges. Based on the mapping, we progressively narrow the visual discrepancy between text and style images by the proposed legibility-preserving structure transfer and texture transfer algorithms, which effectively balance the text legibility and style consistency. Furthermore, we explore a seamless composition of the stylized text and a background image, in which the optimal text layout is determined by a context-aware layout design algorithm utilizing cues for both seamlessness and aesthetics. Given the layout, the text can be seamlessly embedded into the background by texture synthesis under a context-aware boundary constraint. Experimental results demonstrate the effectiveness of the proposed method in automatic artistic typography creation and visual-textual presentation synthesis.
Shuai Yang 0001, Jiaying Liu 0001, Wenhan Yang, Zongming Guo
ACM Multimedia4
2018 Text effects transfer via distribution-aware texture synthesis
Shuai Yang 0001, Jiaying Liu 0001, Zhouhui Lian, Zongming Guo
Comput. Vis. Image Underst.4
2018 Video super-resolution based on spatial-temporal recurrent residual networks
Wenhan Yang, Jiashi Feng, Guosen Xie, Jiaying Liu 0001, Zongming Guo, Shuicheng Yan
Comput. Vis. Image Underst.5
2018 Kernel Wiener filtering model with low-rank approximation for image denoising
Yongqin Zhang, Jinsheng Xiao, Jiaying Liu 0001, Zongming Guo, Xiaopeng Zong
Inf. Sci.6
2018 Blind visual quality assessment for image super-resolution by convolutional neural network
Yuming Fang 0001, Chi Zhang 0027, Wenhan Yang, Jiaying Liu 0001, Zongming Guo
Multim. Tools Appl.5
2018 Decorrelated local binary patterns for efficient texture classification
Xiaolong Li 0001, Zongming Guo
Multim. Tools Appl.3
2018 Efficient large payloads ternary matrix embedding
Guangyuan Yang, Xiaolong Li 0001, Zongming Guo
Multim. Tools Appl.4
2018 Isophote-Constrained Autoregressive Model With Adaptive Window Extension for Image Interpolation
abstract
The autoregressive (AR) model is widely used in image interpolations. Traditional AR models consider utilizing the dependence between pixels to model the image signal. However, they ignore the valuable patch-level information for image modeling. In this paper, we propose to integrate both the pixel-level and patch-level information to depict the relationship between high-resolution and low-resolution pixels and obtain better image interpolation results. In particular, we propose an isophote-constrained AR (ICAR) model to perform AR-flavored interpolation within an identified joint stable region and further develop an AR interpolation with an adaptive window extension. Considering the smoothness along the isophote curve, the ICAR model searches only several successive similar patches along the isophote curve over a large region to construct an adaptive window. These overlapped patches, representing the patch-level structure similarity, are used to construct a joint AR model. To better characterize the piecewise stationarity and determine whether a pixel is suitable for AR estimation, we further propose pixel-level and patch-level similarity metrics and embed them into the ICAR model, introducing a weighted ICAR model. Comprehensive experiments demonstrate that our method can effectively reconstruct the edge structures and suppress jaggy or ringing artifacts. In the objective quality evaluation, our method achieves the best results in terms of both peak signal-to-noise ratio and structural similarity for both simple size doubling (two times) and for arbitrary scale enlargements.
Wenhan Yang, Jiaying Liu 0001, Mading Li, Zongming Guo
IEEE Trans. Circuits Syst. Video Technol.4
2018 Joint-Feature Guided Depth Map Super-Resolution With Face Priors
abstract
In this paper, we present a novel method to super-resolve and recover the facial depth map nicely. The key idea is to exploit the exemplar-based method to obtain the reliable face priors from high-quality facial depth map to improve the depth image. Specifically, a new neighbor embedding (NE) framework is designed for face prior learning and depth map reconstruction. First, face components are decomposed to form specialized dictionaries and then reconstructed, respectively. Joint features, i.e., low-level depth, intensity cues and high-level position cues, are put forward for robust patch similarity measurement. The NE results are used to obtain the face priors of facial structures and smooth maps, which are then combined in an uniform optimization framework to recover high-quality facial depth maps. Finally, an edge enhancement process is implemented to estimate the final high resolution depth map. Experimental results demonstrate the superiority of our method compared to state-of-the-art depth map super-resolution techniques on both synthetic data and real-world data from Kinect.
Shuai Yang 0001, Jiaying Liu 0001, Yuming Fang 0001, Zongming Guo
IEEE Trans. Cybern.4
2018 Structure-Revealing Low-Light Image Enhancement Via Robust Retinex Model
abstract
Low-light image enhancement methods based on classic Retinex model attempt to manipulate the estimated illumination and to project it back to the corresponding reflectance. However, the model does not consider the noise, which inevitably exists in images captured in low-light conditions. In this paper, we propose the robust Retinex model, which additionally considers a noise map compared with the conventional Retinex model, to improve the performance of enhancing low-light images accompanied by intensive noise. Based on the robust Retinex model, we present an optimization function that includes novel regularization terms for the illumination and reflectance. Specifically, we use norm to constrain the piece-wise smoothness of the illumination, adopt a fidelity term for gradients of the reflectance to reveal the structure details in low-light images, and make the first attempt to estimate a noise map out of the robust Retinex model. To effectively solve the optimization problem, we provide an augmented Lagrange multiplier based alternating direction minimization algorithm without logarithmic transformation. Experimental results demonstrate the effectiveness of the proposed method in low-light image enhancement. In addition, the proposed method can be generalized to handle a series of similar problems, such as the image enhancement for underwater or remote sensing and in hazy or dusty conditions.
Mading Li, Jiaying Liu 0001, Wenhan Yang, Xiaoyan Sun 0001, Zongming Guo
IEEE Trans. Image Process.5
2018 Structure-Guided Image Inpainting Using Homography Transformation
abstract
In this paper, we present a novel structure-guided framework for exemplar-based image inpainting to maintain the neighborhood consistence and structure coherence of an inpainted region. The proposed method consists of a data term for pixel validity and boundary continuity, a smoothness term to depict the compatibility of neighboring pixels for contextual continuity, and a coherence term to investigate image inherent regularities to ensure image self-similarity. To better reconstruct image structures, the method utilizes image regularity statistics to extract dominant linear structures of the target image. Guided by these structures, homography transformations are estimated and combined to globally repair the missing region using the Markov random field model. To reduce computational complexity, a hierarchical process is implemented to utilize the regularity effectively. The experimental results demonstrate that our method yields better results for various real-world scenes than existing state-of-the-art image inpainting techniques.
Jiaying Liu 0001, Shuai Yang 0001, Yuming Fang 0001, Zongming Guo
IEEE Trans. Multim.4
2017 Awesome Typography: Statistics-Based Text Effects Transfer
abstract
In this work, we explore the problem of generating fantastic special-effects for the typography. It is quite challenging due to the model diversities to illustrate varied text effects for different characters. To address this issue, our key idea is to exploit the analytics on the high regularity of the spatial distribution for text effects to guide the synthesis process. Specifically, we characterize the stylized patches by their normalized positions and the optimal scales to depict their style elements. Our method first estimates these two features and derives their correlation statistically. They are then converted into soft constraints for texture transfer to accomplish adaptive multi-scale texture synthesis and to make style element distribution uniform. It allows our algorithm to produce artistic typography that fits for both local texture patterns and the global spatial distribution in the example. Experimental results demonstrate the superiority of our method for various text effects over conventional style transfer methods. In addition, we validate the effectiveness of our algorithm with extensive artistic typography library generation.
Shuai Yang 0001, Jiaying Liu 0001, Zhouhui Lian, Zongming Guo
CVPR4
2017 Deep Joint Rain Detection and Removal from a Single Image
abstract
In this paper, we address a rain removal problem from a single image, even in the presence of heavy rain and rain streak accumulation. Our core ideas lie in our new rain image model and new deep learning architecture. We add a binary map that provides rain streak locations to an existing model, which comprises a rain streak layer and a background layer. We create a model consisting of a component representing rain streak accumulation (where individual streaks cannot be seen, and thus visually similar to mist or fog), and another component representing various shapes and directions of overlapping rain streaks, which usually happen in heavy rain. Based on the model, we develop a multi-task deep learning architecture that learns the binary rain streak map, the appearance of rain streaks, and the clean background, which is our ultimate output. The additional binary map is critically beneficial, since its loss function can provide additional strong information to the network. To handle rain streak accumulation (again, a phenomenon visually similar to mist or fog) and various shapes and directions of overlapping rain streaks, we propose a recurrent rain detection and removal network that removes rain streaks and clears up the rain accumulation iteratively and progressively. In each recurrence of our method, a new contextualized dilated network is developed to exploit regional contextual information and to produce better representations for rain detection. The evaluation on real images, particularly on heavy rain, shows the effectiveness of our models and architecture.
Wenhan Yang, Robby T. Tan, Jiashi Feng, Jiaying Liu 0001, Zongming Guo, Shuicheng Yan
CVPR5
2017 General scale interpolation via context-aware autoregressive model and multiplanar constraint
abstract
In this paper, we propose a novel image interpolation algorithm suitable for general scale enlargement. Different from previous AR-based interpolation algorithms which employ predetermined reference configuration to predict pixel values, we consider the context information when building AR models. Optimal references are selected by incorporating nonlocal-based correlation coefficient and the indicator for local edge direction. Furthermore, the multiplanar constraint among similar patches is applied to enhance the correlation within the estimation window and serves as a kind of supplement to data fidelity term in AR model. The experimental results show that our method is effective in several enlargement scales and successfully alleviate the artifacts nearby edges and preserve their sharpness. The comparison experiments demonstrate that the proposed method can obtain desirable performance in terms of both objective and subjective results.
Shihong Deng, Jiaying Liu 0001, Mading Li, Wenhan Yang, Zongming Guo
ICASSP5
2017 1+N fusion: Cascaded self-portrait enhancement
abstract
In this paper, we present a novel cascaded framework to solve a self-portrait enhancement problem we call “1+N” problem, in which a self-portrait is enhanced with the help of N supporting photos that share the same scene and similar shooting time. The key idea is to exploit the extra information of these N photos to expand the field of view of the self-portrait and improve its lighting style. We achieve this by alternatingly optimizing two complementary tasks, namely illumination unification and photo registration. Based on the correspondences extracted in the input 1+N photos, our method estimates and updates the illumination and registration coefficients in a cascaded manner. Then a Markov Random Field formulation is proposed to globally fuse the aligned photos. Experimental results demonstrate the proposed method achieves high-quality results in this novel application scenario.
Shuai Yang 0001, Jiaying Liu 0001, Sifeng Xia, Zongming Guo
ICASSP4
2017 Variation learning guided convolutional network for image interpolation
abstract
In this paper, we propose a variational learning model that effectively exploits the structural similarities for image representation, and construct a deep network based on this model for image interpolation. Based on the local dependency, our learning model represents an image as the three-dimensional features. Besides two coordinate dimensions, an additional neighboring variation dimension is added to encode every pixel as the variation to its nearest low-resolution pixel by the local similarity. This added dimension lowers the risk of over-fitting for learning approaches and constructs abundant structural correspondences for inferring the missing information lost in image degradation. Then, this three-dimensional features are naturally modeled, extracted and refined by an end-to-end trainable recurrent convolutional network for image interpolation. Comprehensive experiments demonstrate that our method leads to a surprisingly superior performance and offers new state-of-the-art benchmark.
Wenhan Yang, Jiaying Liu 0001, Sifeng Xia, Zongming Guo
ICIP4
2017 Dynamic threshold based rate adaptation for HTTP live streaming
abstract
The Dynamic Adaptive Streaming over HTTP (DASH) is specified to cope with the changing network conditions and provide an adaptive bit-rate HTTP-based streaming solution. While there have been many researches of rate adaptation algorithms on adaptive HTTP streaming, much of the work is focused on Video on Demand (VoD) service - which is not same as live streaming. It is generally preferred to minimize the end-to-end delay and make full use of the bandwidth for live services. In this paper, we propose a buffer-based rate adaptation algorithm with dynamic threshold which can decrease the rate transitions and provide a seamless playback under a low latency requirement. The rate adaptation metrics not only take into account the momentary value of bandwidth but also consider its fluctuation as the recognition of bandwidth is crucial over small buffer. Experiments demonstrate that our proposed rate adaptation scheme outperforms the methods using fixed threshold or instant throughput.
Lan Xie, Chao Zhou 0003, Xinggong Zhang, Zongming Guo
ISCAS4
2017 Joint-domain unsupervised stylization for portraits
abstract
People wish to own a portrait painting of themselves by Da Vinci. Unfortunately, it is impossible to make this dream come true; nevertheless, it may give us an opportunity by transferring some artistic features from one single reference painting. To address this issue, we propose a joint-domain image stylization approach, particularly for portrait oil paintings. From the view of artistic appreciation, we analyze an amount of oil painting artworks and summarize three critical factors to depict the figure, i.e. color, structure and texture. First, the tone of the input image is recolored based on semantic regions corresponding to the reference. Those semantic regions are segmented automatically via the color swatch, by considering the constraints of colors and positions. Then, we exploit sparse representation to reconstruct the layout by acquiring the structure from the reference. The paired training set for sparse dictionary learning is built with the guidance of edge features. Third, considering that texture is usually locally stochastic but regularly repetitive in global, a coarse-to-fine texture synthesis is used to enhance the detail pattern. Subjective results demonstrate the proposed method achieves desirable results compared with state-of-art methods while keeping consistent with artist's style.
Saboya Yang, Jiaying Liu 0001, Shuai Yang 0001, Wenhan Yang, Zongming Guo
ISCAS5
2017 Improved Reversible Visible Watermarking Based on Adaptive Block Partition
Guangyuan Yang, Wenfa Qi, Xiaolong Li 0001, Zongming Guo
IWDW4
2017 A Caching Miss Ratio Aware Path Selection Algorithm for Information-Centric Networks
abstract
In Information-Centric Networks (ICN), contents are cached on some intermediary routers. This creates thus a new situation which is totally different from the traditional path-selection paradigm: the source/destination paradigm no longer exists; instead, the new paradigm is how to find a path through a selected group of caches, so that the content is delivered via the shortest way. This paper addresses this issue and proposes a path-selection algorithm taking into account both the caching capability of router and the more traditional link cost between routers. We formulated the problem as a convex optimization problem (named ESP) which aims to get expected shortest path (ESP) by minimizing the transportation cost. By applying the Lagrangian dual theorem, we solved the ESP problem and obtained a criterion for request (and reversely, data) routing. Based on this path-selection criterion, we provide a fully distributed distance-based ESP algorithm that enables routers maintain routes to nearest content, without knowing a network topology and the caching miss ratio of content at other routers. Simulations confirm the efficiency of our approach versus the traditional shortest path algorithm.
Weihong Lin, Xinggong Zhang, Yu Guan 0005, Zongming Guo
LCN5
2017 360ProbDASH: Improving QoE of 360 Video Streaming Using Tile-based HTTP Adaptive Streaming
abstract
Recently, there has been a significant interest towards 360-degree panorama video. However, such videos usually require extremely high bitrate which hinders their widely spread over the Internet. Tile-based viewport adaptive streaming is a promising way to deliver 360-degree video due to its on-request portion downloading. But it is not trivial for it to achieve good Quality of Experience (QoE) because Internet request-reply delay is usually much higher than motion-to-photon latency. In this paper, we leverage a probabilistic approach to pre-fetch tiles countering viewport prediction error, and design a QoE-driven viewport adaptation system, 360ProbDASH. It treats user's head movement as probability events, and constructs a probabilistic model to depict the distribution of viewport prediction error. A QoE-driven optimization framework is proposed to minimize total expected distortion of pre-fetched tiles. Besides, to smooth border effects of mixed-rate tiles, the spatial quality variance is also minimized. With the requirement of short-term viewport prediction under a small buffer, it applies a target-buffer-based rate adaptation algorithm to ensure continuous playback. We implement 360ProbDASH prototype and carry out extensive experiments on a simulation test-bed and real-world Internet with real user's head movement traces. The experimental results demonstrate that 360ProbDASH achieves at almost 39% gains on viewport PSNR, and 46% reduction on spatial quality variance against the existed viewport adaptation methods.
Lan Xie, Zhimin Xu 0001, Yixuan Ban, Xinggong Zhang, Zongming Guo
ACM Multimedia5
2017 A Novel Two-Step Integer-pixel Motion Estimation Algorithm for HEVC Encoding on a GPU
Keji Chen, Jun Sun 0012, Zongming Guo, Dachuan Zhao 0002
MMM (2)3
2017 An optimal spatial-temporal smoothness approach for tile-based 360-degree video streaming
abstract
The world is becoming more and more virtual than we ever thought it would be. Many video service providers have rolled out 360-degree videos which provide immersive experience to users. However, huge bandwidth occupation of 360-degree video hinders its wide spreading over the Internet. Besides, only part of the video is displayed on the screen, transmitting whole video results in waste of bandwidth and computational resources. Tile-based adaptive streaming is regarded as a bandwidth-friendly approach which only delivers specific portion of the whole video. It requires the clients to decide which portion and at which bitrates to deliver. However, due to both space and time partition of 360-degree videos in tile-based adaptive streaming, there still exists a challenge on the quality inconsistence on spatial and temporal domains. In this paper, we propose a optimal spatial-temporal smoothness approach under restricted network for tile-based adaptive streaming. The bitrates of tiles are determined optimally, aiming at maximizing the overall quality while minimizing the spatial and temporal quality variation. By conducting extensive experiments over real bandwidth dataset and user's head movement traces, our approach can get a significant improvement. Specifically, the Viewport-PSNR can be raised by 24.1% compared with traditional delivery of whole 360 video; while the spatial and temporal stability can be improved by 40.5% and 24.6% respectively compared with tile-based streaming.
Yixuan Ban, Lan Xie, Zhimin Xu 0001, Xinggong Zhang, Zongming Guo, Yueyu Hu
VCIP5
2017 Real-time deep image super-resolution via global context aggregation and local queue jumping
abstract
Deep learning-based image super-resolution has provided very impressive reconstruction quality. However, their running time still sets barriers for real-time applications. In this paper, we propose a Global context aggregation and Local queue jumping Network (GLNet) which provides the more effective image SR given a certain number of model parameters. In our GLNet, we reconsider the model design of the real-time image SR paradigm. Then, we construct a deep network with fewer channels but a deeper structure to effectively aggregate the global context. The dilated convolutions are used as parts of basic units of our GLNet, which further enlarges the receptive field. Besides, an additional local queue jumping path is employed to connect the first-layer feature map and the last-layer feature map to better model the local signal structure. Extensive experiments demonstrate the superiority of our GLNet which offers new state-of-the-art performance considering both reconstruction quality and time consumption.
Yueyu Hu, Jiaying Liu 0001, Wenhan Yang, Shihong Deng, Luyao Zhang 0007, Zongming Guo
VCIP6
2017 Improved reversible data hiding based on two-dimensional difference-histogram modification
Xiaolong Li 0001, Zongming Guo
Multim. Tools Appl.4
2017 Objective Quality Assessment of Screen Content Images by Uncertainty Weighting
abstract
In this paper, we propose a novel full-reference objective quality assessment metric for screen content images (SCIs) by structure features and uncertainty weighting (SFUW). The input SCI is first divided into textual and pictorial regions. The visual quality of textual regions is estimated based on perceptual structural similarity, where the gradient information is adopted as the structural feature. To predict the visual quality of pictorial regions in SCIs, we extract the structural features and luminance features for similarity computation between the reference and distorted pictorial patches. To obtain the final visual quality of SCI, we design an uncertainty weighting method by perceptual theories to fuse the visual quality of textual and pictorial regions effectively. Experimental results show that the proposed SFUW can obtain better performance of visual quality prediction for SCIs than other existing ones.
Yuming Fang 0001, Jiebin Yan, Jiaying Liu 0001, Shiqi Wang 0001, Qiaohong Li, Zongming Guo
IEEE Trans. Image Process.6
2017 Deep Edge Guided Recurrent Residual Learning for Image Super-Resolution
abstract
In this paper, we consider the image super-resolution (SR) problem. The main challenge of image SR is to recover high-frequency details of a low-resolution (LR) image that are important for human perception. To address this essentially ill-posed problem, we introduce a Deep Edge Guided REcurrent rEsidual (DEGREE) network to progressively recover the high-frequency details. Different from most of the existing methods that aim at predicting high-resolution (HR) images directly, the DEGREE investigates an alternative route to recover the difference between a pair of LR and HR images by recurrent residual learning. DEGREE further augments the SR process with edge-preserving capability, namely the LR image and its edge map can jointly infer the sharp edge details of the HR image during the recurrent recovery process. To speed up its training convergence rate, by-pass connections across the multiple layers of DEGREE are constructed. In addition, we offer an understanding on DEGREE from the view-point of sub-band frequency decomposition on image signal and experimentally demonstrate how the DEGREE can recover different frequency bands separately. Extensive experiments on three benchmark data sets clearly demonstrate the superiority of DEGREE over the well-established baselines and DEGREE also provides new state-of-the-arts on these data sets. We also present addition experiments for JPEG artifacts reduction to demonstrate the good generality and flexibility of our proposed DEGREE network to handle other image processing tasks.
Wenhan Yang, Jiashi Feng, Jianchao Yang, Fang Zhao 0006, Jiaying Liu 0001, Zongming Guo, Shuicheng Yan
IEEE Trans. Image Process.6
2017 Retrieval Compensated Group Structured Sparsity for Image Super-Resolution
abstract
Sparse representation-based image super-resolution is a well-studied topic; however, a general sparse framework that can utilize both internal and external dependencies remains unexplored. In this paper, we propose a group-structured sparse representation approach to make full use of both internal and external dependencies to facilitate image super-resolution. External compensated correlated information is introduced by a two-stage retrieval and refinement. First, in the global stage, the content-based features are exploited to select correlated external images. Then, in the local stage, the patch similarity, measured by the combination of content and high-frequency patch features, is utilized to refine the selected external data. To better learn priors from the compensated external data based on the distribution of the internal data and further complement their advantages, nonlocal redundancy is incorporated into the sparse representation model to form a group sparsity framework based on an adaptive structured dictionary. Our proposed adaptive structured dictionary consists of two parts: one trained on internal data and the other trained on compensated external data. Both are organized in a cluster-based form. To provide the desired over-completeness property, when sparsely coding a given LR patch, the proposed structured dictionary is generated dynamically by combining several of the nearest internal and external orthogonal subdictionaries to the patch instead of selecting only the nearest one as in previous methods. Extensive experiments on image super-resolution validate the effectiveness and state-of-the-art performance of the proposed method. Additional experiments on contaminated and uncorrelated external data also demonstrate its superior robustness.
Jiaying Liu 0001, Wenhan Yang, Xinfeng Zhang 0001, Zongming Guo
IEEE Trans. Multim.4
2016 MARLow: A Joint Multiplanar Autoregressive and Low-Rank Approach for Image Completion
Mading Li, Jiaying Liu 0001, Zhiwei Xiong, Xiaoyan Sun 0001, Zongming Guo
ECCV (7)5
2016 A new reversible data hiding scheme exploiting high-dimensional prediction-error histogram
abstract
Pairwise prediction-error expansion (pairwise PEE) is an improvement of the conventional PEE and it can provide excellent performance for reversible data hiding (RDH). Unlike PEE in which the prediction-errors are modified individually, the correlation among prediction-errors is exploited in pairwise PEE by jointly modifying each prediction-error pair. In this paper, the idea of pairwise PEE is developed and a new RDH scheme is proposed. A three-dimensional prediction-error histogram (3D-PEH) is generated by counting every non-overlapped prediction-error triple. Then, data embedding is conducted by modifying the 3D-PEH with a specifically designed reversible mapping. By using 3D-PEH and the proposed reversible mapping, the inter-correlation of prediction-errors is better exploited, and the performance of PEE is significantly enhanced. Moreover, the superiority of our method over pairwise PEE and some other state-of-the-art RDH methods is also experimentally verified. The proposed method is an effective extension of PEE towards the direction of high-dimensional histogram modification.
Siren Cai, Xiaolong Li 0001, Jiaying Liu 0001, Zongming Guo
ICIP4
2016 Quality assessment for image super-resolution based on energy change and texture variation
abstract
In this paper, we propose a novel reduced-reference quality assessment metric for image super-resolution (RRIQA-SR) based on the low-resolution (LR) image information. First, we use the Markov Random Field (MRF) to model the pixel correspondence between LR and high-resolution (HR) images. Based on the pixel correspondence, we predict the perceptual similarity between image patches of LR and HR images by two components: the energy change and texture variation. The overall quality of HR images is estimated by the perceptual similarity between local image patches of LR and HR images. Experimental results demonstrate that the proposed method can obtain better performance of quality prediction for HR images than other existing ones, even including some full-reference (FR) metrics.
Yuming Fang 0001, Jiaying Liu 0001, Yabin Zhang 0002, Weisi Lin, Zongming Guo
ICIP5
2016 Local ternary pattern based on path integral for steganalysis
abstract
The least significant bit (LSB) matching is a steganographic method which embeds the stego signal into cover images in the spatial domain. However, the stego signal disturbs the correlation of neighboring pixels in cover image and this can be utilized for steganalysis. Local binary pattern (LBP) is an effective image texture descriptor, and it can summarize the correlation of neighboring pixels. In this paper, a LBP-based steganalyzer is proposed to identify the deviations of the correlation violated by the stego noise. Specifically, our paper proposes the local ternary pattern based on path integral (pi-LTP) to enhance the feature discrimination in large-scale pixels. Moreover, a greedy incremental algorithm is utilized in our method to select the optimal subspace of pi-LTP features. Experimental results show our method has a better performance than the state-of-the-art steganalysis methods.
Qiuyan Lin, Jiaying Liu 0001, Zongming Guo
ICIP3
2016 Robust and automatic video colorization via multiframe reordering refinement
abstract
In this paper, we propose a robust video colorization method automatically through limited color references in a video sequence. The proposed method first estimates motion vectors between a monochrome frame and colored reference frames for initial matching by optical flow. Then it transfers color information to matched points in the monochrome frame and further propagates color information of matched points to other parts of the monochrome frame. Furthermore, we design a multiframe reordering refinement to colorize video sequences robustly. Experimental results demonstrate that the proposed method achieves much better performance in video colorization than state-of-the-art methods.
Sifeng Xia, Jiaying Liu 0001, Yuming Fang 0001, Wenhan Yang, Zongming Guo
ICIP5
2016 Computational modeling of artistic intention: Quantify lighting surprise for painting analysis
abstract
The use of strong lighting contrast to accentuate objects and figures in a painting—called Chiaroscuro—is popular among Renaissance painters such as Caravaggio, La Tour and Rembrandt. In this paper, we propose a new metric called LuCo to quantify the extent to which Chiaroscuro is employed by an artist in a painting. This measurement could be used to assess the capability of any system to fulfill the original artistic intention and consequently ensure minimal disruptions of Quality of Experience. We first argue that Chiaroscuro is a device for artists to draw attention to specific spatial regions; thus it can be understood as a restricted notion of visual saliency computed using only luminance features. Operationally, using a set of local luminance patches we first compute a Bayesian surprise value, where the prior and posterior probabilities are computed assuming a Gaussian Markov Random Field (GMRF) model. Inverse covariance matrices of the GMRF model are estimated via sparse graph learning for robustness. We construct a histogram using the computed surprise values from different local patches in a painting. Finally, we compute a skewness parameter for the constructed histogram as our LuCo score: large skewness means luminance surprises are either very small or very large, meaning that the artist accentuated lighting contrast in the painting. Experimental results show that paintings by Chiaroscuro artists have higher LuCo scores than 19th century French Impressionists, and Rembrandt's self-portraits have increasingly higher LuCo scores as he aged except for his late period—both trends are in agreement with art historians' interpretations.
Saboya Yang, Gene Cheung, Patrick Le Callet, Jiaying Liu 0001, Zongming Guo
QoMEX5
2016 Autoregressive image interpolation via context modeling and multiplanar constraint
abstract
In this paper, we propose a novel image interpolation algorithm by context-aware autoregressive (AR) model and multiplanar constraint. Different from existing AR based methods which employ predetermined reference configuration to predict pixel values, the proposed method considers the anisotropic pixel dependencies in natural images and adaptively chooses the optimal prediction context by utilizing the nonlocal redundancy to interpolate pixels. Furthermore, the multiplanar constraint is applied to enhance the correlations within the estimation window by exploiting the self-similarity property of natural images. Similar patches are collected by the combination of patch-wise pixel values and the gradient information. And the inter-patch dependencies are adopted to improve the interpolation. The experimental results show that our method is effective in image interpolation and successfully decreases the artifacts nearby the sharp edges. The comparison experiments demonstrate that the proposed method can obtain better performance than other related ones in terms of both objective and subjective results.
Shihong Deng, Jiaying Liu 0001, Mading Li, Wenhan Yang, Zongming Guo
VCIP5
2016 Efficient arbitrary ratio downscale transcoding for HEVC
abstract
The arbitrary ratio transcoding usually introduces coding block grid misalignment, which results in difficulties to utilize the decoding information in coding blocks of source videos during the encoding phase. To reduce the large computation load of HEVC downscale transcoding in such situation, we propose an efficient transcoding method that we refer the decoded coding unit (CU) partitioning to accelerate partition decision, which takes up the most complexity in the encoding phase as well as the whole transcoding. First, we predict CU depth in pixel level according to decoded partitioning of source videos. Then, we propose adaptive rules to determine CU partitioning of target videos based on the prediction, so that we can make early CU splitting or pruning decision without complex recursive search. Experiments demonstrate that the proposed method achieves about 74% time reduction on average with acceptable BD-rate increase in the encoding phase compared to the encoder in reference software HM13.0.
Zhenan Lin, Keji Chen, Jun Sun 0012, Zongming Guo
VCIP5
2016 An adaptive intra-frame parallel method based on complexity estimation for HEVC
abstract
Parallelization is an efficient solution for addressing the increased computational complexity in High Efficiency Video Coding (HEVC). To improve the intra-frame parallelism, an adaptive parallel method is proposed based on an encoding complexity model for HEVC. First, by establishing the relationship between encoding complexity and Rate Distortion Optimization (RDO) process, the encoding complexity is measured by the merge skip modes and coding unit partition statistics. Then a greedy algorithm is proposed to partition each frame into several independent regions for parallelism, and the encoding complexity of each region is precisely controlled to achieve computational complexity balancing for better parallelism. Extensive experimental results show that the proposed method can achieve up to 3.34× speedup against wave-front parallel processing (WPP), and 1.19× speedup against tiles with acceptable encoding efficiency loss for low delay video encoding.
Keji Chen, Jun Sun 0012, Xiangyang Ji, Zongming Guo
VCIP5
2016 A Novel Wavefront-Based High Parallel Solution for HEVC Encoding
abstract
With a lot of enhanced coding tools introduced, High Efficiency Video Coding (HEVC) achieves significant improvement in coding efficiency at the cost of increased computational complexity. To efficiently reduce the encoding time of HEVC, a wavefront-based high parallel (WHP) solution integrating novel data-level and task-level methods is proposed in this paper. On data level, optimal single-instruction-multiple-data algorithms are designed for the enhanced coding tools, i.e., replacing the multiplication in motion compensation by add and shift operations with reduced instruction cycles, removing the transpose in transform via realignment of coefficients, and minimizing the memory access in sum of absolute difference/sum of squared differences calculation by fully reusing the registers. On task level, a novel inter-frame wavefront (IFW) method is developed by effectively decreasing the dependence of wavefront parallel processing (WPP). In addition, a coding tree block level parallelism analysis method is presented to prove the superior of IFW method compared with other HEVC representative parallel methods. Besides, a three-level thread management scheme is proposed to best exploit the parallelism of IFW method and achieve corresponding encoding speedup. Extensive experimental results show that, the overall WHP solution can bring up to $57.65\times $ , $65.55\times $ , and $88.17\times $ speedup for HEVC encoding of Wide Video Graphics Array, 720p and 1080p standard test sequences, while maintaining the same coding performance as with WPP. The proposed solution is also applied in several leading video companies in China, providing HEVC video service for more than 1.3 million users everyday.
Keji Chen, Jun Sun 0012, Yizhou Duan, Zongming Guo
IEEE Trans. Circuits Syst. Video Technol.4
2016 Adaptive Video Streaming With Optimized Bitstream Extraction and PID-Based Quality Control
abstract
To cope with the challenges brought about by bandwidth fluctuation and improve the experience of watching online videos, an adaptive video streaming system that can adjust video quality according to actual network conditions is proposed based on the scalable video coding (SVC) extension of H.264/AVC. First, a simple and effective linear error model is proposed and verified for quality scalability of SVC. The model exploits the linear feature of pixel value errors and can be used to accurately estimate the distortion caused by discarding any combination of enhancement data packets in an SVC bitstream. On that basis, a greedy-like algorithm is designed to assign each data packet a priority value according to its rate-distortion (R-D) impact, thus enabling R-D optimized bitstream extraction under certain bitrate constraints. Finally, the proportional-integral-derivative (PID) method is utilized to control the video quality adjustment and determine a suitable bitrate for transmission. By monitoring and predicting the past, current, and future bandwidth information, the PID-based quality control algorithm is able to reduce quality fluctuation, while still preserving a high quality level. Experimental results show that compared with the baseline software, the proposed system that integrates the above algorithms can achieve much lower video quality fluctuation, with PSNR variance reduced from 1.24 to 0.69, and at the same time deliver higher video quality, with the PSNR average increased by 0.83 dB.
Shengbin Meng, Jun Sun 0012, Yizhou Duan, Zongming Guo
IEEE Trans. Multim.4
2016 mDASH: A Markov Decision-Based Rate Adaptation Approach for Dynamic HTTP Streaming
abstract
Dynamic adaptive streaming over HTTP (DASH) has recently been widely deployed in the Internet. It, however, does not impose any adaptation logic for selecting the quality of video fragments requested by clients. In this paper, we propose a novel Markov decision-based rate adaptation scheme for DASH aiming to maximize the quality of user experience under time-varying channel conditions. To this end, our proposed method takes into account those key factors that make a critical impact on visual quality, including video playback quality, video rate switching frequency and amplitude, buffer overflow/underflow, and buffer occupancy. Besides, to reduce computational complexity, we propose a low-complexity sub-optimal greedy algorithm which is suitable for real-time video streaming. Our experiments in network test-bed and real-world Internet all demonstrate the good performance of the proposed method in both objective and subjective visual quality.
Chao Zhou 0003, Chia-Wen Lin, Zongming Guo
IEEE Trans. Multim.3
2015 Image Restoration Based on 3-D Autoregressive Model via Low-Rank Minimization
abstract
Due to all kinds of need of customers and the complicated transmitting environment of digital image and video resources, numerous practical applications emerge, e.g. Image in painting, interpolation, super-resolution and the removal of salt and pepper noise. One thing these cases all have in common is that there are plenty of missing pixels randomly distributed in an image. Existing image restoration methods aiming at solving this problem include kernel regression [1], matrix completion [2] and total variation (TV) model [3]. The 3-D AR model has also been proposed to detect and interpolate the missing data in video sequences. However, the missing rate or missing region in these papers is usually small. With the missing rate increasing, known pixels in a local neighborhood are not going to be enough to form a solvable linear system. Thus, generally speaking, AR model is not suitable for image restoration from high missing rates. Nevertheless, with proper preliminary processing as proposed in this paper, AR models can be well utilized and present good results even in high missing rates. In this paper, we propose a novel method for image restoration. For the first time, the 3-D AR model is utilized in a single image to simultaneously measure correlation within and between similar patches. 2-D AR model combining with a multiscale structure reconstruct the image using its low-resolution versions to preserve important perceptual statistics such as edges. After obtaining the preliminary reconstruction of the reconstructed full size image, similar patches are collected and the 3-D AR model is applied to form a more local-consistent patch set. Then, an iterative singular value thresholding (SVT) method is utilized to solve the low-rank minimization problem. Instead of aggregating all the overlapped patches after each patch set is processed, we perform SVT for each patch set and aggregate all the overlapped patches into an intermediate image, then the iterative regularization is carried out on the image to produce the newly output for next iteration. Experimental results demonstrate that the proposed method achieves higher PSNR and SSIM than state-of-the-art methods [1-3] and the processed images possess a better visual quality especially in edge structures and texture regions.
Mading Li, Jiaying Liu 0001, Zongming Guo
DCC4
2015 Neighborhood regression for edge-preserving image super-resolution
abstract
There have been many proposed works on image super-resolution via employing different priors or external databases to enhance HR results. However, most of them do not work well on the reconstruction of high-frequency details of images, which are more sensitive for human vision system. Rather than reconstructing the whole components in the image directly, we propose a novel edge-preserving super-resolution algorithm, which reconstructs low- and high-frequency components separately. In this paper, a Neighborhood Regression method is proposed to reconstruct high-frequency details on edge maps, and low-frequency part is reconstructed by the traditional bicubic method. Then, we perform an iterative combination method to obtain the estimated high resolution result, based on an energy minimization function which contains both low-frequency consistency and high-frequency adaptation. Extensive experiments evaluate the effectiveness and performance of our algorithm. It shows that our method is competitive or even better than the state-of-art methods.
Yanghao Li, Jiaying Liu 0001, Wenhan Yang, Zongming Guo
ICASSP4
2015 Novel autoregressive model based on adaptive window-extension and patch-geodesic distance for image interpolation
abstract
In this paper, we propose a novel autoregressive (AR) model based on the adaptive window and the patch-geodesic distance for the image interpolation. The model combines the information of inner/inter-patch correlation. To model the inner-patch correlation, we introduce a patch-geodesic distance similarity metric. The proposed metric shows the desirable capacity to depict the piecewise-stationarity of natural images. For the inter-patch correlation, we introduce the inter-patch structure variation and propose an adaptive window-extension AR model. The model extends the interpolation window according to the local structural variation, increasing the adaptation without violating the consistency. Comprehensive experiments demonstrate that the proposed method is better than or competitive with state-of-the-art interpolation methods in both objective and subjective quality evaluations.
Wenhan Yang, Jiaying Liu 0001, Shuai Yang 0001, Zongming Guo
ICASSP4
2015 Unequal error protection for real-time video streaming using expanding window reed-solomon code
abstract
Expanding Window FEC is an emerging scheme for robust real-time video streaming over wireless networks, with low latency and reduction of error propagation. In this work, we focus on the problem of Expanding Window FEC redundancy allocation which has not been adequately addressed in current works. We first analyse the error probability of the adopted Expanding Window Reed-Solomon code (EW-RS), and introduce an equivalent error probability to simplify them. Then we are able to formulate the optimal redundancy allocation into a constrained nonlinear optimization problem, where by allocating the redundancy unequally considering the unequal importance of different frames and their dependency based on the expanding window, unequal error protection (UEP) is achieved and the overall distortion is minimized. Moreover, to reduce the computation complexity, a high-efficiency hill-climbing algorithm is developed to obtain the suboptimal allocation. At last, the experimental results demonstrate the effectiveness of both the proposed allocation scheme and solution algorithm.
Yufeng Geng, Xinggong Zhang, Chao Zhou 0003, Zongming Guo
ICIP4
2015 Multi-pose face hallucination via neighbor embedding for facial components
abstract
In this paper, we propose a novel multi-pose face hallucination method based on Neighbor Embedding for Facial Components (NEFC) to magnify face images with various poses and expressions. To represent the structure of a face, a facial component decomposition is employed on each face image. Then, a neighbor embedding reconstruction method with locality-constraint is performed for each facial component. For the video scenario, we utilize optical flow to locate the position of each patch among the neighboring frames and make use of the Intra and Inter Nonlocal Means method to preserve consistency between neighboring frames. Experimental results evaluate the effectiveness and adaptability of our algorithm. It shows that our method achieves better performance than the state-of-the-art methods, especially on the face images with various poses and expressions.
Yanghao Li, Jiaying Liu 0001, Wenhan Yang, Zongming Guo
ICIP4
2015 A fairness-aware smooth rate adaptation approach for dynamic HTTP streaming
abstract
Recently, Dynamic Adaptive Streaming over HTTP (DASH) has been widely deployed over the Internet. Under time-varying network conditions, it is, however, still a big challenge to provide smooth video bit-rate with high video quality, especially when multiple clients compete for the network resources where the fairness must be considered. In this paper, a fairness-aware smooth rate adaptation approach is designed for DASH under the scenario that multiple clients are competing for the network resources. To avoid the unfair bandwidth estimated by the client induced by the off-intervals during the downloading process, a probe-based bandwidth estimation method is designed which includes a logarithmic law based increase probing scheme and a conservative back-off based decrease probing scheme. Then, with the probed bandwidth, a dual-threshold based video bit-rate switching scheme is designed that buffer overflow/underflow is avoided, and smooth video bit-rate is also provided. The extensive experiments on our network testbed demonstrate that the proposed approach outperforms the existing schemes significantly.
Chao Zhou 0003, Xinggong Zhang, Zongming Guo
ICIP4
2015 Adaptive autoregressive model with window extension via explicit geometry for image interpolation
abstract
In this paper, we propose a novel adaptive autoregressive (AR) model constructed with an explicit geometry based extended window for image interpolation. Geometric features are chosen as criterions to include more useful pixels. These features are estimated explicitly and guide the interpolation window to extend adaptively. To characterize the piecewise stationary of images, the patch-geodesic distance based similarity is proposed and modulated into the adaptive AR model. For increasing the precision of the parameter estimation, a weighted ridge regression based estimation is employed. With the estimation, the multicollinearity between parameters, which occurs in piecewise stationarity conditions, is eliminated. Experimental results demonstrate that the proposed method is better than or competitive with state-of-the-art interpolation methods in both objective and subjective quality evaluations.
Qingyun Wang 0007, Jiaying Liu 0001, Wenhan Yang, Zongming Guo
ICIP4
2015 Software Solution for HEVC Encoding and Decoding
Shengbin Meng, Jun Sun 0012, Zongming Guo
MMM (2)3
2015 Towards Rate-Distortion analysis of general source distributions: Property and principles
abstract
This paper decouples the complex Rate-Distortion (R-D) analysis problem by inspecting the respective influence of source distribution and quantizer design on R-D performance. First, a universal R-D property is theoretically revealed that, for any source distribution can be expressed as the product of a Scaling Factor (SF) and its remaining part, SF does not affect its derivative R-D function. Second, efficient quantizer design principles are deduced for different source distributions, which can be used as convenient R-D performance classifier and indicator when the dead-zone plus uniform threshold scalar quantizer with nearly-uniform reconstruction quantizer (DZ+UTSQ/NURQ) is applied. These two contributions bring new insight and inspiration towards the R-D analysis of various different sources, being solid infrastructure to benefit various video/image applications.
Jun Sun 0012, Yizhou Duan, Jiasi Shen 0001, Zongming Guo
MMSP5
2015 Hierarchical oil painting stylization with limited reference via sparse representation
abstract
Traditional image stylization is enforced by learning the mappings with an external paired training set. But in practice, people usually encounter a specific stylish image and want to transfer its style to their own pictures without the external dataset. Thus, we propose a hierarchical stylization model with limited reference particularly for oil paintings. First, the edge patch based dictionary is trained to build connections between images and limited reference, then reconstruct the structure layer. Due to the highly structured property of saliency regions, the saliency mask is extracted to integrate the structure layer and the texture layer with different weights. Hence, the advantages of both sparse representation based methods and example based methods are integrated. Moreover, the color layer and the surface layer are considered to make the output more consistent with the artist's individual oil painting style. Subjective results demonstrate the proposed method produces desirable results with state-of-art methods while keeping consistent with the artist's oil painting style.
Saboya Yang, Jiaying Liu 0001, Shuai Yang 0001, Sifeng Xia, Zongming Guo
MMSP5
2015 A two-stage fast CU size decision method for HEVC intracoding
abstract
In HEVC coding standard, the encoder employs a quad-tree-based Coding Unit (CU) structure to adapt to various texture characteristics of images. Although it provides better compression performance, the computation load increases drastically. In this paper, a two-stage fast CU size decision method is proposed for HEVC intra coding. In the first stage, we utilize the combination of the weighted variance and the maximum value of the Hadamard costs of the four sub CUs to represent the CU complexity, and classify each CU into three categories: compound, homogeneous, and undetermined. The former two kinds of CUs will be made early split and pruned decisions respectively. In the second stage, an improved SATD-based estimated R-D cost is employed to decide whether the prev undetermined CUs should be early pruned. Experimental results demonstrate that compared with the reference software HM13.0, the proposed fast CU size decision method provides 49% time reduction with slight quality degradation using the HEVC all intra test condition. This proposal has been adopted to provide HEVC video and picture encoding services at the server-side for UC mobile browser which covers over 100 million people in China.
Jun Sun 0012, Yizhou Duan, Zongming Guo
MMSP4
2015 Delay-constrained rate control for real-time video streaming over wireless networks
abstract
Rate control is a big challenge for real-time video streaming on the internet with the needs of low latency, bandwidth-consuming and stable video rate. However, most of the existing Internet congestion control protocols ignore these needs, and some of them use the packet loss event as congestion signal which is deviation especially over error-prone wireless networks. In this paper, we propose a delay-constrained rate control algorithm by locking queueing delay onto a desired objective. The shadow price of video rate is controlled by queueing delay. All flows adapt video rate according to distortion weight and shadow price so as to achieve a distributed bandwidth sharing with low latency, efficient utilization, and distortion fairness. A closed-loop rate control system is designed for the purpose of stable and agile control. The control parameters are analyzed using control-theoretic approach. Additionally, we construct a real-time wireless video streaming test-bed and conduct extensive experiments over it. Compared with the current widely used methods, the experimental results show that the proposed algorithm can achieve 3dB or more gains in PSNR, and better performance on bandwidth utilization, flow stability with well guaranteed multi-flow fairness.
Yufeng Geng, Xinggong Zhang, Chao Zhou 0003, Zongming Guo
VCIP5
2015 Adaptive General Scale Interpolation Based on Weighted Autoregressive Models
abstract
The autoregressive (AR) model has been widely used in signal processing for its effective estimation, especially in image processing. Many dedicated 2× interpolation algorithms adopt the AR model to describe the strong correlation between low-resolution (LR) pixels and high-resolution (HR) pixels. However, these AR model-based methods closely depend on the fixed relative position between LR pixels and HR pixels that are nonexistent in the general scale interpolation. In this paper, we present an adaptive general scale interpolation algorithm that is capable of arbitrary scaling factors considering the nonstationarity of natural images. Different from other dedicated 2× interpolation methods, the proposed AR terms are modeled by pixels with their adjacent unknown HR neighbors. To compensate for the information loss caused by mismatches of AR models, we consider a weighting scheme suitable for general scale situations based on the pixel similarity to increase accuracy of the estimation. Comprehensive experiments demonstrate the effectiveness of the proposed method on general scaling factors. The maximum gain of peak signal-to-noise ratio is 2.07 dB compared with segment adaptive gradient angle in 1.5× enlargements. To evaluate the performance in resolution adaptive video coding, we have also tested our method on Joint Scalable Video Model codec and obtained better subjective quality and rate-distortion performance.
Mading Li, Jiaying Liu 0001, Jie Ren 0012, Zongming Guo
IEEE Trans. Circuits Syst. Video Technol.4
2015 A Novel JSCC Scheme for UEP-Based Scalable Video Transmission Over MIMO Systems
abstract
In this paper, we propose a novel joint source-channel coding (JSCC) scheme for scalable video transmission over multiple-input multiple-output (MIMO) systems. By exploiting the diversity of MIMO antennas and forward error correction (FEC)-based protection, our method aims to provide unequal error protection (UEP) for the video layers, which are mapped to appropriate antennas. Moreover, JSCC is also considered that we extract a proper subset of video layers and allocate suitable FEC redundancy to them. Jointly considering video layer extraction, FEC rate allocation, and video layer scheduling, we are able to achieve UEP so as to minimize end-to-end distortion. We formulate the scheme as a nonlinear integer optimization problem, which is known to be NP-hard. To find a near-optimal solution efficiently, we propose a low-complexity branch-and-bound algorithm, which partitions the original problem into a series of subproblems by a video layer branching technique. In each branch, the upper and lower distortion bounds are derived. In particular, we transform the video layer scheduling subproblem into a 0/1 multiple knapsack problem, which is NP-complete, and employ an evolutionary Lagrangian method to find a solution efficiently. For the FEC allocation subproblem, a Lagrange duality algorithm with fuzzy surrogate subgradient is proposed. The experimental results demonstrate that the proposed method has good efficiency while achieving close performance to the optimal results.
Chao Zhou 0003, Chia-Wen Lin, Xinggong Zhang, Zongming Guo
IEEE Trans. Circuits Syst. Video Technol.4
2015 Image Super-Resolution Based on Structure-Modulated Sparse Representation
abstract
Sparse representation has recently attracted enormous interests in the field of image restoration. The conventional sparsity-based methods enforce sparse coding on small image patches with certain constraints. However, they neglected the characteristics of image structures both within the same scale and across the different scales for the image sparse representation. This drawback limits the modeling capability of sparsity-based super-resolution methods, especially for the recovery of the observed low-resolution images. In this paper, we propose a joint super-resolution framework of structure-modulated sparse representations to improve the performance of sparsity-based image super-resolution. The proposed algorithm formulates the constrained optimization problem for high-resolution image recovery. The multistep magnification scheme with the ridge regression is first used to exploit the multiscale redundancy for the initial estimation of the high-resolution image. Then, the gradient histogram preservation is incorporated as a regularization term in sparse modeling of the image super-resolution problem. Finally, the numerical solution is provided to solve the super-resolution problem of model parameter estimation and sparse representation. Extensive experiments on image super-resolution are carried out to validate the generality, effectiveness, and robustness of the proposed algorithm. Experimental results demonstrate that our proposed algorithm, which can recover more fine structures and details from an input low-resolution image, outperforms the state-of-the-art methods both subjectively and objectively in most cases.
Yongqin Zhang, Jiaying Liu 0001, Wenhan Yang, Zongming Guo
IEEE Trans. Image Process.4
2015 Corrections to "Novel Efficient HEVC Decoding Solution on General-Purpose Processors"
abstract
In the above paper [ibid., vol. 16, no. 7, p. 1915, Nov. 2014], the sentence "has provided HEVC service to over 1500 million people in China via the Xunlei Kankan video client" should have appeared as "has provided HEVC services to over 150 million people in China via the Xunlei Kankan video client." Also. J. Sun should have been noted as the corresponding author.
Yizhou Duan, Jun Sun 0012, Leju Yan, Keji Chen, Zongming Guo
IEEE Trans. Multim.5
2014 An efficient method for no-reference H.264/SVC bitstream extraction
abstract
This paper investigates the no-reference SVC bitstream extraction problem and presents an efficient solution to approximate the “optimal” extracted sub-stream. First, we introduce a linear error model to accurately estimate the distortion caused by discarding any combination of packets, even when the original sequence is not available. Then we propose a greedy algorithm to decide each packet's priority according to its R-D impact. The priority value of packets can be stored in the bitstream and used for R-D optimized extraction. Experimental results show that our bitstream extraction method can achieve a significant PSNR gain compared to the extractors of JSVM, without computational complexity increment. Comparison with other methods also demonstrates the advantage of the proposed method.
Shengbin Meng, Jun Sun 0012, Yizhou Duan, Zongming Guo
ICASSP4
2014 BSIK-SVD: A dictionary-learning algorithm for block-sparse representations
abstract
Sparse dictionary learning has attracted enormous interest in image processing and data representation in recent years. To improve the performance of dictionary learning, we propose an efficient block-structured incoherent K-SVD algorithm for the sparse representation of signals. Without relying on any prior knowledge of the group structure for the input data, we develop a two-stage agglomerative hierarchical clustering method for block sparse representations. This clustering method adaptively identifies the underlying block structure of the dictionary under the restricted conditions of both a maximal block size and a minimal distance between the blocks. Furthermore, to meet the constraints of both the upper bound and the lower bound of the mutual coherence of dictionary atoms, we introduce a regularization term for the objective function to suppress the block coherence of the overcomplete dictionary. The experiments on synthetic data and real images demonstrate that the proposed algorithm has lower representation error, higher visual quality and better reconstructed results than other state-of-the-art methods.
Yongqin Zhang, Jiaying Liu 0001, Mading Li, Zongming Guo
ICASSP4
2014 A PID-based quality control algorithm for SVC video streaming
abstract
Scalable Video Coding (SVC) makes it possible to change video quality dynamically according to real-time bandwidth. For quality control algorithms of SVC video streaming, the biggest challenge is to keep a video quality that is both smooth and as good as possible. In this paper, we first introduce a combined quality level scheme to describe SVC video quality in a unified way. Then an effective and efficient quality control algorithm for SVC video streaming is proposed based on the Proportional-Integral-Derivative (PID) control method. Extensive experiments show that the proposed algorithm improves 8.6% in video quality with 24.8% reduction in quality fluctuation compared with the existing packet delay feedback algorithm. The proposed algorithm has also been implemented in online video website www.7dlive.com and performs well in applications.
Shengbin Meng, Jun Sun 0012, Zongming Guo
ICC4
2014 General scale interpolation based on fine-grained isophote model with consistency constraint
abstract
In this paper, we propose a fine-grained isophote model with consistency constraint to characterize the piecewise-stationarity of image signals. According to this model, we present a novel interpolation algorithm. In this model, the displacement coefficient is used to model the isophote. Then fine-grained pixel intensity information is introduced to correct the displacement calculation and make the isophote estimation more robust. In order to handle the piecewise-stationarity, we force the isophote direction consistent in the local window when an interpolated line is piecewise-stationary. The proposed algorithm can accommodate the general scale enlargement. Experimental results demonstrate that the proposed approach achieves better performances in both objective and subjective quality assessment.
Wenhan Yang, Jiaying Liu 0001, Mading Li, Zongming Guo
ICIP4
2014 Exploiting multi-scale spatial structures for sparsity based single image super-resolution
abstract
To improve the performance of sparsity-based single image super-resolution (SR), we propose a joint SR framework of structure prior based sparse representation (SPSR). The proposed SPSR algorithm exploits the multi-scale spatial structural self-similarities, the gradient prior and nonlocally centralized sparse representation to formulate a constrained optimization problem for high-resolution image recovery. The high-resolution image is firstly initialized by exploiting cross-scale patch redundancy in an image pyramid from single input low-resolution image. Then the sparse modeling of the image SR problem is proposed to refine it further, where the gradient histogram preservation is incorporated as a regularization term. Finally, an iterative solution is provided to solve the problem of model parameter estimation and sparse representation. Experimental results on image super-resolution validate the generality, effectiveness and robustness of the proposed SPSR algorithm.
Yongqin Zhang, Jiaying Liu 0001, Wei Bai 0002, Zongming Guo
ICIP4
2014 Joint multi-CDN and LT-coding for video transport over HTTP
abstract
Video transport over HTTP is becoming more and more popular. Many video service providers construct huge content distribution networks(CDN) to support HTTP streaming service, however, they seldom exploit the benefits of multiple servers to achieve higher bandwidth and reliability by parallel streaming. In this paper, we study the problem of jointing multi-CDN and LT-coding for video transport over HTTP. Using LT coding, a client could download the same video segment from multiple servers without considering data segmentation and server scheduling issue. Thus, we are able to treat all CDN servers as a virtual server with higher bandwidth and reliability. To reduce the ACK overhead, a stochastic model is designed to predict the amount of data to be sent from each server while guaranteeing the decoding probability. Compared with the existing schemes, the experimental results show that our proposed scheme obtains less overhead and fewer number of HTTP requests. Besides, we also achieve better video quality and better robustness to fluctuant bandwidth.
Chao Zhou 0003, Xinggong Zhang, Zongming Guo
ISCAS4
2014 Segmentation-based scale-invariant nonlocal means super resolution
abstract
Zooming in/out appears frequently in video shooting, which makes scale vary between frames. And object motion in videos may cause scale change of the object. It leads to the difficulty in finding similar patches and causes the invalidation of nonlocal means super resolution (NLM SR). In this paper, we propose a novel scale-compensated NLM SR algorithm. First, by considering the parameter model, the image is segmented in order to detect regions with different scales. Then, scale variations in different regions are computed based on SIFT descriptor. And patches extracted from different regions are compensated into the same scale to eliminate the effect of scale change. It is shown by experimental results that our proposed algorithm achieves the average PSNR by up to 0.678dB comparing with the state-of-the-art methods. Subjective results demonstrate the proposed method reduces artifacts and preserves more details.
Saboya Yang, Jiaying Liu 0001, Qiaochu Li, Zongming Guo
ISCAS4
2014 Towards efficient wavefront parallel encoding of HEVC: Parallelism analysis and improvement
abstract
High Efficiency Video Coding (HEVC) is the new generation video coding standard which achieves significant improvement in coding efficiency. Although HEVC is promising in many applications, the increased computational complexity is a serious problem, which makes parallelization necessary in HEVC encoding. To better understand the bottleneck of parallelization and improve the encoding speed, in this paper, we propose a Coding Tree Blocks (CTB) level parallelism analysis method as well as a novel Inter-Frame Wavefront (IFW) parallel encoding method. First, by establishing the relationship between parallelism and dependence, parallelism is precisely described by CTB-level dependence as a criterion to evaluate different parallel methods of HEVC. On this basis, by effectively decreasing the dependence based on Wavefront Parallel Processing (WPP), IFW method is developed. Finally, with the proposed parallelism analysis method, IFW is theoretically proved to be of higher parallelism compared with other HEVC representative parallel methods. Extensive experimental results show that, the proposed method and implementation can bring up to 17.81x, 14.34x and 24.40x speedup for HEVC encoding of WVGA, 720p and 1080p standard test sequences with the same ignorable coding performance degradation as WPP, thus showing a promising technology for future large-scale HEVC video application.
Keji Chen, Yizhou Duan, Jun Sun 0012, Zongming Guo
MMSP4
2014 Highly optimized implementation of HEVC decoder for general processors
abstract
In this paper, we propose a novel design and optimized implementation of the HEVC decoder. First, a novel decoder prototype with refined decoding workflow and efficient memory management is designed. Then on this basis, a series of single-instruction-multiple-data (SIMD) based algorithms are used to speed up several time-consuming modules in HEVC decoding. Finally, a frame-based parallel framework is applied to exploit the multi-threading technology on multicore processors. With the highly optimized HEVC decoder, decoding speed of 246fps on Intel i7-2400 3.4GHz quad-core processor for 1080p videos and 52fps on ARM Cortex-A9 1.2GHz dual-core processor for 720p videos can be achieved in our experiments.
Shengbin Meng, Yizhou Duan, Jun Sun 0012, Zongming Guo
MMSP4
2014 Image transformation using limited reference with application to photo-sketch synthesis
abstract
Image transformation refers to transforming images from a source image space to a target image space. Contemporary image transformation methods achieve this by learning coupled dictionaries from a set of paired images. However, in practical use, such paired training images are not easy to get especially when the target image style is not fixed. Thus in most cases, the reference is limited. In this paper, we propose a sparse representation based framework of transforming images with limited reference, which can be used for the typical image transformation application, photo-sketch synthesis. In the learning stage, the edge features are utilized to map patches between different style images, thus building the coupled database for dictionary learning. In the reconstruction stage, sparse representation can well preserve the basic structure of image contents. In addition, a texture synthesis strategy is introduced to enhance target-like textures in the output image. Experimental results show that the performance of our method is comparable to state-of-the-art methods even with limited reference, which is very efficient and less restrictive for practical use.
Wei Bai 0002, Yanghao Li, Jiaying Liu 0001, Zongming Guo
VCIP4
2014 Patch-based image deblocking using geodesic distance weighted low-rank approximation
abstract
Transform coding based on the discrete cosine transform (DCT) has been widely used in image coding standards. However, the coded images often suffer from severe visual distortions such as blocking artifacts. In this paper, we propose a novel image deblocking method to address the blocking artifacts reduction problem in a patch-based scheme. Image patches are clustered and reconstructed by the low-rank approximation, which is weighted by the geodesic distance. Experimental results show that the proposed method achieves higher PSNR than the state-of-the-art deblocking and denoising methods and the processed images present good visual quality.
Mading Li, Jiaying Liu 0001, Jie Ren 0012, Zongming Guo
VCIP4
2014 Probabilistic chunk scheduling approach in parallel multiple-server DASH
abstract
Recently parallel Dynamic Adaptive Streaming over HTTP (DASH) has emerged as a promising way to supply higher bandwidth, connection diversity and reliability. However, it is still a big challenge to download chunks sequentially in parallel DASH due to heterogeneous and time-varying bandwidth of multiple servers. In this paper, we propose a novel probabilistic chunk scheduling approach considering time-varying bandwidth. Video chunks are scheduled to the servers which consume the least time while with the highest probability to complete downloading before the deadline. The proposed approach is formulated as a constrained optimization problem with the objective to minimize the total downloading time. Using the probabilistic model of time-varying bandwidth, we first estimate the probability of successful downloading chunks before the playback deadline. Then we estimate the download time of chunks. A near-optimal solution algorithm is designed which schedules chunks to the servers with minimal downloading time while the completion probability is under the constraint. Compared with the existing schemes, the experimental results demonstrate that our proposed scheme greatly increases the number of chunks that are received orderly.
Chao Zhou 0003, Xinggong Zhang, Zongming Guo
VCIP4
2014 A low-latency peer-to-peer live and VOD streaming system based on scalable video coding
abstract
This paper demonstrates a peer-to-peer (P2P) live and VOD streaming system called Immedia based on Scalable Video Coding (SVC) with low latency. With features of layered coding in SVC, Immedia system could automatically adapt output video rate to varying network. Especially, the live system combining P2P streaming media transmission techniques works better in this respect and greatly reduces server pressure in a large-scale environment.
Jun Sun 0012, Yanping Zhou, Yizhou Duan, Zongming Guo
VCIP4
2014 Towards simple and smooth rate adaption for VBR video in DASH
abstract
Rate adaption in Dynamic Adaptive Streaming over HTTP (DASH) is widely applied to adapt the transmission rate to varying network capacity. For rate adaption on variable bitrate (VBR) encoded video, it is still a challenge to properly identify and address the dynamics of bandwidth and segment bitrate. In this paper, the trend of client buffer level variation (TBLV) is analyzed to be a more effective metric for detecting the dynamics of bandwidth and segment bitrate compared to previous metrics. Then, a partial-linear trend prediction model is developed to accurately estimate TBLV. Finally, based on the prediction model, a novel simple rate adaption algorithm is designed to achieve efficient and smooth video quality level adjustment. Experimental results show that while maintaining similar average video quality, the proposed algorithm achieves up to 47.3% improvement in rate adaption smoothness compared to the existing work.
Yanping Zhou, Yizhou Duan, Jun Sun 0012, Zongming Guo
VCIP4
2014 Joint image denoising using adaptive principal component analysis and self-similarity
Yongqin Zhang, Jiaying Liu 0001, Mading Li, Zongming Guo
Inf. Sci.4
2014 A Control-Theoretic Approach to Rate Adaption for DASH Over Multiple Content Distribution Servers
abstract
Recently, dynamic adaptive streaming over HTTP (DASH) has been widely deployed on the Internet. However, the research about DASH over multiple content distribution servers (MCDS-DASH) is limited. Compared with traditional single-server DASH, MCDS-DASH is able to offer expanded bandwidth, link diversity, and reliability. It is, however, a challenging problem to smooth video bitrate switching over multiple servers due to their diverse bandwidths. In this paper, we propose a block-based rate adaptation method considering both the diverse bandwidths and feedback buffered video time. In our method, multiple fragments are grouped into a block and the fragments are downloaded in parallel from multiple servers. We propose to adapt video bitrate at the block level rather than at the fragment level. By dynamically adjusting the block length and scheduling fragment requests to multiple servers, the requested video bitrates from the multiple servers are synchronized, making the fragments download in an orderly way. Then, we propose a control-theoretic approach to select an appropriate bitrate for each block. By modeling and linearizing the rate adaption system, we propose a novel proportional-derivative controller to adapt video bitrate with high responsiveness and stability. Theoretical analysis and extensive experiments on our network testbed and the Internet demonstrate the good efficiency of the proposed method.
Chao Zhou 0003, Chia-Wen Lin, Xinggong Zhang, Zongming Guo
IEEE Trans. Circuits Syst. Video Technol.4
2014 Novel Efficient HEVC Decoding Solution on General-Purpose Processors
abstract
Although the emerging video coding standard High Efficiency Video Coding (HEVC) successfully doubles the compression efficiency of H.264/AVC, its growing computational complexity makes real-time decoding of high-definition HEVC videos a very challenging issue for the existing personal computers and mobile devices. In this paper, a systematical, efficient HEVC decoding solution on general processors is provided, consisting of structure-level, data-level, and task-level approaches. First, a redesigned overall structure of a HEVC decoder with data redundancy reduction mechanism is introduced, which cuts down basic data operation cost and achieves an average decoding speedup of 2.37 × compared to the HM 10.0 decoder. On this basis, novel single-instruction multiple-data (SIMD) algorithms such as low-complexity motion compensation, transpose-free transform, symmetric deblocking filter, and parallel-index sample adaptive offset are developed, which further parallelize the data operations of each decoding task and bring another 2.67 × decoding speedup. Finally, a frame-based task-level parallel framework is employed with a flexible entry scheme to efficiently support the simultaneous processing of multiple decoding tasks for different HEVC parallel strategies. The overall solution achieves decoding fps of 40-75 for 4k HEVC videos on the Intel i7-2600 3.4 GHz quad-core processor (4-thread decoding) and 35-55 for 720p videos on the ARM Cortex-A9 1.2 GHz duo-core processor (2-thread decoding). This proposal is the recommended cross-platform HEVC decoding solution of Intel, AMD, and Cisco, and has provided HEVC service to over 1500 million people in China via the Xunlei Kankan video client.
Yizhou Duan, Jun Sun 0012, Leju Yan, Keji Chen, Zongming Guo
IEEE Trans. Multim.5
2013 Single-Pass Dependent Bit Allocation in Temporal Scalability Video Coding
abstract
Summary form only given. In the scalable video coding, we refer to a group-of-pictures (GOP) structure that is composed of hierarchically aligned B-pictures. It employs generalized B-pictures that can be used as a reference to following inter-coded frames. Although it introduces a structural encoding delay of one GOP size, it provides much higher coding efficiency than the conventional GOP structures [2]. Moreover, due to its natural capability of providing the temporal scalability, it is employed as a GOP structure of H.264/SVC [3]. Because of the complex inter-layer dependence of hierarchical B-pictures, the development of an efficient and effective bit allocation algorithm for H.264/SVC is a challenging task. There are several bit allocation algorithms that considered the inter-layer dependence in the literature before. Schwarz et al. proposed the QP cascading scheme that applies a fixed quantization parameter (QP) difference between adjacent temporal layers. Liu et al. introduced constant weights to temporal layers in their H.264/SVC rate control algorithm. Although these algorithms achieve superior coding efficiency, they are limited in two aspects. First, the inter-layer dependence is heuristically addressed. Second, the input video characteristics are not taken into account. For these reasons, the optimality of these bit allocation algorithms cannot be guaranteed. We propose a single-pass dependent bit allocation algorithm for scalable video coding with hierarchical B-pictures in this work. It is generally perceived that dependent bit allocation algorithms cannot be practically employed due to their extremely high complexity requirement. To develop a practical single-pass bit allocation algorithm, we use the number of skipped blocks and the ratio of the mean absolute difference (MAD) as features to measure the inter-layer signal dependence of input video signals. The proposed algorithm performs bit allocation at the target bit rate with two mechanisms: 1) the GOP based rate control and 2) adaptive temporal layer QP decision. The superior performance of the proposed algorithm is demonstrated by experimental results, which is benchmarked by two other single-pass bit allocation algorithms in the literature. The rate and the PSNR coding performance of the proposed scheme and two benchmarks at various target bit rates for GOP-4 and GOP-8, respectively. We see that the proposed rate control algorithm achieves about 0.2-0.3dB improvement in coding efficiency as compared to JSVM. Furthermore, the proposed rate control algorithm outperforms Liu's Algorithm by a significant margin.
Jiaying Liu 0001, Yongjin Cho, Zongming Guo
DCC3
2013 Image Blocking Artifacts Reduction via Patch Clustering and Low-Rank Minimization
abstract
Summary form only given. Block-based Discrete Cosine Transform (BDCT) has been widely used in image and video compression due to its energy compacting property and relative ease of implementation. However, BDCT has a major drawback, which is usually referred to as blocking artifacts. Blocking artifacts appear as grid noise along the block boundaries because each block is transformed and quantized independently. Image deblocking techniques can reduce these distortions and alleviate the conflict between bit rate reduction and visual quality preservation. Many state-of-the-art image deblocking algorithms treated the blocking artifacts reduction of the compressed image as an inverse restoration problem. Natural image prior models are well utilized into the blocking artifacts reduction processing, such as the local sparsity prior model and non-local similarity property of natural images. These two local and non-local models characterize the image prior information in two complementary perspectives. Therefore, it is necessary to combine these two models in a unified framework. In this paper, we propose a novel method to reduce the blocking artifacts of blockcoded images via patch clustering and low-rank minimization, which simultaneously exploits the local and non-local sparse representations in a unified framework. First, the whole compressed image are divided into small patches. For each patch, we perform patch clustering to collect similar patches into a group. Then the whole group are simultaneously reconstructed by a low-rank minimization approach. Singular value thresholding (SVT) algorithm is employed to solve the low-rank minimization problem. To further improve the performance of the proposed algorithm, we adopt an iterative procedure to utilize the newly output data in each iteration and update the noise and signal variance adaptively. Experimental results show that the proposed method achieves higher PSNR and SSIM than the state-of-the-art methods. Comparing to the state-of-theart algorithms and, the proposed algorithm achieves about 0.37dB and 0.11dB improvement on average. For visual quality assessment, the deblocking images produced by the proposed algorithm reveal much more sharp edge structures and richer textures.
Jie Ren 0012, Jiaying Liu 0001, Mading Li, Wei Bai 0002, Zongming Guo
DCC5
2013 Postprocessing of block-coded videos for deflicker and deblocking
abstract
In this paper, we propose a novel postprocessing method to suppress both the flickering and blocking artifacts in block-coded videos. For reducing the flickering effect between adjacent frames, we propose an adaptive multi-scale motion filtering method to maintain the motion coherence of processed video. For blocking artifacts suppression, we adopt a patch-based scheme in which similar patches are grouped in a spatio-temporal domain and each patch group is recovered by solving a low rank matrix completion problem. Experimental results show that the proposed method can significantly reduce the flickering and blocking artifacts in the decoded videos.
Jie Ren 0012, Jiaying Liu 0001, Mading Li, Zongming Guo
ICASSP4
2013 Adaptive general scale interpolation based on similar pixels weighting
abstract
In this paper, we propose an adaptive general scale interpolation algorithm considering the non-stationarity of natural images in local areas. In image 2× enlargement, there are fixed relative positions between low-resolution (LR) pixels and high-resolution (HR) pixels. Unknown HR pixels can be estimated by their available LR neighbors. However, such relative positions are not fixed in the general-scale enlargement situations. The number and position of available LR pixels are indeterminate, therefore HR pixels can not be estimated by LR pixels. To make our method suitable for general scaling factors, we construct autoregressive (AR) models with pixels' neighbors instead of their available LR neighbors. Simultaneously, we introduce the similarity between pixels within a local window, which improves the method's performance by modeling the non-stationarity of image signals. Experimental results demonstrate the effectiveness of the proposed method on general scaling factors.
Mading Li, Jiaying Liu 0001, Jie Ren 0012, Zongming Guo
ISCAS4
2013 Illumination-invariance and nonlocal means based super resolution
abstract
In this paper, we propose a novel algorithm for multi-frame super resolution (SR) with illumination-invariance. Traditional multi-frame SR methods fail to handle images with illumination changes, so in our approach, we adjust the contrast between different search windows and select proper candidate patches to take full advantage of intensity information. We simplify Speed Up Robust Features to get local structure information and incorporate the local structure information into similarity measurement, which does not change significantly in complex illumination situation. By combining intensity and structure information in a proper way, our algorithm Illumination-Invariant Nonlocal Means SR could find more potential similar patches in frames where there are illumination changes than Nonlocal Means SR (NLM SR). Experimental results demonstrate that our algorithm has better performance both in objective and subjective perception with complex illumination conditions and is comparable to NLM SR in stable illumination situation.
Mengyan Wang, Jiaying Liu 0001, Wei Bai 0002, Zongming Guo
ISCAS4
2013 Adaptive channel scheduling for Scalable Video broadcasting over MIMO wireless networks
abstract
Video broadcasting is an efficient way to deliver video content to multiple receivers. However, due to heterogeneous channel conditions of users, it is challenging to minimize all users' video transmission distortion in MIMO broadcasting. In this paper, we investigate the channel scheduling problem, which maps Scalable Video layers to MIMO heterogenous channels to protect video layers unequally, so as to minimize overall received video distortion for all users. We formulate this problem into an integer non-linear optimization problem, which is hard to be solved. An efficient near-optimal algorithm is proposed, which is based on simulated-annealing theory. Experimental results demonstrate the efficiency of our proposed algorithm, and its performance is very close to the optimal results. Compared with the existing MIMO scheduling methods, the proposed scheme significantly improves the overall quality of video broadcasting in MIMO networks.
Chao Zhou 0003, Xinggong Zhang, Zongming Guo
ISCAS3
2013 Multi-frame Super Resolution Using Refined Exploration of Extensive Self-examples
Wei Bai 0002, Jiaying Liu 0001, Mading Li, Zongming Guo
MMM (1)4
2013 Rate-Quantization and Distortion-Quantization Models of Dead-Zone Plus Uniform Threshold Scalar Quantizers for Generalized Gaussian Random Variables
Yizhou Duan, Jun Sun 0012, Zongming Guo
MMM (1)3
2013 Topology-aware content-centric networking
abstract
Making data the first class entity, Information-Centric Networking (ICN) replaces conventional host-to-host model with content sharing model. However, the huge amount of content and the volatility of replicas cached across the Internet pose significant challenges for addressing content only by name. In this paper, we propose a topology-aware name-based routing protocol which combines the benefits of location-oriented routing and content-centric routing together. We adopt a URL-like naming scheme, which defines register locations and content identifier. Node with copies sends Register messages towards a register using location-oriented routing protocols. All en-path routers record forwarding entries in forwarding table (FIB) as the "bread crumb" to this content. Following the bread crumb, routers know the "best" topology path to the available copies. An Interest is either forwarded towards a "known" copy by the content identifier, or towards the register nodes where it would find the bread crumb to the "best" copies. Compared with the existing flooding or name resolution methods, Our design shows a good potential in terms of scalability, availability and overhead.
Xinggong Zhang, Feng Lao, Zongming Guo
SIGCOMM4
2013 Image super resolution using saliency-modulated context-aware sparse decomposition
abstract
This paper presents a novel saliency-modulated sparse representation algorithm for image super resolution. In images, regions salient to human eyes appear to be more organized and structured. This property is utilized in both the dictionary learning and the sparse coding process to capture more structural details for the reconstructed image. Apart from a general dictionary, example patches from the salient regions are extracted to train a salient dictionary. We also incorporate context-aware sparse decomposition to model dependencies between dictionary atoms of adjacent patches, especially in the salient regions. Experiments show the proposed method outperforms state-of-the-art methods with the highest PSNR gain. Subjective results demonstrate the proposed method reduces artifacts and preserves more details.
Wei Bai 0002, Saboya Yang, Jiaying Liu 0001, Jie Ren 0012, Zongming Guo
VCIP5
2013 Joint image denoising using self-similarity based low-rank approximations
abstract
The observed images are usually noisy due to data acquisition and transmission process. Therefore, image denoising is a necessary procedure prior to post-processing applications. The proposed algorithm exploits the self-similarity based low rank technique to approximate the real-world image in the multivariate analysis sense. It consists of two successive steps: adaptive dimensionality reduction of similar patch groups, and the collaborative filtering. For each target patch, the singular value decomposition (SVD) is used to factorize the similar patch group collected in a local search window by block-matching. Parallel analysis automatically selects the principal signal components by discarding the nonsignificant singular values. After the inverse SVD transform, the denoised image is reconstructed by the weighted averaging approach. Finally, the collaborative Wiener filtering is applied to further remove the noise. Experimental results show that the proposed algorithm surpasses the state-of-the-art methods in most cases.
Yongqin Zhang, Jiaying Liu 0001, Saboya Yang, Zongming Guo
VCIP4
2013 A control theory based rate adaption scheme for dash over multiple servers
abstract
Recently, Dynamic Adaptive Streaming over HTTP (DASH) has been widely deployed in the Internet. However, the research about DASH over Multiple Content Distribution Servers (MCDS) is few. Compared with traditional single-server-DASH, MCDS are able to offer expanded bandwidth, link diversity, and reliability. It is, however, a challenging problem to smooth video bitrate switchings over multiple servers due to their diverse bandwidths. In this paper, we propose a block-based rate adaptation method considering both the diverse bandwidths and feedback buffered video time. Multiple fragments are grouped into a block, and the fragments are downloaded in parallel from multiple servers. We propose to adapt video bitrate at the block level rather than at the fragment level. By dynamically adjusting the block length and scheduling fragment requests to multiple servers, the requested video bitrates from the multiple servers are synchronized, making the fragments downloaded orderly. Then, we propose a control-theoretic approach to select an appropriate bitrate for each block. By modeling and linearizing the rate adaption system, we propose a novel Proportional-Derivative (PD) controller to adapt video bitrate with high responsiveness and stability. Theoretical analysis and extensive experiments on the Internet demonstrate the good efficiency of our DASH designs.
Chao Zhou 0003, Xinggong Zhang, Zongming Guo
VCIP3
2013 Optimal adaptive channel scheduling for scalable video broadcasting over MIMO wireless networks
abstract
Video broadcasting is an efficient way to deliver video content to multiple receivers. However, due to heterogeneous channel conditions in MIMO wireless networks, it is challenging for video broadcasting to map scalable video layers to proper MIMO transmit antennas to minimize the average overall video transmission distortion. In this paper, we investigate the channel scheduling problem for broadcasting scalable video content over MIMO wireless networks. An adaptive channel scheduling based unequal error protection (UEP) video broadcasting scheme is proposed. In the scheme, video layers are protected unequally by being mapped to appropriate antennas, and the average overall distortion of all receivers is minimized. We formulate this scheme into a non-linear combinatorial optimization problem. It is not practical to solve the problem by an exhaustive search method with heavy computational complexity. Instead, an efficient branch-and-bound based channel scheduling algorithm, named TBCS, is developed. TBCS finds the global optimal solution with much lower complexity. The complexity is further reduced by relaxing the termination condition of TBCS, which produces a (1 − ε)-optimal solution. Experimental results demonstrate both the effectiveness and efficiency of our proposed scheme and algorithm. As compared with some existing channel scheduling methods, TBCS improves the quality of video broadcasting across all receivers significantly.
Chao Zhou 0003, Xinggong Zhang, Zongming Guo
Comput. Networks3
2013 Guided image filtering using signal subspace projection
abstract
There are various image filtering approaches in computer vision and image processing that are effective for some types of noise, but they invariably make certain assumptions about the properties of the signal and/or noise which lack the generality for diverse image noise reduction. This study describes a novel generalised guided image filtering method with the reference image generated by signal subspace projection (SSP) technique. It adopts refined parallel analysis with Monte Carlo simulations to select the dimensionality of signal subspace in the patch‐based noisy images. The noiseless image is reconstructed from the noisy image projected onto the significant eigenimages by component analysis. Training/test image are utilised to determine the relationship between the optimal parameter value and noise deviation that maximises the output peak signal‐to‐noise ratio (PSNR). The optimal parameters of the proposed algorithm can be automatically selected using noise deviation estimation based on the smallest singular value of the patch‐based image by singular value decomposition (SVD). Finally, we present a quantitative and qualitative comparison of the proposed algorithm with the traditional guided filter and other state‐of‐the‐art methods with respect to the choice of the image patch and neighbourhood window sizes.
Yongqin Zhang, Jiaying Liu 0001, Zongming Guo
IET Image Process.4
2013 Context-Aware Sparse Decomposition for Image Denoising and Super-Resolution
abstract
Image prior models based on sparse and redundant representations are attracting more and more attention in the field of image restoration. The conventional sparsity-based methods enforce sparsity prior on small image patches independently. Unfortunately, these works neglected the contextual information between sparse representations of neighboring image patches. It limits the modeling capability of sparsity-based image prior, especially when the major structural information of the source image is lost in the following serious degradation process. In this paper, we utilize the contextual information of local patches (denoted as context-aware sparsity prior) to enhance the performance of sparsity-based restoration method. In addition, a unified framework based on the markov random fields model is proposed to tune the local prior into a global one to deal with arbitrary size images. An iterative numerical solution is presented to solve the joint problem of model parameters estimation and sparse recovery. Finally, the experimental results on image denoising and super-resolution demonstrate the effectiveness and robustness of the proposed context-aware method.
Jie Ren 0012, Jiaying Liu 0001, Zongming Guo
IEEE Trans. Image Process.3
2013 Rate-Distortion Analysis of Dead-Zone Plus Uniform Threshold Scalar Quantization and Its Application - Part I: Fundamental Theory
abstract
This paper provides a systematic rate-distortion (R-D) analysis of the dead-zone plus uniform threshold scalar quantization (DZ+UTSQ) with nearly uniform reconstruction quantization (NURQ) for generalized Gaussian distribution (GGD), which consists of two aspects: R-D performance analysis and R-D modeling. In R-D performance analysis, we first derive the preliminary constraint of optimum entropy-constrained DZ+UTSQ/NURQ for GGD, under which the property of the GGD distortion-rate (D-R) function is elucidated. Then for the GGD source of actual transform coefficients, the refined constraint and precise conditions of optimum DZ+UTSQ/NURQ are rigorously deduced in the real coding bit rate range, and efficient DZ+UTSQ/NURQ design criteria are proposed to reasonably simplify the utilization of effective quantizers in practice. In R-D modeling, inspired by R-D performance analysis, the D-R function is first developed, followed by the novel rate-quantization (R-Q) and distortion-quantization (D-Q) models derived using analytical and heuristic methods. The D-R, R-Q, and D-Q models form the source model describing the relationship between the rate, distortion, and quantization steps. One application of the proposed source model is the effective two-pass VBR coding algorithm design on an encoder of H.264/AVC reference software, which achieves constant video quality and desirable rate control accuracy.
Jun Sun 0012, Yizhou Duan, Jiaying Liu 0001, Zongming Guo
IEEE Trans. Image Process.5
2013 Rate-Distortion Analysis of Dead-Zone Plus Uniform Threshold Scalar Quantization and Its Application - Part II: Two-Pass VBR Coding for H.264/AVC
abstract
In the first part of this paper, we derive a source model describing the relationship between the rate, distortion, and quantization steps of the dead-zone plus uniform threshold scalar quantizers with nearly uniform reconstruction quantizers for generalized Gaussian distribution. This source model consists of rate-quantization, distortion-quantization (D-Q), and distortion-rate (D-R) models. In this part, we first rigorously confirm the accuracy of the proposed source model by comparing the calculated results with the coding data of JM 16.0. Efficient parameter estimation strategies are then developed to better employ this source model in our two-pass rate control method for H.264 variable bit rate coding. Based on our D-Q and D-R models, the proposed method is of high stability, low complexity and is easy to implement. Extensive experiments demonstrate that the proposed method achieves: 1) average peak signal-to-noise ratio variance of only 0.0658 dB, compared to 1.8758 dB of JM 16.0's method, with an average rate control error of 1.95% and 2) significant improvement in smoothing the video quality compared with the latest two-pass rate control method.
Jun Sun 0012, Yizhou Duan, Jiaying Liu 0001, Zongming Guo
IEEE Trans. Image Process.5
2013 Modeling and Analysis of Skype Video Calls: Rate Control and Video Quality
abstract
Video-conferencing has recently gained its momentum and is widely adopted by end-consumers. But there have been very few studies on the network impacts of video calls and the user Quality-of-Experience (QoE) under different network conditions. In this paper, we study the rate control and video quality of Skype video call, and analyze the network impacts in large-scale networks. We first measure the behaviors of Skype video call on a controlled network testbed. By varying packet loss rate, propagation delay and available network bandwidth, we observe how Skype adjusts its sending rate, FEC redundancy, video rate and frame rate. It is found that Skype is robust against mild packet losses and propagation delays, and can efficiently utilize the available network bandwidth. We also find that it employs an overly aggressive FEC protection strategy. Based on the measurement results, we develop rate control model, FEC model, and video quality model for Skype video calls. Extrapolating from the models, we conduct numerical analysis to study the network impacts. We demonstrate that user back-offs upon quality degradation serve as an effective user-level rate control scheme. We also show that Skype video calls are indeed TCP-friendly and respond to congestion quickly when the network is overloaded. Through a case study of a 4G wireless network, we demonstrate that the proposed models can be used in user-QoE-aware network provisioning.
Xinggong Zhang, Yang Xu 0011, Yong Liu 0013, Zongming Guo, Yao Wang 0001
IEEE Trans. Multim.5
2012 Single-Hop Friends Recommendation and Verification Based Incentive for BitTorrent
Zongming Guo
APWeb2
2012 Rate-Distortion Analysis and Modeling of Dead-Zone Plus Uniform Threshold Scalar Quantization for Generalized Gaussian Random Variables
abstract
This paper provides a systematical rate-distortion (R-D) analysis and modeling of generalized Gaussian distribution (GGD) under the dead-zone plus uniform threshold scalar quantization (DZ+UTSQ) and nearly-uniform reconstruction quantization (NURQ). In R-D analysis, we clearly explain the property of GGD source under efficient DZ+UTSQ/NURQ. On this basis, in R-D modeling, the heuristic D-R model is proposed.
Yizhou Duan, Jun Sun 0012, Zongming Guo
DCC3
2012 Nonlocal based Super Resolution with rotation invariance and search window relocation
abstract
In this paper, we present a novel method for Super Resolution (SR) reconstruction with rotation invariance and search window relocation. To combine complementary information in observed images to generate a higher resolution image, we first relocate search window to involve potential similar patches and then use rotation invariance similarity measure to find accurate similar patches. Comparing with Nonlocal Means SR, our algorithm can find more similar patches for weighted average. Experimental results demonstrate superior performance of the proposed method in terms of both objective measurements and subjective evaluation.
Yue Zhuo 0005, Jiaying Liu 0001, Jie Ren 0012, Zongming Guo
ICASSP4
2012 Single pass dependent bit allocation for H.264 temporal scalability
abstract
In this paper, we propose a single-pass dependent bit allocation algorithm for H.264/SVC hierarchical B-pictures. To develop a practical bit allocation algorithm, we use the number of skipped blocks and the ratio of the mean absolute difference (MAD) as features to measure the inter-layer signal dependence of input video signals. The proposed algorithm performs bit allocation at the target bit rate with two steps: the group-of-picture (GOP) based rate control and adaptive temporal layer quantization parameter (QP) decision. The superior performance of the proposed algorithm is demonstrated by experimental results, which is compared with two other one-pass bit allocation algorithms in the literature.
Jiaying Liu 0001, Yongjin Cho, Zongming Guo
ICIP3
2012 Illumination-invariant non-local means based video denoising
abstract
In this paper, we present a robust illumination-invariant non-local means (NLM) based video denoising algorithm with special illumination handling. Illumination variances pose a challenge to the NLM-based denoising algorithms. To address this issue, we first propose several possible technical improvements, and verify their efficacy of eliminating the influence of illumination changes. Then, by analyzing and comparing these techniques, a histogram processing based technique is integrated into the non-local means denoising framework. Experimental results on synthesis and real video denoising show that the proposed method is able to fully explore the non-local self-similarity property in natural videos under variable illumination conditions.
Jie Ren 0012, Yue Zhuo 0005, Jiaying Liu 0001, Zongming Guo
ICIP4
2012 A novel JSCC scheme for scalable video transmission over MIMO systems
abstract
MIMO recently emerges as one of promising techniques for wireless video streaming. It is still a challenge to provide un-equal error protections by joint source-channel coding (JSCC) over multiple diverse MIMO sub-channels. In this paper, a joint source-channel coding and antenna mapping scheme for scalable video transmission over MIMO systems is proposed. Bandwidth are elaborately allocated between video source and channel protections by layer extracting and FEC coding. For the extracted layers, we determine i) which antenna will they be transmitted over and ii) how much redundancy bits will be added for error protections. We formulate this scheme into a non-linear integer optimization problem, whose complexity is very high. Instead, a low-complexity branch-and-bound algorithm is presented. Source layers are partitioned into subsets of layers, and the selected layer are mapped to antennas using Min-max scheduling algorithm. By branching and pruning, the computation complexity are reduced significantly. We carry out extensive numerical experiments under various network conditions. The results demonstrate our algorithm's efficiency and the overall transmission quality is improved significantly.
Xinggong Zhang, Chao Zhou 0003, Zongming Guo
ICIP3
2012 Cross-entropy based antenna selection for scalable video streaming over MIMO wireless networks
abstract
In this paper, we investigate the antenna selection (AS) problem for scalable video streaming over MIMO wireless networks. By scheduling scalable video layers over MIMO antennas with different signal strength, the video layers are transmitted with un-equal error protections. Considering layer dependencies and various antenna conditions, it is a non-linear combinatorial problem for AS to minimize the overall end-to-end distortion. To find the optimal solution with low complexity, a cross-entropy based solution, named CEBAS, is proposed. All solutions are indexed by unique binary strings, and the primal problem is reformulated to a binary combination problem. Then, random strings are generated using the probability distribution of solutions, which is updated by the cross-entropy optimization method. The feasibility of solution is guaranteed by our proposed projection strategy. CEBAS is iterative in nature and converges to the global optimum in probability. Simulation results reveal both the effectiveness and efficiency of our proposed algorithm. When comparing CEBAS against other existing algorithms, consistent superior performance has been observed.
Chao Zhou 0003, Xinggong Zhang, Zongming Guo
ICIP3
2012 Image super-resolution by structural sparse coding
Jie Ren 0012, Jiaying Liu 0001, Mengyan Wang, Zongming Guo
ICPR4
2012 Profiling Skype video calls: Rate control and video quality
abstract
Video telephony has recently gained its momentum and is widely adopted by end-consumers. But there have been very few studies on the network impacts of video calls and the user Quality-of-Experience (QoE) under different network conditions. In this paper, we study the rate control and video quality of Skype video calls. We first measure the behaviors of Skype video calls on a controlled network testbed. By varying packet loss rate, propagation delay and bandwidth, we observe how Skype adjusts its rates, FEC redundancy and video quality. We find that Skype is robust against mild packet losses and propagation delays, and can efficiently utilize the available network bandwidth. We also find that Skype employs an overly aggressive FEC protection strategy. Based on the measurement results, we develop rate control model, FEC model, and video quality model for Skype. Extrapolating from the models, we conduct numerical analysis to study the network impacts of Skype. We demonstrate that user back-offs upon quality degradation serve as an effective user-level rate control scheme. We also show that Skype video calls are indeed TCP-friendly and respond to congestion quickly when the network is overloaded.
Xinggong Zhang, Yang Xu 0011, Yong Liu 0013, Zongming Guo, Yao Wang 0001
INFOCOM5
2012 Visual-weighted motion compensation frame interpolation with motion vector refinement
abstract
In this paper, we propose a novel frame rate up-conversion algorithm based on joint motion vector refinement and visual-weighted motion compensation interpolation (MCI). It utilizes a hierarchical motion vector refinement to correct inaccurate motion vectors (MVs), which is composed of the global level and the local level. In the global level, distinct inaccurate MVs are detected by global controlling and then corrected by neighborhood information. Afterwards, the local level performs the local controlling to pick out local outliers and re-estimate them with the maximum likelihood method. Finally, plausible weights for each block in the interpolated frame, computed by the similarity index(SSIM), are applied for visual compensation. The experimental results demonstrate that compared with the conventional algorithm EBME, the proposed algorithm achieved the average PSNR by up to 2.7dB while the visual quality improvement is also remarkable.
Wei Bai 0002, Jiaying Liu 0001, Jie Ren 0012, Zongming Guo
ISCAS4
2012 Novel rate-distortion modeling for H.264/AVC and its application in two-pass VBR coding
abstract
In this paper, we first employ rate-distortion (R-D) modeling to derive the source model for H.264 video coding. This source model consists of the distortion-quantization (D-Q) model and distortion-rate (D-R) model, which describe the relationship between rate, distortion and quantization step for generalized Gaussian distribution (GGD) under the quantization scheme of H.264/AVC. The accuracy of the source model is confirmed by comparing the estimated results with the actual data of JM16.0. And the parameter selection criteria are developed to better employ the source model in our two-pass rate control algorithm for H.264 VBR coding. Based on the D-Q and D-R models, the proposed algorithm is of high stability, low complexity and is easy to implement. Extensive experiments covering different resolutions and target rates demonstrates: 1) about 96.7% reduction in PSNR variance compared to the method of JM16.0 with the average rate control error of 1.95%, and 2) significant improvement in smoothing the video quality compared with the latest two-pass rate control method.
Yizhou Duan, Jun Sun 0012, Zongming Guo
ISCAS3
2012 Sparsity estimation in image compressive sensing
abstract
Compressive sensing is an emerging technology which can recover a K-sparse signal vector from M = O(Klog(K=N)) measurements. However, it is a challenge to know exactly how many measurements an image requires to achieve an acceptable recovered visual quality. In this paper, we study the relationship between the image's complexity and its sparsity. We propose a mathematical model to estimate the number of needed measurements by using the image's texture, the edge density and the target reconstruction quality. There exists a linear function between them. The experimental results with a large number of photo pictures show that, quite most reconstructed images using our pre-calculated number of measurements have good enough quality, which confirms our proposed image-complexity-based model well.
Shanzhen Lan, Xinggong Zhang, Zongming Guo
ISCAS4
2012 Parallelizing video transcoding using Map-Reduce-based cloud computing
abstract
Due to the complexity of video coding, fast transcoding is still a challenge. Various parallel coding methods have been proposed. In this paper, we present a parallel transcoding system over Map/Reduce cloud computing architecture. Input video sequences are divided into segments, and mapped to multiple computers. The sub-tasks are launched in parallel with processing results concatenated to the final output sequences. For heterogeneous clips, computing capacity, and task-launching overhead, the task scheduling over cloud is an NP-hard problem. We propose a low-complexity heuristic algorithm, Max-MCT, to find out the optimal solutions for task scheduling. By estimating the low-bound of finish time, we transform the problem into a virtual knapsack problem. But it is not an optimal solution for the original problem therefore we use a minimal complete time (MCT) algorithm to minimize the entire finish time. We carry out extensive experiments on numerical simulations. The results verified that our algorithm outperforms the existing algorithms.
Feng Lao, Xinggong Zhang, Zongming Guo
ISCAS3
2012 Optimized bit extraction of SVC exploiting linear error model
abstract
The Scalable Video Coding (SVC) extension of the H.264/AVC video coding standard supports fidelity or quality (SNR) scalability. The quality enhancement packets would be discarded in case of limited network capacity, which calls for an optimized bit extraction strategy. In this paper, we first analyze the linear feature in H.264/AVC video coding. A linear error model is also constructed using this feature in case of SVC quality scalability. Then based on the linear error model, the rate and distortion (R-D) impact of each quality enhancement packet over the whole sequence is obtained. Finally a new priority assigning algorithm is designed for a more efficient extraction, giving high rank to those with great R-D impacts. Extensive experiments are presented to demonstrate the accuracy of the linear error model and the validity of the priority assigning algorithm. Tests on the set of eight standard video sequences show the quality promotion under any bitrate constraint, and a fidelity gain up to 0.4 dB PSNR is achieved by the proposed strategy, compared to the JSVM reference software with Quality Layer information.
Jun Sun 0012, Jiaying Liu 0001, Zongming Guo
ISCAS4
2012 Implementation of HEVC decoder on x86 processors with SIMD optimization
abstract
High Efficient Video Coding (HEVC) is the next generation video coding standard in progress. Based on the traditional hybrid coding framework, HEVC implements enhanced tools to improve compression efficiency at the cost of far more computational payload than the capacity of real-time video applications. In this paper, we focus on the software implementation of a real-time HEVC decoder over modern Intel x86 processors. First, we identify the most time-consuming modules of HM 4.0 decoder, represented by motion compensation, adaptive loopfilter, deblocking filter and integer transform. Then the single-execution-multiple-data (SIMD) methods are proposed to optimize the computational performance of these modules. Experimental results show that the optimized decoder is more than 4 times faster than the HM 4.0 decoder, with decoding speed of over 40 frames per second for 1920×1080 resolution videos on Intel i5-2400 processor.
Leju Yan, Yizhou Duan, Jun Sun 0012, Zongming Guo
VCIP4
2012 An optimized real-time multi-thread HEVC decoder
abstract
This demonstration illustrates an optimized HEVC decoder, which is compliant with the reference software HM 4.0. The optimized decoder is more than 4 times faster than the reference decoder, and can well meet the real-time decoding demands of the 1080p high-definition (HD) videos. In addition, based on the frame-level multi-thread framework, the optimized decoder can even achieve up to 13.2 times speedup with 4-thread parallel decoding over Intel i5-2400 processor.
Leju Yan, Yizhou Duan, Jun Sun 0012, Zongming Guo
VCIP4
2012 A control-theoretic approach to rate adaptation for dynamic HTTP streaming
abstract
Recently, dynamic adaptive HTTP streaming has been widely used for video content delivery over Internet. However, it is still a challenge how to switch video bitrate under time-varying bandwidth. In this paper, we propose a novel control-theoretic approach to adapt video segments in dynamic HTTP streaming. The rate control is based on a sink-buffer, which has an overflow-threshold and an underflow-threshold. The objective is to maximize the playback quality while keeping the receiver buffer from either overflow or underflow. Using control theory, we formulate this rate control scheme as a proportional (P) control system, which exists oscillations and steady-errors. Furthermore, we design a proportional derivative (PD) controller to improve its adaptation performance. The conditions for stability and settling time of the PD controller are also derived. Numerous experiment results demonstrate the effectiveness of our proposed PD control scheme for dynamic HTTP streaming.
Chao Zhou 0003, Xinggong Zhang, Longshe Huo, Zongming Guo
VCIP4
2011 Joint Spatial-Temporal Layer Bit Allocation with S-Domain Dependent R-D Modeling
abstract
Summary form only given. H.264/SVC, as a scalable extension of H.264/AVC, is finally standardized in 2007. Scalable video stream has achieved great flexibility and adaptability in terms of frame rates, display resolutions and quality levels. With three dimensions of the scalability, each coding unit in an H.264/SVC video is subject to highly complicated inter-dependency, which gives one of major challenges for bit allocating. In this work, we study an optimal solution to joint spatial-temporal (S-T) bit allocation problem with self-domain R-D modeling. The self-domain (ιS-domain) analysis employs the R-D characteristics of the reference layer as the observation domain of those of dependent layers.
Jiaying Liu 0001, Yongjin Cho, Zongming Guo
DCC3
2011 Practical rate control algorithm for temporal scalability in scalable video coding
abstract
A rate control algorithm for hierarchical B-pictures in Scalable Video Coding (SVC) is proposed in this work. The complex inter-frame dependency issue is effectively addressed by the Q-distance policy decision rule while the statistical smoothing effect enables the GOP-based precise bit rate control. The simplicity of the decision processes greatly reduces the encoder complexity providing an efficient and effective rate control algorithm with hierarchical B-pictures. Experimental results verify the significant performance gain by the proposed algorithm.
Jiaying Liu 0001, Yongjin Cho, Zongming Guo
ICIP3
2011 Similarity modulated block estimation for image interpolation
abstract
Modeling the nonstationarity of image signals is one of the challenging issues for image interpolation. In this paper, we propose a similarity probability modeling to faithfully characterize the nonstationarity of image signals, and present a novel image interpolation algorithm based on the proposed model. The missing pixels are estimated in groups by weighted block estimation. The weight of each pixel inside the block is defined as the similarity probability between itself and the centered to-be-interpolated pixel. It is demonstrated by the experimental results that the proposed method preserves the edge structures of the interpolated images better than the state-of-the-art interpolation methods. Annoying artifacts nearby the sharp edges are also greatly reduced.
Jie Ren 0012, Jiaying Liu 0001, Wei Bai 0002, Zongming Guo
ICIP4
2011 Efficient dead-zone plus uniform threshold scalar quantization of generalized Gaussian random variables
abstract
This paper studies the rate-distortion (R-D) performance of entropy-constrained dead-zone plus uniform threshold scalar quantization (DZ+UTSQ) and nearly-uniform reconstruction quantization (NURQ) for generalized Gaussian distribution (GGD). We first derive the preliminary constraint of R-D optimized DZ+UTSQ/NURQ for GGD. Then for GGD source of actual DCT coefficients, the refined constraint and precise conditions of optimum DZ+UTSQ/NURQ are rigorously deduced in the real coding bit rate range. Based on above analysis, efficient DZ+UTSQ/NURQ design criteria are proposed to reasonably simplify the implementation of effective quantizer in practice.
Yizhou Duan, Jun Sun 0012, Jiaying Liu 0001, Zongming Guo
VCIP4
2011 A novel parallel encoding framework for scalable video coding
abstract
In this paper, we first propose a new parallel video coding framework, considering three important factors: parallel strategy, computational complexity and task scheduling. Then combining the characteristics of scalable video coding (SVC), a novel parallel encoding structure for temporal and quality scalabilities is introduced to obtain a high speedup of parallel SVC. Since the data dependencies in SVC are complex and time variant, the scheduling of parallel SVC is extremely difficult. In order to find the optimal scheduling solution, directed acyclic graph (DAG) is exploited to model the dependencies of encoding tasks, and the complexities of the encoding tasks which are accurately estimated by the Kalman filter to weight the scheduling tasks. Finally, two heuristic scheduling algorithms are also proposed to achieve a high encoding speed of parallel SVC. Experimental results show that the speedup of our parallel method was higher (about 60%) than previous work. Using the proposed method, high definition (HD) SVC videos can be encoded in real time.
Jun Sun 0012, Jiaying Liu 0001, Zongming Guo, Longshe Huo
VCIP4
2010 Price Differentiation All-Pay Auction-Based Incentives in BitTorrent
Zongming Guo
GPC2
2010 Collision-detection based rate-adaptation for video multicasting over IEEE 802.11 wireless networks
abstract
Wireless video multicasting/broadcasting is an efficient method for simultaneous transmission of data to a group of users. But the multicasting rates are fixed in current IEEE 802.11 PHYs standard. In this paper, we propose a novel collision-detection based rate-adaptation scheme (CDRA), which fully exploits the potential of rate adaptation capability of wireless physical layer, to improve service qualities of video multicasting. The received signal strength indication (RSSI) and packet error ratio (PER) are comprehensively used to detect collision. The PER-guided rate adjustment algorithm is performed when no collision happens. Otherwise the collision-avoid mechanism works. By detecting the collision, our scheme could adaptively select the maximum data rates for video multicasting. We construct a practical multicasting test-bed in IEEE 802.11b network and carry out extensive experiments. The results show that CDRA achieves throughput gain up to 166% and PSNR gain to 139% compared with existing methods.
Chao Zhou 0003, Xinggong Zhang, Lichuan Lu, Zongming Guo
ICIP4
2010 Time-constrained packet scheduling optimization for video streaming in wireless ad-hoc networks
abstract
Packet schedule is effective to improve the quality of video streaming over time-varying wireless ad-hoc networks. In this paper, we propose a time-constrained packet scheduling algorithm, which minimizes the video distortion by allocating transmission opportunities to packets under the constraint of playback delay. The problem is formulated in a constrained convex optimization framework and solved with Lagrangian method. The packet loss probabilities and transmission time in IEEE802.11 wireless channel are predicted by using a Markov Chain model. The experiments in NS-2 simulator validate that the algorithm achieves a significant improvement on the quality of streaming.
Xinggong Zhang, Zongming Guo
ISCAS2
2010 Efficient Generalized Integer Transform for Reversible Watermarking
abstract
In this letter, an efficient integer transform based reversible watermarking is proposed. We first show that Tian's difference expansion (DE) technique can be reformulated as an integer transform. Then, a generalized integer transform and a payload-dependent location map are constructed to extend the DE technique to the pixel blocks of arbitrary length. Meanwhile, the distortion can be controlled by preferentially selecting embeddable blocks that introduce less distortion. Finally, the superiority of the proposed method is experimental verified by comparing with other existing schemes.
Xiang Wang 0009, Xiaolong Li 0001, Bin Yang 0001, Zongming Guo
IEEE Signal Process. Lett.4
2010 Bit Allocation for Spatial Scalability Coding of H.264/SVC With Dependent Rate-Distortion Analysis
abstract
We propose a model-based spatial layer bit allocation algorithm for H.264/scalable video coding (SVC) in this paper. The challenge of this problem lies in the fact that the rate-distortion (R-D) behavior of an enhancement layer is dependent on its preceding layers because of inter-layer prediction. To solve it, we first focus on the case of two spatial layers, derive the distortion and rate models of the dependent layer analytically, and develop a low-complexity bit allocation algorithm. It is shown by experimental results that the proposed two-layer bit allocation algorithm can achieve the coding performance close to the optimal R-D performance based on the full search method. Then, we extend this result to multilayer bit allocation by performing the two-layer allocation scheme recursively. Finally, we compare the performance of group of pictures-based and frame-based spatial layer bit allocation schemes at a fixed temporal resolution. The superior performance of the proposed spatial layer bit allocation algorithm is demonstrated using Joint Scalable Video Model reference software algorithm and two prior H.264/SVC rate control algorithms as the benchmarks.
Jiaying Liu 0001, Yongjin Cho, Zongming Guo, C.-C. Jay Kuo
IEEE Trans. Circuits Syst. Video Technol.3
2009 Bit allocation for joint spatial-quality scalability in H.264/SVC
abstract
In this work, we propose a model-based layer bit allocation algorithm for joint spatial-quality (S-Q) scalability in H.264/SVC. The complicated inter-layer dependency is decoupled by the proposed spatial and quality rate and distortion (R-D) models. We show that the R-D characteristics of a dependent layer can be represented by a number of independent functions with GOP as a basic coding unit. Then, the joint bit allocation problem is formulated as a two-step optimization problem by the Lagrangian multiplier method, which can be numerically solved using the proposed R-D models. Finally, we develop a low-complexity bit allocation algorithm for the combined spatial and quality scalability in H.264/SVC. It is shown by experimental results that our proposed bit allocation algorithm achieves the coding performance significantly improved from current reference software JSVM.
Jiaying Liu 0001, Zongming Guo, Yongjin Cho
ICIP2
2009 Frame-based bit allocation for spatial scalability in H.264/SVC
abstract
The spatial scalability of H.264/SVC is achieved by a multi-layer approach, where an enhancement layer is dependent on its preceding layers. To address this dependent issue, we propose a model-based spatial layer bit allocation algorithm for H.264/SVC in this work. The inter-layer dependency is decoupled by analyzing the signal flow in the H.264/SVC encoder. We show that the rate and the distortion (R-D) characteristics of a dependent layer with a frame as a basic coding unit. Finally, a low complexity spatial layer bit allocation scheme is developed using the proposed frame-based R-D models. It is shown by experimental results that our proposed bit allocation algorithm can achieve the coding performance close to the optimal R-D performance of full search and is highly improved from current reference codec JSVM.
Jiaying Liu 0001, Yongjin Cho, Zongming Guo
ICME3
2009 Rate-distortion based path selection for video streaming over wireless ad-hoc networks
abstract
In wireless ad-hoc networks with low bandwidth, high bit error, and node mobility, the quality of streaming video is highly dependent on the quality of routing paths. Most of existing ad-hoc routing algorithms select the path according to network parameters. However, the selected path may not be the best path for video applications. This paper proposed a rate-distortion based (RD) path selection algorithm for video streaming over wireless networks. The video distortion on application layer is used as path metric. On each link, the rate-distortion due to transmission error and congestion is estimated using queuing theory. The algorithm selects the path with the minimum expected rate-distortion as the routing path. Extensive experiments are carried out over NS-2 simulation environment. Simulation results demonstrate that RD routing algorithm can improve the quality of video streaming significantly as compared to the conventional shortest-path routing algorithm.
Xinggong Zhang, Zongming Guo
ICME3
2009 A robust content-based watermarking scheme
abstract
Content-based watermarking techniques have been widely developed to resist geometrical attacks. In the content-based watermarking scheme, invariant regions detected by an affine-invariant detector are exploited for watermark embedding. However, many image manipulations, including watermark embedding, will bring noise to the host image. It results in creating new interest points, removing existing points and changing the shape of invariant regions. Consequently, invariant region detection obtains different results between the embedder and the detector of a watermarking scheme. Thus, the watermark fails to be detected because of the loss of synchronization. In this paper, we present a novel content-based watermarking scheme providing robustness to such manipulations. First, the scale and affine invariant regions are extracted on the de-noised image. A new measure is proposed to select non-overlapping invariant regions out of the available invariant regions for watermark embedding. The watermark is embedded surrounding the selected invariant region. The experimental results show that the proposed scheme is robust to various image processing, including geometric attacks, JPEG compression and filtering.
Zongming Guo
MMSP2
2008 Efficient intra-4×4 mode decision based on bit-rate estimation in H.264/AVC
abstract
Rate-distortion optimization (RDO) technique is widely employed by H.264/AVC for the purpose of determining the best mode. However, such technique results in dramatic increase in the computation complexity of the underlying encoder. In this paper, we address this problem by presenting an efficient intra-4×4 mode decision algorithm. The algorithm works by approximating the bit-rate so as to reduce the computational cost of RDO and the main idea is the following: First we have found a quick way to estimate the bit-rate via the number of DCT coefficients to be quantized to 0 and that to be quantized to ±1. The parameters of the estimated function are adaptively obtained by using the Least Squares Fitting method of the above and the left block in the current frame and the co-location one in the previous encoded frame as the feedback. This close loop bit-rate estimation would skip the processes of quantization, inverse transform, entropy coding and reconstruction; we then use the estimated bit-rate and Sum of Absolute Differences (SAD) to simplify the optimization process of R-D cost function. Experimental results show that our scheme decreases the time for intra coding by 50% with negligible loss of PSNR, and that comparing with those fast mode-decision algorithms based-on local edge direction, the optimal prediction mode obtained via our scheme is closer to that obtained via the original RDO in statistic.
Jiaying Liu 0001, Zongming Guo
ISCAS2
2008 Bit allocation for spatial scalability in H.264/SVC
abstract
We propose a model-based spatial layer bit allocation algorithm for H.264/SVC in this work. The spatial scalability of H.264/SVC is achieved by a multi-layer approach, where an enhancement layer is bound by the dependency on its preceding layers. The inter-layer dependency is decoupled in our analysis by a careful examination of the signal flow in the H.264/SVC encoder. We show that the rate and the distortion (R-D) characteristics of a dependent layer can be represented by a number of independent functions with a group of pictures (GOP) as a basic coding unit. Finally, a low complexity spatial layer bit allocation scheme is developed using the proposed GOP-based R-D models. It is shown by experimental results that our proposed bit allocation algorithm can achieve the coding performance close to the optimal R-D performance of full search and is significantly improved from current reference software JSVM.
Jiaying Liu 0001, Yongjin Cho, Zongming Guo, C.-C. Jay Kuo
MMSP3
2006 A Robust Video Watermarking Scheme Via Temporal Segmentation and Middle Frequency Component Adaptive Modification
Liesen Yang, Zongming Guo
IWDW2
2003 Video clip retrieval by maximal matching and optimal matching in graph theory
abstract
In this paper, a novel approach for automatic matching, ranking and retrieval of video clips is proposed. Motivated by the maximal and optimal matching theories in graph analysis, a new similarity measure of video clips is defined based on the representation and modeling of bipartite graph. Four different factors: visual similarity, granularity, interference and temporal order of shots are taken into consideration for similarity ranking. These factors are progressively analyzed in the proposed approach. Maximal matching utilizes the granularity factor to efficiently filter false matches, while optimal matching takes into account the visual, granularity and interference factors for similarity measure. Dynamic programming is also formulated to quantitatively evaluate the temporal order of shots. The final similarity measure is based on the results of optimal matching and dynamic programming. Experimental results indicate that the proposed approach is effective and efficient in retrieving and ranking similar video clips.
Yuxin Peng 0001, Chong-Wah Ngo, Qing-Jie Dong, Zongming Guo, Jianguo Xiao
ICME4