EDBT 2026 Demo / reviewers in the wild / expert
Yueyu Hu
dblp:198/1410
· DBLP profile ↗
28ranked-venue papers
12as first author
11since 2021 · last 2025
0000-0003-4919-4515ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 11 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Bits-to-Photon: End-to-End Learned Scalable Point Cloud Compression for Direct RenderingabstractPoint cloud is a promising 3D representation for volumetric streaming in emerging AR/VR applications. Despite recent advances in point cloud compression, decoding and rendering high-quality images from lossy compressed point clouds is still challenging in terms of quality and complexity, making it a major roadblock to achieve real-time 6-Degree-of-Freedom video streaming. In this paper, we address this problem by developing a point cloud compression scheme which generates a bit stream that can be directly decoded to renderable 3D Gaussians. The encoder and decoder are jointly optimized to consider both bit-rates and rendering quality. It significantly improves the rendering quality while substantially reducing decoding and rendering time, compared to existing point cloud compression methods. Furthermore, the proposed scheme generates a scalable bit stream, allowing multiple levels of details at different bit-rate ranges. Our method supports real-time color decoding and rendering of high quality point clouds, thus paving the way for interactive 3D streaming applications with free view points. The code is available at https://github.com/huzi96/bits2photon. Yueyu Hu, Yao Wang 0001 |
ICIP | 1 |
| 2025 | Spatial Visibility and Temporal Dynamics: Rethinking Field of View Prediction in Adaptive Point Cloud Video StreamingabstractField-of-View (FoV) adaptive streaming significantly reduces bandwidth requirement of immersive point cloud video (PCV) by only transmitting visible points inside a viewer's FoV. The traditional approaches often focus on trajectory-based 6 degree-of-freedom (6DoF) FoV predictions. The predicted FoV is then used to calculate point visibility. Such approaches do not explicitly consider video content's impact on viewer attention, and the conversion from FoV to point visibility is often error-prone and time-consuming. We reformulate the PCV FoV prediction problem from the cell visibility perspective, allowing for precise decision-making regarding the transmission of 3D data at the cell level based on the predicted visibility distribution. We develop a novel spatial visibility and object-aware graph model (CellSight) that leverages the historical 3D visibility data and incorporates spatial perception, occlusion between points, and neighboring cell correlation to predict the cell visibility in the future. We focus on multi-second ahead prediction to enable the use of long pre-fetching buffers in on-demand streaming, critical for enhancing the robustness to network bandwidth fluctuations. CellSight significantly improves the long-term cell visibility prediction, reducing the prediction Mean Squared Error (MSE) loss by up to 50% compared to the state-of-the-art models when predicting 2 to 5 seconds ahead, while maintaining real-time performance (more than 30fps) for point cloud videos with over 1 million points. Chen Li 0043, Tongyu Zong, Yueyu Hu, Yao Wang 0001, Yong Liu 0013 |
MMSys | 3 |
| 2024 | Standard Compatible Efficient Video Coding with Jointly Optimized Neural WrappersabstractWe present a standard-compatible video coding scheme with end-to-end optimized neural wrapper over standard video codecs that achieves significant rate-distortion (R-D) performance gains and is still efficient in decoding. We train a pair of pre- and post-processor using a differential JPEG proxy. The pre-processor applies a learned transform to the video and downsamples the video by a factor of 2. It generates a bottleneck video to be coded by a standard codec as a YUV sequence. The post-processor takes the decoded bottleneck video, does the inverse transform, and upsamples it to the original resolution. We follow the design in [1] , where we configure downsample using a layer of strided convolution. We optimize the post-processor for efficiency by replacing convolutions with kernel size larger than 1×1 to depth-wise convolutions [2] . Yueyu Hu, Onur G. Guleryuz, Debargha Mukherjee, Yao Wang 0001 |
DCC | 1 |
| 2024 | Standard Compliant Video Coding Using Low Complexity, Switchable Neural WrappersabstractThe proliferation of high resolution videos posts great storage and bandwidth pressure on cloud video services, driving the development of next-generation video codecs. Despite great progress made in neural video coding, existing approaches are still far from economical deployment considering the complexity and rate-distortion performance tradeoff. To clear the roadblocks for neural video coding, in this paper we propose a new framework featuring standard compatibility, high performance, and low decoding complexity. We employ a set of jointly optimized neural pre and post-processors, wrapping a standard video codec, to encode videos at different resolutions. The rate-distorion optimal downsampling ratio is signaled to the decoder at the per-sequence level for each target rate. We design a low complexity neural post-processor architecture that can handle different upsampling ratios. The change of resolution exploits the spatial redundancy in high-resolution videos, while the neural wrapper further achieves rate-distortion performance improvement through end-to-end optimization with a codec proxy. Our light-weight post-processor architecture has a complexity of 516 MACs / pixel, and achieves 9.3% BD-Rate reduction over VVC on the UVG dataset, and $6.4 \%$ on AOM CTC Class A1. Our approach has the potential to further advance the performance of the latest video coding standards using neural processing with minimal added complexity. Yueyu Hu, Onur G. Guleryuz, Debargha Mukherjee, Yao Wang 0001 |
ICIP | 1 |
| 2024 | Video Coding for Machines: Compact Visual Representation Compression for Intelligent Collaborative AnalyticsabstractAs an emerging research practice leveraging recent advanced AI techniques, e.g. deep models based prediction and generation, Video Coding for Machines (VCM) is committed to bridging to an extent separate research tracks of video/image compression and feature compression, and attempts to optimize compactness and efficiency jointly from a unified perspective of high accuracy machine vision and full fidelity human vision. With the rapid advances of deep feature representation and visual data compression in mind, in this paper, we summarize VCM methodology and philosophy based on existing academia and industrial efforts. The development of VCM follows a general rate-distortion optimization, and the categorization of key modules or techniques is established including feature-assisted coding, scalable coding, intermediate feature compression/optimization, and machine vision targeted codec, from broader perspectives of vision tasks, analytics resources, etc. From previous works, it is demonstrated that, although existing works attempt to reveal the nature of scalable representation in bits when dealing with machine and human vision tasks, there remains a rare study in the generality of low bit rate representation, and accordingly how to support a variety of visual analytic tasks. Therefore, we investigate a novel visual information compression for the analytics taxonomy problem to strengthen the capability of compact visual representations extracted from multiple tasks for visual analytics. A new perspective of task relationships versus compression is revisited. By keeping in mind the transferability among different machine vision tasks (e.g. high-level semantic and mid-level geometry-related), we aim to support multiple tasks jointly at low bit rates. In particular, to narrow the dimensionality gap between neural network generated features extracted from pixels and a variety of machine vision features/labels (e.g. scene class, segmentation labels), a codebook hyperprior is designed to compress the neural network-generated features. As demonstrated in our experiments, this new hyperprior model is expected to improve feature compression efficiency by estimating the signal entropy more accurately, which enables further investigation of the granularity of abstracting compact features among different tasks. Wenhan Yang, Haofeng Huang, Yueyu Hu, Ling-Yu Duan, Jiaying Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Neural Data-Dependent Transform for Learned Image CompressionabstractLearned image compression has achieved great success due to its excellent modeling capacity, but seldom further considers the Rate-Distortion Optimization (RDO) of each input image. To explore this potential in the learned codec, we make the first attempt to build a neural data-dependent transform and introduce a continuous online mode decision mechanism to jointly optimize the coding efficiency for each individual image. Specifically, apart from the image content stream, we employ an additional model stream to generate the transform parameters at the decoder side. The pres-ence of a model stream enables our model to learn more abstract neural-syntax, which helps cluster the latent repre-sentations of images more compactly. Beyond the transform stage, we also adopt neural-syntax based post-processing for the scenarios that require higher quality reconstructions regardless of extra decoding overhead. Moreover, the in-volvement of the model stream further makes it possible to optimize both the representation and the decoder in an on-line way, i. e. RDO at the testing time. It is equivalent to a continuous online mode decision, like coding modes in the traditional codecs, to improve the coding efficiency based on the individual input image. The experimental results show the effectiveness of the proposed neural-syntax de-sign and the continuous online mode decision mechanism, demonstrating the superiority of our method in coding effi-ciency. Our project is available at: https://dezhao-wang.github.io/Neural-Syntax-Website/. Dezhao Wang, Wenhan Yang, Yueyu Hu, Jiaying Liu 0001 |
CVPR | 3 |
| 2022 | Learning Neural Volumetric Field for Point Cloud Geometry CompressionabstractDue to the diverse sparsity, high dimensionality, and large temporal variation of dynamic point clouds, it remains a challenge to design an efficient point cloud compression method. We propose to code the geometry of a given point cloud by learning a neural volumetric field. Instead of representing the entire point cloud using a single overfit network, we divide the entire space into small cubes and represent each non-empty cube by a neural network and an input latent code. The network is shared among all the cubes in a single frame or multiple frames, to exploit the spatial and temporal redundancy. The neural field representation of the point cloud includes the network parameters and all the latent codes, which are generated by using back-propagation over the network parameters and its input. By considering the entropy of the network parameters and the latent codes as well as the distortion between the original and reconstructed cubes in the loss function, we derive a rate-distortion (R-D) optimal representation. Experimental results show that the proposed coding scheme achieves superior R-D performances compared to the octree-based G-PCC, especially when applied to multiple frames of a point cloud video. The code is available at https://github.com/huzi96/NVFPCC/. Yueyu Hu, Yao Wang 0001 |
PCS | 1 |
| 2022 | Learning End-to-End Lossy Image Compression: A BenchmarkabstractImage compression is one of the most fundamental techniques and commonly used applications in the image and video processing field. Earlier methods built a well-designed pipeline, and efforts were made to improve all modules of the pipeline by handcrafted tuning. Later, tremendous contributions were made, especially when data-driven methods revitalized the domain with their excellent modeling capacities and flexibility in incorporating newly designed modules and constraints. Despite great progress, a systematic benchmark and comprehensive analysis of end-to-end learned image compression methods are lacking. In this paper, we first conduct a comprehensive literature survey of learned image compression methods. The literature is organized based on several aspects to jointly optimize the rate-distortion performance with a neural network, i.e., network architecture, entropy model and rate control. We describe milestones in cutting-edge learned image-compression methods, review a broad range of existing works, and provide insights into their historical development routes. With this survey, the main challenges of image compression methods are revealed, along with opportunities to address the related issues with recent advanced learning methods. This analysis provides an opportunity to take a further step towards higher-efficiency image compression. By introducing a coarse-to-fine hyperprior model for entropy estimation and signal reconstruction, we achieve improved rate-distortion performance, especially on high-resolution images. Extensive benchmark experiments demonstrate the superiority of our model in rate-distortion performance and time complexity on multi-core CPUs and GPUs. Yueyu Hu, Wenhan Yang, Zhan Ma 0001, Jiaying Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Towards Low Light Enhancement With RAW ImagesabstractIn this paper, we make the first benchmark effort to elaborate on the superiority of using RAW images in the low light enhancement and develop a novel alternative route to utilize RAW images in a more flexible and practical way. Inspired by a full consideration on the typical image processing pipeline, we are inspired to develop a new evaluation framework, Factorized Enhancement Model (FEM), which decomposes the properties of RAW images into measurable factors and provides a tool for exploring how properties of RAW images affect the enhancement performance empirically. The empirical benchmark results show that the Linearity of data and Exposure Time recorded in meta-data play the most critical role, which brings distinct performance gains in various measures over the approaches taking the sRGB images as input. With the insights obtained from the benchmark results in mind, a RAW-guiding Exposure Enhancement Network (REENet) is developed, which makes trade-offs between the advantages and inaccessibility of RAW images in real applications in a way of using RAW images only in the training phase. REENet projects sRGB images into linear RAW domains to apply constraints with corresponding RAW images to reduce the difficulty of modeling training. After that, in the testing phase, our REENet does not rely on RAW images. Experimental results demonstrate not only the superiority of REENet to state-of-the-art sRGB-based methods and but also the effectiveness of the RAW guidance and all components. Haofeng Huang, Wenhan Yang, Yueyu Hu, Jiaying Liu 0001, Ling-Yu Duan |
IEEE Trans. Image Process. | 3 |
| 2021 | Bridging the Gap Between Computational Photography and Visual RecognitionabstractWhat is the current state-of-the-art for image restoration and enhancement applied to degraded images acquired under less than ideal circumstances? Can the application of such algorithms as a pre-processing step improve image interpretability for manual analysis or automatic visual recognition to classify scene content? While there have been important advances in the area of computational photography to restore or enhance the visual quality of an image, the capabilities of such techniques have not always translated in a useful way to visual recognition tasks. Consequently, there is a pressing need for the development of algorithms that are designed for the joint problem of improving visual appearance and recognition, which will be an enabling factor for the deployment of visual recognition tools in many real-world scenarios. To address this, we introduce the UG$^2$dataset as a large-scale benchmark composed of video imagery captured under challenging conditions, and two enhancement tasks designed to test algorithmic impact on visual quality and automatic object recognition. Furthermore, we propose a set of metrics to evaluate the joint improvement of such tasks as well as individual algorithmic advances, including a novel psychophysics-based evaluation regime for human assessment and a realistic set of quantitative measures for object recognition performance. We introduce six new algorithms for image restoration or enhancement, which were created as part of the IARPA sponsored UG$^2$Challenge workshop held at CVPR 2018. Under the proposed evaluation regime, we present an in-depth analysis of these algorithms and a host of deep learning-based and classic baseline approaches. From the observed results, it is evident that we are in the early days of building a bridge between computational photography and visual recognition, leaving many opportunities for innovation in this area. Rosaura G. VidalMata, Sreya Banerjee, Brandon RichardWebster, Michael Albright, Pedro Davalos, Scott McCloskey, Ben Miller, Asong Tambo, Sushobhan Ghosh, Sudarshan Nagesh, Ye Yuan 0012, Yueyu Hu, Wenhan Yang, Xiaoshuai Zhang, Jiaying Liu 0001, Zhangyang Wang, Hwann-Tzong Chen, Tzu-Wei Huang, Wen-Chi Chin, Yi-Chun Li, Mahmoud Lababidi, Charles Otto, Walter J. Scheirer |
IEEE Trans. Pattern Anal. Mach. Intell. | 12 |
| 2021 | Towards Coding for Human and Machine Vision: Scalable Face Image CodingabstractThe past decades have witnessed the rapid development of image and video coding techniques in the era of big data. However, the signal fidelity-driven coding pipeline design limits the capability of the existing image/video coding frameworks to fulfill the needs of both machine and human vision. In this paper, we come up with a novel face image coding framework by leveraging both the compressive and the generative models, to support machine vision and human perception tasks jointly. Given an input image, the feature analysis is first applied, and then the generative model is employed to reconstruct image with compact structure and color features, where sparse edges are extracted to connect both kinds of vision and a key reference pixel selection method is proposed to determine the priorities of the reference color pixels for scalable coding. The compact edge map serves as the basic layer for machine vision tasks, and the reference pixels act as an enhanced layer to guarantee signal fidelity for human vision. By introducing advanced generative models, we train a decoding network to reconstruct images from compact structure and color representations, which is flexible to accept inputs in a scalable way and to control the imagery effect of the outputs between signal fidelity and visual realism. Experimental results and comprehensive performance analysis over the face image dataset demonstrate the superiority of our framework in both human vision tasks and machine vision tasks, which provide useful evidence on the emerging standardization efforts on MPEG VCM (Video Coding for Machine). Shuai Yang 0001, Yueyu Hu, Wenhan Yang, Ling-Yu Duan, Jiaying Liu 0001 |
IEEE Trans. Multim. | 2 |
| 2020 | Coarse-to-Fine Hyper-Prior Modeling for Learned Image CompressionabstractApproaches to image compression with machine learning now achieve superior performance on the compression rate compared to existing hybrid codecs. The conventional learning-based methods for image compression exploits hyper-prior and spatial context model to facilitate probability estimations. Such models have limitations in modeling long-term dependency and do not fully squeeze out the spatial redundancy in images. In this paper, we propose a coarse-to-fine framework with hierarchical layers of hyper-priors to conduct comprehensive analysis of the image and more effectively reduce spatial redundancy, which improves the rate-distortion performance of image compression significantly. Signal Preserving Hyper Transforms are designed to achieve an in-depth analysis of the latent representation and the Information Aggregation Reconstruction sub-network is proposed to maximally utilize side-information for reconstruction. Experimental results show the effectiveness of the proposed network to efficiently reduce the redundancies in images and improve the rate-distortion performance, especially for high-resolution images. Our project is publicly available at https://huzi96.github.io/coarse-to-fine-compression.html. Yueyu Hu, Wenhan Yang, Jiaying Liu 0001 |
AAAI | 1 |
| 2020 | Raw-Guided Enhancing Reprocess of Low-Light Image via Deep Exposure Adjustment
Haofeng Huang, Wenhan Yang, Yueyu Hu, Jiaying Liu 0001 |
ACCV (2) | 3 |
| 2020 | Towards Coding For Human And Machine Vision: A Scalable Image Coding ApproachabstractThe past decades have witnessed the rapid development of image and video coding techniques in the era of big data. However, the signal fidelity-driven coding pipeline design limits the capability of the existing image/video coding frameworks to fulfill the needs of both machine and human vision. In this paper, we come up with a novel image coding framework by leveraging both the compressive and the generative models, to support machine vision and human perception tasks jointly. Given an input image, the feature analysis is first applied, and then the generative model is employed to perform image reconstruction with features and additional reference pixels, in which compact edge maps are extracted in this work to connect both kinds of vision in a scalable way. The compact edge map serves as the basic layer for machine vision tasks, and the reference pixels act as a sort of enhanced layer to guarantee signal fidelity for human vision. By introducing advanced generative models, we train a flexible network to reconstruct images from compact feature representations and the reference pixels. Experimental results demonstrate the superiority of our framework in both human visual quality and facial landmark detection, which provide useful evidence on the emerging standardization efforts on MPEG VCM (Video Coding for Machine). Our project website is available at https://williamyang1991.github.io/projects/VCM-Face/. Yueyu Hu, Shuai Yang 0001, Wenhan Yang, Ling-Yu Duan, Jiaying Liu 0001 |
ICME | 1 |
| 2020 | A Benchmark Dataset and Comparison Study for Multi-modal Human Action AnalyticsabstractLarge-scale benchmarks provide a solid foundation for the development of action analytics. Most of the previous activity benchmarks focus on analyzing actions in RGB videos. There is a lack of large-scale and high-quality benchmarks for multi-modal action analytics. In this article, we introduce PKU Multi-Modal Dataset (PKU-MMD), a new large-scale benchmark for multi-modal human action analytics. It consists of about 28,000 action instances and 6.2 million frames in total and provides high-quality multi-modal data sources, including RGB, depth, infrared radiation (IR), and skeletons. To make PKU-MMD more practical, our dataset comprises two subsets under different settings for action understanding, namely Part I and Part II. Part I contains 1,076 untrimmed video sequences with 51 action classes performed by 66 subjects, while Part II contains 1,009 untrimmed video sequences with 41 action classes performed by 13 subjects. Compared to Part I, Part II is more challenging due to short action intervals, concurrent actions and heavy occlusion. PKU-MMD can be leveraged in two scenarios: action recognition with trimmed video clips and action detection with untrimmed video sequences. For each scenario, we provide benchmark performance on both subsets by conducting different methods with different modalities under two evaluation protocols, respectively. Experimental results show that PKU-MMD is a significant challenge to many state-of-the-art methods. We further illustrate that the features learned on PKU-MMD can be well transferred to other datasets. We believe this large-scale dataset will boost the research in the field of action analytics for the community. Jiaying Liu 0001, Sijie Song, Chunhui Liu 0002, Yanghao Li, Yueyu Hu |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2019 | Partition Tree Guided Progressive Rethinking Network for in-Loop Filtering of HEVCabstractIn-Loop filter is a key part in High Efficiency Video Coding (HEVC) which effectively removes the compression artifacts. Recently, many newly proposed methods combine residual learning and dense connection to construct a deeper network for better in-loop filtering performance. However, the long-term dependency between blocks is neglected, and information usually passes between blocks only after dimension compression. To address these issues, we propose the Progressive Rethinking Block (PRB) to deliver long-term memory between the neighboring blocks and allow information to flow without compression, which is similar to human decision mechanism - usually reviewing the complete past memorized experiences to decide in the present, not just based on simple principles summarized before. PRBs further establish the Progressive Rethinking Network (PRN). In addition, we calculate the Multi-scale Mean value of Coding Units (MM-CU) to generate the side information maps which guide the training of the network by novelly telling the network architecture of the entire coding partition tree. Experimental results show that our proposed partition tree guided PRN provides 10.1% BD-rate reduction on average compared to the HEVC baseline. Dezhao Wang, Sifeng Xia, Wenhan Yang, Yueyu Hu, Jiaying Liu 0001 |
ICIP | 4 |
| 2019 | Deep Inter Prediction Via Pixel-Wise Motion Oriented Reference GenerationabstractInter prediction is an important module in video coding for temporal redundancy removal, where the reference blocks are searched from the previously coded frames and employed to predict the block to be coded. However, apart from regular block-wise shift motion, there usually exists inconsistent pixel-wise motion such as rotation and deformation between blocks, which will largely degrade the prediction performance. In this paper, we propose a Multiscale Adaptive Separable Convolutional Neural Network (MASCNN) to generate pixel-wise closer reference frames for inter prediction. A multiscale network is built to interpolate the target frame from coarse to fine. Reconstruction losses are enforced on each scale to make the network infer the main structure at small scales, which improves the interpolation accuracy of the network. Furthermore, a sum of absolute transformed difference (SATD) loss function is proposed to regularize the network training, which further improves the coding performance. Compared with HEVC, our method can obtain on average 5.7% BD-rate saving and up to 9.9% BD-rate saving for the luma component under the random access configuration. Sifeng Xia, Wenhan Yang, Yueyu Hu, Jiaying Liu 0001 |
ICIP | 3 |
| 2019 | Switch Mode Based Deep Fractional Interpolation in Video CodingabstractFractional interpolation is a significant technology in motion compensation of video coding. It generates sub-pixel level reference samples in inter prediction to facilitate temporal redundancy removal between video frames. Recently, some methods explore to introduce the deep learning technique for fractional interpolation and have obtained better compression results. However, existing deep learning based methods still treat fractional interpolation as a traditional interpolation problem but fail to adjust it to the motion compensation scenario. In this paper, we design a switch mode based deep fractional interpolation method to introduce integer pixels of different positions to the interpolation of sub-pixel position samples. By switching between integer pixels of different positions, our method can infer the sub-pixels with smaller variations and achieve better fractional interpolation results. Consequently the motion compensation performance can be further improved. Experimental results have also verified the efficiency of the switch mode based deep fractional interpolation. Compared with High Efficiency Video Coding, our method achieves 2.8% bit saving on average and up to 6.2% bit saving under low-delay P configuration. Sifeng Xia, Wenhan Yang, Yueyu Hu, Wen-Huang Cheng, Jiaying Liu 0001 |
ISCAS | 3 |
| 2019 | Progressive Spatial Recurrent Neural Network for Intra PredictionabstractIntra prediction is an important component of modern video codecs, which is able to efficiently squeeze out the spatial redundancy in video frames. With preceding pixels as the context, traditional intra prediction schemes generate linear predictions based on several predefined directions (i.e., modes) for blocks to be encoded. However, these modes are relatively simple and their predictions may fail when facing blocks with complex textures, which leads to additional bits encoding the residue. In this paper, we design a progressive spatial recurrent neural network (PS-RNN) that learns to conduct intra prediction. Specifically, our PS-RNN consists of three spatial recurrent units and progressively generates predictions by passing information along from preceding contents to blocks to be encoded. To make our network generate predictions considering both distortion and bit rate, we propose using sum of absolute transformed difference (SATD) as the loss function to train PS-RNN since SATD is able to measure rate-distortion cost of encoding a residue block. Moreover, our method supports variable-block-size for intra prediction, which is more practical in real coding conditions. The proposed intra prediction scheme achieves on average 2.5% bit-rate reduction on variable-block-size settings under the same reconstruction quality compared with HEVC. Yueyu Hu, Wenhan Yang, Mading Li, Jiaying Liu 0001 |
IEEE Trans. Multim. | 1 |
| 2018 | Enhanced Intra Prediction with Recurrent Neural Network in Video CodingabstractIntra prediction is one of the important parts in video/image codec. With intra prediction mechanism, spatial redundancy can be largely removed for further bit saving. However, current state-of-the-art intra prediction method does not produce satisfactory prediction result due to its limits in reference samples and modeling ability. To enhance the intra prediction in HEVC, in this paper, a deep neural network featuring spatial RNN, which models the spatial dependency of pixels as sequential dynamics, is proposed to generate better prediction signals. Experimental results show improvement in BD-Rate for the proposed method compared with the original HEVC prediction scheme. Yueyu Hu, Wenhan Yang, Sifeng Xia, Wen-Huang Cheng, Jiaying Liu 0001 |
DCC | 1 |
| 2018 | A Group Variational Transformation Neural Network for Fractional Interpolation of Video CodingabstractMotion compensation is an important technology in video coding to remove the temporal redundancy between coded video frames. In motion compensation, fractional interpolation is used to obtain more reference blocks at sub-pixel level. Existing video coding standards commonly use fixed interpolation filters for fractional interpolation, which are not efficient enough to handle diverse video signals well. In this paper, we design a group variational transformation convolutional neural network (GVTCNN) to improve the fractional interpolation performance of the luma component in motion compensation. GVTCNN infers samples at different sub-pixel positions from the input integer-position sample. It first extracts a shared feature map from the integer-position sample to infer various sub-pixel position samples. Then a group variational transformation technique is used to transform a group of copied shared feature maps to samples at different sub-pixel positions. Experimental results have identified the interpolation efficiency of our GVTCNN. Compared with the interpolation method of High Efficiency Video Coding, our method achieves 1.9% bit saving on average and up to 5.6% bit saving under low-delay P configuration. Sifeng Xia, Wenhan Yang, Yueyu Hu, Siwei Ma 0001, Jiaying Liu 0001 |
DCC | 3 |
| 2018 | Dmcnn: Dual-Domain Multi-Scale Convolutional Neural Network for Compression Artifacts RemovalabstractJPEG is one of the most commonly used standards among lossy image compression methods. However, JPEG compression inevitably introduces various kinds of artifacts, especially at high compression rates, which could greatly affect the Quality of Experience (QoE). Recently, convolutional neural network (CNN) based methods have shown excellent performance for removing the JPEG artifacts. Lots of efforts have been made to deepen the CNN s and extract deeper features, while relatively few works pay attention to the receptive field of the network. In this paper, we illustrate that the quality of output images can be significantly improved by enlarging the receptive fields in many cases. One step further, we propose a Dual-domain Multi-scale CNN (DMCNN) to take full advantage of redundancies on both the pixel and DCT domains. Experiments show that DMCNN sets a new state-of-the-art for the task of JPEG artifact removal. Xiaoshuai Zhang, Wenhan Yang, Yueyu Hu, Jiaying Liu 0001 |
ICIP | 3 |
| 2018 | Optimized Spatial Recurrent Network for Intra Prediction in Video CodingabstractIntra prediction in modern video codecs is able to efficiently reduce spatial redundancy in video frames. With preceding pixels as context, traditional intra prediction schemes generate linear predictions based on several predefined directions (i.e. modes) for the current prediction unit (PU). However, these modes are relatively simple and are not able to handle complex textures, which leads to additional bits encoding the residue. In this paper, we design a convolutional neural network (CNN) guided spatial recurrent neural network (RNN) to improve the intra prediction in High-Efficiency Video Coding (HEVC). By exploring the correlations between pixels, the network learns to generate prediction signal in a progressive manner. The progressive model solves the problem of asymmetry in intra prediction naturally. As the model is designed for global context modeling, no flags for intra prediction modes selection need to be encoded. Our proposed intra prediction scheme achieves on average 1.2% bit-rate saving compared with HEVC. Yueyu Hu, Wenhan Yang, Sifeng Xia, Jiaying Liu 0001 |
VCIP | 1 |
| 2017 | Temporal Perceptive Network for Skeleton-Based Action Recognition
Yueyu Hu, Chunhui Liu 0002, Yanghao Li, Jiaying Liu 0001 |
BMVC | 1 |
| 2017 | Online action detection and forecast via Multitask deep Recurrent Neural NetworksabstractOnline human action detection and forecast on untrimmed 3D skeleton sequences is a novel task based on traditional action recognition and has not been fully studied. Its aim is to localize and recognize one action in a long sequence while doing forecasting task at the same time. In this paper, we propose an online detection algorithm featuring Multi-Task Recurrent Neural Network to solve this problem. First, a deep Long Short Term Memory (LSTM) network is designed for feature extraction and temporal dynamic modeling. Then we utilize a classification subnetwork to classify one action, and predict the status of it at the same time. To forecast the occurrence of actions and estimate the accurate time of occurrence, we incorporate a regression subnetwork to our model. Then we split the action classes to three stages and train the model by optimizing a joint classification regression objective function. Experimental results show that the proposed model achieves satisfactory results on online action detection and forecast. Chunhui Liu 0002, Yanghao Li, Yueyu Hu, Jiaying Liu 0001 |
ICASSP | 3 |
| 2017 | Real-Time Deep Video SpaTial Resolution UpConversion SysTem (STRUCT++ Demo)abstractImage and video super-resolution (SR) has been explored for several decades. However, few works are integrated into practical systems for real-time image and video SR. In this work, we present a real-time deep video SpaTial Resolution UpConversion SysTem (STRUCT++). Our demo system achieves real-time performance (50 fps on CPU for CIF sequences and 45 fps on GPU for HDTV videos) and provides several functions: 1) batch processing; 2) full resolution comparison; 3) local region zooming in. These functions are convenient for super-resolution of a batch of videos (at most 10 videos in parallel), comparisons with other approaches and observations of local details of the SR results. The system is built on a Global context aggregation and Local queue jumping Network (GLNet). It has a thinner and deeper network structure to aggregate global context with an additional local queue jumping path to better model local structures of the signal. GLNet achieves state-of-the-art performance for real-time video SR. Wenhan Yang, Shihong Deng, Yueyu Hu, Junliang Xing, Jiaying Liu 0001 |
ACM Multimedia | 3 |
| 2017 | An optimal spatial-temporal smoothness approach for tile-based 360-degree video streamingabstractThe world is becoming more and more virtual than we ever thought it would be. Many video service providers have rolled out 360-degree videos which provide immersive experience to users. However, huge bandwidth occupation of 360-degree video hinders its wide spreading over the Internet. Besides, only part of the video is displayed on the screen, transmitting whole video results in waste of bandwidth and computational resources. Tile-based adaptive streaming is regarded as a bandwidth-friendly approach which only delivers specific portion of the whole video. It requires the clients to decide which portion and at which bitrates to deliver. However, due to both space and time partition of 360-degree videos in tile-based adaptive streaming, there still exists a challenge on the quality inconsistence on spatial and temporal domains. In this paper, we propose a optimal spatial-temporal smoothness approach under restricted network for tile-based adaptive streaming. The bitrates of tiles are determined optimally, aiming at maximizing the overall quality while minimizing the spatial and temporal quality variation. By conducting extensive experiments over real bandwidth dataset and user's head movement traces, our approach can get a significant improvement. Specifically, the Viewport-PSNR can be raised by 24.1% compared with traditional delivery of whole 360 video; while the spatial and temporal stability can be improved by 40.5% and 24.6% respectively compared with tile-based streaming. Yixuan Ban, Lan Xie, Zhimin Xu 0001, Xinggong Zhang, Zongming Guo, Yueyu Hu |
VCIP | 6 |
| 2017 | Real-time deep image super-resolution via global context aggregation and local queue jumpingabstractDeep learning-based image super-resolution has provided very impressive reconstruction quality. However, their running time still sets barriers for real-time applications. In this paper, we propose a Global context aggregation and Local queue jumping Network (GLNet) which provides the more effective image SR given a certain number of model parameters. In our GLNet, we reconsider the model design of the real-time image SR paradigm. Then, we construct a deep network with fewer channels but a deeper structure to effectively aggregate the global context. The dilated convolutions are used as parts of basic units of our GLNet, which further enlarges the receptive field. Besides, an additional local queue jumping path is employed to connect the first-layer feature map and the last-layer feature map to better model the local signal structure. Extensive experiments demonstrate the superiority of our GLNet which offers new state-of-the-art performance considering both reconstruction quality and time consumption. Yueyu Hu, Jiaying Liu 0001, Wenhan Yang, Shihong Deng, Luyao Zhang 0007, Zongming Guo |
VCIP | 1 |