Wen Gao 0001

dblp:g/WenGao · DBLP profile ↗
← Back
71ranked-venue papers in the field
2as first author
17since 2021 · last 2025
0000-0001-8894-1806ORCID · conflict

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 51Other / Interdisciplinary · 6 (1 first)Data Mining & Knowledge Discovery · 5Information Retrieval & Web Search · 5 (1 first)Knowledge Engineering, Semantic Web & Information Systems · 3Database Systems & Data Management · 1
YearPublicationVenuePosition
2025 STACO: Spatio-Temporal Adaptive Context Optimization for Neural Video Compression
abstract
This paper introduces the Spatio-Temporal Adaptive Context Optimization (STACO) method, which enhances the quality of contextual prediction across various resolutions, essential for subsequent compression. The STACO takes predicted contexts$C_t^{\{1,2,3\}}$as input and improves their quality by aligning them better with decoded features$f_{t}$, thus boosting coding efficiency. The STACO comprises Quality Perception Units and Consistency Synergy Modules, arranged in a hierarchical stacked architecture. This multi-scale design enables simultaneous processing of contexts at different spatial resolutions and facilitates information exchange through upsampling and downsampling. Enhanced contexts$\tilde{C}_{t}^{\{1,2,3\}}$are output after passing through residual connections, ensuring better alignment with reconstructed features. Using VTM-11.0 as anchor, the STACO significantly improves compression efficiency on common test condition (CTC) in HEVC, achieving an average BD-rate reduction of 17.99% for PSNR and 43.84% for MS-SSIM. By incorporating spatial quality mapping and temporal propagation, STACO offers a significant advancement in video compression.
Kexiang Feng, Shuhong Liao, Zhimeng Huang, Chuanmin Jia, Siwei Ma 0001, Wen Gao 0001
DCC8
2025 Rethinking Bjøntegaard Delta for Compression Efficiency Evaluation: Are we Calculating it Precisely and Reliably?
abstract
For decades, the Bjøntegaard Delta (BD) has been the metric for evaluating codec Rate-Distortion (R-D) performance. Yet, in most studies, BD is determined using just 4–5 R-D data points, could this be sufficient? As codecs and quality metrics advance, does the conventional BD estimation still hold up? Crucially, are the performance improvements of new codecs and tools genuine, or merely artifacts of estimation flaws? We address these concerns by reevaluating BD estimation. We have established a large-scale, high-precision R-D dataset to verify the accuracy of existing BD estimation algorithms. Moreover, we propose a robust method for high-precision BD estimation across diverse compression scenarios, enhanced by a reliability assessment to determine the probability distribution of BD values from R-D sample points. This approach both assesses the reliability of BD calculations and serves as a precise BD estimator. Our method's validity is confirmed through extensive testing on a dataset we constructed. Our findings advocate for the adoption of rigorous R-D sampling and reliability metrics in future compression research to ensure the validity and reliability of results. Our code and additional experimental details are publicly accessible at https://github.com/fgvfgfg564/BDCI.
Xinyu Hang, Shenpeng Song, Zhimeng Huang, Chuanmin Jia, Siwei Ma 0001, Wen Gao 0001
DCC6
2025 Image Coding for Machine with Visual-Language Mimic Feature Learning
abstract
This paper propose a Image Coding for Machine (ICM) framework with Visual-Language Mimic Feature Learning (VLM-ICM). VLM-ICM decouples the position and semantic information into language modality and extracts universal features from the input image. Language, inherently more semantically compact, helps reduce the bitrate. Meanwhile, the universal features in VLM-ICM, guided by the language at the decoder side, allow for flexible domain adaptation, thereby enhancing versatility and practicality.
Zhimeng Huang, Junlong Gao, Jiaqi Zhang 0007, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001, Chuanmin Jia
DCC6
2025 FAPC: Frequency-Based Adaptive Pixel Correction for Compressed Screen Content
abstract
Screen content is an important category of video. The statistical distribution of pixels in screen content exhibits substantial differences compared to camera-captured video, leading to different compression needs and challenges. Hitherto, most Screen Content Coding (SCC) tools are designed for block-based hybrid coding frameworks, which may not be suitable for emerging wavelet-based and learning-based coding frameworks. Consequently, a plug-and-play SCC tool independent of coding frameworks is lacking in the current video coding landscape. In this paper, an out-loop coding method, Frequency-based Adaptive Pixel Correction (FAPC), is proposed to improve the SCC performance for arbitrary codecs. First, a True Color Value (TCV) table is established based on the most frequently occurring pixel values. Then, the reconstructed pixels are corrected according to the TCV table. To realize precise pixel correction, an adaptive threshold derivation method is meticulously designed to control the pixel correction process. Furthermore, an inheritance coding strategy is proposed to reduce the overhead of parameter transmission. The proposed method has been integrated into three different coding frameworks. Simulation results demonstrate that the proposed method can achieve 10.66%, 3.55% and 10.67% luma component BD-BR gains on the three coding frameworks, respectively. These results prove the superiority and universality of the proposed method.
Zetian Song, Jiaqi Zhang 0007, Chuanmin Jia, Siwei Ma 0001, Wen Gao 0001
DCC7
2025 MoRLACS: A Monocular RGBD-based Locomotion Approach for CAVE Systems
abstract
Navigation within Cave Automatic Virtual Environment (CAVE) systems often faces challenges due to limited physical space and the necessity for seamless user interaction. Traditional solutions typically rely on multi-view tracking systems or constrained locomotion techniques, which can interrupt immersion and hinder usability. In this paper, we introduce MoRLACS, a novel locomotion approach for CAVE systems that leverages a single RGBD camera. This hybrid framework integrates small-scale physical walking with controller-based large-scale exploration through a tailored guidance method. By accurately tracking the user's head position in the real world and synchronizing it with the virtual camera, MoRLACS enables natural walking within confined CAVE spaces and supports extended interaction in larger virtual environments. Preliminary user experiments demonstrate the approach's effectiveness, revealing improvements in usability and a heightened sense of presence. These findings underscore the potential of MoRLACS to enrich user experiences in immersive CAVE settings and offer valuable design insights for integrating 3D sensor data into multimedia interaction frameworks.
Haopeng Lu, Qian Yin 0002, Li Song 0001, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001
ICMR8
2025 Peng Cheng Cloud Brain and Mind Series of Large Model
abstract
As a revolutionary pre-training model, ChatGPT has already had a huge impact on global economic. It is the strong foundation of computing power that enables large models to continuously improve in the process of understanding massive data, resulting in breakthrough innovations. Based on the Peng Cheng Cloud Brain II E-level intelligent computing platform, Peng Cheng Laboratory is training PCL Mind Series of Large Model. Mind is the first fully autonomous, controllable, safe, open-source pre-training foundation model in China, where the performance of the 200 billion parameter base model reaches the international advanced, and the output content conforms to the Chinese core values. Peng Cheng Laboratory is opening up PCL Mind cooperation and work with external partners to continuously build a large model open-source consortium for domestic large model ecosystem. The next generation-Peng Cheng Cloud Brain III will break through key technologies such as high computing power chips, large-scale networking communication, high-performance software stacks, and large-scale parallel training, and support 10,000 chip-level parallel training of trillion-level parameter AI.
Wen Gao 0001
WWW1
2024 Lightweight super resolution network for point cloud geometry compression
abstract
We present an approach for compressing point cloud geometry by leveraging a lightweight super-resolution network. It involves decomposing a point cloud into a base point cloud and the interpolation patterns for reconstructing the original point cloud. While the base point cloud can be efficiently compressed using any lossless codec, such as Geometry-based Point Cloud Compression, a distinct strategy is employed for handling the interpolation patterns. Rather than directly compressing the interpolation patterns, a lightweight super-resolution network is utilized to learn this information through overfitting. Subsequently, the network parameter is transmitted to assist in point cloud reconstruction at the decoder side. Our approach differentiates itself from lookup table-based methods, allowing us to obtain more accurate interpolation patterns by accessing a broader range of neighboring voxels at an acceptable computational cost. Experiments on MPEG Cat1 (Solid) and Cat2 datasets demonstrate the remarkable compression performance achieved by our method.
Wei Zhang 0072, Dingquan Li, Ge Li 0002, Wen Gao 0001
DCC4
2024 IME: Integrating Multi-curvature Shared and Specific Embedding for Temporal Knowledge Graph Completion
abstract
Temporal Knowledge Graphs (TKGs) incorporate a temporal dimension, allowing for a precise capture of the evolution of knowledge and reflecting the dynamic nature of the real world. Typically, TKGs contain complex geometric structures, with various geometric structures interwoven. However, existing Temporal Knowledge Graph Completion (TKGC) methods either model TKGs in a single space or neglect the heterogeneity of different curvature spaces, thus constraining their capacity to capture these intricate geometric structures. In this paper, we propose a novel Integrating Multi-curvature shared and specific Embedding (IME) model for TKGC tasks. Concretely, IME models TKGs into multi-curvature spaces, including hyperspherical, hyperbolic, and Euclidean spaces. Subsequently, IME incorporates two key properties, namely space-shared property and space-specific property. The space-shared property facilitates the learning of commonalities across different curvature spaces and alleviates the spatial gap caused by the heterogeneous nature of multi-curvature spaces, while the space-specific property captures characteristic features. Meanwhile, IME proposes an Adjustable Multi-curvature Pooling (AMP) approach to effectively retain important information. Furthermore, IME innovatively designs similarity, difference, and structure loss functions to attain the stated objective. Experimental results clearly demonstrate the superior performance of IME over existing state-of-the-art TKGC models.
Jiapu Wang, Boyue Wang, Shirui Pan, Junbin Gao, Wen Gao 0001
WWW7
2023 Rate-Distortion Optimization for Cross Modal Compression
abstract
Recently, cross modal compression (CMC) is proposed to compress highly redundant visual data into a compact, common, human-comprehensible domain (such as text) to preserve semantic fidelity for semantic-related applications. However, CMC only achieves a certain level of semantic fidelity at a constant rate, and the model aims to optimize the probability of the ground truth text but not directly semantic fidelity. To tackle the problems, we propose a novel scheme named rate-distortion optimized CMC (RDO-CMC). Specifically, we model the text generation process as a Markov decision process and propose rate-distortion reward which is used in reinforcement learning to optimize text generation. In rate-distortion reward, the distortion measures both the semantic fidelity and naturalness of the encoded text. The rate for the text is estimated by the sum of the amount of information of all the tokens in the text since the amount of information of each token is a lower bound of coding bits. Experimentally, RDO-CMC effectively controls the rate in the CMC framework and achieves competitive performance on MSCOCO dataset.
Junlong Gao, Chuanmin Jia, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001
DCC5
2023 Learning to Compress Unmanned Aerial Vehicle (UAV) Captured Video: Benchmark and Analysis
abstract
In this paper, we propose to build a novel benchmark and neural video coding task named learning based Unmanned Aerial Vehicle (UAV) video coding. We collect the UAV videos with different content variations, including in-door and out-door scenes, object-scale variations and viewpoint distance, different climate condition etc. Then we encode those properly-selected videos using popular end-to-end optimized video codecs and conventional hybrid codecs, to form a comprehensive benchmark for learned drone video compression. We also provide a detailed analysis and envision the challenge of such task for future research. The main contributions of this paper are three folds. First, we construct a comprehensive benchmark for the task of drone video compression which consists of the rate-distortion (R-D) behavior of both hybrid and learned video codecs. To our knowledge, it is the first attempt in end-to-end optimized solution to compress drone videos. Second, we provide the review and analysis of the learned drone video compression schemes and further discuss the challenges of encoding UAV videos. Third, this benchmark and related research is accomplished as a milestone MPAI End-to-end Video (EEV) coding project. The proposed benchmark has constructed a solid baseline for compressing UAV videos and facilitates the future research works for related task.
Chuanmin Jia, Huifang Sun, Siwei Ma 0001, Wen Gao 0001
DCC5
2023 An Efficient Rate Control Scheme for Video Compression in Low-latency Interoperable Interfaces
abstract
Lightweight video compression has effectively alleviated the tension between growing transmission demands and expensive integration upgrades. Effective rate control algorithms are believed to be the crucial bottleneck for quality improvement during those ultra-high throughput coding processes. This paper proposes a novel rate control (RC) scheme that constructs a contextual adaptive bit estimation model through clustering historical compression information into block-gradient complexity categories. A buffer-aware tuning method and a flexible quantization parameter (QP) mapping algorithm are designed to determine the Luma/Chroma QP distribution where a simplified Lagrangian multiplier is further defined to preserve the stability of the overall compression process. As a result, the constant-bitrate compression towards low-latency interoperable ASICs is implemented with a promising RC performance.
Huiwen Ren, Zetian Song, Yan Wang 0011, Shanshe Wang, Fangdong Chen, Shiliang Pu, Siwei Ma 0001, Wen Gao 0001
DCC10
2023 An Adaptive Intra-frame Quantization Parameter Derivation Model Jointing with Inter-frame Analysis
abstract
This paper proposes a novel quantization parameter (QP) derivation module that constructs several spatiotemporal characteristics into key-frame QP determination through an efficient pre-analysis progress. A series of simplified prediction modes and a histogram statistic are employed to model the reference quality that key-frames provide to subsequent frames. An adaptive delta-QP value is generated to address the conflict between the low compression efficiency of intra-only frames and the critical predictive basis of temporal-underlying frames. The experimental result shows that the proposed method reduces 41.91% of peak-to-valley bitrate difference while leading to a 0.02% BDBR performance change, indicates that high-quality key-frames may not be indispensable in nowadays video compression frameworks. The proposed method has shown that the pre-analysis based QP optimization for intra-only frames is promising for the enhancement of transmission bandwidth utilization, which may hopefully provide new inspiration for bit allocation and rate control designs.
Huiwen Ren, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001
DCC4
2022 Rate Distortion Characteristic Modeling for Neural Image Compression
abstract
End-to-end optimized neural image compression (NIC) has obtained superior lossy compression performance recently. In this paper, we consider the problem of rate-distortion (R-D) characteristic analysis and modeling for NIC. We make efforts to formulate the essential mathematical functions to describe the R-D behavior of NIC using deep networks. Thus arbitrary bit-rate points could be elegantly realized by leveraging such model via a single trained network. We propose a plugin-in module to learn the relationship between the target bit-rate and the binary representation for the latent variable of auto-encoder. The proposed scheme resolves the problem of training distinct models to reach different points in the R-D space. Furthermore, we model the rate and distortion characteristic of NIC as a function of the coding parameter$\lambda$respectively. Our experiments show our proposed method is easy to adopt and realizes state-of-the-art continuous bit-rate coding performance, which implies that our approach would benefit the practical deployment of NIC.
Chuanmin Jia, Ziqing Ge, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001
DCC5
2022 High-Order Intra Prediction for Future Video Coding
abstract
Intra prediction acts a significant role in removing the spatial redundancy in the hybrid coding framework. Versatile Video Coding (VVC) employs a set of angular intra modes to generate directional contents based on the linear projection hypothesis. However, the linear based predictor is not expressive enough for generating high-fidelity patterns with non-linear structure. To compensate for that, we propose a high-order intra prediction (HOIP) for future video coding in this paper. In particular, the HOIP is modeled by a quadratic extrapolation function. To be compatible with the present intra prediction mechanism, the quadratic function can be formulated by two angular intra modes. To reduce the encoding complexity, we further propose a search pruning strategy to find the most appropriate pair-wise modes, and it can be flexibly extended for higher coding performance. The extensive experimental results demonstrate the effectiveness of the proposed method. Up to 0.6% BD-rate saving is obtained with the moderate complexity increment.
Jiaqi Zhang 0007, Chuanmin Jia, Wen Gao 0001
DCC5
2022 Fast Partition Mode Decision via a Plug-in Fully Connected Network for Video Coding
abstract
Flexible coding unit partitioning such as quad-tree nested binary-tree and ternary-tree adopted by the emerging enhanced compression model (ECM) brings promising coding performance improvement. Meanwhile, the computational complexity increases dramatically, which may block the exploration and validation of new coding tools. This paper investigates a partition mode early pruning scheme via a fully connected network to reduce the encoding complexity for the ECM. In particular, we carefully select features and devise the fully connected network, which could seamlessly cooperate with the encoder, revealing promising learning and inference capability. Experimental results demonstrate that the proposed method achieves 15%~50% encoding time savings with moderate bit-rate increasing on the ECM, and the extra complexity regarding the fully connected network and feature extraction is negligible.
Jiaqi Zhang 0007, Meng Wang 0017, Chuanmin Jia, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001
DCC7
2021 Intra Block Partition Structure Prediction via Convolutional Neural Network
abstract
In video coding, block partition segments images into non-overlap blocks for individual coding, the structure of which is becoming more and more flexible along with the development of video coding standards. Multiple types of tree structures have been proposed recently, which extensively improved the complexity of the encoding process due to the recursive rate-distortion search for the optimal partition. In this paper, a two-stage Convolutional Neural Network (CNN) based partition structure prediction method is proposed to bypass the decision process of the block size in intra frame coding. Specifically, the Coding Unit (CU) partition is first represented in sub-block granularity and predicted by the end-to-end trained CNNs. Then, the final partition structure compatible with the coding standard is derived from the prediction results directly. Experimental results show that the proposed CNN based partition method achieves about 56 times speedup (97% time-saving) with 9% BD-rate degradation against the reference software of the latest AVS3 coding standard (IEEE Standard 1857.10).
Shanshe Wang, Siwei Ma 0001, Wen Gao 0001
DCC5
2021 Flow-Grounded Dynamic Texture Synthesis for Video Compression
abstract
The basic ingredients of modern video coding standards are block-based prediction and transforms. However, when dealing with video contents containing dynamic textures (DT), the existing prediction schemes usually failed due to temporal variability and randomness of DT, which results in more bit cost on residual coding compared with other contents. In view of this point, a novel video compression scheme for DT is proposed in this work. In particular, wavelet-based analysis on motion characteristics of DT is firstly presented and based on the analysis, we introduce a flow-grounded texture synthesis method for video compression. Instead of conventional inter prediction, synthesized DT contents are used for reconstruction at the decoder. The proposed scheme has been fully integrated into the test model of Versatile Video Coding standard, VTM-10.0, for validation and a subjective test has also been carried out. Experimental results show that bitrate savings can be achieved by 40% on average at comparable visual quality.
Suhong Wang, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001
DCC5
2020 Sub-Sampled Cross-Component Prediction for Chroma Component Coding
abstract
Cross-component prediction, which takes advantage of inter-channel correlations, predicts the chroma block with the luma reconstructed block according to associated linear model. Instead of involving all available reference samples in building the linear model, in this paper, we propose a sub-sampled approach that utilizes at most four neighboring chroma samples and their corresponding down-sampled luma samples, leading to significantly reduced operations in the derivation of model parameters at both encoder and decoder. The proposed scheme is hardware friendly in terms of the overheads of memory access and clock cycles, and greatly benefits the practical implementations of the emerging video coding standard in real applications. Extensive experiments reveal that the proposed sub-sampled method provides simple operations and robust coding performance, leading to the adoption by Versatile Video Coding (VVC) Standard and the third generation Audio Video Coding Standard (AVS3).
Meng Wang 0017, Li Zhang 0006, Kai Zhang 0007, Shiqi Wang 0001, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001
DCC8
2020 Implicit Geometry Partition for Point Cloud Compression
abstract
Octree (OT) geometry partitioning has been acknowledged as an efficient representation in state-of-the-art point cloud compression (PCC). In this work, a new geometry partition and coding scheme is proposed to improve the OT based coding framework, in which the quad-tree (QT) and binary-tree (BT) partitions are introduced. It brings in and harmonizes the asymmetric kd-tree like concept with the symmetric OT-based geometry coding framework. It also enables asymmetric bounding boxes that can better fit the shape of 3D scenes. Bit savings can be obtained by skipping encoding unnecessary bits from implicit QT and BT partitions. Parameters are introduced to specify the conditions on which implicit geometry partitions will be applied. Experimental results have shown significant coding gains over the OT-only coding scheme in the state-of-the-art test model of MPEG Geometry based PCC (G-PCC) standard. For the dynamically acquired point clouds, the average coding gains are 8.3% for lossy geometry coding and 4.8% for lossless geometry coding without significant increase in coding complexity.
Xiang Zhang 0004, Wen Gao 0001, Shan Liu 0001
DCC2
2020 Linear Model Based Geometry Coding for Lidar Acquired Point Clouds
abstract
In light of above observations, in this work, we propose a new model-based method for geometry coding in G-PCC. More specifically, it is a linear model that takes the explicit line structures as a prior knowledge to improve the coding efficiency. The encoding of the linear model can be expressed by two parts, including the principle component along the line direction and the offsets from the line. Compact representation and high-efficiency coding methods are presented by encoding the parameters of linear model with appropriate quantization step-sizes (QS). To maximize the coding performance, encoder optimization techniques are employed to find the optimal trade-off between coding bits and errors, involving the Lagrangian multiplier method, where the rate-distortion behavior in terms of QS and multiplier is analyzed. We implement our method on top of the MPEG G-PCC reference software, and the results have shown that the proposed method is effective in coding point clouds with explicit line structures, such as the Lidar acquired data for autonomous driving. About 20% coding gains can be achieved on lossy geometry coding.
Xiang Zhang 0004, Wen Gao 0001, Shan Liu 0001
DCC2
2019 Separable KLT for Intra Coding in Versatile Video Coding (VVC)
abstract
After the works on the state-of-the-art High Efficiency Video Coding (HEVC) standard, the standard organizations continued to study the potential video coding technologies for the next generation of video coding standard, named Versatile Video Coding (VVC). Transform is a key technique for compression efficiency, and core experiment 6 (CE6) is carried out to explore the transform related coding tools. In this paper, we propose a novel separable transform based on Karhunen-Loève Transform (KLT) to eliminate the horizontal and vertical correlations in the residual samples of intra coding. In the proposed method, the weaknesses of the traditional KLT are addressed. The separable KLT is developed as an alternative transform type in addition to DCT-II, and the transform matrices from 4×4 to 64×64 are trained from intra residual samples. Experimental results show the proposed method can achieve 2.7% bitrate saving averagely on top of the reference software of VVC (VTM-1.1), and the consistent performance improvement on test set also validates the strong generalization capacity of the proposed separable KLT.
Kui Fan, Ronggang Wang, Weisi Lin, Jong-Uk Hou, Ling-Yu Duan, Ge Li 0002, Wen Gao 0001
DCC7
2019 Adaptive Wavelet Domain Filter for Versatile Video Coding (VVC)
abstract
Owing to the ability of removing compression artifacts, extensive in-loop filters have been proposed for video coding standards. They are performed after the reconstruction of all coding units (CUs), however, none of them has been taken into account in the mode decision when coding each CU. To address this issue and make the rate-distortion optimization (RDO) more precise for each CU, we introduce a low-pass filter when checking the rate-distortion cost after the reconstruction of each CU. Specifically, based on Haar wavelet, the reconstructed block is transformed to the frequency domain, and then an adaptive wavelet domain filter (AWF) is proposed to suppress the quantization noises in coded blocks. To be adaptive, the filter strength varies from CU to CU according to the texture complexity and quantization parameters (QPs). Experimental results show that the proposed method can reduce the compression artifacts and improve both the objective and subjective quality.
Suhong Wang, Xiang Zhang 0004, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001
DCC5
2018 Compressed Image Restoration via External-Image Assisted Band Adaptive PCA Model Learning
abstract
Visually annoying compression artifacts frequently appear in block-based transform coding at low bit rates, due to coarse and independent quantization of transform coefficients in coding blocks. This paper presents a subband adaptive modeling framework for reducing quantization artifacts. In this framework, each patch is jointly regularized by bandwise distribution priors adaptively learned in its PCA transform domain together with a quantization constraint prior in the DCT domain. Since the compression artifacts influence the covariance statistics of coded image patches remarkably, external images are utilized to provide more robust PCA domains for patch sparse modeling. Instead of using a global distribution model for all patches, the distribution prior of each patch is adaptively learned from similar patches within the compressed image itself to address the non-stationarity of image signals. The coefficients in different PCA bands are regularized unequally according to the learned priors. Experimental results show that the proposed scheme outperforms existing schemes in terms of both the objective and the perceptual qualities.
Ruiqin Xiong, Xiaopeng Fan 0001, Xianming Liu 0005, Tiejun Huang 0001, Wen Gao 0001
DCC6
2018 Fast and Robust Image Upsampling by Local Adaptive Gradient Field Sharpening Transform
abstract
This paper proposes an image upsampling scheme by introducing a new gradient field sharpening transform that converts the blurry gradient field of upsampled low-resolution (LR) image to a much sharper gradient field of original high-resolution (HR) image. Different from the existing methods that need to figure out the whole gradient profile structure and locate the edge points, we derive a new approach that sharpens the gradient field adaptively only based on the pixels in a small neighborhood. To maintain image contrast, image gradient is adaptively scaled to keep the integral of gradient field stable. Finally the HR image is reconstructed by fusing the LR image with the sharpened HR gradient field. Experimental results demonstrate that the proposed algorithm can generate more accurate gradient field and produce super-resolved images with better objective and visual qualities. Another advantage is that the proposed gradient sharpening transform is very fast and suitable for low-complexity applications.
Ruiqin Xiong, Dong Liu 0002, Zhiwei Xiong, Feng Wu 0001, Wen Gao 0001
DCC6
2018 Local patch encoding-based method for single image super-resolution
Yang Zhao 0002, Ronggang Wang, Wei Jia 0001, Jianchao Yang, Wenmin Wang 0001, Wen Gao 0001
Inf. Sci.6
2017 Wireless Image SoftCast Using Compressive Gradient
abstract
Summary form only given: Based on observations that the visual quality has strong correlation with image gradients, gradient based image SoftCast (G-Cast) [1] advocates to convey visual information by delivering image gradients. In G-Cast, both horizontal gradients and vertical gradients needs to be transmitted, even if the channel bandwidth is insufficient. This paper propose to send out the random projection measurements of the gradients instead of delivering gradients directly, so that data size can be reduced to an arbitrary ratio and channel bandwidth occupation can be lowered. We name this scheme as compressive gradient based SoftCast (CG-Cast). At CG-Cast sender, after generated by gradient transform, the gradients are down-sampled by random projection, sample rate of which is set according to the channel bandwidth condition. Then the produced measurements are sent out for raw OFDM transmission. A few lowfrequency components are also transmitted to tell the global luminance. At CG-Cast receiver, the received noisy measurements are used for compressive gradient based reconstruction procedure, which utilizes sparsity in gradient domain and non-local similarity in spatial domain [2]. The proposed method is compared with SoftCast [3] and compressive sensing (CS) in bandwidth limited and power constrained scenarios. To make fair comparison, these three schemes are tested under the same channel signal-to-noise ratio (CSNR) conditions to transmit equal amount of data for reconstruction, using equivalent power and bandwidth. CG-Cast outperforms SoftCast and CS in terms of SSIM and gradient signal-to-noise ratio (GSNR) at different bandwidth ratios. Comparing with SoftCast in different channel conditions, the average SSIM gain of all the tested images varies from 0.04 to 0.13, and the average GSNR gain ranges from 1.5dB to 2.9dB. CS is rather unstable in noisy conditions. SoftCast performs better than CS because of its power allocation.
Hangfan Liu, Ruiqin Xiong, Xiaopeng Fan 0001, Siwei Ma 0001, Wen Gao 0001
DCC5
2017 Compact Deep Invariant Descriptors for Video Retrieval
abstract
With emerging demand for large-scale video analysis, the Motion Picture Experts Group (MPEG) initiated the Compact Descriptor for Video Analysis (CDVA) standardization in 2014. In this work, we develop novel deep-learning features and incorporate them into the well-established CDVA evaluation framework to study its effectiveness in video analysis. In particular, we propose a Nested Invariance Pooling (NIP) method to obtain compact and robust Convolutional Neural Network (CNNs) descriptors. The CNNs descriptors are generated by applying three different pooling operations to the feature maps of CNNs in a nested way towards rotation and scale invariant feature representation. In particular, the rational, advantages and performance on the combination of CNNs and handcrafted descriptors are provided to better investigate the complementary effects of deep learnt and handcrafted features. Extensive experimental results show that the proposed CNNs descriptors outperform both state-of-the-art CNNs descriptors and canonical handcrafted descriptors adopted in CDVA Experimental Model (CXM) with significant mAP gains of 11.3% and 4.7%, respectively. Moreover, the combination of NIP derived deep invariant descriptors and handcrafted descriptors not only fulfills the lowest bitrate budget of CDVA, but also significantly advances the performance of CDVA core techniques.
Yihang Lou, Jie Lin 0001, Shiqi Wang 0001, Jie Chen 0006, Vijay Chandrasekhar 0001, Ling-Yu Duan, Tiejun Huang 0001, Alex Chichung Kot, Wen Gao 0001
DCC10
2017 Globally Variance-Constrained Sparse Representation for Rate-Distortion Optimized Image Representation
abstract
Sparse representation is efficient to approximately recover signals by a linear composition of a few bases from an over-complete dictionary. However, in the scenario of data compression, its efficiency and popularity are hindered due to the extra overhead for encoding the sparse coefficients. Therefore, how to establish an accurate rate model in sparse coding and dictionary learning becomes meaningful, which has been not fully exploited in the context of sparse representation. According to the Shannon entropy inequality, the variance of data source can bound its entropy, thus can reflect the actual coding bits. Therefore, a Globally Variance-Constrained Sparse Representation (GVCSR) model is proposed, where a variance-constrained rate term is introduced to the conventional sparse representation. To solve the non-convex optimization problem, we employ the Alternating Direction Method of Multipliers (ADMM) for sparse coding and dictionary learning, both of which have shown state-of-the-art rate-distortion performance in image representation.
Xiang Zhang 0004, Siwei Ma 0001, Zhouchen Lin, Jian Zhang 0018, Shiqi Wang 0001, Wen Gao 0001
DCC6
2016 Structure-driven Adaptive Non-local Filter for High Efficiency Video Coding (HEVC)
abstract
Deblocking filter (DF) Is High Efficiency Video Coding (HEVC) is Only Applied to all Samples Adjacent to prediction units (PU), or transform units (TU), which actually exists two issues. The first one is that DF in HEVC does not fully exploit nonlocal similarity structure information in video. The second one is that DF is HEVC does not consider the inside pixels, which often suffer from quantization distrotion. To alleviate these issues, in this paper, a structure-driven adaptive non-local filter (SANF) Is Proposed By Simultaneously Enforcing The Intrinsic Local Sparsity And The Non-Local Self-Similarity Of Each Frame. Not only SANF deals with the boundary pixels, but also the inside area, which is able to effectively reduce block artifacts while enhancing the quality of the deblocked frames. Applying SANF to luma and chroma components after DF, simulation results demonstrate that the proposed SANF can save BD-rate reduction up to 10.3% with ALF off. For luma component, SANF achieves 4.1%. 3.3%, 4.4% BD-rate saving for all intra, low delay B and random access configurations, respectively with ALF off. furthermore, the performance with ALF on is also discussed.
Jian Zhang 0018, Chuanmin Jia, Nan Zhang 0015, Siwei Ma 0001, Wen Gao 0001
DCC5
2016 From Visual Search to Video Compression: A Compact Representation Framework for Video Feature Descriptors
abstract
Visual feature descriptors have been successfully deployed in a wide range of applications, e.g. visual retrieval and analysis. To transmit these descriptors over bandwidth-limited networks, a high efficiency feature coding technique is highly desired to maximize compression capability and achieve compact feature representations. In this paper, a hybrid visual feature descriptor compression framework is presented and implemented in the encoding and decoding loops of texture videos. In particular, the multiple-hypothesis prediction is employed to effectively remove redundancies originated not only from spatial and temporal similarities, but also from reconstructed video frames. As the ultimate purpose of the transmitted descriptors is retrieval, the rate-accuracy optimization (RAO) technique is proposed to obtain the best tradeoff between the rate and retrieval performance. Such paradigm enables the conventional video stream to achieve high efficient retrieval/analysis with very low bitrate consumption. Moreover, we also demonstrate that texture video compression can also benefit from the additional information provided by the transmitted descriptors, leading to significantly improvement of coding efficiency on top of the high efficiency video coding (HEVC) standard. Extensive simulations have shown that the proposed method can offer significant bitrate reduction in representing both the descriptors and texture video frames, and meanwhile providing desirable retrieval performance.
Xiang Zhang 0004, Siwei Ma 0001, Shiqi Wang 0001, Shanshe Wang, Xinfeng Zhang 0001, Wen Gao 0001
DCC6
2016 Nonconvex Lp Nuclear Norm based ADMM Framework for Compressed Sensing
abstract
Compressed Sensing (CS) has drawn quite an amount of attention as a joint sampling and compression methodology. Recent studies further show that image prior models play an important role in image CS recovery. By exploiting the non-local self-similarity of natural images and clustering similar patches, low-rank prior model is adopted in this paper. Different from traditional nuclear norm, we extend thelp(0plpnuclear norm prior model for image CS recovery, which is able to more accurately enforce image structural sparsity and self-similarity at the same time. The proposed optimization problem is efficiently solved within the alternative direction multiplier method (ADMM) framework. Experimental results demonstrate that the proposedlpnuclear norm based ADMM framework for image CS recovery framework exhibits good convergence and achieves significant performance improvements over the current state-of-the-art methods.
Chen Zhao 0002, Jian Zhang 0018, Siwei Ma 0001, Wen Gao 0001
DCC4
2016 Compressive-Sensed Image Coding via Stripe-based DPCM
abstract
These years have seen the advances of compressive sensing (CS), but efficient coding of sensed measurements is still an issue. In this paper, we propose an image coding system based on the compressive sensing paradigm via stripe-based differential pulse-code modulation (DPCM). In the system, we sample and encode an image in a unit of multiple rows, which we call a stripe. Through extensive experiments, we observe that the correlation between measurements of adjacent stripes are much higher than that of the neighboring blocks. Based on this, we combine the stripe-based CS acquisition with the DPCM framework and design a mechanism that predicts a stripe of measurements from its preceding stripe of measurements. The produced measurement residuals are then quantized and entropy-encoded into binary coding bits, which are tremendously reduced compared to the traditional block-based framework. Furthermore, we provide an image CS reconstruction algorithm corresponding to the stripe-based acquisition. Experiments verify that the reconstruction quality is no worse or even better than the block-based case when much lower bitrate is consumed. In a rate-distortion point of view, the proposed system also outperforms the methods using block-based sampling and achieves the state-of-the-art performance for compressive-sensed image coding.
Chen Zhao 0002, Jian Zhang 0018, Siwei Ma 0001, Wen Gao 0001
DCC4
2016 Multimodal Deep Convolutional Neural Network for Audio-Visual Emotion Recognition
abstract
Emotion recognition is a challenging task because of the emotional gap between subjective emotion and the low-level audio-visual features. Inspired by the recent success of deep learning in bridging the semantic gap, this paper proposes to bridge the emotional gap based on a multimodal Deep Convolution Neural Network (DCNN), which fuses the audio and visual cues in a deep model. This multimodal DCNN is trained with two stages. First, two DCNN models pre-trained on large-scale image data are fine-tuned to perform audio and visual emotion recognition tasks respectively on the corresponding labeled speech and face data. Second, the outputs of these two DCNNs are integrated in a fusion network constructed by a number of fully-connected layers. The fusion network is trained to obtain a joint audio-visual feature representation for emotion recognition. Experimental results on the RML audio-visual database demonstrates the promising performance of the proposed method. To the best of our knowledge, this is an early work fusing audio and visual cues in DCNN for emotion recognition. Its success guarantees further research in this direction.
Shiqing Zhang, Shiliang Zhang, Tiejun Huang 0001, Wen Gao 0001
ICMR4
2016 Efficient Generalized Fused Lasso and Its Applications
abstract
Generalized fused lasso (GFL) penalizes variables with l 1 norms based both on the variables and their pairwise differences. GFL is useful when applied to data where prior information is expressed using a graph over the variables. However, the existing GFL algorithms incur high computational costs and do not scale to high-dimensional problems. In this study, we propose a fast and scalable algorithm for GFL. Based on the fact that fusion penalty is the Lovász extension of a cut function, we show that the key building block of the optimization is equivalent to recursively solving graph-cut problems. Thus, we use a parametric flow algorithm to solve GFL in an efficient manner. Runtime comparisons demonstrate a significant speedup compared to existing GFL algorithms. Moreover, the proposed optimization framework is very general; by designing different cut functions, we also discuss the extension of GFL to directed graphs. Exploiting the scalability of the proposed algorithm, we demonstrate the applications of our algorithm to the diagnosis of Alzheimer’s disease (AD) and video background subtraction (BS). In the AD problem, we formulated the diagnosis of AD as a GFL regularized classification. Our experimental evaluations demonstrated that the diagnosis performance was promising. We observed that the selected critical voxels were well structured, i.e., connected, consistent according to cross validation, and in agreement with prior pathological knowledge. In the BS problem, GFL naturally models arbitrary foregrounds without predefined grouping of the pixels. Even by applying simple background models, e.g., a sparse linear combination of former frames, we achieved state-of-the-art performance on several public datasets.
Bo Xin, Yoshinobu Kawahara, Yizhou Wang 0001, Lingjing Hu, Wen Gao 0001
ACM Trans. Intell. Syst. Technol.5
2015 Overview of the MPEG CDVS Standard
abstract
Towards mobile visual search, compact visual descriptors have been well advocated in both academic and industry endeavors. Moving Picture Experts Group (MPEG) initiated the remarkable Compact Descriptors for Visual Search (CDVS) standard activity in Jan. 2010 to push forward the frontiers of compact descriptors in mobile internet industry. In Oct. 2014, MPEG CDVS successfully entered the Final Draft of International Standard. CDVS made a series of significant breakthroughs in high performance and low complexity compact descriptors. In this paper, we give an overview of the MPEG CDVS standard, with emphasis on the development of the core techniques and their technical merits.
Ling-Yu Duan, Tiejun Huang 0001, Wen Gao 0001
DCC3
2015 Optimizing Binary Fisher Codes for Visual Search
abstract
Fisher vectors (FV) aggregated from local invariant features (e.g., SIFT) is one of the state-of-the-art descriptors for visual search, due to high discriminability but small visual vocabulary. Nevertheless, a high-dimensional FV needs to be compressed into a compact descriptor for light storage and high matching eficiency. In this paper, we formulate the FV compression as a resource-constrained optimization problem. Our goal is to maximize search performance subject to the constraints of descriptor compactness, compression complexity in terms of memory usage and time cost. Accordingly, we present a selective binary Fisher codes (SBFC) to compress the raw FV. Firstly, to fulfill the constraint of compression complexity, we binarize the FV by a sign function, Secondly, we propose to select discriminative bits from the binarized FV (BFC) to maximize search performance, subject to the constraint of descriptor compactness. Extensive experiments over MPEG Compact Descriptor for Visual Search (CDVS) benchmark datasets have shown that S-BFC significantly improves search performance at a smaller descriptor size as well as much lower complexity, compared with the state-of-the-art FV compression algorithms like Hashing and Product Quantziation (PQ). A simplified version of SBFC, SBFC LS has been adopted by the MPEG CDVS standard. In the CDVS evaluation framework, SBFC LS has achieved promising performance mean Average Precision (mAP) 83% on average at much lower memory cost of 40KB.
Zhe Wang 0019, Ling-Yu Duan, Jie Lin 0001, Jie Chen 0006, Tiejun Huang 0001, Wen Gao 0001
DCC6
2014 G-CAST: Gradient Based Image SoftCast for Perception-Friendly Wireless Visual Communication
abstract
Conventional image and video communication systems are usually designed with the objective being to maximize the fidelity of reconstructed images measured by mean square errors (MSE). It is well known that the fidelity metric MSE may not reflect the visual quality perceived by human eyes. Recent advancements in image quality assessment tell us that the structural similarity (SSIM), especially the gradient similarity, reveals the perceptual fidelity of images more reliably. Inspired by this observation, this paper proposes a new image communication approach, which conveys the visual information in an image by transmitting the image gradients and recovers the image from the received gradient data at decoder side using statistical image prior knowledge. In particular, we designed a gradient-based image SoftCast scheme for wireless scenarios. Experimental results show that the proposed scheme can produce reconstruction images with much better perceptual quality. The advantage in perceptual quality is verified by the quality improvement measured by the metrics SSIM and gradient signal-to-noise ratio (GSNR).
Ruiqin Xiong, Hangfan Liu, Siwei Ma 0001, Xiaopeng Fan 0001, Feng Wu 0001, Wen Gao 0001
DCC6
2014 Superimage: Packing Semantic-Relevant Images for Indexing and Retrieval
abstract
As an important procedure in image retrieval, off-line indexing focuses on organizing relevant images together and making them easy to access. However, most of existing indexing strategies view database images individually and only consider partial relevance, i.e., either visual or semantic relevance among them. To overcome these issues and design better indexing strategy, we propose to package semantically relevant images into superimages, and then index superimages instead of single images. Superimage effectively packages multiple images into one new unit, hence significantly decreases the number of images to be indexed. This naturally saves the memory cost and retrieval time. To make the final index file discriminative to both visual and semantic relevances, we extract local descriptors from superimages and index them with inverted file. During online retrieval, we only need to extract local descriptors from queries, but could get semantic-aware retrieval results. This is because during our off-line indexing stage, both the semantically and visually relevant images are organized together. Therefore, our approach is superior to many online retrieval fusion algorithms. Experimental results on UKbench, Holidays, and one large-scale dataset all manifest the promising performance of our approach, i.e., competitive precision, better efficiency, and only about 1/2 memory consumption compared with state-of-the-arts.
Qingjun Luo, Shiliang Zhang, Tiejun Huang 0001, Wen Gao 0001, Qi Tian 0001
ICMR4
2013 Image Super-Resolution via Hierarchical and Collaborative Sparse Representation
abstract
In this paper, we propose an efficient image super-resolution algorithm based on hierarchical and collaborative sparse representation (HCSR). Motivated by the observation that natural images typically exhibit multi-modal statistics, we propose a hierarchical sparse coding model which includes two layers: the first layer encodes individual patches, and the second layer jointly encodes the set of patches that belong to the same homogeneous subset of image space. We further present a simple alternative to achieve such target by identifying optimal sparse representation that is adaptive to specific statistics of images. Specially, we cluster images from the offline training set into regions of similar geometric structure, and model each region (cluster) by learning adaptive bases describing the patches within that cluster using principal component analysis (PCA). This cluster-specific dictionary is then exploited to optimally estimate the underlying HR pixel values using the idea of collaborative sparse coding, in which the similarity between patches in the same cluster is further considered. It conceptually and computationally remedies the limitation of many existing algorithms based on standard sparse coding, in which patches are independently encoded. Experimental results demonstrate the proposed method appears to be competitive with state-of-the-art algorithms.
Xianming Liu 0005, Deming Zhai, Debin Zhao, Wen Gao 0001
DCC4
2013 Low Complexity Rate Distortion Optimization for HEVC
abstract
The emerging High Efficiency Video Coding (HEVC) standard has improved the coding efficiency drastically, and can provide equivalent subjective quality with more than 50% bit rate reduction compared to its predecessor H.264/AVC. As expected, the improvement on coding efficiency is obtained at the expense of more intensive computation complexity. In this paper, based on an overall analysis of computation complexity in HEVC encoder, a low complexity rate distortion optimization (RDO) coding scheme is proposed by reducing the number of available candidates for evaluation in terms of the intra prediction mode decision, reference frame selection and CU splitting. With the proposed scheme, the RDO technique of HEVC can be implemented in a low-complexity way for complexity-constrained encoders. Experimental results demonstrate that, compared with the original HEVC reference encoder implementation, the proposed algorithms can achieve about 30% reduced encoding time on average with ignorable coding performance degradation (0.8%).
Siwei Ma 0001, Shiqi Wang 0001, Shanshe Wang, Liang Zhao 0007, Qin Yu 0003, Wen Gao 0001
DCC6
2013 Progressive Image Restoration through Hybrid Graph Laplacian Regularization
abstract
In this paper, we propose a unified framework to perform progressive image restoration based on hybrid graph Laplacian regularized regression. We first construct a multi-scale representation of the target image by Laplacian pyramid, then progressively recover the degraded image in the scale space from coarse to fine so that the sharp edges and texture can be eventually recovered. On one hand, within each scale, a graph Laplacian regularization model represented by implicit kernel is learned which simultaneously minimizes the least square error on the measured samples and preserves the geometrical structure of the image data space by exploring non-local self-similarity. In this procedure, the intrinsic manifold structure is considered by using both measured and unmeasured samples. On the other hand, between two scales, the proposed model is extended to the parametric manner through explicit kernel mapping to model the inter-scale correlation, in which the local structure regularity is learned and propagated from coarser to finer scales. Experimental results on benchmark test images demonstrate that the proposed method achieves better performance than state-of-the-art image restoration algorithms.
Deming Zhai, Xianming Liu 0005, Debin Zhao, Hong Chang 0001, Wen Gao 0001
DCC5
2013 Hierarchical-and-Adaptive Bit-Allocation with Selective Background Prediction for High Efficiency Video Coding (HEVC)
abstract
Summary form only given. Recently, a low-delay and high-efficiency hierarchical prediction structure (HPS) has been proposed for the forthcoming HEVC. Actually, frames and coding units (CUs) at different HPS positions have different importance to predict following frames and CUs. This paper firstly analyzes what frames and CUs should be quantified less. Based on the analysis, we propose a Hierarchical-and-Adaptive BIT-allocation method with Selective background prediction (HABITS) to optimize the video performance of HEVC. Extensive experiments on HM8.0 show that, HABITS saves 13.3% and 35.5% of the total bit rate for eight HEVC conference videos and eight common used surveillance videos. Even for the normal videos in HEVC's Class B and C, there is still 2.2% bit-saving.
Xianguo Zhang, Tiejun Huang 0001, Yonghong Tian 0001, Wen Gao 0001
DCC4
2013 Structural Group Sparse Representation for Image Compressive Sensing Recovery
abstract
Compressive Sensing (CS) theory shows that a signal can be decoded from many fewer measurements than suggested by the Nyquist sampling theory, when the signal is sparse in some domain. Most of conventional CS recovery approaches, however, exploited a set of fixed bases (e.g. DCT, wavelet, contour let and gradient domain) for the entirety of a signal, which are irrespective of the nonstationarity of natural signals and cannot achieve high enough degree of sparsity, thus resulting in poor rate-distortion performance. In this paper, we propose a new framework for image compressive sensing recovery via structural group sparse representation (SGSR) modeling, which enforces image sparsity and self-similarity simultaneously under a unified framework in an adaptive group domain, thus greatly confining the CS solution space. In addition, an efficient iterative shrinkage/thresholding algorithm based technique is developed to solve the above optimization problem. Experimental results demonstrate that the novel CS recovery strategy achieves significant performance improvements over the current state-of-the-art schemes and exhibits nice convergence.
Jian Zhang 0018, Debin Zhao, Feng Jiang 0001, Wen Gao 0001
DCC4
2012 Distributed Soft Video Broadcast (DCAST) with Explicit Motion
abstract
Video broadcasting is a popular application of wireless network. However, the existing layered approaches can hardly accommodate users with diverse channel conditions as analog communication can do. The newly emerged `soft cast' approach, utilizing soft broadcast, provides smooth multicast performance but is not very efficient in inter frame compression. In this work, we propose a motion-aligned wireless video multicast scheme DCAST. Instead of using conventional close loop prediction (CLP), DCAST is based on distributed source coding (DSC) theory. This helps DCAST to avoid error propagation but still achieve high compression efficiency in inter frame coding. DCAST outperforms soft cast 5dB in video PSNR while maintaining the similar graceful degradation feature as soft cast.
Xiaopeng Fan 0001, Feng Wu 0001, Debin Zhao, Oscar C. Au, Wen Gao 0001
DCC5
2012 Multi-scale Spatial Error Concealment via Hybrid Bayesian Regression
abstract
In this paper, we propose a novel multi-scale spatial error concealment algorithm to combine the modeling strengthes of the parametric and nonparametric Bayesian regression. We progressively recover missing blocks in the scale space from coarse to fine so that the sharp edges and texture in the finest scale can be eventually recovered. On one hand, in each scale, the nonparametric part of the methodology is used to exploit the intra-scale correlation, which relies on the data itself to dictate the structure of the model. In this procedure, the non-local self-similarity property is utilized as a fruitful resource for abstracting a priori knowledge of images. On the other hand, the parametric part is used to explicitly model the inter-scale correlation, in which the local structure regularity is thoroughly explored to recover the sharp edges and major texture features of images. It is not respected if only the nonparametric modeling is considering. We achieve the best of both worlds within a multi-scale framework. Experimental results on benchmark test images demonstrate that the proposed method achieves very competitive performance with the state-of-the-art error concealment algorithms.
Xianming Liu 0005, Deming Zhai, Guangtao Zhai, Debin Zhao, Ruiqin Xiong, Wen Gao 0001
DCC6
2012 A Compact Stereoscopic Video Representation for 3D Video Generation and Coding
abstract
We propose a novel compact representation for stereoscopic videos - a 2D video and its depth cues. Depth cues are derived from an interactive labeling process during 2D-to-3D video conversion, they are contour points of foreground objects and a background geometric model. By using such cues and image features of 2D video frames, depth maps of the frames can be recovered. Compared with traditional 3D video representation, the proposed one is more compact. We also design algorithms to encode and decode the depth cues. The representation benefits both 3D video generation and coding. Experimental results demonstrate that the bit rate can be saved about 10%-50% in coding 3D videos compared with multi-view video coding and 2D+depth methods. A system coupling 2D-to-3D video conversion and coding (CVCC) is proposed to verify advantages of the representation.
Zhebin Zhang, Ronggang Wang, Yizhou Wang 0001, Wen Gao 0001
DCC5
2012 Compressed Sensing Recovery via Collaborative Sparsity
abstract
Compressed Sensing (CS) has drawn quite an amount of attention as a joint sampling and compression approach. Its theory shows that a signal can be decoded from many fewer measurements than suggested by the Nyquist sampling theory, when the signal is sparse in some domain. So one of the most significant challenges in CS is to seek a domain where a signal can exhibit a high degree of sparsity and hence be recovered faithfully. Most of conventional CS recovery approaches, however, exploited a set of fixed bases (e.g. DCT, wavelet and gradient domain) for the entirety of a signal, which are irrespective of the nonstationarity of natural signals and cannot achieve high enough degree of sparsity, thus resulting in poor rate-distortion performance. In this paper, we propose a new framework for compressed sensing recovery via collaborative sparsity (RCoS), which enforces local two-dimensional sparsity and nonlocal three-dimensional sparsity simultaneously in an adaptive hybrid space-transform domain, thus substantially utilizing intrinsic sparsities of natural images and greatly confining the CS solution space. In addition, an efficient augmented Lagrangian based technique is developed to solve the above optimization problem. Experimental results on a wide range of natural images are presented to demonstrate the efficacy of the new CS recovery strategy.
Jian Zhang 0018, Debin Zhao, Chen Zhao 0002, Ruiqin Xiong, Siwei Ma 0001, Wen Gao 0001
DCC6
2012 Multiview Metric Learning with Global Consistency and Local Smoothness
abstract
In many real-world applications, the same object may have different observations (or descriptions) from multiview observation spaces, which are highly related but sometimes look different from each other. Conventional metric-learning methods achieve satisfactory performance on distance metric computation of data in a single-view observation space, but fail to handle well data sampled from multiview observation spaces, especially those with highly nonlinear structure. To tackle this problem, we propose a new method calledMultiview Metric Learning with Global consistency and Local smoothness(MVML-GL) under a semisupervised learning setting, which jointly considers global consistency and local smoothness. The basic idea is to reveal the shared latent feature space of the multiview observations by embodying global consistency constraints and preserving local geometric structures. Specifically, this framework is composed of two main steps. In the first step, we seek a global consistent shared latent feature space, which not only preserves the local geometric structure in each space but also makes those labeled corresponding instances as close as possible. In the second step, the explicit mapping functions between the input spaces and the shared latent space are learned via regularized locally linear regression. Furthermore, these two steps both can be solved by convex optimizations in closed form. Experimental results with application to manifold alignment on real-world datasets of pose and facial expression demonstrate the effectiveness of the proposed method.
Deming Zhai, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001, Wen Gao 0001
ACM Trans. Intell. Syst. Technol.5
2012 A Generic Approach for Systematic Analysis of Sports Videos
abstract
Various innovative and original works have been applied and proposed in the field of sports video analysis. However, individual works have focused on sophisticated methodologies with particular sport types and there has been a lack of scalable and holistic frameworks in this field. This article proposes a solution and presents a systematic and generic approach which is experimented on a relatively large-scale sports consortia. The system aims at the event detection scenario of an input video with an orderly sequential process. Initially, domain knowledge-independent local descriptors are extracted homogeneously from the input video sequence. Then the video representation is created by adopting a bag-of-visual-words (BoW) model. The video’s genre is first identified by applying the k-nearest neighbor (k-NN) classifiers on the initially obtained video representation, and various dissimilarity measures are assessed and evaluated analytically. Subsequently, an unsupervised probabilistic latent semantic analysis (PLSA)-based approach is employed at the same histogram-based video representation, characterizing each frame of video sequence into one of four view groups, namely closed-up-view, mid-view, long-view, and outer-field-view. Finally, a hidden conditional random field (HCRF) structured prediction model is utilized for interesting event detection. From experimental results, k-NN classifier using KL-divergence measurement demonstrates the best accuracy at 82.16% for genre categorization. Supervised SVM and unsupervised PLSA have average classification accuracies at 82.86% and 68.13%, respectively. The HCRF model achieves 92.31% accuracy using the unsupervised PLSA based label input, which is comparable with the supervised SVM based input at an accuracy of 93.08%. In general, such a systematic approach can be widely applied in processing massive videos generically.
Ning Zhang 0023, Ling-Yu Duan, Lingfang Li, Qingming Huang, Wen Gao 0001, Ling Guan
ACM Trans. Intell. Syst. Technol.6
2011 Transductive Regression with Local and Global Consistency for Image Super-Resolution
abstract
In this paper, we propose a novel image super-resolution algorithm, referred to as interpolation based on transductive regression with local and global consistency (TRLGC). Our algorithm first constructs a set of local interpolation models which can predict the intensity labels of all image samples, and a loss term will be minimized to keep the predicted labels of available low-resolution (LR) samples sufficiently close to the original ones. Then, all of the losses evaluated in local neighborhoods are accumulated together to measure the global consistency on all samples. Furthermore, a graph-Laplacian based manifold regularization term is incorporated to penalize the global smoothness of intensity labels, such smoothing can alleviate the insufficient training of the local models and make them more robust. Finally, we construct a unified objective function to combine together the accumulated loss of the locally linear regression, square error of prediction bias on the available LR samples and the manifold regularization term, which could be solved with a closed-form solution as a convex optimization problem. In this way, a transductive regression algorithm with local and global consistency is developed. Experimental results on benchmark test images demonstrate that the proposed image super-resolution method achieves very competitive performance with the state-of-the-art algorithms.
Xianming Liu 0005, Debin Zhao, Ruiqin Xiong, Siwei Ma 0001, Wen Gao 0001, Huifang Sun
DCC5
2010 Tanner Graph Based Image Interpolation
abstract
This paper interprets image interpolation as a channel decoding problem and proposes a tanner graph based interpolation framework, which regards each pixel in an image as a variable node and the local image structure around each pixel as a check node. The pixels available from low-resolution image are "received" whereas other missing pixels of highresolution image are "erased", through an imaginary channel. Local image structures exhibited by the low-resolution image provide information on the joint distribution of pixels in a small neighborhood, and thus play the same role as parity symbols in the classic channel coding scenarios. We develop an efficient solution for the sum-product algorithm of belief propagation in this framework, based on a gaussian auto-regressive image model. Initial experiments show up to 3dB gain over other methods with the same image model. The proposed framework is flexible in message processing at each node and provides much room for incorporating more sophisticated image modelling techniques.
Ruiqin Xiong, Wen Gao 0001
DCC2
2010 Auto Regressive Model and Weighted Least Squares Based Packet Video Error Concealment
abstract
In this paper, auto regressive (AR) model is applied to error concealment for block-based packet video encoding. Each pixel within the corrupted block is restored as the weighted summation of corresponding pixels within the previous frame in a linear regression manner. Two novel algorithms using weighted least squares method are proposed to derive the AR coefficients. First, we present a coefficient derivation algorithm under the spatial continuity constraint, in which the summation of the weighted square errors within the available neighboring blocks is minimized. The confident weight of each sample is inversely proportional to the distance between the sample and the corrupted block. Second, we provide a coefficient derivation algorithm under the temporal continuity constraint, where the summation of the weighted square errors around the target pixel within the previous frame is minimized. The confident weight of each sample is proportional to the similarity of geometric proximity as well as the intensity gray level. The regression results generated by the two algorithms are then merged to form the ultimate restorations. Various experimental results demonstrate that the proposed error concealment strategy is able to increase the peak signal-to-noise ratio (PSNR) compared to other methods.
Yongbing Zhang 0002, Xinguang Xiang, Siwei Ma 0001, Debin Zhao, Wen Gao 0001
DCC5
2009 Compression-Induced Rendering Distortion Analysis for Texture/Depth Rate Allocation in 3D Video Compression
abstract
In 3D video applications, the virtual view is generally rendered by the compressed texture and depth. The texture and depth compression with different bit-rate overheads can lead to different virtual view rendering qualities. In this paper, we analyze the compression-induced rendering distortion for the virtual view. Based on the 3D warping principle, we first address how the texture and depth compression affects the virtual view quality, and then derive an upper bound for the compression-induced rendering distortion. The derived distortion bound depends on the compression-induced depth error and texture intensity error. Simulation results demonstrate that the theoretical upper bound is an approximate indication of the rendering quality and can be used to guide sequence-level texture/depth rate allocation for 3D video compression.
Yanwei Liu 0001, Siwei Ma 0001, Qingming Huang, Debin Zhao, Wen Gao 0001, Nan Zhang 0015
DCC5
2008 Performance Analysis of Dual Frame Motion Compensation
abstract
In dual frame video coding, one short-term reference frame (STR) and one long-term reference frame (LTR) are available for motion compensation. The STR is the previous frame of current frame. The LTR remains static for a few frames, and then jump forward. In this paper, for different GOP length and bits allocation of the LTR, the coding performance of dual frame motion compensation is analyzed. The rate-distortion modeling of multi-hypothesis motion compensated prediction is employed to analyze the performance of dual frame motion compensation.
Xiangyang Ji, Debin Zhao, Zhi Bian, Wen Gao 0001
DCC6
2007 An Enhanced Robust Entropy Coder for Video Codecs Based on Context-Adaptive Reversible VLC
abstract
This paper proposes an enhanced RVLC coder, context-adaptive reversible variable length coder (CRVLC), for DCT coefficients by using the techniques of data sub-partitioning and context modeling. The data sub-partitioning means that the data part of DCT coefficients is split into several small sub-partitions. As each sub-partition can be reversibly decoded by RVLC, more data as well as higher error resilience can be obtained. The context modeling exploits the correlation of DCT coefficients for further compression. This modeling defines the contexts by hierarchical-dependent information. The information is also available in the backward decoding, so that it supports the reversible decoding. And with it the data outputted by CRVLC can be naturally placed into multiple sub-partitions.
Qiang Wang 0011, Debin Zhao, Siwei Ma 0001, Wen Gao 0001
DCC4
2006 Semantic Scoring Based on Small-World Phenomenon for Feature Selection in Text Mining
Chong Huang 0006, Yonghong Tian 0001, Tiejun Huang 0001, Wen Gao 0001
ADMA4
2006 Robust Collective Classification with Contextual Dependency Network Models
Yonghong Tian 0001, Tiejun Huang 0001, Wen Gao 0001
ADMA3
2006 Latent linkage semantic kernels for collective classification of link data
Yonghong Tian 0001, Tiejun Huang 0001, Wen Gao 0001
J. Intell. Inf. Syst.3
2006 Learning Contextual Dependency Network Models for Link-Based Classification
abstract
Links among objects contain rich semantics that can be very helpful in classifying the objects. However, many irrelevant links can be found in real-world link data such as Web pages. Often, these noisy and irrelevant links do not provide useful and predictive information for categorization. It is thus important to automatically identify which links are most relevant for categorization. In this paper, we present a contextual dependency network (CDN) model for classifying linked objects in the presence of noisy and irrelevant links. The CDN model makes use of a dependency function that characterizes the contextual dependencies among linked objects. In this way, CDNs can differentiate the impacts of the related objects on the classification and consequently reduce the effect of irrelevant links on the classification. We show how to learn the CDN model effectively and how to use the Gibbs inference framework over the learned model for collective classification of multiple linked objects. The experiments show that the CDN model demonstrates relatively high robustness on data sets containing irrelevant links.
Yonghong Tian 0001, Qiang Yang 0001, Tiejun Huang 0001, Charles Ling 0001, Wen Gao 0001
IEEE Trans. Knowl. Data Eng.5
2005 Bandwidth Adaptive Quality Smoothing for Unequal Error Protected Scalable Video Streaming
abstract
Summary form only given. We address the problem of inter-GOF bit allocation for FEC-MDC protected scalable video sequences. The objective is to minimize the variation in quality while streaming over packet-loss channels with time-varying bandwidth. We present an online heuristic algorithm to adaptively allocate bits for every GOF. Before transmitting a GOF, we first estimate the current available network bandwidth, and calculate the spare channel bit rate available to be used for current GOF. Then we propose a novel AIMD (additive increase/multiplicative decrease) quality control mechanism to regulate the changing behavior of target quality: (a) if the spare channel bit rate is greater than a certain value, then we increase the target quality gracefully in a linear mode; (b) else if it is less than another certain value, then we decrease the target quality aggressively. Finally, once the target quality is determined, we allocate bits for current GOF to meet the target quality requirement by iteratively increasing or decreasing the packet length used by FEC-MDC packetization. Since this procedure is time costly, we propose a fast approximate-approaching based dual-stage iteration technique to accelerate it. Experimental results show that our techniques can achieve near constant or graceful increasing quality in a segment by segment scheme, and the improved algorithm is very efficient and can be used in real-time online computation. Besides, we also analyze the impacts of some algorithm parameters to the quality smoothing results.
Longshe Huo, Wen Gao 0001, Qingming Huang
DCC2
2004 An Ontology-based Approach to Retrieve Digitized Art Images
abstract
Although much progress has been made, current low-level based visual information retrieval technology does not allow users to formulate queries through high-level semantics. More and more digitized art images appear on the Internet, and techniques need to be established on how to organize and retrieve them. In this work, a framework for retrieving art images using an ontology-based method is introduced. The proposed ontology describes images in various aspects. Non-objectionable semantics are first introduced, and how to express these semantics is given. Concepts in the ontology could be automatically derived. The retrieval scheme makes users more naturally find visual information and experimental implementation demonstrates good potential on retrieving art images in a human-centered manner.
Shuqiang Jiang, Tiejun Huang 0001, Wen Gao 0001
Web Intelligence3
2003 IISM: An Image Internal Semantic Model for Image Database Based on Relevance Feedback
abstract
A semantic model - IISM (image internal semantic model) is introduced. Unlike other semantic extracting methods, IISM extracts the semantic information not by image segmentation and image understanding, but by analyzing relevance feedback image retrieval results. For relevance feedback image retrieval system, the images relevant to query are pointed as positive example, otherwise the images irrelevant to query are pointed as negative examples. It is assumed that these positive examples are related in semantic content. IISM computes comprehensive pair-wise mutual information for all images through analyzing the results of relevance feedback image retrieval. An association with a high mutual information means that one image is semantically associated with another. Semantic retrieval and clustering is carried out based on these association relationships.
Lijuan Duan, Wen Gao 0001
Web Intelligence2
2003 Two-Phase Web Site Classification Based on Hidden Markov Tree Models
abstract
With the exponential growth of both the amount and diversity of the information that the Web encompasses, automatic classification of topic-specific Web sites is highly desirable. We propose a novel approach for Web site classification based on the content, structure and context information of Web sites. In our approach, the site structure is represented as a two-layered tree in which each page is modeled as a DOM (document object model) tree and a site tree is used to hierarchically link all pages within the site. Two context models are presented to capture the topic dependences in the site. Then the hidden Markov tree (HMT) model is utilized as the statistical model of the site tree and the DOM tree, and an HMT-based classifier is presented for their classification. Moreover, for reducing the download size of Web sites but still keeping high classification accuracy, an entropy-based approach is introduced to dynamically prune the site trees. On these bases, we employ the two-phase classification system for classifying Web sites through a fine-to-coarse recursion. The experiments show our approach is able to offer high accuracy and efficient process performance.
Yonghong Tian 0001, Tiejun Huang 0001, Wen Gao 0001, PingBo Kang
Web Intelligence3
2002 Learning Prosodic Patterns for Mandarin Speech Synthesis
Yiqiang Chen 0001, Wen Gao 0001, Tingshao Zhu, Charles Ling 0001
J. Intell. Inf. Syst.2
2001 Morphological Representation of DCT Data for Image Coding
Debin Zhao, Wen Gao 0001
Data Compression Conference3
1999 Extending DACLIC for Near Lossless Compression with Postprocessing of Greyscale Images
abstract
Summary form only given. A lossless/near lossless coding scheme, DACLIC, is presented. The proposed scheme attempts to remove redundancy in a given image in the spatial domain. The redundancy removal is achieved by block direction prediction and context-based error modeling. The block direction operation in DACLIC first partitions an image into blocks. Pixels within each incoming block are analyzed resulting in a best directional prediction for that block. The best direction is chosen from a given set that results in the minimum prediction error. Removal of redundancy by the block direction technique is not possible for removing all possible redundancy in a given image. Another decorrelation part of the DACLIC scheme is the context-based error modeling which exploits context-dependent DPCM error structures. The DACLIC scheme is primarily used as a lossless image compression technique. However, the scheme can be easily extended to near-lossless compression applications by introducing a small quantization loss. This small quantization loss is restricted to an absolute error not exceeding a prescribed value n for all pixels in a given image. Application of block direction and context modeling reduces a given image into residuals. This residual typically has a lower entropy than the given image. A quadtree Rice coder (QRC) is proposed as an entropy coder of DACLIC. An arithmetic coder is also given as an option. The QRC operates on residual blocks with low computing complexity that compares favorably with the residual coding method used by LOCO-I as the proposed QRC is two-dimensional in nature. For near-lossless compression with a larger value of n, banding artifacts are visible in the decoded image. In the DACLIC system, a postprocessing technique is proposed to remove the banding artifacts.
Debin Zhao, Wen Gao 0001
Data Compression Conference3
1999 A Robust Method for Unknown Forms Analysis
abstract
This paper proposes a strategy for analyzing unknown, filled forms. First, horizontal and vertical line segments are detected, extracted and filtered. A recursive splitting and merging algorithm eliminates overlapping segments, filters false segments, and groups the segments into lines. Based on the extracted lines, an algorithm for rectangle extraction is proposed. We define the constraints between rectangles and edges. In a process of scanning the horizontal and vertical lines, candidate edges are validated and rectangles are generated if its surrounding edges and their combination are all valid. The process is recursively applied. It can tolerate large breaks in form lines, ignore irrelevant segments and deal with embedded rectangles. Experiments on a collection of forms show that our approach works well on poor quality images.
Xingyuan Li 0003, Wen Gao 0001, David S. Doermann, Weon-Geun Oh
ICDAR2
1998 On-Line Sprite Encoding with Large Global Motion Estimation
abstract
Summary form only given. A sprite which is an image composed of pixels belonging to a video object visible throughout a video segment is a very important concept proposed by MPEG4. Because of the search region limitation in the global motion estimation, the performance of traditional sprite coding technology is not satisfactory in the case of fast camera motion. Only enlarging the search region is difficult to ensure the right motion estimation. An improved algorithm is proposed with enlarging the search region, predicting the motion of the current VOP (video object plane) and shortening the iterative time. Three main techniques are adopted in the new algorithm. They are: (1) enlarging search region; (2) weighting sum of absolute difference (SAD); and (3) shortening iteration. Two group experiments present a comparison between the original algorithm and the improved algorithm. The first group experiment shows the improved algorithm gets the same performance as the original algorithm in coding the general motion sequences. In the second group experiment, the Stefan background sequences which frame rates of 15 Hz, 10 Hz, 7.5 Hz, and 6 Hz are encoded in various transformations. The results given in a table show the coding performances of the algorithm are significantly improved.
Feng Wu 0001, Wen Gao 0001, Yangzhou Xiang, Datong Chen
Data Compression Conference2
1998 FACOLA - Face Coder Based on Location and Attention
abstract
Summary form only given. FACOLA (face coder based on location and attention) is proposed for potential applications such as compression of face pictures used in IC and ID cards. The face locator locates a face in an image using template matching and eigenface techniques. The attention detector detects high attention and low attention in the located face according to their different frequency characteristics. For high attention, usually high quality or lossless coding is required. The DCT is not suitable for such a case because it will not lead to an efficient compaction of the image energy and variable length coding (VLC) cannot be done efficiently. So a DPCM coder with three directional predictors is presented instead of the DCT. The prediction difference is quantized and the entropy coded DPCM coder supports lossless and lossy compression. The DCT coder is adopted for low attention and non-face area compression using different quantization factors.
Debin Zhao, Wen Gao 0001
Data Compression Conference3
1997 Recognizing components of handwritten characters by attributed relational graphs with stable features
abstract
We present a method for Chinese character component recognition. We use an attributed relational graph model to describe a component. This model allows us to express the knowledge of the component shape and can include stable features of the component in various writing styles and various characters. We then describe a method to extract graph representation of an input character. The graph models are used to recognize components from a whole character by subgraph isomorphism. Our method need not segment component from character and can tolerate links and overlaps between a component and the other part. Experiments for handwritten Chinese characters show its efficiency.
Xingyuan Li 0003, Weon-Geun Oh, Jiarong Hong, Wen Gao 0001
ICDAR4
1995 A system for automatic Chinese seal imprint verification
abstract
Chinese seal imprint verification by computer is very difficult, but is much needed by the application. An improved method for seal imprint verification based on stroke edge matching combined with image difference analysis is proposed. Experimental results show that the proposed approach is excellent in consistency, reliability and adaptability and is feasible for practical applications.
Wen Gao 0001, Shengfu Dong, Xilin Chen 0001
ICDAR1