EDBT 2026 Demo / reviewers in the wild / expert
Xiulian Peng
dblp:88/3644
· DBLP profile ↗
49ranked-venue papers
8as first author
21since 2021 · last 2025
0000-0001-8213-4878ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 44 · 6 first-author · 20 since 2021Artificial intelligence and machine learning · 13 · 10 since 2021Systems, architecture and hardware · 3 · 2 first-authorComputer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Bitrate-Controlled Diffusion for Disentangling Motion and Content in VideoabstractWe propose a novel and general framework to disentangle video data into its dynamic motion and static content components. Our proposed method is a self-supervised pipeline with less assumptions and inductive biases than previous works: it utilizes a transformer-based architecture to jointly generate flexible implicit features for frame-wise motion and clip-wise content, and incorporates a low-bitrate vector quantization as an information bottleneck to promote disentanglement and form a meaningful discrete motion space. The bitrate-controlled latent motion and content are used as conditional inputs to a denoising diffusion model to facilitate self-supervised representation learning. We validate our disentangled representation learning framework on real-world talking head videos with motion transfer and auto-regressive motion generation tasks. Furthermore, we also show that our method can generalize to other types of video data, such as pixel sprites of 2D cartoon characters. Our work presents a new perspective on self-supervised learning of disentangled video representations, contributing to the broader field of video analysis and generation. Xiao Li 0030, Qi Chen 0009, Xiulian Peng, Kai Yu 0004, Xie Chen 0001, Yan Lu 0001 |
ICCV | 3 |
| 2024 | QDFormer: Towards Robust Audiovisual Segmentation in Complex Environments with Quantization-based Semantic DecompositionabstractAudiovisual segmentation (AVS) is a challenging task that aims to segment visual objects in videos according to their associated acoustic cues. With multiple sound sources and background disturbances involved, establishing robust correspondences between audio and visual contents poses unique challenges due to (1) complex entanglement across sound sources and (2) frequent changes in the occurrence of distinct sound events. Assuming sound events occur in- dependently, the multi-source semantic space can be rep- resented as the Cartesian product of single-source sub- spaces. We are motivated to decompose the multi-source audio semantics into single-source semantics for more ef- fective interactions with visual content. We propose a se- mantic decomposition method based on product quanti- zation, where the multi-source semantics can be decom- posed and represented by several disentangled and noise- suppressed single-source semantics. Furthermore, we in- troduce a global-to-local quantization mechanism, which distills knowledge from stable global (clip-level) features into local (frame-level) ones, to handle frequent changes in audio semantics. Extensive experiments demonstrate that our semantically decomposed audio representation signifi- cantly improves AVS performance, e.g., +21.2% mIoU on the challenging AVS-Semantic benchmark with ResNet50 backbone. Xiang Li 0106, Jinglu Wang, Xiaohao Xu, Xiulian Peng, Rita Singh, Yan Lu 0001, Bhiksha Raj |
CVPR | 4 |
| 2024 | Low-Latency Speech Enhancement via Speech Token GenerationabstractExisting deep learning based speech enhancement mainly employ a data-driven approach, which leverage large amounts of data with a variety of noise types to achieve noise removal from noisy signal. However, the high dependence on the data limits its generalization on the unseen complex noises in real-life environment. In this paper, we focus on the low-latency scenario and regard speech enhancement as a speech generation problem conditioned on the noisy signal, where we generate clean speech instead of identifying and removing noises. Specifically, we propose a conditional generative framework for speech enhancement, which models clean speech by acoustic codes of a neural speech codec and generates the speech codes conditioned on past noisy frames in an auto-regressive way. Moreover, we propose an explicitalignment approach to align noisy frames with the generated speech tokens to improve the robustness and scalability to different input lengths. Different from other methods that leverage multiple stages to generate speech codes, we leverage a single-stage speech generation approach based on the TF-Codec neural codec to achieve high speech quality with low latency. Extensive results on both synthetic and real-recorded test set show its superiority over data-driven approaches in terms of noise robustness and temporal speech coherence. Huaying Xue, Xiulian Peng, Yan Lu 0001 |
ICASSP | 2 |
| 2024 | Convert and Speak: Zero-shot Accent Conversion with Minimum SupervisionabstractLow resource of parallel data is the key challenge of accent conversion(AC) problem in which both the pronunciation units and prosody pattern need to be converted. We propose a two-stage generative framework "convert-and-speak" in which the conversion is only operated on the semantic token level and the speech is synthesized conditioned on the converted semantic token with a speech generative model in target accent domain. The decoupling design enables the "speaking" module to use massive amount of target accent speech and relieves the parallel data required for the "conversion" module. Conversion with the bridge of semantic token also relieves the requirement for the data with text transcriptions and unlocks the usage of language pre-training technology to further efficiently reduce the need of parallel accent speech data. To reduce the complexity and latency of "speaking", a single-stage AR generative model is designed to achieve good quality as well as lower computation cost. Experiments on Indian-English to general American-English conversion show that the proposed framework achieves state-of-the-art performance in accent similarity, speech quality, and speaker maintenance with only 15 minutes of weakly parallel data which is not constrained to the same speaker. Extensive experimentation with diverse accent types suggests that this framework possesses a high degree of adaptability, making it readily scalable to accommodate other accents with low-resource data. Audio samples are available at https://www.microsoft.com/en-us/research/project/convert-and-speak-zero-shot-accent-conversion-with-minimumsupervision/. Zhijun Jia, Huaying Xue, Xiulian Peng, Yan Lu 0001 |
ACM Multimedia | 3 |
| 2023 | Disentangled Feature Learning for Real-Time Neural Speech CodingabstractRecently end-to-end neural audio/speech coding has shown its great potential to outperform traditional signal analysis based audio codecs. This is mostly achieved by following the VQ-VAE paradigm where blind features are learned, vector-quantized and coded. In this paper, instead of blind end-to-end learning, we propose to learn disentangled features for real-time neural speech coding. Specifically, more global-like speaker identity and local content features are learned with disentanglement to represent speech. Such a compact feature decomposition not only achieves better coding efficiency by exploiting bit allocation among different features but also provides the flexibility to do audio editing in embedding space, such as voice conversion in real-time communications. Both subjective and objective results demonstrate its coding efficiency and we find that the learned disentangled features show comparable performance on any-to-any voice conversion with modern self-supervised speech representation learning models with far less parameters and low latency, showing the potential of our neural coding framework. Xiulian Peng, Yuan Zhang 0013, Yan Lu 0001 |
ICASSP | 2 |
| 2023 | Dasformer: Deep Alternating Spectrogram Transformer For Multi/Single-Channel Speech SeparationabstractFor the task of speech separation, previous study usually treats multi-channel and single-channel scenarios as two research tracks with specialized solutions developed respectively. Instead, we propose a simple and unified architecture - DasFormer (Deep alternating spectrogram transFormer) to handle both of them in the challenging reverberant environments. Unlike frame-wise sequence modeling, each TF-bin in the spectrogram is assigned with an embedding encoding spectral and spatial information. With such input, DasFormer is then formed by multiple repetition of simple blocks each of which integrates 1) two multi-head self-attention (MHSA) modules alternately processing within each frequency bin & temporal frame of the spectrogram 2) MBConv before each MHSA for modeling local features on the spectrogram. Experiments show that DasFormer has a powerful ability to model the time-frequency representation, whose performance far exceeds the current SOTA models in multi-channel speech separation, and also achieves single-channel SOTA in the more challenging yet realistic reverberation scenario. Xiulian Peng, Hesam Movassagh, Vinod Prakash, Yan Lu 0001 |
ICASSP | 3 |
| 2023 | Improving Speech Enhancement via Event-Based QueryabstractExisting deep learning based speech enhancement (SE) methods either use blind end-to-end training or explicitly incorporate speaker embedding or phonetic information into the SE network to enhance speech quality. In this paper, we perceive speech and noises as different types of sound events and propose an event-based query method for SE. Specifically, speech embeddings that can discriminate speech from noises are first pre-trained with the sound event detection (SED) task. The embeddings are then clustered into fixed golden speech queries, i.e., general but representative speech embeddings, on a diverse clean speech dataset to assist the SE network. The golden speech queries can be obtained offline and generalizable to different SE datasets and networks. Therefore, little extra complexity is introduced and no enrollment is needed for each speaker. Experimental results show that the proposed method yields significant gains compared with baselines and the golden queries are well generalized to different datasets. Yifei Xin, Xiulian Peng, Yan Lu 0001 |
ICASSP | 2 |
| 2023 | Contrast-PLC: Contrastive Learning for Packet Loss ConcealmentabstractPacket loss concealment (PLC) is challenging in concealing missing contents both plausibly and naturally when there are only limited available context to use. Recently deep-learning based PLC algorithms have demonstrated their superiority over traditional counterparts; but their concealment ability is still mostly limited to a maximum of 120ms loss. Even with strong GAN-based generative models, it is still very challenging to predict long burst losses that could happen within/in-between phonemes. In this paper, we propose to use contrastive learning to learn a loss-robust semantic representation for PLC. A hybrid neural PLC architecture combining the semantic prediction and GAN-based generative model is designed to verify its effectiveness. Results on the blind test set of Interspeech2022 PLC Challenge show its superiority over commonly used UNet-style framework and the one without contrastive learning, especially for the longer burst loss at (120, 220]ms. Huaying Xue, Xiulian Peng, Yan Lu 0001 |
ICASSP | 2 |
| 2023 | Real-Time Speech Enhancement with Dynamic Attention SpanabstractFor real-time speech enhancement (SE) including noise suppression, dereverberation and acoustic echo cancellation, the time-variance of the audio signals becomes a severe challenge. The causality and memory usage limit that only the historical information can be used for the system to capture the time-variant characteristics. We propose to adaptively change the receptive field according to the input signal in deep neural network based SE model. Specifically, in an encoder-decoder framework, a dynamic attention span mechanism is introduced to all the attention modules for controlling the size of historical content used for processing the current frame. Experimental results verify that this dynamic mechanism can better track time-variant factors and capture speech-related characteristics, benefiting to both interference removing and speech quality retaining. Xiulian Peng, Yuan Zhang 0013, Yan Lu 0001 |
ICASSP | 3 |
| 2023 | ABC-KD: Attention-Based-Compression Knowledge Distillation for Deep Learning-Based Noise SuppressionabstractNoise suppression (NS) models have been widely applied to enhance speech quality.Recently, Deep Learning-Based NS, which we denote as Deep Noise Suppression (DNS), became the mainstream NS method due to its excelling performance over traditional ones.However, DNS models face 2 major challenges for supporting the real-world applications.First, highperforming DNS models are usually large in size, causing deployment difficulties.Second, DNS models require extensive training data, including noisy audios as inputs and clean audios as labels.It is often difficult to obtain clean labels for training DNS models.We propose the use of knowledge distillation (KD) to resolve both challenges.Our study serves 2 main purposes.To begin with, we are among the first to comprehensively investigate mainstream KD techniques on DNS models to resolve the two challenges.Furthermore, we propose a novel Attention-Based-Compression KD method that outperforms all investigated mainstream KD frameworks on DNS task. Yixin Wan, Xiulian Peng, Kai-Wei Chang 0001, Yan Lu 0001 |
INTERSPEECH | 3 |
| 2023 | Masked Audio Modeling with CLAP and Multi-Objective Learning
Yifei Xin, Xiulian Peng, Yan Lu 0001 |
INTERSPEECH | 2 |
| 2023 | Latent-Domain Predictive Neural Speech CodingabstractNeural audio/speech coding has recently demonstrated its capability to deliver high quality at much lower bitrates than traditional methods. However, existing neural audio/speech codecs employ either acoustic features or learned blind features with a convolutional neural network for encoding, by which there are still temporal redundancies within encoded features. This article introduces latent-domain predictive coding into the VQ-VAE framework to fully remove such redundancies and proposes the TF-Codec for low-latency neural speech coding in an end-to-end manner. Specifically, the extracted features are encoded conditioned on a prediction from past quantized latent frames so that temporal correlations are further removed. Moreover, we introduce a learnable compression on the time-frequency input to adaptively adjust the attention paid to main frequencies and details at different bitrates. A differentiable vector quantization scheme based on distance-to-soft mapping and Gumbel-Softmax is proposed to better model the latent distributions with rate constraint. Subjective results on multilingual speech datasets show that, with low latency, the proposed TF-Codec at 1 kbps achieves significantly better quality than Opus at 9 kbps, and TF-Codec at 3 kbps outperforms both EVS at 9.6 kbps and Opus at 12 kbps. Numerous studies are conducted to demonstrate the effectiveness of these techniques. Xiulian Peng, Huaying Xue, Yuan Zhang 0013, Yan Lu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Text Image Super-Resolution Guided by Text Structure and Embedding PriorsabstractWe aim to super-resolve text images from unrecognizable low-resolution inputs. Existing super-resolution methods mainly learn a direct mapping from low-resolution to high-resolution images by exploring low-level features, which usually generate blurry outputs and suffer from severe structure distortion for text parts, especially when the resolution is quite low. Both the visual quality and the readability will suffer. To tackle these issues, we propose a new text super-resolution paradigm by recovering with understanding. Specifically, we extract a text-embedding prior and a text-structure prior from the upsampled image by learning to understand the text. The two priors with rich structure information and text-embedding information are then used as auxiliary information to recover the clear text structure. In addition, we introduce a text-feature loss to guide the training for better text recognizability. Extensive evaluations on both screen and scene text image datasets show that our method largely outperforms the state-of-the-art in both visual quality and recognition accuracy. Xiulian Peng, Dong Liu 0002, Yan Lu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2022 | Reference-Based Speech Enhancement via Feature Alignment and Fusion NetworkabstractSpeech enhancement aims at recovering a clean speech from a noisy input, which can be classified into single speech enhancement and personalized speech enhancement. Personalized speech enhancement usually utilizes the speaker identity extracted from the noisy speech itself (or a clean reference speech) as a global embedding to guide the enhancement process. Different from them, we observe that the speeches of the same speaker are correlated in terms of frame-level short-time Fourier Transform (STFT) spectrogram. Therefore, we propose reference-based speech enhancement via a feature alignment and fusion network (FAF-Net). Given a noisy speech and a clean reference speech spoken by the same speaker, we first propose a feature level alignment strategy to warp the clean reference with the noisy speech in frame level. Then, we fuse the reference feature with the noisy feature via a similarity-based fusion strategy. Finally, the fused features are skipped connected to the decoder, which generates the enhanced results. Experimental results demonstrate that the performance of the proposed FAF-Net is close to state-of-the-art speech enhancement methods on both DNS and Voice Bank+DEMAND datasets. Our code is available at https://github.com/HieDean/FAF-Net. Huanjing Yue, Wenxin Duo, Xiulian Peng, Jing-Yu Yang 0002 |
AAAI | 3 |
| 2022 | End-to-End Neural Speech Coding for Real-Time CommunicationsabstractDeep-learning based methods have shown their advantages in audio coding over traditional ones but limited attention has been paid on real-time communications (RTC). This paper proposes the TFNet, an end-to-end neural speech codec with low latency for RTC. It takes an encoder-temporal filtering-decoder paradigm that has seldom been investigated in audio coding. An interleaved structure is proposed for temporal filtering to capture both short-term and long-term temporal dependencies. Furthermore, with end-to-end optimization, the TFNet is jointly optimized with speech enhancement and packet loss concealment, yielding a one-for-all network for three tasks. Both subjective and objective results demonstrate the efficiency of the proposed TFNet. Xiulian Peng, Huaying Xue, Yuan Zhang 0013, Yan Lu 0001 |
ICASSP | 2 |
| 2022 | Cross-Scale Vector Quantization for Scalable Neural Speech CodingabstractBitrate scalability is a desirable feature for audio coding in real-time communications.Existing neural audio codecs usually enforce a specific bitrate during training, so different models need to be trained for each target bitrate, which increases the memory footprint at the sender and the receiver side and transcoding is often needed to support multiple receivers.In this paper, we introduce a cross-scale scalable vector quantization scheme (CSVQ), in which multi-scale features are encoded progressively with stepwise feature fusion and refinement.In this way, a coarse-level signal is reconstructed if only a portion of the bitstream is received, and progressively improves the quality as more bits are available.The proposed CSVQ scheme can be flexibly applied to any neural audio coding network with a mirrored auto-encoder structure to achieve bitrate scalability.Subjective results show that the proposed scheme outperforms the classical residual VQ (RVQ) with scalability.Moreover, the proposed CSVQ at 3 kbps outperforms Opus at 9 kbps and Lyra at 3kbps and it could provide a graceful quality boost with bitrate increase. Xiulian Peng, Huaying Xue, Yuan Zhang 0013, Yan Lu 0001 |
INTERSPEECH | 2 |
| 2022 | Multi-Modal Multi-Correlation Learning for Audio-Visual Speech SeparationabstractIn this paper we propose a multi-modal multi-correlation learning framework targeting at the task of audio-visual speech separation.Although previous efforts have been extensively put on combining audio and visual modalities, most of them solely adopt a straightforward concatenation of audio and visual features.To exploit the real useful information behind these two modalities, we define two key correlations which are: (1) identity correlation (between timbre and facial attributes); (2) phonetic correlation (between phoneme and lip motion).These two correlations together comprise the complete information, which shows a certain superiority in separating target speaker's voice especially in some hard cases, such as the same gender or similar content.For implementation, contrastive learning or adversarial training approach is applied to maximize these two correlations.Both of them work well, while adversarial training shows its advantage by avoiding some limitations of contrastive learning.Compared with previous research, our solution demonstrates clear improvement on experimental metrics without additional complexity.Further analysis reveals the validity of the proposed architecture and its good potential for future extension. Xiulian Peng, Yan Lu 0001 |
INTERSPEECH | 3 |
| 2022 | Towards Error-Resilient Neural Speech CodingabstractNeural audio coding has shown very promising results recently in the literature to largely outperform traditional codecs but limited attention has been paid on its error resilience.Neural codecs trained considering only source coding tend to be extremely sensitive to channel noises, especially in wireless channels with high error rate.In this paper, we investigate how to elevate the error resilience of neural audio codecs for packet losses that often occur during real-time communications.We propose a feature-domain packet loss concealment algorithm (FD-PLC) for real-time neural speech coding.Specifically, we introduce a self-attention-based module on the received latent features to recover lost frames in the feature domain before the decoder.A hybrid segment-level and frame-level frequencydomain discriminator is employed to guide the network to focus on both the generative quality of lost frames and the continuity with neighbouring frames.Experimental results on several error patterns show that the proposed scheme can achieve better robustness compared with the corresponding error-free and error-resilient baselines.We also show that feature-domain concealment is superior to waveform-domain counterpart as postprocessing. Huaying Xue, Xiulian Peng, Yan Lu 0001 |
INTERSPEECH | 2 |
| 2022 | Time-Variance Aware Dynamic Kernel Generation for Real-Time Acoustic Echo CancellationabstractTime-variant factors including dynamic delay and varying echo path often occur in real-world acoustic echo cancellation (AEC) applications. Current end-to-end deep neural network (DNN) based methods usually model the time-variant components implicitly and can hardly handle the unpredictable time-variance in real-time AEC. To explicitly capture the time-variant components, we propose a dynamic kernel generation (DKG) module that can be introduced as a learnable plug-in to a DNN-based end-to-end pipeline. Specifically, the DKG module generates a convolutional kernel regarding to each input audio frame, so that the DNN model is able to dynamically adjust its weights according to the input signal during inference. Experimental results verify that DKG module improves the AEC performance of the model under time-variant scenarios, especially in the double-talk cases. Xiulian Peng, Yuan Zhang 0013, Yan Lu 0001 |
IEEE Signal Process. Lett. | 3 |
| 2021 | Interactive Speech and Noise Modeling for Speech EnhancementabstractSpeech enhancement is challenging because of the diversity of background noise types. Most of the existing methods are focused on modelling the speech rather than the noise. In this paper, we propose a novel idea to model speech and noise simultaneously in a two-branch convolutional neural network, namely SN-Net. In SN-Net, the two branches predict speech and noise, respectively. Instead of information fusion only at the final output layer, interaction modules are introduced at several intermediate feature domains between the two branches to benefit each other. Such an interaction can leverage features learned from one branch to counteract the undesired part and restore the missing component of the other and thus enhance their discrimination capabilities. We also design a feature extraction module, namely residual-convolution-and-attention (RA), to capture the correlations along temporal and frequency dimensions for both the speech and the noises. Evaluations on public datasets show that the interaction module plays a key role in simultaneous modeling and the SN-Net outperforms the state-of-the-art by a large margin on various evaluation metrics. The proposed SN-Net also shows superior performance for speaker separation. Xiulian Peng, Yuan Zhang 0013, Sriram Srinivasan 0003, Yan Lu 0001 |
AAAI | 2 |
| 2021 | Phoneme-Based Distribution Regularization for Speech EnhancementabstractExisting speech enhancement methods mainly separate speech from noises at the signal level or in the time-frequency domain. They seldom pay attention to the semantic information of a corrupted signal. In this paper, we aim to bridge this gap by extracting phoneme identities to help speech enhancement. Specifically, we propose a phoneme-based distribution regularization (PbDr) for speech enhancement, which incorporates frame-wise phoneme information into speech enhancement network in a conditional manner. As different phonemes always lead to different feature distributions in frequency, we propose to learn a parameter pair, i.e. scale and bias, through a phoneme classification vector to modulate the speech enhancement network. The modulation parameter pair includes not only frame-wise but also frequency-wise conditions, which effectively map features to phoneme-related distributions. In this way, we explicitly regularize speech enhancement features by recognition vectors. Experiments on public datasets demonstrate that the proposed PbDr module can not only boost the perceptual quality for speech enhancement but also the recognition accuracy of an ASR system on the enhanced speech. This PbDr module could be readily incorporated into other speech enhancement networks as well. Xiulian Peng, Zhiwei Xiong, Yan Lu 0001 |
ICASSP | 2 |
| 2020 | Convolutional Neural Network-Based Arithmetic Coding for HEVC Intra-Predicted ResiduesabstractEntropy coding is a fundamental technology in video coding that removes statistical redundancy among syntax elements. In high efficiency video coding (HEVC), context-adaptive binary arithmetic coding (CABAC) is adopted as the primary entropy coding method. The CABAC consists of three steps: binarization, context modeling, and binary arithmetic coding. As the binarization processes and context models are both manually designed in CABAC, the probability of the syntax elements may not be estimated accurately, which restricts the coding efficiency of CABAC. To address the problem, we propose a convolutional neural network-based arithmetic coding (CNNAC) method and apply it to compress the syntax elements of the intra-predicted residues in HEVC. Instead of manually designing the binarization processes and context models, we propose directly estimating the probability distribution of the syntax elements with a convolutional neural network (CNN), as CNNs can adaptively build complex relationships between inputs and outputs by training with a lot of data. Then, the values of the syntax elements, together with their estimated probability distributions, are fed into a multi-level arithmetic codec to perform entropy coding. In this paper, we have utilized the CNNAC to code the syntax elements of the DC coefficient; the lowest frequency AC coefficient; the second, third, fourth, and fifth lowest frequency AC coefficients; and the position of the last non-zero coefficient in the HEVC intra-predicted residues. The experimental results show that our proposed method achieves up to 6.7% BD-rate reduction and an average of 4.7% BD-rate reduction compared to the HEVC anchor under all intra (AI) configuration. Changyue Ma, Dong Liu 0002, Xiulian Peng, Li Li 0040, Feng Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | Neural Network-Based Arithmetic Coding for Inter Prediction Information in HEVCabstractEntropy coding is a fundamental technique in video coding to remove the statistical redundancy in syntax elements. Currently, context-adaptive binary arithmetic coding (CABAC) is used as the entropy coding tool in HEVC. Considering that the manually designed binarization and context models are not flexible to estimate the probability of the syntax elements, we use neural networks to estimate the probability of the syntax elements, then the estimated probabilities together with the values of the syntax elements are fed into an arithmetic coding engine to fulfill entropy coding. In this paper, we focus on the syntax elements of inter prediction information that consists of merge flag, merge index, reference index, motion vector difference and motion vector prediction index in HEVC under low-delay P (LDP) setting. Compared with the previous work on neural network-based arithmetic coding for intra prediction modes and intra DC coefficients, there are three new characteristics in this paper. First, surrounding syntax elements are directly fed into the neural network without converting to reconstructed pixels. Second, unified neural networks are designed for different prediction block sizes. Finally, dependency among the syntax elements in current prediction unit is omitted to improve parallelism. Experimental results show that compared with HEVC, our proposed method achieves up to 0.5% and on average 0.3% BD-rate reduction in LDP configuration. Changyue Ma, Dong Liu 0002, Xiulian Peng, Zhengjun Zha, Feng Wu 0001 |
ISCAS | 3 |
| 2019 | Traffic surveillance video coding with libraries of vehicles and background
Changyue Ma, Dong Liu 0002, Xiulian Peng, Li Li 0040, Feng Wu 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2019 | Reference Clip for Inter Prediction in Video CodingabstractInter prediction is a fundamental technology in video coding to remove the temporal redundancy between video frames. Traditionally, the reconstructed frames are directly put into a reference frame buffer to serve as references for inter prediction. Using multiple reference frames increases the accuracy of inter prediction, but also incurs a waste of memory of the buffer since the content of reference frames is highly similar. To address this problem, we propose to organize the references at clip level in addition to frame level, i.e. the reference buffer stores not only reference frames, but also reference clips that are cropped regions selected from the reconstructed frames. Using clip-level references, we can manage the reference content more economically, since the content of multiple reference frames is divided into the singular content of each frame as well as the repetitive content that appears in multiple frames. For the repetitive content, only one copy is stored in reference clips so as to avoid duplicate. Moreover, using reference clips also facilitates the bit-rate allocation among reference content, i.e. the quality of each clip can be decided adaptively to achieve the rate-distortion optimization. In this paper, we propose a complete video coding framework using reference clips, and investigate the problems including how to generate reference clips as either singular content clips or repetitive content clips, how to manage the clips, how to utilize the clips for inter prediction, and how to allocate bit-rate among clips, in a systematic manner. The proposed video coding framework is implemented upon the state-of-the-art video coding scheme, High Efficiency Video Coding (HEVC). Experimental results show that our scheme achieves on average 5.1% and 5.0% BD-rate reduction than the HEVC anchor, in low-delay B and low-delay P settings, respectively. We believe that reference clip opens up a new dimension for optimizing inter prediction in video coding, and thus is worthy of further study. Changyue Ma, Dong Liu 0002, Xiulian Peng, Feng Wu 0001, Houqiang Li, Tingting Wang 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | End-to-End United Video Dehazing and DetectionabstractThe recent development of CNN-based image dehazing has revealed the effectiveness of end-to-end modeling. However, extending the idea to end-to-end video dehazing has not been explored yet. In this paper, we propose an End-to-End Video Dehazing Network (EVD-Net), to exploit the temporal consistency between consecutive video frames. A thorough study has been conducted over a number of structure options, to identify the best temporal fusion strategy. Furthermore, we build an End-to-End United Video Dehazing and Detection Network (EVDD-Net), which concatenates and jointly trains EVD-Net with a video object detection model. The resulting augmented end-to-end pipeline has demonstrated much more stable and accurate detection results in hazy video. Boyi Li 0001, Xiulian Peng, Zhangyang Wang, Jizheng Xu, Dan Feng 0001 |
AAAI | 2 |
| 2018 | Convolutional Neural Network-Based Arithmetic Coding of DC Coefficients for HEVC Intra CodingabstractIn the state-of-the-art video coding standard-High Efficiency Video Coding (HEVC), context-adaptive binary arithmetic coding (CABAC) is adopted as the entropy coding tool. In CABAC, the binarization processes are manually designed, and the context models are empirically crafted, both of which incur that the probability distribution of the syntax elements may not be estimated accurately, and restrict the coding efficiency. In this paper, we adopt a convolutional neural network-based arithmetic coding (CNNAC) strategy, and conduct studies on the coding of the DC coefficients for HEVC intra coding. Instead of manually designing binarization process and context model, we propose to directly estimate the probability distribution of the value of the DC coefficient using densely connected convolutional networks. The estimated probability together with the real DC coefficient are then input into a multi-level arithmetic codec to fulfill entropy coding. Simulation results show that our proposed CNNAC leads to on average 22.47% bits saving compared with CABAC for the bits of DC coefficients, which corresponds to 1.6% BD-rate reduction than the HEVC anchor. Changyue Ma, Dong Liu 0002, Xiulian Peng, Feng Wu 0001 |
ICIP | 3 |
| 2018 | Frequency-Domain Dynamic Pruning for Convolutional Neural NetworksabstractDeep convolutional neural networks have demonstrated their powerfulness in a variety of applications. However, the storage and computational requirements have largely restricted their further extensions on mobile devices. Recently, pruning of unimportant parameters has been used for both network compression and acceleration. Considering that there are spatial redundancy within most filters in a CNN, we propose a frequency-domain dynamic pruning scheme to exploit the spatial correlations. The frequency-domain coefficients are pruned dynamically in each iteration and different frequency bands are pruned discriminatively, given their different importance on accuracy. Experimental results demonstrate that the proposed scheme can outperform previous spatial-domain counterparts by a large margin. Specifically, it can achieve a compression ratio of 8.4x and a theoretical inference speed-up of 9.2x for ResNet-110, while the accuracy is even better than the reference model on CIFAR-110. Zhenhua Liu 0003, Jizheng Xu, Xiulian Peng, Ruiqin Xiong |
NeurIPS | 3 |
| 2018 | Unequal Error Protection for Scalable Video Storage in the CloudabstractRedundancy is necessary for a storage system to achieve reliability. Frequent errors in large-scale storage systems, for example, cloud, make it desirable to reduce the cost of recovery. Among all types of data in cloud storage, videos generally occupy significant amounts of space due to high volumes and the rapid development of video sharing and video-on-demand services. Unlike general data, videos can tolerate a certain level of quality degradation. This paper investigates multilayer video representations, such as scalable videos and simulcast streaming, and proposes an unequal error protection scheme based on local reconstruction codes (LRC) for video storage. By providing less protection for less important layers or video copies, a better tradeoff between storage and repair cost is achieved. Both theoretical and simulation results show that such a tradeoff can be achieved over the LRC with equal error protection, though the recovered video quality might be slightly lower in rare cases. Xiaodan Song, Xiulian Peng, Jizheng Xu, Guangming Shi, Feng Wu 0001 |
IEEE Trans. Multim. | 2 |
| 2017 | AOD-Net: All-in-One Dehazing NetworkabstractThis paper proposes an image dehazing model built with a convolutional neural network (CNN), called All-in-One Dehazing Network (AOD-Net). It is designed based on a re-formulated atmospheric scattering model. Instead of estimating the transmission matrix and the atmospheric light separately as most previous models did, AOD-Net directly generates the clean image through a light-weight CNN. Such a novel end-to-end design makes it easy to embed AOD-Net into other deep models, e.g., Faster R-CNN, for improving high-level tasks on hazy images. Experimental results on both synthesized and natural hazy image datasets demonstrate our superior performance than the state-of-the-art in terms of PSNR, SSIM and the subjective visual quality. Furthermore, when concatenating AOD-Net with Faster R-CNN, we witness a large improvement of the object detection performance on hazy images. Boyi Li 0001, Xiulian Peng, Zhangyang Wang, Jizheng Xu, Dan Feng 0001 |
ICCV | 2 |
| 2017 | Surveillance video coding with vehicle libraryabstractInter prediction in video coding is very efficient to remove temporal redundancy. However, due to the limitation of short-term references, inter prediction can work only within a very short time interval. In surveillance videos, we observe that there are always similar vehicles passing through one static camera, but the time intervals of similar vehicles are usually several seconds to minutes, exceeding the time interval that short-term references can handle. To solve this problem, we propose to build a vehicle library, and to put high-quality copies of the similar vehicles into the vehicle library. During encoding, vehicles are detected from the current frame, and for each vehicle we can retrieve similar vehicles from the vehicle library, and take the retrieved vehicle picture as additional references for inter prediction. Preliminary experimental results show that the proposed vehicle library based method achieves as high as 10.1% bit-rate saving for surveillance video coding, compared to HEVC anchor. Changyue Ma, Dong Liu 0002, Xiulian Peng, Feng Wu 0001 |
ICIP | 3 |
| 2017 | Distributed Compressive Sensing for Cloud-Based Wireless Image TransmissionabstractWe consider efficient image transmission via time-varying channels. To improve the performance, we propose a new distributed compressive sensing (CS) scheme that can leverage similar images in the cloud. It is featured by channel SNR and bandwidth scalability, high efficiency, and low encoding complexity. For each image, a compressed thumbnail is first transmitted after forward error correction (FEC) and modulation to retrieve similar images and generate a side information (SI) in the cloud. The residual image after subtracting the decompressed thumbnail is then coded and transmitted by CS through a very dense constellation without FEC. The linearly and ratelessly generated CS measurements make it capable of achieving both graceful quality degradation (GD) with the channel SNR and bandwidth scalability in a universal scheme. A mode decision and transform-domain power allocation are introduced for better bandwidth usage and protection against channel errors. At the decoder, a two-step CS decoding is performed to recover the residual signal, where both the local and nonlocal correlations within the image and that with the SI are exploited. Simulations on landmark images and an AWGN channel show that the received image quality gracefully increases with the channel SNR and bandwidth. Furthermore, it outperforms existing schemes both subjectively and objectively by up to 11 dB gains compared with the state-of-the-art transmission scheme with GD, i.e. SoftCast. Xiaodan Song, Xiulian Peng, Jizheng Xu, Guangming Shi, Feng Wu 0001 |
IEEE Trans. Multim. | 2 |
| 2016 | Hash-Based Line-by-Line Template Matching for Lossless Screen Image CodingabstractTemplate matching (TM) was proposed in the literature a decade ago to efficiently remove non-local redundancies within an image without transmitting any overhead of displacement vectors. However, the large computational complexity introduced at both the encoder and the decoder, especially for a large search range, limits its widespread use. This paper proposes a hash-based line-by-line template matching (hLTM) for lossless screen image coding, where the non-local redundancy commonly exists in text and graphics parts. By hash-based search, it can largely reduce the search complexity of template matching without an accuracy degradation. Besides, the line-by-line template matching increases prediction accuracy by using a fine granularity. Experimental results show that the hLTM can significantly reduce both the encoding and decoding complexities by 68 and 23 times, respectively, compared with the traditional TM with a search radius of 128. Moreover, when compared with High Efficiency Video Coding screen content coding test model SCM-1.0, it can largely improve coding efficiency by up to 12.68% bits saving on screen contents with rich texts/graphics. Xiulian Peng, Jizheng Xu |
IEEE Trans. Image Process. | 1 |
| 2015 | Unequal error protection for scalable video storage in the cloudabstractRedundancy is necessary for a storage system to recover from errors. The frequent errors in large-scale systems, e.g. cloud, make it desired to reduce the recovery cost. Among all kinds of data stored in the cloud, video takes a large portion due to its large data volume. The other characteristic of video is that a certain distortion can be tolerated. This paper investigates using scalable video representation and unequal error protection scheme to reduce the storage and recovery costs in the cloud. By introducing more protection for the base layer and less on the enhancement layers, it can achieve a better tradeoff between storage and reconstruction costs although the reliability for the enhancement layer sacrifices a little. Simulation results based on local reconstruction codes (LRC) show that comparing with the existing (12, 2, 2) LRC code in Windows Azure Storage, the reconstruction cost can be reduced from 6x to 3x at the same storage cost at the expense of possible video quality loss. Xiaodan Song, Xiulian Peng, Jizheng Xu, Guangming Shi, Feng Wu 0001 |
ICME | 2 |
| 2015 | Compressive sensing based image transmission with side information at the decoderabstractThis paper proposes a distributed compressive sensing (CS) scheme for robust image transmission over unknown or time-varying channels with highly correlated images at the decoder. A compressed thumbnail is first transmitted after digital forward error correction (FEC) and modulation to retrieve highly correlated images and generate a side information (SI) at the decoder. The current residual image after subtracting the decompressed thumbnail is then coded and transmitted by CS through a very dense constellation without FEC. The linear representation of the residual signal by CS measurements and rateless sampling makes it able to achieve graceful degradation and bandwidth scalability without channel feedback. Moreover, a transform-domain power allocation is employed before random sampling to protect against channel errors. At the decoder, both the nonlocal correlations within the original image and the correlation with the SI are exploited in CS decoding via a low-rank regulation on similar patches. After CS decoding, a block-wise minimum-mean-square-error (MMSE) reconstruction using the SI is further performed in the spatial domain to enhance the reconstruction quality. Simulations on landmark images and an unknown Gaussian channel show that an up to 10 dB gain is achieved at low channel SNRs compared with the state-of-the-art uncoded image transmission scheme, i.e. SoftCast, when highly correlated images are available at the decoder. Xiaodan Song, Xiulian Peng, Jizheng Xu, Guangming Shi, Feng Wu 0001 |
VCIP | 2 |
| 2015 | Cloud-Based Distributed Image CodingabstractWith multimedia flourishing on the Web, it is easy to find similar images for a query, especially landmark images. Traditional image coding, such as JPEG, cannot exploit correlations with external images. Existing vision-based approaches are able to exploit such correlations by reconstructing from local descriptors but cannot ensure the pixel-level fidelity of the reconstruction. In this paper, a cloud-based distributed image coding (Cloud-DIC) scheme is proposed to exploit external correlations for mobile photo uploading. For each input image, a thumbnail is transmitted to retrieve correlated images and reconstruct it in the cloud by geometrical and illumination registrations. Such a reconstruction serves as the side information (SI) in the Cloud-DIC. The image is then compressed by a transform-domain syndrome coding to correct the disparity between the original image and the SI. Once a bitplane is received in the cloud, an iterative refinement process is performed between the final reconstruction and the SI. Moreover, a joint encoder/decoder mode decision at block, frequency, and bitplane levels is proposed to adapt to different correlations. Experimental results on a landmark image database show that the Cloud-DIC can largely enhance the coding efficiency both subjectively and objectively, with up to 5-dB gains and 70% bits saving over JPEG with arithmetic coding, and perform comparably at low bitrates with the intra coding of the High Efficiency Video Coding standard with a much lower encoder complexity. Xiaodan Song, Xiulian Peng, Jizheng Xu, Guangming Shi, Feng Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2014 | Cloud-based distributed image codingabstractThis paper proposes a cloud-based distributed image coding scheme (Cloud-DIC) to exploit the strong correlations with external partial-duplicate images in the cloud. It features both high coding efficiency and low encoder complexity, which makes it suitable for photo sharing on mobile devices. To get the side information in the cloud, a thumbnail of the current image is transmitted to retrieve highly correlated images and reconstruct through geometrical registration and adaptive patched-based stitching. The current image is then compressed by a transform-domain syndrome coding, bitplane by bitplane. Once a bitplane is received, the decoded high-quality image is further used to refine the side information in the cloud, which will benefit the coding of following bitplanes and the reconstruction. Experimental results on a landmark image database show that it can largely enhance the coding efficiency both subjectively and objectively with up to 5 dB gains and 58% bits saving over JPEG. Xiaodan Song, Xiulian Peng, Jizheng Xu, Feng Wu 0001 |
ICIP | 2 |
| 2014 | Screen content coding for HEVC by improved line-based intra block copyabstractThe line-based intra block copy (IntraBC) technique is newly proposed in the High Efficient Video Coding (HEVC) Range Extensions to deal with repeated patterns within a picture for screen content coding. One challenge for line-based IntraBC is the large overhead by transmitting displacement vectors (DV) for each line. In this paper, a DV prediction is proposed to reduce such an overhead. Moreover, a flipping copy based prediction is proposed to better exploit the correlations within screen content. Experimental results show that our scheme can provide a BD-rate reduction of 11.0% compare with the HEVC Range Extension anchor with the encoding time increased by 23% and no decoding time increase. Xiulian Peng, Jizheng Xu |
ICIP | 3 |
| 2014 | LineCast: Line-Based Distributed Coding and Transmission for Broadcasting Satellite ImagesabstractIn this paper, we propose a novel coding and transmission scheme, called LineCast, for broadcasting satellite images to a large number of receivers. The proposed LineCast matches perfectly with the line scanning cameras that are widely adopted in orbit satellites to capture high-resolution images. On the sender side, each captured line is immediately compressed by a transform-domain scalar modulo quantization. Without syndrome coding, the transmission power is directly allocated to quantized coefficients by scaling the coefficients according to their distributions. Finally, the scaled coefficients are transmitted over a dense constellation. This line-based distributed scheme features low delay, low memory cost, and low complexity. On the receiver side, our proposed line-based prediction is used to generate side information from previously decoded lines, which fully utilizes the correlation among lines. The quantized coefficients are decoded by the linear least square estimator from the received data. The image line is then reconstructed by the scalar modulo dequantization using the generated side information. Since there is neither syndrome coding nor channel coding, the proposed LineCast can make a large number of receivers reach the qualities matching their channel conditions. Our theoretical analysis shows that the proposed LineCast can achieve Shannon's optimum performance by using a high-dimensional modulo-lattice quantization. Experiments on satellite images demonstrate that it achieves up to 1.9-dB gain over the state-of-the-art 2D broadcasting scheme and a gain of more than 5 dB over JPEG 2000 with forward error correction. Feng Wu 0001, Xiulian Peng, Jizheng Xu |
IEEE Trans. Image Process. | 2 |
| 2013 | Scene-aware perceptual video codingabstractThe mean-square-error (MSE) distortion criterion used in the state-of-the-art video coding standards, e.g. H.264/AVC and the High Efficiency Video Coding (HEVC) under standardization, is widely criticized for poor measurement of perceived visual quality. Existing research on perceptual video coding mainly employs low-level features of images/video, which cannot take into account the big picture people see. This paper proposes a scene-aware perceptual video coding scheme (SAPC), which accommodates human visual perception of the scene by reconstructing the scene from video and perform scene-based bits allocation. To be specific, more bits are allocated to the foreground object and its boundaries considering that people tend to pay more attention to the foreground and object boundaries are prone to blur at low bitrates for object occlusion. The structure from motion (SFM) technology is employed for scene reconstruction. Experiments taking HEVC as the benchmark show that our algorithm can give better visual quality than the original HEVC encoder at the same bitrate. Xiulian Peng, Jizheng Xu |
VCIP | 2 |
| 2013 | A light-weight HEVC encoder for image codingabstractHigh Efficiency Video Coding (HEVC), not only provides a much better coding efficiency than previous video coding standards, but also shows significantly superior performance than other image coding schemes when applied to image coding. However, the improvement is at the cost of significant increase of encoding complexity. In this paper, we focus on retaining the high coding efficiency provided by HEVC while largely reducing its encoding complexity for image coding. By applying various techniques including optimized coding structure parameters, coding unit early termination, fast intra prediction and transform skip mode decision, we significantly reduce the complexity of HEVC intra coding while keeping most of its coding efficiency. Experimental results show that our light-weight HEVC encoder can save about 82% coding time compared with original HEVC encoder. With a slight loss to the HEVC reference software, the proposed scheme still gains about 19% in BD-BR compared with H.264/AVC. Xiulian Peng, Jizheng Xu |
VCIP | 2 |
| 2012 | Highly Parallel Line-Based Image Coding for Many CoresabstractComputers are developing along with a new trend from the dual-core and quad-core processors to ones with tens or even hundreds of cores. Multimedia, as one of the most important applications in computers, has an urgent need to design parallel coding algorithms for compression. Taking intraframe/image coding as a start point, this paper proposes a pure line-by-line coding scheme (LBLC) to meet the need. In LBLC, an input image is processed line by line sequentially, and each line is divided into small fixed-length segments. The compression of all segments from prediction to entropy coding is completely independent and concurrent at many cores. Results on a general-purpose computer show that our scheme can get a 13.9 times speedup with 15 cores at the encoder and a 10.3 times speedup at the decoder. Ideally, such near-linear speeding relation with the number of cores can be kept for more than 100 cores. In addition to the high parallelism, the proposed scheme can perform comparatively or even better than the H.264 high profile above middle bit rates. At near-lossless coding, it outperforms H.264 more than 10 dB. At lossless coding, up to 14% bit-rate reduction is observed compared with H.264 lossless coding at the high 4:4:4 profile. Xiulian Peng, Jizheng Xu, Feng Wu 0001 |
IEEE Trans. Image Process. | 1 |
| 2011 | Highly parallel image coding for many coresabstractThis paper proposes a pure line-by-line image coding scheme (LBC) for many-core computers. The proposed scheme can simultaneously use tens and even hundreds of cores to compress one HD image. In LBC, an input image is coded line by line sequentially and each line is divided into many short segments at equal length that are coded independently. Experimental results show that the proposed scheme can speed the encoding up to 14 times through 16 cores, whereas parallel H.264 intra-frame coding can only speed up less than 5 times. In terms of coding performance, efficient intra predictions are proposed to make full use of the strong correlations among neighboring lines. Experiments results also show that the proposed scheme can have a comparable and even better performance than H.264 high profile at middle bit rates. At high bit rates, the proposed scheme can outperform H.264 high profile more than 10dB. Xiulian Peng, Jizheng Xu, Feng Wu 0001 |
ISCAS | 1 |
| 2011 | Exploiting inter-frame correlations in compound video codingabstractCompound images/videos are a mixture of text, graphics and natural images/video. Despite extensive research on compound image coding, rare research on compound video coding has been reported in the literature. This paper proposes three approaches to exploit inter-frame correlations in compound video. One is the motion compensation aided base color (MCA-BC) approach, which handles a special type of motion on text/graphics parts: foreground/background color change. The other two are motion-compensated residue base color and index map (MC-BCIM) and the motion-compensated residue scalar quantization (MC-RSQ) approach, both of which exploit spatial correlations among motion compensated residues. Experiments based on the High-Efficiency Video Coding (HEVC) reference software HM0.9 show that the proposed scheme largely improves the coding efficiency for more than 3dB on compound video with text/graphics and still keep a comparable performance on natural video content. Xiulian Peng, Jizheng Xu, Feng Wu 0001 |
VCIP | 1 |
| 2010 | Line-based image coding using adaptive prediction filtersabstractThis paper proposes a line-based approach for image coding. Unlike conventional block-based image coding, the proposed method takes one line as a basic unit in coding and reconstruction. Such a structure allows us to perform better predictions between neighboring units and affords us more flexibility when coding each unit. We also present an efficient prediction method using adaptive prediction filters for this structure. Based on the neighboring information, we apply the best prediction filter on each line to get high-quality predictions. We show that the proposed line-based coding can perform comparatively or even better than the state-of-the-art block-based coding. When combing line-based and block-based coding together, the state-of-the-art image coding can be much improved. Xiulian Peng, Jizheng Xu, Feng Wu 0001 |
ISCAS | 1 |
| 2010 | Improved line-based image coding by exploiting long-distance correlationsabstractLine-based coding has shown its potential in improving the coding efficiency of intra-frame/image coding due to its flexibility in prediction. In our previous work we proposed an efficient line-based image coding method (LIC) by adaptive line-by-line prediction (ALP) and adaptive residue coding. In this paper we further improve this line-based coding scheme by exploiting long-distance correlations in prediction. Line-by-line template matching (LTM) is introduced to perform the long-distance prediction. Experiments in the KTA software show that the LTM scheme can effectively improve the performance of LIC at low bitrates and also brings some improvement at high bitrates. Up to 1db gain is achieved compared to previous LIC and 1.5db compared to KTA on images with regular patterns or strong edges. Up to 20% rate reduction over KTA is also achieved on some standard video sequences with different resolutions. Xiulian Peng, Jizheng Xu, Feng Wu 0001 |
VCIP | 1 |
| 2010 | Directional Filtering Transform for Image/Intra-Frame CompressionabstractWhile directional adaption is introduced into traditional transforms, different orders of two 1-D transforms will result in different results of one 2-D transform. Based upon an anisotropic image model, this paper analyzes the effect of transform orders in terms of theoretical coding gain. Our results reveal that the transform orders have little effect on the coding gain with full decomposition, good directional modes and good interpolation. However, in practical compression schemes, since high-pass bands are not decomposed fully because of the consideration on complexity, different transform orders have different coding performances, which can be solved by an adaptive transform order. Motivated by our analyzed results, a directional filtering transform (dFT, in order to distinguish from the common usage on DFT) is proposed in this paper to better exploit correlations among samples in H.264 intraframe coding. It provides an evenly distributed set of prediction modes with an adaptive transform order. Both interblock and intrablock correlations are exploited in this scheme. Experimental results in H.264 intraframe coding demonstrate its superiority both objectively and subjectively. Xiulian Peng, Jizheng Xu, Feng Wu 0001 |
IEEE Trans. Image Process. | 1 |
| 2009 | How Can Intra Correlation Be Exploited Better?abstractSummary form only given. This paper studies how to better exploit intra correlation to compress images. In general, edge and texture areas of images exhibit strong anisotropic property. The correlation among samples is determined by not only their distance but also the link orientation. Traditional transforms are not efficient on handling this anisotropic correlation. Therefore, in this paper we propose a directional filtering transform (dFT, in order to distinguish from the common usage on DFT) to exploit local anisotropic correlation among samples. Similar to directional prediction in H.264 intra-frame coding, but it adopts the hierarchal structure to decrease the distance between samples to be predicted and that are used for prediction. From another viewpoint, the dFT prediction resembles the directional wavelet transform without update, which can take both intra-block and inter-block correlations into account. Feng Wu 0001, Xiulian Peng, Jizheng Xu, Shipeng Li 0001 |
DCC | 2 |
| 2009 | Directional filtering transformabstractThis paper proposes the directional filtering transform (dFT, in order to distinguish from the common usage on DFT) to better exploit intra-frame correlation in H.264 intra-frame coding. It consists of a directional filtering and an optional DCT transform. In the proposed directional filtering, there are two different approaches. One is the uni-directional filtering (UDF) that is similar to H.264 directional intra prediction. In this approach, only samples from neighboring blocks can be used in prediction. Another is bidirectional filtering (BDF) that exploits the correlations among samples from not only neighboring blocks but also the current block. The prediction structure in this approach is hierarchical multi-layer. In this paper, we present mathematical analyses on UDF and BDF and show the advantage to combine them together. The proposed dFT is integrated into H.264 intra-frame coding too. The preliminary experimental results in H.264 demonstrate its superiority. Xiulian Peng, Feng Wu 0001, Jizheng Xu |
ICME | 1 |