Chen Chen 0015

dblp:65/4423-15 · DBLP profile ↗
← Back
30ranked-venue papers
6as first author
17since 2021 · last 2026
0000-0003-3498-2527ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 23 · 6 first-author · 10 since 2021Artificial intelligence and machine learning · 15 · 1 first-author · 14 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 X2Edit: Revisiting Arbitrary-Instruction Image Editing Through Self-Constructed Data and Task-Aware Representation Learning
abstract
Existing open-source datasets for arbitrary-instruction image editing remain suboptimal, while a plug-and-play editing module compatible with community-prevalent generative models is notably absent. In this paper, we first introduce the X2Edit Dataset, a comprehensive dataset covering 14 diverse editing tasks, including subject-driven generation. We utilize the industry-leading unified image generation models and expert models to construct the data. Meanwhile, we design reasonable editing instructions with the VLM and implement various scoring mechanisms to filter the data. As a result, we construct 3.7 million high-quality data with balanced categories. Second, to better integrate seamlessly with community image generation models, we design task-aware MoE-LoRA training based on FLUX.1, with only 8% of the parameters of the full model. To further improve the final performance, we utilize the internal representations of the diffusion model and define positive/negative samples based on image editing types to introduce contrastive learning. Extensive experiments demonstrate that the model's editing performance is competitive among many excellent models. Additionally, the constructed dataset exhibits substantial advantages over existing open-source datasets.
Jian Ma 0010, Xujie Zhu, Qirong Peng, Chen Chen 0015, Haonan Lu
AAAI6
2025 SCott: Accelerating Diffusion Models with Stochastic Consistency Distillation
abstract
The iterative sampling procedure employed by diffusion models (DMs) often leads to significant latency. To address this, we propose Stochastic Consistency Distillation (SCott) to enable accelerated text-to-image generation, where high-quality generations can be achieved with just 2-4 sampling steps or even1 step, and further improvements can be obtained by additional cost, e.g., 4 steps. In contrast to vanilla consistency distillation (CD) which distills the ordinary differential equation solvers-based sampling process of a pre-trained teacher model into a student, SCott explores the possibility and validates the efficacy of integrating stochastic differential equation (SDE) solvers into CD to fully unleash the potential of the teacher. SCott is augmented with elaborate strategies to control the noise strength and sampling process of the SDE solver. An adversarial loss is further incorporated to strengthen the sample quality with rare sampling steps. Empirically, on the MSCOCO-2017 5K dataset with a Stable Diffusion-V1.5 teacher, SCott achieves an FID of 21.9, surpassing that of the 1-step InstaFlow (23.4) and the 4-step UFOGen (22.1). Moreover, SCott can yield more diverse samples than other consistency models for high-resolution image generation, with up to 16% improvement in a qualified metric.
Hongjian Liu, Qingsong Xie, Tianxiang Ye, Zhijie Deng, Chen Chen 0015, Shixiang Tang, Xueyang Fu, Haonan Lu, Zhengjun Zha
AAAI5
2025 GlyphDraw2: Automatic Generation of Complex Glyph Posters with Diffusion Models and Large Language Models
abstract
Posters serve an essential function in marketing and advertising by improving visual communication and brand visibility, thus significantly contributing to industrial design. With the latest developments in controllable T2I diffusion models, research interest has surged in text rendering within synthesized images. Although text rendering accuracy has seen advancements, automatic poster generation remains a relatively untapped area. This paper presents an automatic poster generation framework featuring text rendering capabilities through the use of LLMs. Our framework employs a triple-cross attention mechanism based on alignment learning to achieve precise text placement within detailed contextual backgrounds. Moreover, it supports adjustable fonts, varying image resolutions, and poster rendering with textual prompts in both English and Chinese. Additionally, we present a comprehensive bilingual image-text dataset, GlyphDraw-3M, comprising 3 million image-text pairs, each with OCR annotations and resolutions exceeding 1024. Our method utilizes the SDXL architecture, and extensive experiments confirm its ability to generate posters with intricate and context-rich backgrounds.
Jian Ma 0010, Yonglin Deng, Chen Chen 0015, Nanyang Du, Haonan Lu
AAAI3
2025 X2i: Seamless Integration of Multimodal Understanding Into Diffusion Transformer Via Attention Distillation
Jian Ma 0010, Qirong Peng, Chen Chen 0015, Haonan Lu
ICCV4
2025 Learning Mutual Excitation for Hand-to-Hand and Human-to-Human Interaction Recognition
abstract
Recognizing interactive actions, including hand-to-hand interaction and human-to-human interaction, has attracted increasing attention for various applications in the field of video analysis and human–robot interaction. Considering the success of graph convolution in modeling topology-aware features from skeleton data, recent methods commonly operate graph convolution on separate entities and use late fusion for interactive action recognition, which can barely model the mutual semantic relationships between pairwise entities. To this end, we propose a mutual excitation graph convolutional network (me-GCN) by stacking mutual excitation graph convolution (me-GC) layers. Specifically, me-GC uses a mutual topology excitation module to firstly extract adjacency matrices from individual entities and then adaptively model the mutual constraints between them. Moreover, me-GC extends the above idea and further uses a mutual feature excitation module to extract and merge deep features from pairwise entities. Compared with graph convolution, our proposed me-GC gradually learns mutual information in each layer and each stage of graph convolution operations. Extensive experiments on a challenging hand-to-hand interaction dataset, i.e., the Assembely101 dataset, and two large-scale human-to-human interaction datasets, i.e., NTU60-Interaction and NTU120-Interaction consistently verify the superiority of our proposed method, which outperforms the state-of-the-art GCN-based and Transformer-based methods.
Mengyuan Liu 0001, Chen Chen 0015, Songtao Wu, Fanyang Meng, Hong Liu 0008
IEEE Trans. Hum. Mach. Syst.2
2024 Compositional Text-to-Image Synthesis with Attention Map Control of Diffusion Models
abstract
Recent text-to-image (T2I) diffusion models show outstanding performance in generating high-quality images conditioned on textual prompts. However, they fail to semantically align the generated images with the prompts due to their limited compositional capabilities, leading to attribute leakage, entity leakage, and missing entities. In this paper, we propose a novel attention mask control strategy based on predicted object boxes to address these issues. In particular, we first train a BoxNet to predict a box for each entity that possesses the attribute specified in the prompt. Then, depending on the predicted boxes, a unique mask control is applied to the cross- and self-attention maps. Our approach produces a more semantically accurate synthesis by constraining the attention regions of each token in the prompt to the image. In addition, the proposed method is straightforward and effective and can be readily integrated into existing cross-attention-based T2I generators. We compare our approach to competing methods and demonstrate that it can faithfully convey the semantics of the original text to the generated content and achieve high availability as a ready-to-use plugin. Please refer to https://github.com/OPPO-Mente-Lab/attention-mask-control.
Zekang Chen, Chen Chen 0015, Jian Ma 0010, Haonan Lu, Xiaodong Lin 0004
AAAI3
2024 A Dual-Augmentor Framework for Domain Generalization in 3D Human Pose Estimation
abstract
3D human pose data collected in controlled laboratory settings present challenges for pose estimators that generalize across diverse scenarios. To address this, domain generalization is employed. Current methodologies in do-main generalization for 3D human pose estimation typically utilize adversarial training to generate synthetic poses for training. Nonetheless, these approaches exhibit several limitations. First, the lack of prior information about the target domain complicates the application of suitable augmentation through a single pose augmentor, affecting generalization on target domains. Moreover, adversarial training's discriminator tends to enforce similarity between source and synthesized poses, impeding the exploration of out-of-source distributions. Furthermore, the pose estimator's op-timization is not exposed to domain shifts, limiting its over-all generalization ability. To address these limitations, we propose a novel frame-work featuring two pose augmentors: the weak and the strong augmentors. Our framework employs differential strategies for generation and discrimination processes, facilitating the preservation of knowledge related to source poses and the exploration of out-of-source distributions without prior information about target poses. Besides, we leverage meta-optimization to simulate domain shifts in the optimization process of the pose estimator, thereby improving its generalization ability. Our proposed approach significantly outperforms existing methods, as demonstrated through comprehensive experiments on various benchmark datasets. Our code will be released at https://github.com/davidpengucf/DAF-DG.
Qucheng Peng, Chen Chen 0015
CVPR3
2024 PEA-Diffusion: Parameter-Efficient Adapter with Knowledge Distillation in Non-english Text-to-Image Generation
Jian Ma 0010, Chen Chen 0015, Qingsong Xie, Haonan Lu
ECCV (68)2
2024 InsCL: A Data-efficient Continual Learning Paradigm for Fine-tuning Large Language Models with Instructions
abstract
Yifan Wang, Yafei Liu, Chufan Shi, Haoling Li, Chen Chen, Haonan Lu, Yujiu Yang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Chufan Shi, Haoling Li, Chen Chen 0015, Haonan Lu, Yujiu Yang 0001
NAACL-HLT5
2024 Delving into Identify-Emphasize Paradigm for Combating Unknown Bias
Bowen Zhao 0003, Chen Chen 0015, Qian-Wei Wang, Anfeng He, Shutao Xia
Int. J. Comput. Vis.2
2024 PatchNet: Maximize the Exploration of Congeneric Semantics for Weakly Supervised Semantic Segmentation
abstract
With the increase in the number of image data and the lack of corresponding labels, weakly supervised learning has drawn a lot of attention recently in computer vision tasks, especially in the fine-grained semantic segmentation problem. To alleviate human efforts from expensive pixel-by-pixel annotations, our method focuses on weakly supervised semantic segmentation (WSSS) with image-level labels, which are much easier to obtain. As a considerable gap exists between pixel-level segmentation and image-level labels, how to reflect the image-level semantic information on each pixel is an important question. To explore the congeneric semantic regions from the same class to the maximum, we construct the patch-level semantic augmentation network (PatchNet) based on the self-detected patches from different images that contain the same class labels. Patches can frame the objects as much as possible and include as little background as possible. The patch-level semantic augmentation network that is established with patches as the nodes can maximize the mutual learning of similar objects. We regard the embedding vectors of patches as nodes and use a transformer-based complementary learning module to construct weighted edges according to the embedding similarity between different nodes. Moreover, to better supplement semantic information, we propose softcomplementary loss functions matched with the whole network structure. We conduct experiments on the popular PASCAL VOC 2012 and MS COCO 2014 benchmarks, and our model yields the state-of-the-art performance.
Ke Zhang 0046, Chen Chen 0015, Chun Yuan 0003, Xinfeng Wang
IEEE Trans. Neural Networks Learn. Syst.2
2024 Dream360: Diverse and Immersive Outdoor Virtual Scene Creation via Transformer-Based 360° Image Outpainting
abstract
360° images, with a field-of-view (FoV) of $180^{\circ}\times 360^{\circ}$, provide immersive and realistic environments for emerging virtual reality (VR) applications, such as virtual tourism, where users desire to create diverse panoramic scenes from a narrow FoV photo they take from a viewpoint via portable devices. It thus brings us to a technical challenge: 'How to allow the users to freely create diverse and immersive virtual scenes from a narrow FoV image with a specified viewport?' To this end, we propose a transformer-based 360° image outpainting framework called Dream360, which can generate diverse, high-fidelity, and high-resolution panoramas from user-selected viewports, considering the spherical properties of 360° images. Compared with existing methods, e.g., [3], which primarily focus on inputs with rectangular masks and central locations while overlooking the spherical property of 360° images, our Dream360 offers higher outpainting flexibility and fidelity based on the spherical representation. Dream360 comprises two key learning stages: (I) codebook-based panorama outpainting via Spherical-VQGAN (S-VQGAN), and (II) frequency-aware refinement with a novel frequency-aware consistency loss. Specifically, S-VQGAN learns a sphere-specific codebook from spherical harmonic (SH) values, providing a better representation of spherical data distribution for scene modeling. The frequency-aware refinement matches the resolution and further improves the semantic consistency and visual fidelity of the generated results. Our Dream360 achieves significantly lower Frechet Inception Distance (FID) scores and better visual fidelity than existing methods. We also conducted a user study involving 15 participants to interactively evaluate the quality of the generated results in VR, demonstrating the flexibility and superiority of our Dream360 framework.
Hao Ai, Zidong Cao, Haonan Lu, Chen Chen 0015, Jian Ma 0010, Peng Yuan Zhou, Tae-Kyun Kim 0001, Pan Hui 0001, Lin Wang 0025
IEEE Trans. Vis. Comput. Graph.4
2023 Combating Unknown Bias with Effective Bias-Conflicting Scoring and Gradient Alignment
abstract
Models notoriously suffer from dataset biases which are detrimental to robustness and generalization. The identify-emphasize paradigm shows a promising effect in dealing with unknown biases. However, we find that it is still plagued by two challenges: A, the quality of the identified bias-conflicting samples is far from satisfactory; B, the emphasizing strategies just yield suboptimal performance. In this work, for challenge A, we propose an effective bias-conflicting scoring method to boost the identification accuracy with two practical strategies --- peer-picking and epoch-ensemble. For challenge B, we point out that the gradient contribution statistics can be a reliable indicator to inspect whether the optimization is dominated by bias-aligned samples. Then, we propose gradient alignment, which employs gradient statistics to balance the contributions of the mined bias-aligned and bias-conflicting samples dynamically throughout the learning process, forcing models to leverage intrinsic features to make fair decisions. Experiments are conducted on multiple datasets in various settings, demonstrating that the proposed solution can alleviate the impact of unknown biases and achieve state-of-the-art performance.
Bowen Zhao 0003, Chen Chen 0015, Qian-Wei Wang, Anfeng He, Shutao Xia
AAAI2
2023 Delta: Degradation-Free Fully Test-Time Adaptation
Bowen Zhao 0003, Chen Chen 0015, Shutao Xia
ICLR2
2023 A Single 2D Pose with Context is Worth Hundreds for 3D Human Pose Estimation
abstract
The dominant paradigm in 3D human pose estimation that lifts a 2D pose sequence to 3D heavily relies on long-term temporal clues (i.e., using a daunting number of video frames) for improved accuracy, which incurs performance saturation, intractable computation and the non-causal problem. This can be attributed to their inherent inability to perceive spatial context as plain 2D joint coordinates carry no visual cues. To address this issue, we propose a straightforward yet powerful solution: leveraging the $\textit{readily available}$ intermediate visual representations produced by off-the-shelf (pre-trained) 2D pose detectors -- no finetuning on the 3D task is even needed. The key observation is that, while the pose detector learns to localize 2D joints, such representations (e.g., feature maps) implicitly encode the joint-centric spatial context thanks to the regional operations in backbone networks. We design a simple baseline named $\textbf{Context-Aware PoseFormer}$ to showcase its effectiveness. $\textit{Without access to any temporal information}$, the proposed method significantly outperforms its context-agnostic counterpart, PoseFormer, and other state-of-the-art methods using up to $\textit{hundreds of}$ video frames regarding both speed and precision. $\textit{Project page:}$ https://qitaozhao.github.io/ContextAware-PoseFormer
Qitao Zhao, Mengyuan Liu 0001, Chen Chen 0015
NeurIPS4
2022 Energy Alignment for Bias Rectification in Class Incremental Learning
abstract
In class incremental learning (CIL), models are expected to be able to learn new categories continuously. However, the standard DNNs suffer from catastrophic forgetting. Recent studies show class imbalance is an essential factor that causes catastrophic forgetting in CIL. In this paper, from the perspective of energy-based model, we demonstrate that the free energies of categories are aligned with the label distribution theoretically, thus the energies of different classes are expected to be close to each other when aiming for "balanced" performance. However, we discover a severe energy-bias phenomenon in the models trained in CIL. To eliminate the bias, we propose a simple and effective method named Energy Alignment by merely adding the calculated shift scalars onto the output logits, which does not require to (i) modify the network architectures, (ii) intervene the standard learning paradigm. Experimental results show that energy alignment can achieve good performance on several CIL benchmarks.
Bowen Zhao 0003, Chen Chen 0015, Xi Xiao 0001, Qi Ju 0002, Shutao Xia
ICASSP2
2022 Towards a category-extended object detector with limited data
Bowen Zhao 0003, Chen Chen 0015, Xi Xiao 0001, Shutao Xia
Pattern Recognit.2
2019 Improving Image Captioning with Conditional Generative Adversarial Nets
abstract
In this paper, we propose a novel conditional-generativeadversarial-nets-based image captioning framework as an extension of traditional reinforcement-learning (RL)-based encoder-decoder architecture. To deal with the inconsistent evaluation problem among different objective language metrics, we are motivated to design some “discriminator” networks to automatically and progressively determine whether generated caption is human described or machine generated. Two kinds of discriminator architectures (CNN and RNNbased structures) are introduced since each has its own advantages. The proposed algorithm is generic so that it can enhance any existing RL-based image captioning framework and we show that the conventional RL training method is just a special case of our approach. Empirically, we show consistent improvements over all language evaluation metrics for different state-of-the-art image captioning models. In addition, the well-trained discriminators can also be viewed as objective image captioning evaluators.
Chen Chen 0015, Wanpeng Xiao, Zexiong Ye, Liesi Wu, Qi Ju 0002
AAAI1
2019 High-Quality Color Image Compression by Quantization Crossing Color Spaces
abstract
Coding of a color image usually happens in the YCbCr space so that the rate-distortion optimization is conducted in this space. Due to the use of a non-unitary matrix in the RGB-to-YCbCr conversion, an optimal coding performance achieved in the YCbCr space does not guarantee an optimal quality in the RGB space, which would impact most display devices that need RGB signals as the inputs. In this paper, we first study the relationship between the coding distortions of the compressed RGB signals and the quantization errors occurred in the coded YCbCr signals. Then, we design a new quantization scheme crossing the RGB and YCbCr spaces to achieve a high-quality color image compression with the YCbCr 4:4:4 format. Although our proposed quantization takes place in the YCbCr space, it aims at reducing the coding distortion in the RGB space as much as possible. Experimental results demonstrate that our proposed method offers a significant quality gain over the existing block-based coding methods for various images.
Shuyuan Zhu, Zhiying He, Chen Chen 0015, Shuaicheng Liu, Jiantao Zhou 0001, Yuanfang Guo, Bing Zeng 0001
IEEE Trans. Circuits Syst. Video Technol.3
2018 A New HEVC In-Loop Filter Based on Multi-channel Long-Short-Term Dependency Residual Networks
abstract
In this paper, we propose a new HEVC in-loop filter based on a multi-channel long-short-term dependency residual network (MLSDRN). Inspired by the information storage and information update function of human memory cell, our MLSDRN introduces an update cell to adaptively store and select the long-term and short-term dependency information through an adaptive learning process. In addition, we leverage the block boundary information that recorded in the bit-streams to improve the filter performance, which also makes our MLSDRN to unequally treat the video content. Meanwhile, the multi-channel is introduced to solve the illumination discrepancy problem. We integrate the novel in-loop filter into HM reference software, and applying it to luma and chroma components, simulation results demonstrate that the proposed in-loop filter can save BD-rate reduction up to 15.9% with ALF off. For luma component, the novel in-loop filter achieves 6.0%, 8.1%, 7.4% BD-rate saving for all intra, low delay and random access configurations, respectively.
Xiandong Meng, Chen Chen 0015, Shuyuan Zhu, Bing Zeng 0001
DCC2
2018 DC Coefficient Estimation of Intra-Predicted Residuals in HEVC
abstract
This paper presents a DC coefficient estimation algorithm for intra-predicted residual blocks in the High-Efficiency Video Coding (HEVC) standard. Discarding the DC coefficient directly in each transform block leads to substantial bit-saving, but at the same time produces strong discontinuities between neighboring blocks. To overcome this problem, we propose an estimation algorithm for the DC coefficient, which solves an optimal offset in a closed-form to recover the corresponding block edges. Then, we embed this algorithm into HEVC in its rate-distortion optimized quantization and sign bit hiding steps. Furthermore, a flag is signaled to decide whether the DC estimation strategy is used for each transform block. Test results under the common test condition show that our algorithm achieves 1.5% and 1.6% BD-rate reduction on average for luma and chroma, respectively, under all intra configuration. In the meantime, our simulation results show that both encoding time and decoding time increase only slightly (about 10%, without any special optimization on programming our proposed algorithm). When testing the proposed DC estimation algorithm on inter coding configurations, including low delay with P pictures, low delay with B pictures, and random access, we can also achieve 0.5%-1.1% bit-rate savings on average, while nearly no extra encoding and decoding time is needed.
Chen Chen 0015, Zexiang Miao, Xiandong Meng, Shuyuan Zhu, Bing Zeng 0001
IEEE Trans. Circuits Syst. Video Technol.1
2018 Cross-Space Distortion Directed Color Image Compression
abstract
Traditional color image compression is usually conducted in the YCbCr space but many color displayers only accept RGB signals as inputs. Due to the use of a non-unitary matrix in the YCbCr-RGB conversion, low distortion achieved in the YCbCr space cannot guarantee low distortion for the RGB signals. To solve this problem, we propose a novel compression scheme for color images through defining a cross-space distortion so as to reduce as much as possible the distortion in the RGB space. To this end, we first derive the relationship between the distortions in the YCbCr space and RGB space. Then, we develop two solutions to implement color image compression for the most popular 4:2:0 chroma format. The first solution focuses on the design of a new spatial downsampling method to generate the 4:2:0 YCbCr image for a high-efficiency compression. The second one provides a novel way to reduce the distortion of the compressed color image by controlling the quantization error of the 4:2:0 YCbCr image, especially the one generated by using the traditional spatial downsampling. Experimental results show that both proposed solutions offer a remarkable quality gain over some state-of-the-art approaches when tested on various textured color images.
Shuyuan Zhu, Chen Chen 0015, Shuaicheng Liu, Bing Zeng 0001
IEEE Trans. Multim.3
2017 A New Block-Based Method for HEVC Intra Coding
abstract
This paper presents a new block-based method for the High Efficiency Video Coding (HEVC) intra coding. First, we have found through analysis and test that the prediction errors on some pixels in each prediction block (PB) that are neighboring to the reference pixels would be no bigger than the corresponding coding errors. Based on this observation, the pixels in each PB are divided into two parts: half pixels are coded via a novel padding technique together with a constrained quantization algorithm (leading to around 3 dB gain under the same bit rate), whereas the other half are reconstructed by linear interpolations along a prediction direction by utilizing the neighboring reference pixels and the first half coded pixels. In the final implementation, a competition mechanism is employed between this new method and the original HEVC intra coding in order to choose the best mode for each PB. Experimental results show that about 2% BD-rate reduction has been achieved both for luma and chroma with respect to the original HEVC intra coding, whereas the encoder complexity increases by 130%, but the decoding time remains nearly unchanged.
Chen Chen 0015, Shuyuan Zhu, Bing Zeng 0001, Moncef Gabbouj
IEEE Trans. Circuits Syst. Video Technol.1
2017 A Hierarchical Approach for Rain or Snow Removing in a Single Color Image
abstract
In this paper, we propose an efficient algorithm to remove rain or snow from a single color image. Our algorithm takes advantage of two popular techniques employed in image processing, namely, image decomposition and dictionary learning. At first, a combination of rain/snow detection and a guided filter is used to decompose the input image into a complementary pair: 1) the low-frequency part that is free of rain or snow almost completely and 2) the high-frequency part that contains not only the rain/snow component but also some or even many details of the image. Then, we focus on the extraction of image's details from the high-frequency part. To this end, we design a 3-layer hierarchical scheme. In the first layer, an overcomplete dictionary is trained and three classifications are carried out to classify the high-frequency part into rain/snow and non-rain/snow components in which some common characteristics of rain/snow have been utilized. In the second layer, another combination of rain/snow detection and guided filtering is performed on the rain/snow component obtained in the first layer. In the third layer, the sensitivity of variance across color channels is computed to enhance the visual quality of rain/snow-removed image. The effectiveness of our algorithm is verified through both subjective (the visual quality) and objective (through rendering rain/snow on some ground-truth images) approaches, which shows a superiority over several state-of-the-art works.
Yinglong Wang 0002, Shuaicheng Liu, Chen Chen 0015, Bing Zeng 0001
IEEE Trans. Image Process.3
2016 Low bit-rate intra coding scheme based on constrained quantization and median-type filter
abstract
This paper presents a new intra coding scheme for low bit-rate video compression. We first propose an improved codec architecture based on HEVC encoder. Then we divide pixels in each prediction block into two parts: three quarters of pixels are coded via a smart padding technique together with a constrained quantization algorithm (leading to a significantly improved quality); whereas the other quarter are reconstructed according to median-type filtering by utilizing the 8-neighboring reference samples after all blocks have been encoded and reconstructed. Experimental results show that about 3% BD-rate reduction has been achieved both for luma and chroma components without apparent increase of encoding complexity with respect to the original HEVC intra coding.
Chen Chen 0015, Bing Zeng 0001
ICASSP1
2016 A framework of single-image deraining method based on analysis of rain characteristics
abstract
In this paper, we propose an algorithm to remove rain streaks from single color image. Firstly, the guided filter, cooperated with rain pixels detection are used to separate a color image into low-frequency and high-frequency parts so that most rain components exist in the high-frequency part. Then, we focus on the high-frequency part to extract the non-rain details according to the characteristics of the rain in which a dictionary learning method is used. Meanwhile, to enhance the quality of the rain-removed image, the proposed principal direction of an image patch (PDIP) and the sensitivity of variance of color channels (SVCC) are employed in our work to help extract more non-rain details. Compared with the state-of-the-art works, our proposed method can remove the rain (especially heavy rain) from color images more efficiently.
Yinglong Wang 0002, Chen Chen 0015, Shuyuan Zhu, Bing Zeng 0001
ICIP2
2016 DC coefficient estimation of intra-predicted residuals in high efficiency video coding
abstract
This paper proposes a DC coefficient estimation algorithm for intra-predicted residual blocks in the High Efficiency Video Coding (HEVC) standard. Discarding the DC coefficient in the current coding block leads to a substantial bit-saving but produces at the same time strong discontinuities between this block and its neighboring reconstructed blocks. To overcome this problem, we propose an estimation algorithm for the DC coefficient, which solves an optimal offset in a closed-form in the pixel domain to recover the corresponding block edges. Test results show that our algorithm achieves 1.0% and 1.4% BD-rate reduction on average for luma and chroma as compared with HM-16.6, respectively, when the sign-bit-hiding (SBH) technique is disabled. When SDH is set on, namely under the common test condition (CTC), the BD-rate reduction drops slightly to 0.7% and 1.1% for luma and chroma, respectively. In the meantime, the test results show that both encoding time and decoding time increase only slightly (about 10%, without any special optimization on programming our proposed algorithm).
Chen Chen 0015, Zexiang Miao, Xiandong Meng, Shuyuan Zhu, Bing Zeng 0001
VCIP1
2016 Mode-dependent transforms based on elliptical model for high efficiency video coding
abstract
High efficiency video coding (HEVC) defines 35 prediction modes in its intra prediction stage to signal the direction information of residual blocks. Traditionally, separable two-dimension (2-D) transforms (integer DCT and DST) are utilized in a similar manner as in the previous H.264/AVC standards. However, such 2-D transforms cannot yield the best energy compaction for a 2-D directional source where the dominating directional information is other than the horizontal or vertical one. In order to overcome this drawback, we build an elliptical model with directionality and design some non-separable transforms based on the Karhunen-Loeve transform in this paper. Specifically, we derive a non-separable transform in closed-form for each intra-prediction mode and replace the default transform in HEVC. Simulation results reveal that 1.7% and 2.0% on average and up to 7.7% and 8.1% BD-rate reduction can be achieved for luma and chroma component, respectively. In the meantime, the test results show that both the encoding time and decoding time increase only about 5%.
Kaiyuan Jia, Chen Chen 0015, Xiandong Meng, Shuyuan Zhu, Bing Zeng 0001
VCIP2
2016 Interpolation-directed transform domain downward conversion for block-based image compression
abstract
In this paper, we design an interpolation-directed transform domain downward conversion (ITDDC) to build up a new block-based image compression scheme. This ITDDC is derived from our proposed 2-D padding and performed on each 16×16 macro-block of pixels to convert it into an 8×8 coefficient block, leading to a downward image conversion in the transform domain. More interestingly, the further compression is just performed on the down-sized coefficient block and the reconstruction for an entire macro-block is achieved via the interpolation by using the decoded pixels only locating in some specific positions of it. To make the interpolation more efficient, the pixels participating in the interpolation will be optimized before the compression. The ITDDC-based coding is used competitively with the JPEG baseline coding to compress each macro-block in our proposed compression scheme according to a simple but efficient rate-distortion optimization based criterion. Experimental results demonstrate that our proposed method gets a remarkable quality gain over the existing approaches.
Shuyuan Zhu, Jinglin Yu, Chen Chen 0015, Liaoyuan Zeng, Bing Zeng 0001
VCIP4
2015 Adaptive guided image filter for improved in-loop filtering in video coding
abstract
This paper proposes a new adaptive sharpening filter based on guided image filter and improves HEVC's in-loop filter architecture by embedding sharpening filter between deblocking filter and SAO. The proposed algorithm classifies pixels of a frame into several groups according to uniform quantization of each pixel's Sum-Modified-Laplacian value and assigns identical optimal filtering parameters to the pixels belonging to the same group based on rate-distortion optimization. Simulation results show that our proposed algorithm achieves 0.7% on average and up to 8% BD-rate reduction with respect to the original HEVC in-loop filtering method. Encoding time increases slightly by about 15% and decoding time increases by 70% on average without special optimization of C++ program integrated in HM-16.5.
Chen Chen 0015, Zexiang Miao, Bing Zeng 0001
MMSP1