VLDB 2026 Research / reviewers in the wild / expert
Jiaying Liu 0001
dblp:32/197
· DBLP profile ↗
247ranked-venue papers
22as first author
85since 2021 · last 2026
0000-0002-0468-9576ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 196 · 20 first-author · 55 since 2021Artificial intelligence and machine learning · 78 · 2 first-author · 40 since 2021Systems, architecture and hardware · 17 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 9 · 2 first-author · 1 since 2021Computer networks · 7 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Self-Supervised Skeleton-Based Action Representation Learning: A Benchmark and Beyond
Jiahang Zhang 0001, Lilang Lin, Shuai Yang 0001, Jiaying Liu 0001 |
Int. J. Comput. Vis. | 4 |
| 2026 | QuadPrior++: Multi-Dimension Augmented Physical Prior for Zero-Reference Illumination EnhancementabstractExisting low-light enhancement methods typically rely on fitting data mappings (pixel-wise mappings through fully supervised methods or distribution-wise mappings through weakly supervised or self-supervised methods). However, their performance is heavily dependent on specific scenes and fails to adequately model the intrinsic prior of natural images, resulting in poor generalization. To tackle this challenge, we leverage the strengths of powerful generative diffusion models, conditioned on a thoughtfully designed prior, and propose a novel zero-reference low-light enhancement framework that gets rid of dependence on the distribution of low-light images. In detail, we address the most fundamental core by proposing an illumination-invariant prior derived from the theory of physical light transfer, bridging the gap between normal and low-light domains, and enabling zero-shot enhancement without the need for low-light-specific training. A prior-to-image restoration framework is built upon generative diffusion models, pre-trained on normal-light data. During inference, the framework extracts the illumination-invariant prior from low-light inputs and maps them back to high-quality images, naturally for low-light enhancement. Additionally, such intrinsic properties of illumination-invariant prior open up opportunities for distilling diffusion models into compact CNN-based networks. We propose a novel prior-injected distillation paradigm incorporating intensity, frequency, and gradient domain-augmented regularization comprehensively. This distillation framework not only reduces computational costs but also maintains high fidelity and perceptual quality in enhanced outputs, making it more efficient and practical for real-world applications. The approach further extends seamlessly to handle over-exposure scenarios, demonstrating its versatility in addressing complex lighting conditions. Extensive experiments demonstrate the superiority of our framework in various scenarios, as well as its strong interpretability, robustness, and efficiency. Haofeng Huang, Wenjing Wang 0001, Wenhan Yang, Ling-Yu Duan, Jiaying Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2026 | Syntax-Driven Multi-Realism Image Compression With Consistency Guided Diffusion ModelabstractGiven the challenge of balancing high fidelity with perceptual quality, multi-realism image compression is developed to adapt flexibly to varying requirements. It allows images with different levels of realism to be decoded from the same bit stream. Diffusion models are known for generating images with high perceptual quality. However, their inherent process of adding noise and denoising is often difficult to control and will bring more distortion. This limits their direct application in image compression, especially in multi-realism image compression which requires precise control to adapt to different requirements. To address this issue, we propose aConsistency Guided Diffusion Modelas a post-processing network for multi-realism image compression, aiming to control the addition of detail representations, thereby adjusting the trade-off between subjective quality and fidelity. In detail, our proposed novel method is crafted to introduce an additional consistency guided feature branch into the diffusion model to constrain the deviation caused by randomness in the diffusion process to ensure fidelity. Furthermore, a syntax-driven feature fusion module is constructed to guide the information adaptive fusion of two branches with an input extra ultra-low stream, which contains the context information and trade-off control information. In addition, we design a warm-up based training strategy and adopt a continuous online optimization method to improve coding efficiency and trade-off control precision. Extensive experiments validate the superiority of our method over existing compression techniques, as well as the effectiveness of each component. Haowei Kuang, Wenhan Yang, Zongming Guo, Jiaying Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Prompting Rain Off: Evolving Compact Dual Prompts for Continual De-RainingabstractIn recent years, there has been notable progress in single-image rain removal, particularly focusing on static data distributions in these approaches. When dealing with data that constantly changes, the challenge of catastrophic forgetting arises, which is quite common and critical in real-world scenarios. To address this, we propose Evolving COmpact Dual Prompt Learning (EcoDPL), an efficient rehearsal-free continual learning deraining framework designed specifically for low-level vision tasks. Specifically, we design two prompt pools at both image and feature levels and insert these prompts into images and embedding tokens, for better knowledge transfer across tasks. Our adaptive weight generation module, P-Fuser, attaches an attention map to each prompt, to adaptively pay attention to different inputs, and get different weights to fuse prompts, making the inserted prompts more flexible with various inputs. Also, we introduce Grad-Tuner, a dictionary learning strategy, to compress knowledge into fewer prompts. This makes the knowledge more compact and provides more space for new prompts to learn new tasks. Our method stands out by leveraging small, learnable prompts for efficient knowledge retention across tasks, not increasing training time or parameters. Furthermore, we present an augmented method that upgrades the distance function $\gamma $ from simple cosine distance to a more advanced weight generation network. We also employ a fine-tuned dictionary learning technique, compressing knowledge into a more compact form, and enhancing the ability of prompts to learn new tasks. With our new designs, the model becomes more flexible with various inputs and it compresses knowledge into fewer prompts to free up spaces to learn new tasks. Through extensive experiments on various rain removal datasets, our EcoDPL method consistently outperforms previous continual learning techniques. Notably, although EcoDPL is designed for continual learning with changing data, it also performs well with stationary data, proving its robustness and versatility. Our website is available at: https://starymoon.github.io/Prompting-Rain-Off. Minghao Liu 0019, Wenhan Yang, Jiaying Liu 0001 |
IEEE Trans. Image Process. | 3 |
| 2026 | Local Dimension Enhancement Representation Learning for Skeleton-Based Action SegmentationabstractMost existing self-supervised learning methods for skeleton-based temporal action segmentation (TAS) fail to capture the short-term motion semantics essential for dense frame-level prediction, as they typically learn representations that are either too coarse or motion-insensitive. This issue is reflected in local dimension collapse, which highlights the limitations of current approaches and suggests directions for improvement. Specifically, to address the issue of local dimension collapse for self-supervised learning in TAS, we propose the Local Dimension Enhancement (LoDE) framework, which introduces the local effective rank (LER) as a metric to measure and a learning objective to reduce this collapse. A new fine-grained representation scale, termed a motion unit, is defined as a temporal clip of consecutive skeleton frames to model skeleton data. Centered on this representation scale, we analyze existing methods (sequence-scale and frame-scale learning) with the tool of LER and theoretically demonstrate that introducing motion unit-scale learning is essential to alleviate local dimension collapse. Inspired by our theoretical insights, we design a multi-scale semantics module that integrates frame-, sequence-, and motion unit-scale learning, with LER-based regularization to enrich local representation diversity. These designs effectively alleviate local dimension collapse and lead to significant improvements in TAS, as evidenced by LoDE's superior performance over state-of-the-art methods on three large-scale untrimmed datasets: PKUMMD, TSU, and BABEL. Our project website is available at https://carefreesun.github.io/LoDE_TIP_2026/. Shaofan Sun, Lilang Lin, Jiahang Zhang 0001, Ling-Yu Duan, Jiaying Liu 0001 |
IEEE Trans. Image Process. | 5 |
| 2026 | Seeing in the Dark with Ambient GuidanceabstractA low-light image taken in a dark scene usually suffers from severe distortions, which does not accurately characterize the ambient lighting. Long exposure is an accustomed way to capture more supplementary light and alleviate the degradation, but sometimes it induces other distortions, e.g., blurriness. To address this issue, we propose a new paradigm that introduces additional captured ambient guidance, i.e., a long-exposure image to steer the low-light enhancement. In practice, this long-exposure image can be obtained conveniently, but usually suffers from blurriness and misalignment. To effectively extract and fuse information from degraded and misaligned low-light and guidance image pairs, we propose a Long Exposure Compensation Network (LECNet). Adaptive Band Regression is introduced to disentangle the image into multi-scale representations and coarse-to-fine aggregate them with an attention mechanism. For stable image-guidance registration and artifact suppression, we propose a Bounded Cross-domain Deformable Alignment to warp the guidance based on extracted feature pyramids step by step. To integrate knowledge about the degradation into our LECNet for better fidelity, a dual learned back projection is enforced between the predicted result and the paired inputs in illumination and texture detail consistency, serving the model training for both offline training and online sample-adaptive finetuning. For training and evaluation of this new paradigm, we build a dataset with both synthetic and real-captured image triplets of long/short exposure pairs and extra blurry guidance. The experimental evaluation demonstrates the significance of our new paradigm, as well as the superiority of our LECNet and its usability in the real world. Haofeng Huang, Wenhan Yang, Mengnan Wang, Ling-Yu Duan, Jiaying Liu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2025 | UP-Restorer: When Unrolling Meets Prompts for Unified Image RestorationabstractAll-in-one restoration needs to implicitly distinguish between different degradation conditions and apply specific prior constraints accordingly. To fulfill this goal, our work makes the first effort to create an all-in-one restoration via unrolling from the typical maximum a-posterior optimization function. This unrolling framework naturally leads to the construction of progressively solving models, which are equivalent to a diffusion enhancer taking as input dynamically generated prompts. Under a score-based diffusion model, the prompts are integrated for propogating and updating several context-related variables, i.e. transmission map, atmospheric light map and noise or rain map progressively. Such learned prompt generation process, which simulates the nonlinear operations in the unrolled solution, is combined with linear operations owning clear physics implications to make the diffusion models well reguarlized and more effective in learning degradation-related visual priors. Experimental results demonstrate that our method achieves significant performance improvements across various image restoration tasks, realizing true all-in-one image restoration. Minghao Liu 0019, Wenhan Yang, Jinyi Luo, Jiaying Liu 0001 |
AAAI | 4 |
| 2025 | PTDiffusion: Free Lunch for Generating Optical Illusion Hidden Pictures with Phase-Transferred Diffusion ModelabstractOptical illusion hidden picture is an interesting visual perceptual phenomenon where an image is cleverly integrated into another picture. Established on the off-the-shelf text-to-image (T2I) diffusion model, we propose a novel text-guided image-to-image (I2I) translation framework dubbed as Phase-Transferred Diffusion Model (PTDiffusion) for hidden art syntheses, which harmoniously embeds an input reference image into arbitrary scenes described by the text prompts. At the heart of our method is a plug-and-play phase transfer mechanism that dynamically and progressively transplants diffusion features’ phase spectrum from the denoising process to reconstruct the reference image into the one to sample the generated illusion image, realizing deep fusion of the reference structural information and the textual semantic information. Furthermore, we propose asynchronous phase transfer to enable flexible control over the degree of hidden content discernability. Our method bypasses any model training and fine-tuning process, all while substantially outperforming related methods in image quality, text fidelity, visual discernibility, and contextual naturalness for illusion picture synthesis, as demonstrated by extensive qualitative and quantitative experiments. Our project is publically available at this web page. Xiang Gao 0014, Shuai Yang 0001, Jiaying Liu 0001 |
CVPR | 3 |
| 2025 | JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and GenerationabstractWe present JanusFlow, a powerful framework that unifies image understanding and generation in a single model. JanusFlow introduces a minimalist architecture that integrates autoregressive language models with rectified flow, a state-of-the-art method in generative modeling. Our key finding demonstrates that rectified flow can be straightforwardly trained within the large language model framework, eliminating the need for complex architectural modifications. To further improve the performance of our unified model, we adopt two key strategies: (i) decoupling the understanding and generation encoders, and (ii) aligning their representations during unified training. Extensive experiments show that JanusFlow achieves comparable or superior performance to specialized models in their respective domains, while significantly outperforming existing unified approaches. This work represents a step toward more efficient and versatile vision-language models. Yiyang Ma, Xingchao Liu, Xiaokang Chen, Wen Liu 0011, Chengyue Wu, Zhiyu Wu, Zizheng Pan, Zhenda Xie, Xingkai Yu, Liang Zhao 0026, Jiaying Liu 0001, Chong Ruan |
CVPR | 13 |
| 2025 | Cross-Granularity Online Optimization with Masked Compensated Information for Learned Image Compression
Haowei Kuang, Wenhan Yang, Zongming Guo, Jiaying Liu 0001 |
ICCV | 4 |
| 2025 | Splice, Focus and Relife: High-Resolution Periodic Pattern GenerationabstractThe printing and dyeing industry requires periodic and high-resolution patterns to ensure seamless designs on large fabric sections and high-quality final products. However, current manual approaches to pattern creation are time-consuming and labor-intensive. Leveraging powerful image generative models, such as Latent Diffusion Models (LDMs), offers a promising alternative, but challenges persist in generating strictly periodic and high-resolution patterns due to the inherent randomness and high computational demands of LDMs. In this paper, we propose a novel text-driven framework for generating periodic and high-resolution patterns. We introduce a new training-free Splice-and-Focus Mechanism, which enhances the model by constraining latent features and modifying the attention mechanism to produce natural and strictly periodic patterns. Additionally, we present a ReLife Pipeline, which integrates super-resolution and guided image synthesis to enhance pattern resolution while eliminating artifacts and distortions. Experimental results demonstrate that our framework produces patterns of superior quality. Xicheng Lan, Wenshuo Gao, Luyao Zhang 0007, Jiaying Liu 0001, Shuai Yang 0001 |
ISCAS | 4 |
| 2025 | SGAR: Structural Generative Augmentation for 3D Human Motion Retrievalabstract3D human motion-text retrieval is essential for accurate motion understanding, targeted at cross-modal alignment learning. Existing methods typically align the global motion-text concepts directly, suffering from sub-optimal generalization due to the uncertainty of correspondence learning between multiple motion concepts coupled in a single motion/text sequence. Therefore, we study the explicit fine-grained concept decomposition for alignment learning and present a novel framework, Structural Generative Augmentation for 3D Human Motion Retrieval (SGAR), to enable generation-augmented retrieval. Specifically, relying on the strong priors of existing large language model (LLM) assets, we effectively decompose human motions structurally into subtler semantic units, \ie, body parts, for fine-grained motion modeling. Based on this, we develop part-mixture learning to better decouple the local motion concept learning, boosting part-level alignment. Moreover, a directional relation alignment strategy exploiting the correspondence between full-body and part motions is incorporated to regularize feature manifold for better consistency. Extensive experiments on three benchmarks, including motion-text retrieval as well as recognition and generation applications, demonstrate the superior performance and promising transferability of our method. Jiahang Zhang 0001, Lilang Lin, Shuai Yang 0001, Jiaying Liu 0001 |
NeurIPS | 4 |
| 2025 | Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation
Wenjing Wang 0001, Huan Yang 0005, Zixi Tuo, Huiguo He, Junchen Zhu, Jianlong Fu, Jiaying Liu 0001 |
Int. J. Comput. Vis. | 7 |
| 2025 | Self-Supervised Skeleton Representation Learning Via Actionlet Contrast and ReconstructabstractContrastive learning has shown remarkable success in the domain of skeleton-based action recognition. However, the design of data transformations, which is crucial for effective contrastive learning, remains a challenging aspect in the context of skeleton-based action recognition. The difficulty lies in creating data transformations that capture rich motion patterns while ensuring that the transformed data retains the same semantic information. To tackle this challenge, we introduce an innovative framework called ActCLR+ (Actionlet-Dependent Contrastive Learning), which explicitly distinguishes between static and dynamic regions in a skeleton sequence. We begin by introducing the concept of actionlet, connecting self-supervised learning quantitatively with downstream tasks. Actionlets represent regions in the skeleton where features closely align with action prototypes, highlighting dynamic sequences as distinct from static ones. We propose an anchor-based method for unsupervised actionlet discovery, establishing a motion-adaptive data transformation approach based on this discovery. This motion-adaptive data transformation strategy tailors data transformations for actionlet and non-actionlet regions, respectively, introducing more diverse motion patterns while preserving the original motion semantics. Additionally, we incorporate a semantic-aware masked motion modeling technique to enhance the learning of actionlet representations. Our comprehensive experiments on well-established benchmark datasets such as NTU RGB+D and PKUMMD validate the effectiveness of our proposed method. Lilang Lin, Jiahang Zhang 0001, Jiaying Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Facial Image Compression via Neural Image Manifold CompressionabstractAlthough the recent learning-based image and video coding techniques achieve rapid development, the signal fidelity-driven target in these methods leads to the divergence to a highly effective and efficient coding framework for both human and machine. In this paper, we aim to address the issue by making use of the power of generative models to bridge the gap between full fidelity (for human vision) and high discrimination (for machine vision). Therefore, relying on existing pretrained generative adversarial networks (GAN), we build a GAN inversion framework that projects the image into a low-dimensional natural image manifold. In this manifold, the feature is highly discriminative and also encodes the appearance information of the image, named aslatent code. Taking a variational bit-rate constraint with a hyperprior model to model/suppress the entropy of image manifold code, our method is capable of fulfilling the needs of both machine and human visions at very low bit-rates. To improve the visual quality of image reconstruction, we further proposemultiple latent codesandscalable inversion. The former gets several latent codes in the inversion, while the latter additionally compresses and transmits a shallow compact feature to support visual reconstruction. Experimental results demonstrate the superiority of our method in both human vision tasks,i.e. image reconstruction, and machine vision tasks, including semantic parsing and attribute prediction. Wenhan Yang, Haofeng Huang, Jiaying Liu 0001, Alex Chichung Kot |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Frequency-Controlled Diffusion Model for Versatile Text-Guided Image-to-Image TranslationabstractRecently, text-to-image diffusion models have emerged as a powerful tool for image-to-image translation (I2I), allowing flexible image translation via user-provided text prompts. This paper proposes frequency-controlled diffusion model (FCDiffusion), an end-to-end diffusion-based framework contributing a novel solution to text-guided I2I from a frequency-domain perspective. At the heart of our framework is a feature-space frequency-domain filtering module based on Discrete Cosine Transform, which extracts image features carrying different DCT spectral bands to control the text-to-image generation process of the Latent Diffusion Model, realizing versatile I2I applications including style-guided content creation, image semantic manipulation, image scene translation, and image style translation. Different from related methods, FCDiffusion establishes a unified text-driven I2I framework suiting diverse I2I application scenarios simply by switching among different frequency control branches. The effectiveness and superiority of our method for text-guided I2I are demonstrated with extensive experiments both qualitatively and quantitatively. Our project is publicly available at: https://xianggao1102.github.io/FCDiffusion/. Xiang Gao 0014, Zhengbo Xu, Junhan Zhao, Jiaying Liu 0001 |
AAAI | 4 |
| 2024 | Seeing Dark Videos via Self-Learned Bottleneck Neural RepresentationabstractEnhancing low-light videos in a supervised style presents a set of challenges, including limited data diversity, misalignment, and the domain gap introduced through the dataset construction pipeline. Our paper tackles these challenges by constructing a self-learned enhancement approach that gets rid of the reliance on any external training data. The challenge of self-supervised learning lies in fitting high-quality signal representations solely from input signals. Our work designs a bottleneck neural representation mechanism that extracts those signals. More in detail, we encode the frame-wise representation with a compact deep embedding and utilize a neural network to parameterize the video-level manifold consistently. Then, an entropy constraint is applied to the enhanced results based on the adjacent spatial-temporal context to filter out the degraded visual signals, e.g. noise and frame inconsistency. Last, a novel Chromatic Retinex decomposition is proposed to effectively align the reflectance distribution temporally. It benefits the entropy control on different components of each frame and facilitates noise-to-noise training, successfully suppressing the temporal flicker. Extensive experiments demonstrate the robustness and superior effectiveness of our proposed method. Our project is publicly available at: https://huangerbai.github.io/SLBNR/. Haofeng Huang, Wenhan Yang, Ling-Yu Duan, Jiaying Liu 0001 |
AAAI | 4 |
| 2024 | Zero-Reference Low-Light Enhancement via Physical Quadruple PriorsabstractUnderstanding illumination and reducing the need for supervision pose a significant challenge in low-light enhancement. Current approaches are highly sensitive to data usage during training and illumination-specific hyper-parameters, limiting their ability to handle unseen scenarios. In this paper, we propose a new zero-reference low-light enhancement framework trainable solely with normal light images. To accomplish this, we devise an illumination-invariant prior inspired by the theory of physical light transfer. This prior serves as the bridge between normal and low-light images. Then, we develop a prior-to-image framework trained without low-light data. During testing, this frame-work is able to restore our illumination-invariant prior back to images, automatically achieving low-light enhancement. Within this framework, we leverage a pretrained generative diffusion model for model ability, introduce a by-pass decoder to handle detail distortion, as well as offer a lightweight version for practicality. Extensive experiments demonstrate our framework's superiority in various scenarios as well as good interpretability, robustness, and efficiency. Code is available on our project homepage. Wenjing Wang 0001, Huan Yang 0005, Jianlong Fu, Jiaying Liu 0001 |
CVPR | 4 |
| 2024 | Idempotent Unsupervised Representation Learning for Skeleton-Based Action Recognition
Lilang Lin, Lehong Wu, Jiahang Zhang 0001, Jiaying Liu 0001 |
ECCV (26) | 4 |
| 2024 | MacDiff: Unified Skeleton Modeling with Masked Conditional Diffusion
Lehong Wu, Lilang Lin, Jiahang Zhang 0001, Yiyang Ma, Jiaying Liu 0001 |
ECCV (26) | 5 |
| 2024 | Solving Diffusion ODEs with Optimal Boundary Conditions for Better Image Super-ResolutionabstractDiffusion models, as a kind of powerful generative model, have given impressive results on image super-resolution (SR) tasks. However, due to the randomness introduced in the reverse process of diffusion models, the performances of diffusion-based SR models are fluctuating at every time of sampling, especially for samplers with few resampled steps. This inherent randomness of diffusion models results in ineffectiveness and instability, making it challenging for users to guarantee the quality of SR results. However, our work takes this randomness as an opportunity: fully analyzing and leveraging it leads to the construction of an effective plug-and-play sampling method that owns the potential to benefit a series of diffusion-based SR methods. More in detail, we propose to steadily sample high-quality SR images from pre-trained diffusion-based SR models by solving diffusion ordinary differential equations (diffusion ODEs) with optimal boundary conditions (BCs) and analyze the characteristics between the choices of BCs and their corresponding SR results. Our analysis shows the route to obtain an approximately optimal BC via an efficient exploration in the whole space. The quality of SR results sampled by the proposed method with fewer steps outperforms the quality of results sampled by current methods with randomness from the same pre-trained diffusion-based SR model, which means that our sampling method ''boosts'' current diffusion-based SR models without any additional training. Yiyang Ma, Huan Yang 0005, Wenhan Yang, Jianlong Fu, Jiaying Liu 0001 |
ICLR | 5 |
| 2024 | Correcting Diffusion-Based Perceptual Image Compression with Privileged End-to-End DecoderabstractThe images produced by diffusion models can attain excellent perceptual quality. However, it is challenging for diffusion models to guarantee distortion, hence the integration of diffusion models and image compression models still needs more comprehensive explorations. This paper presents a diffusion-based image compression method that employs a privileged end-to-end decoder model as correction, which achieves better perceptual quality while guaranteeing the distortion to an extent. We build a diffusion model and design a novel paradigm that combines the diffusion model and an end-to-end decoder, and the latter is responsible for transmitting the privileged information extracted at the encoder side. Specifically, we theoretically analyze the reconstruction process of the diffusion models at the encoder side with the original images being visible. Based on the analysis, we introduce an end-to-end convolutional decoder to provide a better approximation of the score function $\nabla_{\mathbf{x}_t}\log p(\mathbf{x}_t)$ at the encoder side and effectively transmit the combination. Experiments demonstrate the superiority of our method in both distortion and perception compared with previous perceptual compression methods. Yiyang Ma, Wenhan Yang, Jiaying Liu 0001 |
ICML | 3 |
| 2024 | Shap-Mix: Shapley Value Guided Mixing for Long-Tailed Skeleton Based Action Recognition
Jiahang Zhang 0001, Lilang Lin, Jiaying Liu 0001 |
IJCAI | 3 |
| 2024 | FBSDiff: Plug-and-Play Frequency Band Substitution of Diffusion Features for Highly Controllable Text-Driven Image TranslationabstractLarge-scale text-to-image diffusion models have been a revolutionary milestone in the evolution of generative AI, allowing wonderful image generation with natural-language text prompt. However, the issue of lacking controllability of such models restricts their practical applicability for real-life content creation. Thus, attention has been focused on leveraging a reference image to control text-to-image synthesis, which is also regarded as manipulating (or editing) a reference image as per a text prompt, namely, text-driven image-to-image translation. This paper contributes a novel, concise, and efficient approach that adapts pre-trained large-scale text-to-image (T2I) diffusion model to the image-to-image (I2I) paradigm in a plug-and-play manner, realizing high-quality and versatile text-driven I2I translation without model training, fine-tuning, or online optimization process. To guide T2I generation with a reference image, we propose to decompose diverse guiding factors with different frequency bands of diffusion features in the DCT spectral space, and accordingly devise a novel frequency band substitution layer which realizes dynamic control of the reference image to the T2I generation result in a plug-and-play manner. We demonstrate that our method allows flexible control over both guiding factor and guiding intensity of the reference image simply by tuning the type and bandwidth of the substituted frequency band, respectively. Extensive qualitative and quantitative experiments verify superiority of our approach over related methods in I2I translation visual quality, versatility, and controllability. Our project is publicly available at: https://xianggao1102.github.io/FBSDiff_webpage/. Xiang Gao 0014, Jiaying Liu 0001 |
ACM Multimedia | 2 |
| 2024 | Consistency Guided Diffusion Model with Neural Syntax for Perceptual Image CompressionabstractDiffusion models show impressive performances in image generation with excellent perceptual quality. However, its tendency to introduce additional distortion prevents its direct application in image compression. To address the issue, this paper introduces a Consistency Guided Diffusion Model (CGDM) tailored for perceptual image compression, which integrates an end-to-end image compression model with a diffusion-based post-processing network, aiming to learn richer detail representations with less fidelity loss. In detail, the compression and post-processing networks are cascaded and a branch of consistency guided features is added to constrain the deviation in the diffusion process for better reconstruction quality. Furthermore, a Syntax driven Feature Fusion (SFF) module is constructed to take an extra ultra-low bitstream from the encoding end as input, guiding the adaptive fusion of information from the two branches. In addition, we design a globally uniform boundary control strategy with overlapped patches and adopt a continuous online optimization mode to improve both coding efficiency and global consistency. Extensive experiments validate the superiority of our method to existing perceptual compression techniques. Our project is publicly available at: https://ellisonkuang.github.io/CGDM.github.io/. Haowei Kuang, Yiyang Ma, Wenhan Yang, Zongming Guo, Jiaying Liu 0001 |
ACM Multimedia | 5 |
| 2024 | COCO-LC: Colorfulness Controllable Language-based ColorizationabstractLanguage-based image colorization aims to convert grayscale images to plausible and visually pleasing color images with language guidance, enjoying wide applications in historical photo restoration and film industry. Existing methods mainly leverage large language models and diffusion models to incorporate language guidance into the colorization process. However, it is still a great challenge to build accurate correspondence between the gray image and the semantic instructions, leading to mismatched, overflowing and under-saturated colors. In this paper, we introduce a novel coarse-to-fine framework, COlorfulness COntrollable Language-based Colorization (COCO-LC), that effectively reinforces the image-text correspondence with a coarsely colorized results. In addition, a multi-level condition that leverages both low-level and high-level cues of the gray image is introduced to realize accurate semantic-aware colorization without color overflows. Furthermore, we condition COCO-LC with a scale factor to determine the colorfulness of the output, flexibly meeting the different needs of users. We validate the superiority of COCO-LC over state-of-the-art image colorization methods in accurate, realistic and controllable colorization through extensive experiments. The code and demo will be released at https://lyf1212.github.io/COCO-LC. Shuai Yang 0001, Jiaying Liu 0001 |
ACM Multimedia | 4 |
| 2024 | CoolColor: Text-guided COherent OLd film COLORization
Zichuan Huang, Shuai Yang 0001, Jiaying Liu 0001 |
MMAsia | 4 |
| 2024 | Unsupervised Illumination Adaptation for Low-Light VisionabstractInsufficient lighting poses challenges to both human and machine visual analytics. While existing low-light enhancement methods prioritize human visual perception, they often neglect machine vision and high-level semantics. In this paper, we make pioneering efforts to build an illumination enhancement model for high-level vision. Drawing inspiration from camera response functions, our model could enhance images from the machine vision perspective despite being lightweight in architecture and simple in formulation. We also introduce two approaches that leverage knowledge from base enhancement curves and self-supervised pretext tasks to train for different downstream normal-to-low-light adaptation scenarios. Our proposed framework overcomes the limitations of existing algorithms without requiring access to labeled data in low-light conditions. It facilitates more effective illumination restoration and feature alignment, significantly improving the performance of downstream tasks in a plug-and-play manner. This research advances the field of low-light machine analytics and broadly applies to various high-level vision tasks, including classification, face detection, optical flow estimation, and video action recognition. Wenjing Wang 0001, Rundong Luo, Wenhan Yang, Jiaying Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Video Coding for Machines: Compact Visual Representation Compression for Intelligent Collaborative AnalyticsabstractAs an emerging research practice leveraging recent advanced AI techniques, e.g. deep models based prediction and generation, Video Coding for Machines (VCM) is committed to bridging to an extent separate research tracks of video/image compression and feature compression, and attempts to optimize compactness and efficiency jointly from a unified perspective of high accuracy machine vision and full fidelity human vision. With the rapid advances of deep feature representation and visual data compression in mind, in this paper, we summarize VCM methodology and philosophy based on existing academia and industrial efforts. The development of VCM follows a general rate-distortion optimization, and the categorization of key modules or techniques is established including feature-assisted coding, scalable coding, intermediate feature compression/optimization, and machine vision targeted codec, from broader perspectives of vision tasks, analytics resources, etc. From previous works, it is demonstrated that, although existing works attempt to reveal the nature of scalable representation in bits when dealing with machine and human vision tasks, there remains a rare study in the generality of low bit rate representation, and accordingly how to support a variety of visual analytic tasks. Therefore, we investigate a novel visual information compression for the analytics taxonomy problem to strengthen the capability of compact visual representations extracted from multiple tasks for visual analytics. A new perspective of task relationships versus compression is revisited. By keeping in mind the transferability among different machine vision tasks (e.g. high-level semantic and mid-level geometry-related), we aim to support multiple tasks jointly at low bit rates. In particular, to narrow the dimensionality gap between neural network generated features extracted from pixels and a variety of machine vision features/labels (e.g. scene class, segmentation labels), a codebook hyperprior is designed to compress the neural network-generated features. As demonstrated in our experiments, this new hyperprior model is expected to improve feature compression efficiency by estimating the signal entropy more accurately, which enables further investigation of the granularity of abstracting compact features among different tasks. Wenhan Yang, Haofeng Huang, Yueyu Hu, Ling-Yu Duan, Jiaying Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Learning to Remove Rain in Video With Self-SupervisionabstractIn heavy rain video, rain streak and rain accumulation are the most common causes of degradation. They occlude background information and can significantly impair the visibility. Most existing methods rely heavily on the synthetic training data, and thus raise the domain gap problem that prevents the trained models from performing adequately in real testing cases. Unlike these methods, we introduce a self-learning method to remove both rain streaks and rain accumulation without using any ground-truth clean images in training our model, which consequently can alleviate the domain gap issue. The main idea is based on the assumptions that (1) adjacent clean frames can be aligned or warped from one frame to another frame, (2) rain streaks are distributed randomly in the temporal domain, (3) the rain streak/accumulation related variables/priors can be inferred reliably from the information within the images/sequences. Based on these assumptions, we construct an augmented Self-Learned Deraining Network (SLDNet+) to remove both rain streaks and rain accumulation by utilizing temporal correlation, consistency, and rain-related priors. For the temporal correlation, our SLDNet+ takes rain degraded adjacent frames as its input, aligns them, and learns to predict the clean version of the current frame. For the temporal consistency, a new loss is designed to build a robust mapping between the predicted clean frame and non-rain regions from the adjacent rain frames. For the rain-streak-related prior, the rain streak removal network is optimized jointly with motion estimation and rain region detection; while for the rain-accumulation-related prior, a novel non-local video rain accumulation removal method is developed to estimate the accumulation-lines from the whole input video and to offer better color constancy and temporal smoothness. Extensive experiments show the effectiveness of our approach, which provides superior results compared with the existing state of the art methods both quantitatively and qualitatively. The source code will be made publicly available at: https://github.com/flyywh/CVPR-2020-Self-Rain-Removal-Journal. Wenhan Yang, Robby T. Tan, Shiqi Wang 0001, Alex Chichung Kot, Jiaying Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Diffusion Enhancement for Cloud Removal in Ultra-Resolution Remote Sensing ImageryabstractThe presence of cloud layers severely compromises the quality and effectiveness of optical remote sensing (RS) images. However, existing deep-learning (DL)-based cloud removal (CR) techniques, which usually take the fidelity-driven losses as constraints, e.g.,$L_{1}$or$L_{2}$losses, tend to generate smooth results, often failing to reconstruct visually pleasing results and cause semantic loss. To tackle this challenge, this work proposes to encompass enhancements at the data and methodology fronts. On the data side, an ultra-resolution benchmark named CUHK cloud removal (CUHK-CR) of 0.5 m spatial resolution is established. This benchmark incorporates rich detailed textures and diverse cloud coverage, serving as a robust foundation for designing and assessing CR models. From the methodology perspective, a novel diffusion-based framework for CR named diffusion enhancement (DE) is introduced. This framework aims to gradually recover texture details, leveraging a reference visual prior providing foundational structure of the images to enhance inference accuracy. Additionally, a weight allocation (WA) network is developed to dynamically adjust the weights for feature fusion, thereby further improving performance, particularly in the context of ultra-resolution image generation. Furthermore, a coarse-to-fine training strategy is applied to effectively expedite training convergence while reducing the computational complexity required to handle ultra-resolution images. Extensive experiments on the newly established CUHK-CR and existing datasets such as RICE confirm that the proposed DE framework outperforms existing DL-based methods in terms of both perceptual quality and signal fidelity. Jialu Sui, Yiyang Ma, Wenhan Yang, Man-On Pun, Jiaying Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | S⁵Mars: Semi-Supervised Learning for Mars Semantic SegmentationabstractDeep learning has become a powerful tool for Mars exploration. Mars terrain semantic segmentation is an important Martian vision task, which is the base of rover autonomous planning and safe driving. However, there is a lack of sufficient detailed and high-confidence data annotations, which are exactly required by most deep learning methods to obtain a good model. To address this problem, we propose our solution from the perspective of joint data and method design. We first present a new dataset S5Mars for Semi-SuperviSed learning on Mars Semantic Segmentation, which contains 6K high-resolution images and is sparsely annotated based on confidence, ensuring the high quality of labels. Then to learn from this sparse data, we propose a semi-supervised learning (SSL) framework for Mars image semantic segmentation, to learn representations from limited labeled data. Different from the existing SSL methods which are mostly targeted at the Earth image data, our method takes into account Mars data characteristics. Specifically, we first investigate the impact of current widely used natural image augmentations on Mars images. Based on the analysis, we then proposed two novel and effective augmentations for SSL of Mars segmentation,AugINandSAM-Mix, which serve as strong augmentations to boost the model performance. Meanwhile, to fully leverage the unlabeled data, we introduce a soft-to-hard consistency learning strategy, learning from different targets based on prediction confidence. Experimental results show that our method can outperform state-of-the-art SSL approaches remarkably. Our proposed dataset is available at https://jhang2020.github.io/S5Mars.github.io/. Jiahang Zhang 0001, Lilang Lin, Zejia Fan, Wenjing Wang 0001, Jiaying Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Toward Real-World Super Resolution With Adaptive Self-Similarity MiningabstractDespite efforts to construct super-resolution (SR) training datasets with a wide range of degradation scenarios, existing supervised methods based on these datasets still struggle to consistently offer promising results due to the diversity of real-world degradation scenarios and the inherent complexity of model learning. Our work explores a new route: integrating the sample-adaptive property learned through image intrinsic self-similarity and the universal knowledge acquired from large-scale data. We achieve this by uniting internal learning and external learning by an unrolled optimization process. With the merits of both, the tuned fully-supervised SR models can be augmented to broadly handle the real-world degradation in a plug-and-play style. Furthermore, to promote the efficiency of combining internal/external learning, we apply an attention-based weight-updating method to guide the mining of self-similarity, and various data augmentations are adopted while applying the exponential moving average strategy. We conduct extensive experiments on real-world degraded images and our approach outperforms other methods in both qualitative and quantitative comparisons. Our project is available at: https://github.com/ZahraFan/AdaSSR/. Zejia Fan, Wenhan Yang, Zongming Guo, Jiaying Liu 0001 |
IEEE Trans. Image Process. | 4 |
| 2024 | Mutual Information Driven Equivariant Contrastive Learning for 3D Action Representation LearningabstractSelf-supervised contrastive learning has proven to be successful for skeleton-based action recognition. For contrastive learning, data transformations are found to fundamentally affect the learned representation quality. However, traditional invariant contrastive learning is detrimental to the performance on the downstream task if the transformation carries important information for the task. In this sense, it limits the application of many data transformations in the current contrastive learning pipeline. To address these issues, we propose to utilize equivariant contrastive learning, which extends invariant contrastive learning and preserves important information. By integrating equivariant and invariant contrastive learning into a hybrid approach, the model can better leverage the motion patterns exposed by data transformations and obtain a more discriminative representation space. Specifically, a self-distillation loss is first proposed for transformed data of different intensities to fully utilize invariant transformations, especially strong invariant transformations. For equivariant transformations, we explore the potential of skeleton mixing and temporal shuffling for equivariant contrastive learning. Meanwhile, we analyze the impacts of different data transformations on the feature space in terms of two novel metrics proposed in this paper, namely, consistency and diversity. In particular, we demonstrate that equivariant learning boosts performance by alleviating the dimensional collapse problem. Experimental results on several benchmarks indicate that our method outperforms existing state-of-the-art methods. Lilang Lin, Jiahang Zhang 0001, Jiaying Liu 0001 |
IEEE Trans. Image Process. | 3 |
| 2024 | Prompt-Based Modality Bridging for Unified Text-to-Face Generation and ManipulationabstractText-driven face image generation and manipulation are significant tasks. However, such tasks are quite challenging due to the gap between text and image modalities. It is difficult to utilize current methods to deal with both of the two problems because these methods are usually designed for one certain task, limiting their application in real scenarios. To address the two problems in one framework, we propose a Unified Prompt-based Cross-Modal Framework (UPCM-Frame) to bridge the gap between the text modality and image modality with CLIP and StyleGAN, which are two large-scale pre-trained models. The proposed framework is combined with two main modules: a Text Embedding-to-Image Embedding projection module based on a special prompt embedding pair, and a projection module mapping Image Embeddings to semantically aligned StyleGAN Embeddings which can be used in both image generation and manipulation. The proposed framework is able to handle complicated descriptions and generate impressive results with high quality due to the utilization of large-scale pre-trained models. In order to demonstrate the effectiveness of the proposed method in the two tasks, we conduct experiments to evaluate the results of our method both quantitatively and qualitatively. Yiyang Ma, Haowei Kuang, Huan Yang 0005, Jianlong Fu, Jiaying Liu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2023 | Hierarchical Consistent Contrastive Learning for Skeleton-Based Action Recognition with Growing AugmentationsabstractContrastive learning has been proven beneficial for self-supervised skeleton-based action recognition. Most contrastive learning methods utilize carefully designed augmentations to generate different movement patterns of skeletons for the same semantics. However, it is still a pending issue to apply strong augmentations, which distort the images/skeletons’ structures and cause semantic loss, due to their resulting unstable training. In this paper, we investigate the potential of adopting strong augmentations and propose a general hierarchical consistent contrastive learning framework (HiCLR) for skeleton-based action recognition. Specifically, we first design a gradual growing augmentation policy to generate multiple ordered positive pairs, which guide to achieve the consistency of the learned representation from different views. Then, an asymmetric loss is proposed to enforce the hierarchical consistency via a directional clustering operation in the feature space, pulling the representations from strongly augmented views closer to those from weakly augmented views for better generalizability. Meanwhile, we propose and evaluate three kinds of strong augmentations for 3D skeletons to demonstrate the effectiveness of our method. Extensive experiments show that HiCLR outperforms the state-of-the-art methods notably on three large-scale datasets, i.e., NTU60, NTU120, and PKUMMD. Our project is publicly available at: https://jhang2020.github.io/Projects/HiCLR/HiCLR.html. Jiahang Zhang 0001, Lilang Lin, Jiaying Liu 0001 |
AAAI | 3 |
| 2023 | Actionlet-Dependent Contrastive Learning for Unsupervised Skeleton-Based Action RecognitionabstractThe self-supervised pretraining paradigm has achieved great success in skeleton-based action recognition. However, these methods treat the motion and static parts equally, and lack an adaptive design for different parts, which has a negative impact on the accuracy of action recognition. To realize the adaptive action modeling of both parts, we propose an Actionlet-Dependent Contrastive Learning method (ActCLR). The actionlet, defined as the discriminative subset of the human skeleton, effectively decomposes motion regions for better action modeling. In detail, by contrasting with the static anchor without motion, we extract the motion region of the skeleton data, which serves as the actionlet, in an unsupervised manner. Then, centering on actionlet, a motion-adaptive data transformation method is built. Different data transformations are applied to action let and non-actionlet regions to introduce more diversity while maintaining their own characteristics. Meanwhile, we propose a semantic-aware feature pooling method to build feature representations among motion and static regions in a distinguished manner. Extensive experiments on NTU RGB+D and PKUMMD show that the proposed method achieves remarkable action recognition performance. More visualization and quantitative experiments demonstrate the effectiveness of our method. Our project website is available at https://langlandslin.github.io/projects/ActCLR/ Lilang Lin, Jiahang Zhang 0001, Jiaying Liu 0001 |
CVPR | 3 |
| 2023 | Similarity Min-Max: Zero-Shot Day-Night Domain AdaptationabstractLow-light conditions not only hamper human visual experience but also degrade the model’s performance on downstream vision tasks. While existing works make remarkable progress on day-night domain adaptation, they rely heavily on domain knowledge derived from the task-specific nighttime dataset. This paper challenges a more complicated scenario with border applicability, i.e., zero-shot day-night domain adaptation, which eliminates reliance on any nighttime data. Unlike prior zero-shot adaptation approaches emphasizing either image-level translation or model-level adaptation, we propose a similarity min-max paradigm that considers them under a unified framework. On the image level, we darken images towards minimum feature similarity to enlarge the domain gap. Then on the model level, we maximize the feature similarity between the darkened images and their normal-light counterparts for better model adaptation. To the best of our knowledge, this work represents the pioneering effort in jointly optimizing both levels, resulting in a significant improvement of model generalizability. Extensive experiments demonstrate our method’s effectiveness and broad applicability on various nighttime vision tasks, including classification, semantic segmentation, visual place recognition, and video action recognition. Our project page is available at https://red-fairy.github.io/ZeroShotDayNightDA-Webpage/ Rundong Luo, Wenjing Wang 0001, Wenhan Yang, Jiaying Liu 0001 |
ICCV | 4 |
| 2023 | Flash Compensated Low-Light Enhancement Via Hierarchical Network PredictionabstractPhotography in low-light conditions suffers from dense noise and insufficient light. Flash photography, introducing extra light sources, performs better at suppressing noise and revealing details, while being interrupted by unnatural ambient illumination. This paper offers an analysis of the pros and cons to utilize low-light and flash images for enhancement, which inspires us to design a unified sample-adaptive CNN to capture diverse focuses from different inputs in a complementary way. Specifically, a Flash Compensated Dynamic Filtering Network is proposed to utilize the revealed details of flash images to compensate for fine structure reconstruction in low-light enhancement. To adaptively fuse information from misaligned low-light and flash image pairs, our network is designed with three distinctive features. Firstly, we adopt a layer-wise regression strategy, where results are predicted from the single input first and then fused to sufficiently leverage complementary information. Secondly, we employ a sample-adaptive mechanism, where each pixel is estimated with its distinctive parameters augmented by weighted residual connections. Finally, we utilize a coarse-to-fine architecture, where features are extracted by diversified receptive fields to utilize hierarchical contextual information. Experimental results demonstrate that the three design principles lead to the significant superiority of the proposed method over state-of-the-art methods. Haowei Kuang, Haofeng Huang, Wenhan Yang, Jiaying Liu 0001 |
ICIP | 4 |
| 2023 | Content-Adaptive Parallel Entropy Coding for End-to-End Image CompressionabstractState-of-the-art entropy models, e.g. autoregressive context models, utilize spatial correlation among latent representations, leading to more accurate entropy estimation. However, this autoregressive design naturally results in serial decoding and the infeasibility of parallelization, which makes the decoding procedure slow and less practical. To address the issue, we propose a Content-Adaptive Parallel Entropy Model (CAPEM) that takes a two-pass context calculation with dynamically generated patterns. Our CAPEM relaxes the strict coding order while the dynamic context mechanism still promotes flexibility in capturing latent dependency. This design greatly improves the parallelism of the context model, leading to higher coding efficiency while maintaining the same rate-distortion performance. We test it on the widely used Kodak and CLIC image datasets. Experimental results show that the proposed model outperforms the recent works with less complexity. Shujia Li, Dezhao Wang, Zejia Fan, Jiaying Liu 0001 |
ICIP | 4 |
| 2023 | Collaborative Spatial-Temporal Distillation for Efficient Video DerainingabstractIn this paper, we propose a novel knowledge distillation framework to improve the efficiency of deep networks for video deraining. The knowledge is transferred from a large-scale powerful teacher network to a compact efficient student network via the proposed collaborative spatial-temporal distillation framework. The framework is equipped with three collaboration schemes of different granularities that make use of spatial-temporal redundancy in a complementary way for better distillation performance. First, the spatial alignment module applies distillation constraints at different spatial scales to achieve better scale invariance in transferred knowledge. Second, the temporal alignment module traces both temporal status between teacher and student separately and collaboratively, to comprehensively utilize inter-frame information. Third, these two alignment modules interact through a spatial-temporal adaptor, where spatial-temporal knowledge is transferred in a unified framework. Extensive experiments demonstrate the superiority of our distillation framework as well as the effectiveness of each module. Our code is available at: https://github.com/HuYuzhang/Knowledge-Distillation. Yuzhang Hu, Minghao Liu 0019, Wenhan Yang, Jiaying Liu 0001, Zongming Guo |
ICME | 4 |
| 2023 | Bayesian Contrastive Learning with Manifold Regularization for Self-Supervised Skeleton Based Action RecognitionabstractIn this paper, we address skeleton-based action recognition under the self-supervised setting. We propose a novel framework Bayesian Contrastive Learning with Manifold Regularization (BCLR). In Bayesian contrastive learning, we employ Monte Carlo Dropout sampling on the adjacency matrix of the skeleton data to obtain positive/negative samples for model robustness. A novel entropy-based memory bank updating strategy is further proposed to take full advantage of hard negative samples for better separability. The feature manifold regularization, including projection-based data reconstruction and similarity-based feature decoupling, on the other hand, is designed to extract comprehensive information to avoid overfitting and increase feature diversity to prevent a collapse of the model. With Bayesian contrastive learning and feature manifold regularization, our model learns stronger and more discriminative features. Extensive experiments on NTU RGB+D and PKUMMD show that the proposed method achieves remarkable action recognition performance. Lilang Lin, Jiahang Zhang 0001, Jiaying Liu 0001 |
ISCAS | 3 |
| 2023 | Temporal Consistent Oil Painting Video StylizationabstractThe automatic rendering of oil painting style video has great artistic and commercial application value. Temporal consistency is the bottleneck of video rendering. However, existing translation methods are either designed for images, or have high training/inference costs on videos due to the estimation of optical flows. This paper explores how to render videos in oil painting styles without video training data. We adopt a motion-based regularization in the training phase and a feature statistics sharing strategy in the inference phase. Experiments show that our model can render vivid and temporally smooth oil painting videos. Luyao Zhang 0007, Wenjing Wang 0001, Jiaying Liu 0001 |
ISCAS | 3 |
| 2023 | Prompted Contrast with Masked Motion Modeling: Towards Versatile 3D Action Representation LearningabstractSelf-supervised learning has proved effective for skeleton-based human action understanding, which is an important yet challenging topic. Previous works mainly rely on contrastive learning or masked motion modeling paradigm to model the skeleton relations. However, the sequence-level and joint-level representation learning cannot be effectively and simultaneously handled by these methods. As a result, the learned representations fail to generalize to different downstream tasks. Moreover, combining these two paradigms in a naive manner leaves the synergy between them untapped and can lead to interference in training. To address these problems, we propose Prompted Contrast with Masked Motion Modeling, PCM 3, for versatile 3D action representation learning. Our method integrates the contrastive learning and masked prediction tasks in a mutually beneficial manner, which substantially boosts the generalization capacity for various downstream tasks. Specifically, masked prediction provides novel training views for contrastive learning, which in turn guides the masked prediction training with high-level semantic information. Moreover, we propose a dual-prompted multi-task pretraining strategy, which further improves model representations by reducing the interference caused by learning the two different pretext tasks. Extensive experiments on five downstream tasks under three large-scale datasets are conducted, demonstrating the superior generalization capacity of PCM3 compared to the state-of-the-art works. Our project is publicly available at: https://jhang2020.github.io/Projects/PCM3/PCM3.html. Jiahang Zhang 0001, Lilang Lin, Jiaying Liu 0001 |
ACM Multimedia | 3 |
| 2023 | Unsupervised Face Detection in the DarkabstractLow-light face detection is challenging but critical for real-world applications, such as nighttime autonomous driving and city surveillance. Current face detection models rely on extensive annotations and lack generality and flexibility. In this paper, we explore how to learn face detectors without low-light annotations. Fully exploiting existing normal light data, we propose adapting face detectors from normal light to low light. This task is difficult because the gap between brightness and darkness is too large and complicated at the object level and pixel level. Accordingly, the performance of current low-light enhancement or adaptation methods is unsatisfactory. To solve this problem, we propose a joint High-Low Adaptation (HLA) framework. We design bidirectional low-level adaptation and multitask high-level adaptation. For low-level, we enhance the dark images and degrade the normal-light images, making both domains move toward each other. For high-level, we combine context-based and contrastive learning to comprehensively close the features on different domains. Experiments show that our HLA-Face v2 model obtains superior low-light face detection performance even without the use of low-light annotations. Moreover, our adaptation scheme can be extended to a wide range of applications, such as improving supervised learning and generic object detection. Project publicly available at: https://daooshee.github.io/HLA-Face-v2-Website/. Wenjing Wang 0001, Wenhan Yang, Jiaying Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Automatic Model-Based Dataset Generation for High-Level Vision Tasks of Autonomous Driving in Haze WeatherabstractImproving the performance of high-level computer vision tasks in adverse weather (e.g., haze) is highly critical for autonomous driving safety. However, collecting and annotating training sets for various high-level tasks in haze weather are expensive and time-consuming. To address this issue, we propose a novel haze generation model called HazeGEN by coupling the variational autoencoder and the generative adversarial network to automatically generate annotated datasets. The proposed HazeGEN leverages a shared latent space assumption based on an optimized encoder–decoder architecture, which guarantees high fidelity in the cross-domain image translations. To ensure that the generated image can truly facilitate high-level vision task performance, a semisupervised learning strategy is developed for HazeGEN to efficiently learn the useful knowledge from both the real-world images (with unsupervised losses) and the synthetic images generated following the atmosphere scattering model (with supervised losses). Extensive experiments and ablation studies demonstrate that training the model with our generated haze dataset greatly improves accuracy in high-level tasks such as semantic segmentation and object detection. Furthermore, one important but under-exploited issue is investigated to find out whether the developed dataset can be a good substitute for the real ones. Results show that the generated dataset has the most similar performance to the real-world collected haze dataset on multiple challenging industrial scenarios compared with prior works. Tianqi Su, Siyi Chen 0004, Wenhan Yang, Jiaying Liu 0001, Zhongfeng Wang 0001 |
IEEE Trans. Ind. Informatics | 5 |
| 2023 | Intelligent Typography: Artistic Text Style Transfer for Complex Texture and StructureabstractText style transfer is an important task to render artistic texts from a reference image or style, and is widely desired in many visual creations. Previous works have brought some efficient methods for text style transfer, which facilitate users to design various artistic texts automatically. However, these works mainly focus on relatively simple text effects, and do not perform well on complex reference styles. In this paper, we propose a coarse-to-fine framework to generate exquisite texts with complex texture and structure in an unsupervised way, achieving real-time control of style scales (i.e., text stylistic degree or deformation degree). The key idea is to decouple the overall task into two steps, prototype generation and detail refinement, and explore delicate networks for each step to imitate the features at different levels. Based on this idea, in the first step, we present a novel pro-gen GAN to generate prototypes of artistic texts using the reference style, and develop a deformable module to empower the pro-gen GAN to continuously characterize the multi-scale shape features without network retraining. Furthermore, we propose a mix-attention training scheme for text style transfer, which can avoid artifacts and retain a clear text background. In the second step, we introduce two optimized networks for detail refinements. Experimental results show that the proposed method can synthesize exquisite stylized texts with complex reference styles, and surpass the state of the arts in texture reconstruction, contour imitation, and text image quality drastically. Wendong Mao, Shuai Yang 0001, Huihong Shi, Jiaying Liu 0001, Zhongfeng Wang 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | Semi-supervised Learning for Mars Imagery Classification and SegmentationabstractWith the progress of Mars exploration, numerous Mars image data are being collected and need to be analyzed. However, due to the severe train-test gap and quality distortion of Martian data, the performance of existing computer vision models is unsatisfactory. In this article, we introduce a semi-supervised framework for machine vision on Mars and try to resolve two specific tasks: classification and segmentation. Contrastive learning is a powerful representation learning technique. However, there is too much information overlap between Martian data samples, leading to a contradiction between contrastive learning and Martian data. Our key idea is to reconcile this contradiction with the help of annotations and further take advantage of unlabeled data to improve performance. For classification, we propose to ignore inner-class pairs on labeled data as well as neglect negative pairs on unlabeled data, forming supervised inter-class contrastive learning and unsupervised similarity learning. For segmentation, we extend supervised inter-class contrastive learning into an element-wise mode and use online pseudo labels for supervision on unlabeled areas. Experimental results show that our learning strategies can improve the classification and segmentation models by a large margin and outperform state-of-the-art approaches. Wenjing Wang 0001, Lilang Lin, Zejia Fan, Jiaying Liu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | Deep Inter Prediction with Error-Corrected Auto-Regressive Network for Video CodingabstractModern codecs remove temporal redundancy of a video via inter prediction, i.e., searching previously coded frames for similar blocks and storing motion vectors to save bit-rates. However, existing codecs adopt block-level motion estimation, where a block is regressed by reference blocks linearly and is doomed to fail to deal with non-linear motions. In this article, we generate virtual reference frames (VRFs) with previously reconstructed frames via deep networks to offer an additional candidate, which is not constrained to linear motion structure and further significantly improves coding efficiency. More specifically, we propose a novel deep Auto-Regressive Moving-Average (ARMA) model, Error-Corrected Auto-Regressive Network (ECAR-Net), equipped with the powers of the conventional statistic ARMA models and deep networks jointly for reference frame prediction. Similar to conventional ARMA models, the ECAR-Net consists of two stages: Auto-Regression (AR) stage and Error-Correction (EC) stage, where the first part predicts the signal at the current time-step based on previously reconstructed frames, while the second one compensates for the output of the AR stage to obtain finer details. Different from the statistic AR models only focusing on short-term temporal dependency, the AR model of our ECAR-Net is further injected with the long-term dynamics mechanism, where long temporal information is utilized to help predict motions more accurately. Furthermore, ECAR-Net works in a configuration-adaptive way, i.e., using different dynamics and error definitions for the Low Delay B and Random Access configurations, which helps improve the adaptivity and generality in diverse coding scenarios. With the well-designed network, our method surpasses HEVC on average 5.0% and 6.6% BD-rate saving for the luma component under the Low Delay B and Random Access configurations and also obtains on average 1.54% BD-rate saving over VVC. Furthermore, ECAR-Net works in a configuration-adaptive way, i.e., using different dynamics and error definitions for the Low Delay B and Random Access configurations, which helps improve the adaptivity and generality in diverse coding scenarios. Yuzhang Hu, Wenhan Yang, Jiaying Liu 0001, Zongming Guo |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2022 | Neural Data-Dependent Transform for Learned Image CompressionabstractLearned image compression has achieved great success due to its excellent modeling capacity, but seldom further considers the Rate-Distortion Optimization (RDO) of each input image. To explore this potential in the learned codec, we make the first attempt to build a neural data-dependent transform and introduce a continuous online mode decision mechanism to jointly optimize the coding efficiency for each individual image. Specifically, apart from the image content stream, we employ an additional model stream to generate the transform parameters at the decoder side. The pres-ence of a model stream enables our model to learn more abstract neural-syntax, which helps cluster the latent repre-sentations of images more compactly. Beyond the transform stage, we also adopt neural-syntax based post-processing for the scenarios that require higher quality reconstructions regardless of extra decoding overhead. Moreover, the in-volvement of the model stream further makes it possible to optimize both the representation and the decoder in an on-line way, i. e. RDO at the testing time. It is equivalent to a continuous online mode decision, like coding modes in the traditional codecs, to improve the coding efficiency based on the individual input image. The experimental results show the effectiveness of the proposed neural-syntax de-sign and the continuous online mode decision mechanism, demonstrating the superiority of our method in coding effi-ciency. Our project is available at: https://dezhao-wang.github.io/Neural-Syntax-Website/. Dezhao Wang, Wenhan Yang, Yueyu Hu, Jiaying Liu 0001 |
CVPR | 4 |
| 2022 | Self-Learned Video Super-Resolution with Augmented Spatial and Temporal ContextabstractVideo super-resolution methods typically rely on paired training data, in which the low-resolution frames are usually synthetically generated under predetermined degradation conditions (e.g., Bicubic downsampling). However, in real applications, it is labor-consuming and expensive to obtain this kind of training data, which limits the practical performance of these methods. To address the issue and get rid of the synthetic paired data, in this paper, we make exploration in utilizing the internal self-similarity redundancy within the video to build a Self-Learned Video Super-Resolution (SLVSR) method, which only needs to be trained on the input testing video itself. We employ a series of data augmentation strategies to make full use of the spatial and temporal context of the target video clips. The idea is applied to two branches of mainstream SR methods: frame fusion and frame recurrence methods. Since the former takes advantage of the short-term temporal consistency and the latter of the long-term one, our method can satisfy different practical situations. The experimental results show the superiority of our proposed method, especially in addressing the video super-resolution problems in real applications. Zejia Fan, Jiaying Liu 0001, Wenhan Yang, Zongming Guo |
ICASSP | 2 |
| 2022 | Rain-Prior Injected Knowledge Distillation for Single Image DerainingabstractThis paper makes efforts in improving the efficiency of deep networks for single image deraining with a newly proposed knowledge distillation framework. Specifically, we propose a rain-prior injected distillation scheme to transfer the knowledge from a large-scale teacher network to a more compact student network. Previous works directly calculate the distillation loss between the features extracted from the student and teacher networks. Differently, our distillation scheme adaptively removes the noisy background patterns by calculating the distillation loss based on the residual feature, which is inferred from the features extracted from the rain and ground truth images. This residual operation makes the student network focus on transferring only the knowledge on the rain streaks instead of the background, which facilitates more effective distillation results. Furthermore, our method can be applied to reduce both the network size and the deraining recurrence stage, which makes it a plug-and-play module that can be integrated into diverse existing deraining methods. Experimental results prove the efficiency of our method to build an efficient deraining network and the superiority over existing distillation methods. Yuzhang Hu, Wenhan Yang, Jiaying Liu 0001, Zongming Guo |
ICIP | 3 |
| 2022 | On the Connection between Local Attention and Dynamic Depth-wise Convolution
Qi Han 0007, Zejia Fan, Qi Dai 0001, Ming-Ming Cheng, Jiaying Liu 0001, Jingdong Wang 0001 |
ICLR | 6 |
| 2022 | Self-supervised Learning and Adaptation for Single Image DehazingabstractExisting deep image dehazing methods usually depend on supervised learning with a large number of hazy-clean image pairs which are expensive or difficult to collect. Moreover, dehazing performance of the learned model may deteriorate significantly when the training hazy-clean image pairs are insufficient and are different from real hazy images in applications. In this paper, we show that exploiting large scale training set and adapting to real hazy images are two critical issues in learning effective deep dehazing models. Under the depth guidance estimated by a well-trained depth estimation network, we leverage the conventional atmospheric scattering model to generate massive hazy-clean image pairs for the self-supervised pre-training of dehazing network. Furthermore, self-supervised adaptation is presented to adapt pre-trained network to real hazy images. Learning without forgetting strategy is also deployed in self-supervised adaptation by combining self-supervision and model adaptation via contrastive learning. Experiments show that our proposed method performs favorably against the state-of-the-art methods, and is quite efficient, i.e., handling a 4K image in 23 ms. The codes are available at https://github.com/DongLiangSXU/SLAdehazing. Yudong Liang, Bin Wang 0071, Wangmeng Zuo, Jiaying Liu 0001, Wenqi Ren |
IJCAI | 4 |
| 2022 | Collaborative Scalable Visual Compression for Human-Centered VideosabstractMachine intelligence systems have been increasingly widely deployed in real-world circumstances, while the conventional human-vision oriented video coding schemes are inefficient to be embedded in large-scale systems and further support a wide range of applications. There have been urgent demands for a new generation of compression framework to efficiently encodes visual data, where the compression and analytics for machine vision and human perception can be jointly optimized. To this end, we propose a novel visual compression framework to provide visual contents with different granularity for both human and machine vision tasks collaboratively. The proposed scalable compression framework maintains the critical semantic information in a basic layer, so that it is capable of supporting the accurate machine vision analysis under a tight bit-rate constraint. It is scalable to provide visual representations of different granularity to support various kinds of tasks, including video reconstruction that serves human vision examination. Experimental results on the human-centered videos have demonstrated the promising functionality of scalable visual coding with improved efficiency for high-performance machine analysis and human perception. Haofeng Huang, Wenhan Yang, Jiaying Liu 0001, Ling-Yu Duan |
ISCAS | 4 |
| 2022 | Meta-Interpolation: Time-Arbitrary Frame Interpolation via Dual Meta-LearningabstractExisting video frame interpolation methods can only interpolate the frame at a given intermediate time-step, e.g. 1/2. In this paper, we aim to explore a more generalized kind of video frame interpolation, that at an arbitrary time-step. To this end, we consider processing different time-steps with adaptively generated convolutional kernels in a unified way with the help of meta-learning. Specifically, we develop a dual meta-learned frame interpolation framework to synthesize intermediate frames with the guidance of context information and optical flow as well as taking the time-step as side information. First, a content-aware meta-learned flow refinement module is built to improve the accuracy of the optical flow estimation based on the down-sampled version of the input frames. Second, with the refined optical flow and the time-step as the input, a motion-aware meta-learned frame interpolation module generates the convolutional kernels for every pixel used in the convolution operations on the feature map of the coarse warped version of the input frames to generate the predicted frame. Extensive qualitative and quantitative evaluations, as well as ablation studies, demonstrate that, via introducing meta-learning in our framework in such a well-designed way, our method not only achieves superior performance to state-of-the-art frame interpolation approaches but also owns an extended capacity to support the interpolation at an arbitrary time-step. Shixing Yu, Yiyang Ma, Wenhan Yang, Jiaying Liu 0001 |
ISCAS | 5 |
| 2022 | Self-Aligned Concave Curve: Illumination Enhancement for Unsupervised AdaptationabstractLow light conditions not only degrade human visual experience, but also reduce the performance of downstream machine analytics. Although many works have been designed for low-light enhancement or domain adaptive machine analytics, the former considers less on high-level vision, while the latter neglects the potential of image-level signal adjustment. How to restore underexposed images/videos from the perspective of machine vision has long been overlooked. In this paper, we are the first to propose a learnable illumination enhancement model for high-level vision. Inspired by real camera response functions, we assume that the illumination enhancement function should be a concave curve, and propose to satisfy this concavity through discrete integral. With the intention of adapting illumination from the perspective of machine vision without task-specific annotated data, we design an asymmetric cross-domain self-supervised training strategy. Our model architecture and training designs mutually benefit each other, forming a powerful unsupervised normal-to-low light adaptation framework. Comprehensive experiments demonstrate that our method surpasses existing low-light enhancement and adaptation methods and shows superior generalization on various low-light vision tasks, including classification, detection, action recognition, and optical flow estimation. All of our data, code, and results will be available online upon publication of the paper. Wenjing Wang 0001, Zhengbo Xu, Haofeng Huang, Jiaying Liu 0001 |
ACM Multimedia | 4 |
| 2022 | Learning Hierarchical Dynamics with Spatial Adjacency for Image EnhancementabstractIn various real-world image enhancement applications, the degradations are always non-uniform or non-homogeneous and diverse, which challenges most deep networks with fixed parameters during the inference phase. Inspired by the dynamic deep networks that adapt the model structures or parameters conditioned on the inputs, we propose a DCP-guided hierarchical dynamic mechanism for image enhancement to adapt the model parameters and features from local to global as well as to keep spatial adjacency within the region. Specifically, channel-spatial-level, structure-level, and region-level dynamic components are sequentially applied. Channel-spatial-level dynamics obtain channel- and spatial-wise representation variations, and structure-level dynamics enable modeling geometric transformations and augment sampling locations for the varying local features to better describe the structures. In addition, a novel region-level dynamic is proposed to generate spatially continuous masks for dynamic features which capitalizes on the Dark Channel Priors (DCP). The proposed region-level dynamics benefit from exploiting the statistical differences between distorted and undistorted images. Moreover, the DCP-guided region generations are inherently spatial coherent which facilitates capturing local coherence of the images. The proposed method achieves state-of-the-art performance and generates visually pleasing images for multiple enhancement tasks,i.e. , image dehazing, image deraining and low-light image enhancement. The codes are available at https://github.com/DongLiangSXU/HDM. Yudong Liang, Bin Wang 0071, Wenqi Ren, Jiaying Liu 0001, Wangmeng Zuo |
ACM Multimedia | 4 |
| 2022 | AI Illustrator: Translating Raw Descriptions into Images by Prompt-based Cross-Modal GenerationabstractAI illustrator aims to automatically design visually appealing images for books to provoke rich thoughts and emotions. To achieve this goal, we propose a framework for translating raw descriptions with complex semantics into semantically corresponding images. The main challenge lies in the complexity of the semantics of raw descriptions, which may be hard to be visualized e.g., "gloomy" or "Asian"). It usually poses challenges for existing methods to handle such descriptions. To address this issue, we propose a Prompt-based Cross-Modal Generation Framework (PCM-Frame) to leverage two powerful pre-trained models, including CLIP and StyleGAN. Our framework consists of two components: a projection module from Text Embeddings to Image Embeddings based on prompts, and an adapted image generation module built on StyleGAN which takes Image Embeddings as inputs and is trained by combined semantic consistency losses. To bridge the gap between realistic images and illustration designs, we further adopt a stylization model as post-processing in our framework for better visual effects. Benefiting from the pre-trained models, our method can handle complex descriptions and does not require external paired data for training. Furthermore, we have built a benchmark that consists of 200 descriptions from literature books or online resources. We conduct a user study to demonstrate our superiority over the competing methods of text-to-image translation with complicated semantics. Yiyang Ma, Huan Yang 0005, Bei Liu 0001, Jianlong Fu, Jiaying Liu 0001 |
ACM Multimedia | 5 |
| 2022 | Learning-Based Video Coding with Joint Deep Compression and EnhancementabstractEnd-to-end learning-based video coding has attracted substantial attentions by compressing video signals as stacked visual features. This paper proposes an end-to-end deep video codec with jointly optimized compression and enhancement modules (JCEVC). First, we propose a dual-path generative adversarial network (DPEG) to reconstruct video details after compression. An α-path and a β-path concurrently reconstruct the structure information and local textures. Second, we reuse the DPEG network in both motion compensation and quality enhancement modules, which are further combined with other necessary modules to formulate our JCEVC framework. Third, we employ a joint training of deep video compression and enhancement that further improves the rate-distortion (RD) performance of compression. Compared with x265 LDP very fast mode, our JCEVC reduces the average bit-per-pixel (bpp) by 39.39%/54.92% at the same PSNR/MS-SSIM, which outperforms the state-of-the-art deep video codecs by a considerable margin. Sourcecode is available at: https://github.com/fwz1021/JCEVC. Tiesong Zhao, Weize Feng, Hongji Zeng, Yuzhen Niu, Jiaying Liu 0001 |
ACM Multimedia | 6 |
| 2022 | Learning End-to-End Lossy Image Compression: A BenchmarkabstractImage compression is one of the most fundamental techniques and commonly used applications in the image and video processing field. Earlier methods built a well-designed pipeline, and efforts were made to improve all modules of the pipeline by handcrafted tuning. Later, tremendous contributions were made, especially when data-driven methods revitalized the domain with their excellent modeling capacities and flexibility in incorporating newly designed modules and constraints. Despite great progress, a systematic benchmark and comprehensive analysis of end-to-end learned image compression methods are lacking. In this paper, we first conduct a comprehensive literature survey of learned image compression methods. The literature is organized based on several aspects to jointly optimize the rate-distortion performance with a neural network, i.e., network architecture, entropy model and rate control. We describe milestones in cutting-edge learned image-compression methods, review a broad range of existing works, and provide insights into their historical development routes. With this survey, the main challenges of image compression methods are revealed, along with opportunities to address the related issues with recent advanced learning methods. This analysis provides an opportunity to take a further step towards higher-efficiency image compression. By introducing a coarse-to-fine hyperprior model for entropy estimation and signal reconstruction, we achieve improved rate-distortion performance, especially on high-resolution images. Extensive benchmark experiments demonstrate the superiority of our model in rate-distortion performance and time complexity on multi-core CPUs and GPUs. Yueyu Hu, Wenhan Yang, Zhan Ma 0001, Jiaying Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Recurrent Multi-Frame Deraining: Combining Physics Guidance and Adversarial LearningabstractExisting video rain removal methods mainly focus on rain streak removal and are solely trained based on the synthetic data, which neglect more complex degradation factors, e.g., rain accumulation, and the prior knowledge in real rain data. Thus, in this paper, we build a more comprehensive rain model with several degradation factors and construct a novel two-stage video rain removal method that combines the power of synthetic videos and real data. Specifically, a novel two-stage progressive network is proposed: recovery guided by a physics model, and further restoration by adversarial learning. The first stage performs an inverse recovery process guided by our proposed rain model. An initially estimated background frame is obtained based on the input rain frame. The second stage employs adversarial learning to refine the result, i.e., recovering the overall color and illumination distributions of the frame, the background details that are failed to be recovered in the first stage, and removing the artifacts generated in the first stage. Furthermore, we also introduce a more comprehensive rain model that includes degradation factors, e.g., occlusion and rain accumulation, which appear in real scenes yet ignored by existing methods. This model, which generates more realistic rain images, will train and evaluate our models better. Extensive evaluations on synthetic and real videos show the effectiveness of our method in comparisons to the state-of-the-art methods. Our datasets, results and code are available at: https://github.com/flyywh/Recurrent-Multi-Frame-Deraining. Wenhan Yang, Robby T. Tan, Jiashi Feng, Shiqi Wang 0001, Bin Cheng 0001, Jiaying Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | Shape-Matching GAN++: Scale Controllable Dynamic Artistic Text Style TransferabstractDynamic artistic text style transfer aims to migrate the style in terms of both the appearance and motion patterns from a reference style video to the target text to create artistic text animation. Recent researches have improved the usability of transfer models by introducing texture control. However, it remains an important open challenge to investigate the control of the stylistic degree with respect to shape deformation. In this paper, we explore a new problem of dynamic artistic text style transfer with glyph stylistic degree control. The key idea is to build multi-scale glyph-style shape mappings through a novel bidirectional shape matching framework. Following this idea, we first introduce a scale-ware Shape-Matching GAN to learn such mappings to simultaneously model the style shape features at multiple scales and transfer them onto the target glyph. Furthermore, an advanced Shape-Matching GAN++ is proposed to animate a static text image based on the reference style video. Our Shape-Matching GAN++ characterizes the short-term consistency of motion patterns via shape matchings within consecutive frames, which are propagated to achieve effective long-term consistency. Experiments show that the proposed method outperforms previous state-of-the-arts both qualitatively and quantitatively, and generate high-quality and controllable artistic text. Shuai Yang 0001, Zhangyang Wang, Jiaying Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Towards Low Light Enhancement With RAW ImagesabstractIn this paper, we make the first benchmark effort to elaborate on the superiority of using RAW images in the low light enhancement and develop a novel alternative route to utilize RAW images in a more flexible and practical way. Inspired by a full consideration on the typical image processing pipeline, we are inspired to develop a new evaluation framework, Factorized Enhancement Model (FEM), which decomposes the properties of RAW images into measurable factors and provides a tool for exploring how properties of RAW images affect the enhancement performance empirically. The empirical benchmark results show that the Linearity of data and Exposure Time recorded in meta-data play the most critical role, which brings distinct performance gains in various measures over the approaches taking the sRGB images as input. With the insights obtained from the benchmark results in mind, a RAW-guiding Exposure Enhancement Network (REENet) is developed, which makes trade-offs between the advantages and inaccessibility of RAW images in real applications in a way of using RAW images only in the training phase. REENet projects sRGB images into linear RAW domains to apply constraints with corresponding RAW images to reduce the difficulty of modeling training. After that, in the testing phase, our REENet does not rely on RAW images. Experimental results demonstrate not only the superiority of REENet to state-of-the-art sRGB-based methods and but also the effectiveness of the RAW guidance and all components. Haofeng Huang, Wenhan Yang, Yueyu Hu, Jiaying Liu 0001, Ling-Yu Duan |
IEEE Trans. Image Process. | 4 |
| 2022 | CLAST: Contrastive Learning for Arbitrary Style TransferabstractArbitrary style transfer aims at migrating the style of a reference style painting to a target content image. Existing methods find it challenging to achieve good content fidelity and style migration at the same time. Moreover, they all rely on manually defined content and style, which is of limited universality and robustness. In this paper, we propose to introduce contrastive learning into style transfer, instructing the network to automatically learn to model the structural content and artistic style based on natural contrastive relationships in style transfer. Compared with existing methods, our learned modeling of content and style is more robust and universal. In addition, we further propose instance-wise contrastive style losses and a patch-wise contrastive content loss to guide style transfer. Combining the proposed contrastive losses and two self-reconstruction strategies, we develop a new style transfer framework, which is pluggable and can be flexibly applied to various style transfer modules. Experimental results demonstrate that our method has strong flexibility and synthesizes stylized images with higher quality. Wenjing Wang 0001, Shuai Yang 0001, Jiaying Liu 0001 |
IEEE Trans. Image Process. | 4 |
| 2022 | Recurrent Exposure Generation for Low-Light Face DetectionabstractFace detection from low-light images is challenging due to limited photons and inevitable noise, which, to make the task even harder, are often spatially unevenly distributed. A natural solution is to borrow the idea frommulti-exposure, which captures multiple shots to obtain well-exposed images under challenging conditions. High-quality implementation/approximation of multi-exposure from a single image is however nontrivial. Fortunately, as shown in this paper, neither is such high-quality necessary since our task isface detectionrather thanimage enhancement. Specifically, we propose a novelRecurrent Exposure Generation (REG)module and couple it seamlessly with aMulti-Exposure Detection (MED)module, and thus significantly improve face detection performance by effectively inhibiting non-uniform illumination and noise issues. REG produces progressively and efficiently intermediate images corresponding to various exposure settings, and such pseudo-exposures are then fused by MED to detect faces across different lighting conditions. The proposed method, namedREGDet, is the first ‘detection-with-enhancement’ framework for low-light face detection. It not only encourages rich interaction and feature fusion across different illumination levels, but also enables effective end-to-end learning of the REG component to be better tailored for face detection. Moreover, as clearly shown in our experiments, REG can be flexibly coupled with different face detectors without extra low/normal-light image pairs for training. We tested REGDet on the DARK FACE low-light face benchmark with thorough ablation study, where REGDet outperforms previous state-of-the-arts by a significant margin, with only negligible extra parameters. Jinxiu Liang, Jingwen Wang 0003, Yuhui Quan, Jiaying Liu 0001, Haibin Ling, Yong Xu 0007 |
IEEE Trans. Multim. | 5 |
| 2022 | Learning to Recognize Human Actions From Noisy Skeleton Data Via Noise AdaptationabstractRecent studies have made great progress on skeleton-based action recognition. However, most of them are developed with relatively clean skeletons without the presence of intensive noise. We argue that the models learned from relatively clean data are not well generalizable to handle noisy skeletons commonly appeared in the real world. In this paper, we address the challenge of recognizing human actions from noisy skeletons, which is seldom explored by previous methods. Beyond exploring the new problem, we further take a new perspective to address it, \textit{i.e.}, noise adaptation, which gets rid of explicit skeleton noise modeling and reliance on skeleton ground truths. Specifically, we develop regression-based and generation-based adaptation models according to whether pairs of noisy skeletons are available. The regression-based model aims to learn noise-suppressed intrinsic feature representations by mapping pairs of noisy skeletons into a noise-robust space. When only unpaired skeletons are accessible, the generation-based model aims to adapt the features from noisy skeletons to a low-noise space by adversarial learning. To verify our proposed model and facilitate research on noisy skeletons, we collect a new dataset Noisy Skeleton Dataset (NSD), the skeletons of which are with much noise and more similar to daily-life data than previous datasets. Extensive experiments are conducted on the NSD, VV-RGBD and N-UCLA datasets, and results consistently show the outstanding performance of our proposed model. Sijie Song, Jiaying Liu 0001, Lilang Lin, Zongming Guo |
IEEE Trans. Multim. | 2 |
| 2022 | Template-Free Try-On Image Synthesis via Semantic-Guided OptimizationabstractThe virtual try-on task is so attractive that it has drawn considerable attention in the field of computer vision. However, presenting the 3-D physical characteristic (e.g., pleat and shadow) based on a 2-D image is very challenging. Although there have been several previous studies on 2-D-based virtual try-on work, most: 1) required user-specified target poses that are not user-friendly and may not be the best for the target clothing and 2) failed to address some problematic cases, including facial details, clothing wrinkles, and body occlusions. To address these two challenges, in this article, we propose an innovative template-free try-on image synthesis (TF-TIS) network. The TF-TIS first synthesizes the target pose according to the user-specified in-shop clothing. Afterward, given an in-shop clothing image, a user image, and a synthesized pose, we propose a novel model for synthesizing a human try-on image with the target clothing in the best fitting pose. The qualitative and quantitative experiments both indicate that the proposed TF-TIS outperforms the state-of-the-art methods, especially for difficult cases. Chien-Lung Chou, Chieh-Yun Chen, Chia-Wei Hsieh, Hong-Han Shuai, Jiaying Liu 0001, Wen-Huang Cheng |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2021 | Co-Grounding Networks With Semantic Attention for Referring Expression Comprehension in VideosabstractIn this paper, we address the problem of referring expression comprehension in videos, which is challenging due to complex expression and scene dynamics. Unlike previous methods which solve the problem in multiple stages (i.e., tracking, proposal-based matching), we tackle the problem from a novel perspective, co-grounding, with an elegant one-stage framework. We enhance the single-frame grounding accuracy by semantic attention learning and improve the cross-frame grounding consistency with co-grounding feature learning. Semantic attention learning explicitly parses referring cues in different attributes to reduce the ambiguity in the complex expression. Co-grounding feature learning boosts visual feature representations by integrating temporal correlation to reduce the ambiguity caused by scene dynamics. Experiment results demonstrate the superiority of our framework on the video grounding datasets VID and LiOTB in generating accurate and stable results across frames. Our model is also applicable to referring expression comprehension in images, illustrated by the improved performance on the RefCOCO dataset. Our project is available at https://sijiesong.github.io/co-grounding. Sijie Song, Xudong Lin 0003, Jiaying Liu 0001, Zongming Guo, Shih-Fu Chang |
CVPR | 3 |
| 2021 | HLA-Face: Joint High-Low Adaptation for Low Light Face DetectionabstractFace detection in low light scenarios is challenging but vital to many practical applications, e.g., surveillance video, autonomous driving at night. Most existing face detectors heavily rely on extensive annotations, while collecting data is time-consuming and laborious. To reduce the burden of building new datasets for low light conditions, we make full use of existing normal light data and explore how to adapt face detectors from normal light to low light. The challenge of this task is that the gap between normal and low light is too huge and complex for both pixel-level and object-level. Therefore, most existing low-light enhancement and adaptation methods do not achieve desirable performance. To address the issue, we propose a joint High-Low Adaptation (HLA) framework. Through a bidirectional low-level adaptation and multi-task high-level adaptation scheme, our HLA-Face outperforms state-of-the-art methods even without using dark face labels for training. Our project is publicly available at: https://daooshee.github.io/HLA-Face-Website/. Wenjing Wang 0001, Wenhan Yang, Jiaying Liu 0001 |
CVPR | 3 |
| 2021 | Semi-Supervised Learning for Mars Imagery ClassificationabstractWith the progress of Mars exploration, numerous Mars image data are collected and need to be analyzed. However, because of the imbalance and distortion in Mars data, the performance of existing classification models is unsatisfactory. In this paper, we design a new framework based on semi-supervised contrastive learning for Mars rover image classification. The redundancy of Mars data can disable the effectiveness of contrastive learning. To strip out problematic learning samples, we propose to ignore inner-class pairs on labeled data as well as neglect negative pairs on unlabeled data. Experimental results show that our learning strategies can improve the classification model by a large margin and outperform state-of-the-art methods. Wenjing Wang 0001, Lilang Lin, Zejia Fan, Jiaying Liu 0001 |
ICIP | 4 |
| 2021 | Instance-Aware Coherent Video Style Transfer for Chinese Ink Wash PaintingabstractRecent researches have made remarkable achievements in fast video style transfer based on western paintings. However, due to the inherent different drawing techniques and aesthetic expressions of Chinese ink wash painting, existing methods either achieve poor temporal consistency or fail to transfer the key freehand brushstroke characteristics of Chinese ink wash painting. In this paper, we present a novel video style transfer framework for Chinese ink wash paintings. The two key ideas are a multi-frame fusion for temporal coherence and an instance-aware style transfer. The frame reordering and stylization based on reference frame fusion are proposed to improve temporal consistency. Meanwhile, the proposed method is able to adaptively leave the white spaces in the background and to select proper scales to extract features and depict the foreground subject by leveraging instance segmentation. Experimental results demonstrate the superiority of the proposed method over state-of-the-art style transfer methods in terms of both temporal coherence and visual quality. Our project website is available at https://oblivioussy.github.io/InkVideo/. Shuai Yang 0001, Wenjing Wang 0001, Jiaying Liu 0001 |
IJCAI | 4 |
| 2021 | Edit Like A Designer: Modeling Design Workflows for Unaligned Fashion EditingabstractFashion editing has drawn increasing research interest with its extensive application prospect. Instead of directly manipulating the real fashion item image, it is more intuitive for designers to modify it via the design draft. In this paper, we model design workflows for a novel task of unaligned fashion editing, allowing the user to edit a fashion item through manipulating its corresponding design draft. The challenge lies in the large misalignment between the real fashion item and the design draft, which could severely degrade the quality of editing results. To address this issue, we propose an Unaligned Fashion Editing Network (UFE-Net). A coarsely rendered fashion item is firstly generated from the edited design draft via a translation module. With this as guidance, we align and manipulate the original unedited fashion item via a novel alignment-driven fashion editing module, and then optimize the details and shape via a reference-guided refinement module. Furthermore, a joint training strategy is introduced to exploit the synergy between the alignment and editing tasks. Our UFE-Net enables the edited fashion item to have semantically consistent geometric shape and realistic details to the edited draft in the edited region, as well as to keep the unedited region intact. Experiments demonstrate our superiority over the competing methods on unaligned fashion editing. Qiyu Dai, Shuai Yang 0001, Wenjing Wang 0001, Jiaying Liu 0001 |
ACM Multimedia | 5 |
| 2021 | Benchmarking Low-Light Image Enhancement and Beyond
Jiaying Liu 0001, Dejia Xu, Wenhan Yang, Minhao Fan, Haofeng Huang |
Int. J. Comput. Vis. | 1 |
| 2021 | Mask-guided GAN for robust text editing in the scene
Boxi Yu, Yong Xu 0007, Shuai Yang 0001, Jiaying Liu 0001 |
Neurocomputing | 5 |
| 2021 | TE141K: Artistic Text Benchmark for Text Effect TransferabstractText effects are combinations of visual elements such as outlines, colors and textures of text, which can dramatically improve its artistry. Although text effects are extensively utilized in the design industry, they are usually created by human experts due to their extreme complexity; this is laborious and not practical for normal users. In recent years, some efforts have been made toward automatic text effect transfer; however, the lack of data limits the capabilities of transfer models. To address this problem, we introduce a new text effects dataset, TE141K,11.Project page:https://daooshee.github.io/TE141K/.with 141,081 text effect/glyph pairs in total. Our dataset consists of 152 professionally designed text effects rendered on glyphs, including English letters, Chinese characters, and Arabic numerals. To the best of our knowledge, this is the largest dataset for text effect transfer to date. Based on this dataset, we propose a baseline approach called text effect transfer GAN (TET-GAN), which supports the transfer of all 152 styles in one model and can efficiently extend to new styles. Finally, we conduct a comprehensive comparison in which 14 style transfer models are benchmarked. Experimental results demonstrate the superiority of TET-GAN both qualitatively and quantitatively and indicate that our dataset is effective and challenging. Shuai Yang 0001, Wenjing Wang 0001, Jiaying Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Unpaired Person Image Generation With Semantic Parsing TransformationabstractIn this paper, we tackle the problem of pose-guided person image generation with unpaired data, which is a challenging problem due to non-rigid spatial deformation. Instead of learning a fixed mapping directly between human bodies as previous methods, we propose a new pathway to decompose a single fixed mapping into two subtasks, namely, semantic parsing transformation and appearance generation. First, to simplify the learning for non-rigid deformation, a semantic generative network is developed to transform semantic parsing maps between different poses. Second, guided by semantic parsing maps, we render the foreground and background image, respectively. A foreground generative network learns to synthesize semantic-aware textures, and another background generative network learns to predict missing background regions caused by pose changes. Third, we enable pseudo-label training with unpaired data, and demonstrate that end-to-end training of the overall network further refines the semantic map prediction and final results accordingly. Moreover, our method is generalizable to other person image generation tasks defined on semantic maps, e.g., clothing texture transfer, controlled image manipulation, and virtual try-on. Experimental results on DeepFashion and Market-1501 datasets demonstrate the superiority of our method, especially in keeping better body shapes and clothing attributes, as well as rendering structure-coherent backgrounds. Sijie Song, Wei Zhang 0031, Jiaying Liu 0001, Zongming Guo, Tao Mei 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Bridging the Gap Between Computational Photography and Visual RecognitionabstractWhat is the current state-of-the-art for image restoration and enhancement applied to degraded images acquired under less than ideal circumstances? Can the application of such algorithms as a pre-processing step improve image interpretability for manual analysis or automatic visual recognition to classify scene content? While there have been important advances in the area of computational photography to restore or enhance the visual quality of an image, the capabilities of such techniques have not always translated in a useful way to visual recognition tasks. Consequently, there is a pressing need for the development of algorithms that are designed for the joint problem of improving visual appearance and recognition, which will be an enabling factor for the deployment of visual recognition tools in many real-world scenarios. To address this, we introduce the UG$^2$dataset as a large-scale benchmark composed of video imagery captured under challenging conditions, and two enhancement tasks designed to test algorithmic impact on visual quality and automatic object recognition. Furthermore, we propose a set of metrics to evaluate the joint improvement of such tasks as well as individual algorithmic advances, including a novel psychophysics-based evaluation regime for human assessment and a realistic set of quantitative measures for object recognition performance. We introduce six new algorithms for image restoration or enhancement, which were created as part of the IARPA sponsored UG$^2$Challenge workshop held at CVPR 2018. Under the proposed evaluation regime, we present an in-depth analysis of these algorithms and a host of deep learning-based and classic baseline approaches. From the observed results, it is evident that we are in the early days of building a bridge between computational photography and visual recognition, leaving many opportunities for innovation in this area. Rosaura G. VidalMata, Sreya Banerjee, Brandon RichardWebster, Michael Albright, Pedro Davalos, Scott McCloskey, Ben Miller, Asong Tambo, Sushobhan Ghosh, Sudarshan Nagesh, Ye Yuan 0012, Yueyu Hu, Wenhan Yang, Xiaoshuai Zhang, Jiaying Liu 0001, Zhangyang Wang, Hwann-Tzong Chen, Tzu-Wei Huang, Wen-Chi Chin, Yi-Chun Li, Mahmoud Lababidi, Charles Otto, Walter J. Scheirer |
IEEE Trans. Pattern Anal. Mach. Intell. | 16 |
| 2021 | Single Image Deraining: From Model-Based to Data-Driven and BeyondabstractThe goal of single-image deraining is to restore the rain-free background scenes of an image degraded by rain streaks and rain accumulation. The early single-image deraining methods employ a cost function, where various priors are developed to represent the properties of rain and background layers. Since 2017, single-image deraining methods step into a deep-learning era, and exploit various types of networks, i.e., convolutional neural networks, recurrent neural networks, generative adversarial networks, etc., demonstrating impressive performance. Given the current rapid development, in this paper, we provide a comprehensive survey of deraining methods over the last decade. We summarize the rain appearance models, and discuss two categories of deraining approaches: model-based and data-driven approaches. For the former, we organize the literature based on their basic models and priors. For the latter, we discuss the developed ideas related to architectures, constraints, loss functions, and training datasets. We present milestones of single-image deraining methods, review a broad selection of previous works in different categories, and provide insights on the historical development route from the model-based to data-driven methods. We also summarize performance comparisons quantitatively and qualitatively. Beyond discussing the technicality of deraining methods, we also discuss the future possible directions. Wenhan Yang, Robby T. Tan, Shiqi Wang 0001, Yuming Fang 0001, Jiaying Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2021 | Combining Progressive Rethinking and Collaborative Learning: A Deep Framework for In-Loop FilteringabstractIn this paper, we aim to address issues of (1) joint spatial-temporal modeling and (2) side information injection for deep-learning based in-loop filter. For (1), we design a deep network with both progressive rethinking and collaborative learning mechanisms to improve quality of the reconstructed intra-frames and inter-frames, respectively. For intra coding, a Progressive Rethinking Network (PRN) is designed to simulate the human decision mechanism for effective spatial modeling. Our designed block introduces an additional inter-block connection to bypass a high-dimensional informative feature before the bottleneck module across blocks to review the complete past memorized experiences and rethinks progressively. For inter coding, the current reconstructed frame interacts with reference frames (peak quality frame and the nearest adjacent frame) collaboratively at the feature level. For (2), we extract both intra-frame and inter-frame side information for better context modeling. A coarse-to-fine partition map based on HEVC partition trees is built as the intra-frame side information. Furthermore, the warped features of the reference frames are offered as the inter-frame side information. Our PRN with intra-frame side information provides 9.0% BD-rate reduction on average compared to HEVC baseline under All-intra (AI) configuration. While under Low-Delay B (LDB), Low-Delay P (LDP) and Random Access (RA) configuration, our PRN with inter-frame side information provides 9.0%, 10.6% and 8.0% BD-rate reduction on average respectively. Our project webpage is https://dezhao-wang.github.io/PRN-v2/. Dezhao Wang, Sifeng Xia, Wenhan Yang, Jiaying Liu 0001 |
IEEE Trans. Image Process. | 4 |
| 2021 | Band Representation-Based Semi-Supervised Low-Light Image Enhancement: Bridging the Gap Between Signal Fidelity and Perceptual QualityabstractIt has been widely acknowledged that under-exposure causes a variety of visual quality degradation because of intensive noise, decreased visibility, biased color, etc. To alleviate these issues, a novel semi-supervised learning approach is proposed in this paper for low-light image enhancement. More specifically, we propose a deep recursive band network (DRBN) to recover a linear band representation of an enhanced normal-light image based on the guidance of the paired low/normal-light images. Such design philosophy enables the principled network to generate a quality improved one by reconstructing the given bands based upon another learnable linear transformation which is perceptually driven by an image quality assessment neural network. On one hand, the proposed network is delicately developed to obtain a variety of coarse-to-fine band representations, of which the estimations benefit each other in a recursive process mutually. On the other hand, the extracted band representation of the enhanced image in the recursive band learning stage of DRBN is capable of bridging the gap between the restoration knowledge of paired data and the perceptual quality preference to high-quality images. Subsequently, the band recomposition learns to recompose the band representation towards fitting perceptual regularization of high-quality images with the perceptual guidance. The proposed architecture can be flexibly trained with both paired and unpaired data. Extensive experiments demonstrate that our method produces better enhanced results with visually pleasing contrast and color distributions, as well as well-restored structural details. Wenhan Yang, Shiqi Wang 0001, Yuming Fang 0001, Yue Wang 0032, Jiaying Liu 0001 |
IEEE Trans. Image Process. | 5 |
| 2021 | Sparse Gradient Regularized Deep Retinex Network for Robust Low-Light Image EnhancementabstractDue to the absence of a desirable objective for low-light image enhancement, previous data-driven methods may provide undesirable enhanced results including amplified noise, degraded contrast and biased colors. In this work, inspired by Retinex theory, we design an end-to-end signal prior-guided layer separation and data-driven mapping network with layer-specified constraints for single-image low-light enhancement. A Sparse Gradient Minimization sub-Network (SGM-Net) is constructed to remove the low-amplitude structures and preserve major edge information, which facilitates extracting paired illumination maps of low/normal-light images. After the learned decomposition, two sub-networks (Enhance-Net and Restore-Net) are utilized to predict the enhanced illumination and reflectance maps, respectively, which helps stretch the contrast of the illumination map and remove intensive noise in the reflectance map. The effects of all these configured constraints, including the signal structure regularization and losses, combine together reciprocally, which leads to good reconstruction results in overall visual quality. The evaluation on both synthetic and real images, particularly on those containing intensive noise, compression artifacts and their interleaved artifacts, shows the effectiveness of our novel models, which significantly outperforms the state-of-the-art methods. Wenhan Yang, Wenjing Wang 0001, Haofeng Huang, Shiqi Wang 0001, Jiaying Liu 0001 |
IEEE Trans. Image Process. | 5 |
| 2021 | Controllable Sketch-to-Image Translation for Robust Face SynthesisabstractIn this paper, we propose a novel controllable sketch-to-image translation framework that allows users to interactively and robustly synthesize and edit face images with hand-drawn sketches. Inspired by the coarse-to-fine painting process of human artists, we propose a novel dilation-based sketch refinement method to refine sketches at varied coarse levels without the need for real sketch training data. We further investigate multi-level refinement that enables users to flexibly define how "reliable" the input sketch should be considered for the final output through a refinement level control parameter, which helps balance between the realism of the output and its structural consistency with the input sketch. It is realized by leveraging scale-aware style transfer to model and adjust the style features of sketches at different coarse levels. Moreover, advanced user controllability in terms of the editing region control, facial attribute editing, and spatially non-uniform refinement is further explored for fine-grained and semantic editing. We demonstrate the effectiveness of the proposed method in terms of visual quality and user controllability through extensive experiments including qualitative and quantitative comparison with state-of-the-art methods, ablation studies and various applications. Shuai Yang 0001, Zhangyang Wang, Jiaying Liu 0001, Zongming Guo |
IEEE Trans. Image Process. | 3 |
| 2021 | Towards Coding for Human and Machine Vision: Scalable Face Image CodingabstractThe past decades have witnessed the rapid development of image and video coding techniques in the era of big data. However, the signal fidelity-driven coding pipeline design limits the capability of the existing image/video coding frameworks to fulfill the needs of both machine and human vision. In this paper, we come up with a novel face image coding framework by leveraging both the compressive and the generative models, to support machine vision and human perception tasks jointly. Given an input image, the feature analysis is first applied, and then the generative model is employed to reconstruct image with compact structure and color features, where sparse edges are extracted to connect both kinds of vision and a key reference pixel selection method is proposed to determine the priorities of the reference color pixels for scalable coding. The compact edge map serves as the basic layer for machine vision tasks, and the reference pixels act as an enhanced layer to guarantee signal fidelity for human vision. By introducing advanced generative models, we train a decoding network to reconstruct images from compact structure and color representations, which is flexible to accept inputs in a scalable way and to control the imagery effect of the outputs between signal fidelity and visual realism. Experimental results and comprehensive performance analysis over the face image dataset demonstrate the superiority of our framework in both human vision tasks and machine vision tasks, which provide useful evidence on the emerging standardization efforts on MPEG VCM (Video Coding for Machine). Shuai Yang 0001, Yueyu Hu, Wenhan Yang, Ling-Yu Duan, Jiaying Liu 0001 |
IEEE Trans. Multim. | 5 |
| 2021 | Introduction to the Special Issue on Explainable AI on Multimedia ComputingabstractNo abstract available. Wen-Huang Cheng, Jiaying Liu 0001, Nicu Sebe, Junsong Yuan 0001, Hong-Han Shuai |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2020 | Coarse-to-Fine Hyper-Prior Modeling for Learned Image CompressionabstractApproaches to image compression with machine learning now achieve superior performance on the compression rate compared to existing hybrid codecs. The conventional learning-based methods for image compression exploits hyper-prior and spatial context model to facilitate probability estimations. Such models have limitations in modeling long-term dependency and do not fully squeeze out the spatial redundancy in images. In this paper, we propose a coarse-to-fine framework with hierarchical layers of hyper-priors to conduct comprehensive analysis of the image and more effectively reduce spatial redundancy, which improves the rate-distortion performance of image compression significantly. Signal Preserving Hyper Transforms are designed to achieve an in-depth analysis of the latent representation and the Information Aggregation Reconstruction sub-network is proposed to maximally utilize side-information for reconstruction. Experimental results show the effectiveness of the proposed network to efficiently reduce the redundancies in images and improve the rate-distortion performance, especially for high-resolution images. Our project is publicly available at https://huzi96.github.io/coarse-to-fine-compression.html. Yueyu Hu, Wenhan Yang, Jiaying Liu 0001 |
AAAI | 3 |
| 2020 | Consistent Video Style Transfer via Compound RegularizationabstractRecently, neural style transfer has drawn many attentions and significant progresses have been made, especially for image style transfer. However, flexible and consistent style transfer for videos remains a challenging problem. Existing training strategies, either using a significant amount of video data with optical flows or introducing single-frame regularizers, have limited performance on real videos. In this paper, we propose a novel interpretation of temporal consistency, based on which we analyze the drawbacks of existing training strategies; and then derive a new compound regularization. Experimental results show that the proposed regularization can better balance the spatial and temporal performance, which supports our modeling. Combining with the new cost formula, we design a zero-shot video style transfer framework. Moreover, for better feature migration, we introduce a new module to dynamically adjust inter-channel distributions. Quantitative and qualitative results demonstrate the superiority of our method over other state-of-the-art style transfer methods. Our project is publicly available at: https://daooshee.github.io/CompoundVST/. Wenjing Wang 0001, Jizheng Xu, Li Zhang 0006, Yue Wang 0032, Jiaying Liu 0001 |
AAAI | 5 |
| 2020 | Towards Scale-Free Rain Streak Removal via Self-Supervised Fractal Band LearningabstractData-driven rain streak removal methods, which most of rely on synthesized paired data, usually come across the generalization problem when being applied in real cases. In this paper, we propose a novel deep-learning based rain streak removal method injected with self-supervision to improve the ability to remove rain streaks in various scales. To realize this goal, we made efforts in two aspects. First, considering that rain streak removal is highly correlated with texture characteristics, we create a fractal band learning (FBL) network based on frequency band recovery. It integrates commonly seen band feature operations with neural modules and effectively improves the capacity to capture discriminative features for deraining. Second, to further improve the generalization ability of FBL for rain streaks in various scales, we add cross-scale self-supervision to regularize the network training. The constraint forces the extracted features of inputs in different scales to be equivalent after rescaling. Therefore, FBL can offer similar responses based on solely image content without the interleave of scale and is capable to remove rain streaks in various scales. Extensive experiments in quantitative and qualitative evaluations demonstrate the superiority of our FBL for rain streak removal, especially for the real cases where very large rain streaks exist, and prove the effectiveness of its each component. Our code will be public available at: https://github.com/flyywh/AAAI-2020-FBL-SS. Wenhan Yang, Shiqi Wang 0001, Dejia Xu, Jiaying Liu 0001 |
AAAI | 5 |
| 2020 | Spherical Criteria for Fast and Accurate 360° Object DetectionabstractWith the advance of omnidirectional panoramic technology, 360◦ imagery has become increasingly popular in the past few years. To better understand the 360◦ content, many works resort to the 360◦ object detection and various criteria have been proposed to bound the objects and compute the intersection-over-union (IoU) between bounding boxes based on the common equirectangular projection (ERP) or perspective projection (PSP). However, the existing 360◦ criteria are either inaccurate or inefficient for real-world scenarios. In this paper, we introduce a novel spherical criteria for fast and accurate 360◦ object detection, including both spherical bounding boxes and spherical IoU (SphIoU). Based on the spherical criteria, we propose a novel two-stage 360◦ detector, i.e., Reprojection R-CNN, by combining the advantages of both ERP and PSP, yielding efficient and accurate 360◦ object detection. To validate the design of spherical criteria and Reprojection R-CNN, we construct two unbiased synthetic datasets for training and evaluation. Experimental results reveal that compared with the existing criteria, the two-stage detector with spherical criteria achieves the best mAP results under the same inference speed, demonstrating that the spherical criteria can be more suitable for 360◦ object detection. Moreover, Reprojection R-CNN outperforms the previous state-of-the-art methods by over 30% on mAP with competitive speed, which confirms the efficiency and accuracy of the design. Ansheng You, Yuanxing Zhang, Jiaying Liu 0001, Kaigui Bian, Yunhai Tong |
AAAI | 4 |
| 2020 | Raw-Guided Enhancing Reprocess of Low-Light Image via Deep Exposure Adjustment
Haofeng Huang, Wenhan Yang, Yueyu Hu, Jiaying Liu 0001 |
ACCV (2) | 4 |
| 2020 | From Fidelity to Perceptual Quality: A Semi-Supervised Approach for Low-Light Image EnhancementabstractUnder-exposure introduces a series of visual degradation, i.e. decreased visibility, intensive noise, and biased color, etc. To address these problems, we propose a novel semi-supervised learning approach for low-light image enhancement. A deep recursive band network (DRBN) is proposed to recover a linear band representation of an enhanced normal-light image with paired low/normal-light images, and then obtain an improved one by recomposing the given bands via another learnable linear transformation based on a perceptual quality-driven adversarial learning with unpaired data. The architecture is powerful and flexible to have the merit of training with both paired and unpaired data. On one hand, the proposed network is well designed to extract a series of coarse-to-fine band representations, whose estimations are mutually beneficial in a recursive process. On the other hand, the extracted band representation of the enhanced image in the first stage of DRBN (recursive band learning) bridges the gap between the restoration knowledge of paired data and the perceptual quality preference to real high-quality images. Its second stage (band recomposition) learns to recompose the band representation towards fitting perceptual properties of high-quality images via adversarial learning. With the help of this two-stage design, our approach generates enhanced results with well-reconstructed details and visually promising contrast and color distributions. Qualitative and quantitative evaluations demonstrate the superiority of our DRBN. Wenhan Yang, Shiqi Wang 0001, Yuming Fang 0001, Yue Wang 0032, Jiaying Liu 0001 |
CVPR | 5 |
| 2020 | Self-Learning Video Rain Streak Removal: When Cyclic Consistency Meets Temporal CorrespondenceabstractIn this paper, we address the problem of rain streaks removal in video by developing a self-learned rain streak removal method, which does not require any clean groundtruth images in the training process. The method is inspired by fact that the adjacent frames are highly correlated and can be regarded as different versions of identical scene, and rain streaks are randomly distributed along the temporal dimension. With this in mind, we construct a two-stage Self-Learned Deraining Network (SLDNet) to remove rain streaks based on both temporal correlation and consistency. In the first stage, SLDNet utilizes the temporal correlations and learns to predict the clean version of the current frame based on its adjacent rain video frames. In the second stage, SLDNet enforces the temporal consistency among different frames. It takes both the current rain frame and adjacent rain video frames to recover structural details. The first stage is responsible for reconstructing main structures, and the second stage is responsible for extracting structural details. We build our network architecture with two sub-tasks, i.e. motion estimation, and rain region detection, and optimize them jointly. Our extensive experiments demonstrate the effectiveness of our method, offering better results both quantitatively and qualitatively. Wenhan Yang, Robby T. Tan, Shiqi Wang 0001, Jiaying Liu 0001 |
CVPR | 4 |
| 2020 | Deep Plastic Surgery: Robust and Controllable Image Editing with Human-Drawn Sketches
Shuai Yang 0001, Zhangyang Wang, Jiaying Liu 0001, Zongming Guo |
ECCV (15) | 3 |
| 2020 | Towards Coding For Human And Machine Vision: A Scalable Image Coding ApproachabstractThe past decades have witnessed the rapid development of image and video coding techniques in the era of big data. However, the signal fidelity-driven coding pipeline design limits the capability of the existing image/video coding frameworks to fulfill the needs of both machine and human vision. In this paper, we come up with a novel image coding framework by leveraging both the compressive and the generative models, to support machine vision and human perception tasks jointly. Given an input image, the feature analysis is first applied, and then the generative model is employed to perform image reconstruction with features and additional reference pixels, in which compact edge maps are extracted in this work to connect both kinds of vision in a scalable way. The compact edge map serves as the basic layer for machine vision tasks, and the reference pixels act as a sort of enhanced layer to guarantee signal fidelity for human vision. By introducing advanced generative models, we train a flexible network to reconstruct images from compact feature representations and the reference pixels. Experimental results demonstrate the superiority of our framework in both human visual quality and facial landmark detection, which provide useful evidence on the emerging standardization efforts on MPEG VCM (Video Coding for Machine). Our project website is available at https://williamyang1991.github.io/projects/VCM-Face/. Yueyu Hu, Shuai Yang 0001, Wenhan Yang, Ling-Yu Duan, Jiaying Liu 0001 |
ICME | 5 |
| 2020 | An Emerging Coding Paradigm Vcm: A Scalable Coding Approach Beyond Feature And SignalabstractIn this paper, we study a new problem arising from the emerging MPEG standardization effort Video Coding for Machine (VCM)1, which aims to bridge the gap between visual feature compression and classical video coding. VCM is committed to address the requirement of compact signal representation for both machine and human vision in a more or less scalable way. To this end, we make endeavors in leveraging the strength of predictive and generative models to support advanced compression techniques for both machine and human vision tasks simultaneously, in which visual features serve as a bridge to connect signal-level and task-level compact representations in a scalable manner. Specifically, we employ a conditional deep generation network to reconstruct video frames with the guidance of learned motion pattern. By learning to extract sparse motion pattern via a predictive model, the network elegantly leverages the feature representation to generate the appearance of to-be-coded frames via a generative model, relying on the appearance of the coded key frames. Meanwhile, the sparse motion pattern is compact and highly effective for high-level vision tasks, e.g. action recognition. Experimental results demonstrate that our method yields much better reconstruction quality compared with the traditional video codecs (0.0063 gain in SSIM), as well as state-of-the-art action recognition performance over highly compressed videos (9.4% gain in recognition accuracy), which showcases a promising paradigm of coding signal for both human and machine vision. Sifeng Xia, Kunchangtai Liang, Wenhan Yang, Ling-Yu Duan, Jiaying Liu 0001 |
ICME | 5 |
| 2020 | Multitask Attentive Network For Text Effects Quality AssessmentabstractAlong with the fast development of image style transfer, large amounts of style transfer algorithms were proposed. However, not enough attention has been paid to assess the quality of stylized images, which is of great value in allowing users to efficiently search for high quality images as well as guiding the designing of style transfer algorithms. In this paper, we focus on artistic text stylization and build a novel deep neural network equipped with multitask learning and attention mechanism for text effects quality assessment. We first select stylized images from TE141K [1] dataset and then collect the corresponding visual scores from users. Then through multitask learning, the network learns to extract features related to both style and content information. Furthermore, we employ an attention module to simulate the process of human high-level visual judgement. Experimental results demonstrate the superiority of our network in achieving a high judgement accuracy over the state-of-the-art methods. Our project website is available at https://ykq98.github.io/projects/TEA/. Keqiang Yan, Shuai Yang 0001, Wenjing Wang 0001, Jiaying Liu 0001 |
ICME | 4 |
| 2020 | Memory-Augmented Auto-Regressive Network for Frame Recurrent Inter PredictionabstractInter prediction is quite important for the modern codecs to remove temporal redundancy. In this paper, we make endeavors in generating artificial reference frames with previous reconstructed frames for inter prediction, to offer a better choice when the traditional block-wise motion estimation fails to find a good reference block. Long-term temporal dynamics are tracked during the whole coding process to generate more accurate and realistic artificial reference frames. Specifically, we propose a Memory-Augmented Auto-Regressive Network (MAAR-Net) for frame prediction in video coding. MAAR-Net regresses the current frame with two nearest frames via an auto-regressive (AR) model to better capture the main spatial and temporal structures. The AR regression coefficients are generated based on adjacent frame information as well as the long-term motion dynamics accumulated and propagated by a convolutional Long Short-Term Memory (LSTM). To generate the target frame with higher quality, a quality attention mechanism is introduced for the temporal regularization between different reconstructed frames. With the well-designed network, our method surpasses HEVC on average 4.0% BD-rate saving and up to 10.6% BD-rate saving for the luma component under the low-delay configuration. Yuzhang Hu, Sifeng Xia, Wenhan Yang, Jiaying Liu 0001 |
ISCAS | 4 |
| 2020 | Coping with Pandemics: Opportunities and Challenges for AI Multimedia in the "New Normal"abstractTheworld iswelcoming the newnormal - the coronavirus pandemic has significantly changed the way people live, work, communicate and learn. Almost everyone now is wearing a face mask when they go in public. People are working from home, some taking care of children at the same time. Bars and restaurants are limited to carry-out and delivery only. Meetings and conferences go online. Schools are closed and educators are instead holding video conference classes regularly. All these become the new normal as our ways of life. The panel thus provides a valuable opportunity for people from a variety of backgrounds to exchange views on opportunities and challenges for AI multimedia in the current and post pandemics era. Jiaying Liu 0001, Wen-Huang Cheng, Klara Nahrstedt, Ramesh Jain 0001, Elisa Ricci 0001, Hyeran Byun |
ACM Multimedia | 1 |
| 2020 | Integrating Semantic Segmentation and Retinex Model for Low-Light Image EnhancementabstractRetinex model is widely adopted in various low-light image enhancement tasks. The basic idea of the Retinex theory is to decompose images into reflectance and illumination. The ill-posed decomposition is usually handled by hand-crafted constraints and priors. With the recently emerging deep-learning based approaches as tools, in this paper, we integrate the idea of Retinex decomposition and semantic information awareness. Based on the observation that various objects and backgrounds have different material, reflection and perspective attributes, regions of a single low-light image may require different adjustment and enhancement regarding contrast, illumination and noise. We propose an enhancement pipeline with three parts that effectively utilize the semantic layer information. Specifically, we extract the segmentation, reflectance as well as illumination layers, and concurrently enhance every separate region, i.e. sky, ground and objects for outdoor scenes. Extensive experiments on both synthetic data and real world images demonstrate the superiority of our method over current state-of-the-art low-light enhancement algorithms. Minhao Fan, Wenjing Wang 0001, Wenhan Yang, Jiaying Liu 0001 |
ACM Multimedia | 4 |
| 2020 | From Design Draft to Real Attire: Unaligned Fashion Image TranslationabstractFashion manipulation has attracted growing interest due to its great application value, which inspires many researches towards fashion images. However, little attention has been paid to fashion design draft. In this paper, we study a new unaligned translation problem between design drafts and real fashion items, whose main challenge lies in the huge misalignment between the two modalities. We first collect paired design drafts and real fashion item images without pixel-wise alignment. To solve the misalignment problem, our main idea is to train a sampling network to adaptively adjust the input to an intermediate state with structure alignment to the output. Moreover, built upon the sampling network, we present design draft to real fashion item translation network (D2RNet), where two separate translation streams that focus on texture and shape, respectively, are combined tactfully to get both benefits. D2RNet is able to generate realistic garments with both texture and shape consistency to their design drafts. We show that this idea can be effectively applied to the reverse translation problem and present R2DNet accordingly. Extensive experiments on unaligned fashion design translation demonstrate the superiority of our method over state-of-the-art methods. Our project website is available at: https://victoriahy.github.io/MM2020/. Yu Han 0008, Shuai Yang 0001, Wenjing Wang 0001, Jiaying Liu 0001 |
ACM Multimedia | 4 |
| 2020 | MS2L: Multi-Task Self-Supervised Learning for Skeleton Based Action RecognitionabstractIn this paper, we address self-supervised representation learning from human skeletons for action recognition. Previous methods, which usually learn feature presentations from a single reconstruction task, may come across the overfitting problem, and the features are not generalizable for action recognition. Instead, we propose to integrate multiple tasks to learn more general representations in a self-supervised manner. To realize this goal, we integrate motion prediction, jigsaw puzzle recognition, and contrastive learning to learn skeleton features from different aspects. Skeleton dynamics can be modeled through motion prediction by predicting the future sequence. And temporal patterns, which are critical for action recognition, are learned through solving jigsaw puzzles. We further regularize the feature space by contrastive learning. Besides, we explore different training strategies to utilize the knowledge from self-supervised tasks for action recognition. We evaluate our multi-task self-supervised learning approach with action classifiers trained under different configurations, including unsupervised, semi-supervised and fully-supervised settings. Our experiments on the NW-UCLA, NTU RGB+D, and PKUMMD datasets show remarkable performance for action recognition, demonstrating the superiority of our method in learning more discriminative and general features. Our project website is available at https://langlandslin.github.io/projects/MSL/. Lilang Lin, Sijie Song, Wenhan Yang, Jiaying Liu 0001 |
ACM Multimedia | 4 |
| 2020 | Sensitivity-Aware Bit Allocation for Intermediate Deep Feature CompressionabstractIn this paper, we focus on compressing and transmitting deep intermediate features to support the prosperous applications at the cloud side efficiently, and propose a sensitivity-aware bit allocation algorithm for the deep intermediate feature compression. Considering that different channels' contributions to the final inference result of the deep learning model might differ a lot, we design a channel-wise bit allocation mechanism to maintain the accuracy while trying to reduce the bit-rate cost. The algorithm consists of two passes. In the first pass, only one channel is exposed to compression degradation while other channels are kept as the original ones in order to test this channel's sensitivity to the compression degradation. This process will be repeated until all channels' sensitivity is obtained. Then, in the second pass, bits allocated to each channel will be automatically decided according to the sensitivity obtained in the first pass to make sure that the channel with higher sensitivity can be allocated with more bits to maintain accuracy as much as possible. With the well-designed algorithm, our method surpasses state-of-the-art compression tools with on average 6.4% BD-rate saving. Yuzhang Hu, Sifeng Xia, Wenhan Yang, Jiaying Liu 0001 |
VCIP | 4 |
| 2020 | Joint Rain Detection and Removal from a Single Image with Contextualized Deep NetworksabstractRain streaks, particularly in heavy rain, not only degrade visibility but also make many computer vision algorithms fail to function properly. In this paper, we address this visibility problem by focusing on single-image rain removal, even in the presence of dense rain streaks and rain-streak accumulation, which is visually similar to mist or fog. To achieve this, we introduce a new rain model and a deep learning architecture. Our rain model incorporates a binary rain map indicating rain-streak regions, and accommodates various shapes, directions, and sizes of overlapping rain streaks, as well as rain accumulation, to model heavy rain. Based on this model, we construct a multi-task deep network, which jointly learns three targets: the binary rain-streak map, rain streak layers, and clean background, which is our ultimate output. To generate features that can be invariant to rain steaks, we introduce a contextual dilated network, which is able to exploit regional contextual information. To handle various shapes and directions of overlapping rain streaks, our strategy is to utilize a recurrent process that progressively removes rain streaks. Our binary map provides a constraint and thus additional information to train our network. Extensive evaluation on real images, particularly in heavy rain, shows the effectiveness of our model and architecture. Wenhan Yang, Robby T. Tan, Jiashi Feng, Zongming Guo, Shuicheng Yan, Jiaying Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2020 | A lightweight face detector by integrating the convolutional neural network with the image pyramid
Jiapeng Luo, Jiaying Liu 0001, Jun Lin 0001, Zhongfeng Wang 0001 |
Pattern Recognit. Lett. | 2 |
| 2020 | Video Coding for Machines: A Paradigm of Collaborative Compression and Intelligent AnalyticsabstractVideo coding, which targets to compress and reconstruct the whole frame, and feature compression, which only preserves and transmits the most critical information, stand at two ends of the scale. That is, one is with compactness and efficiency to serve for machine vision, and the other is with full fidelity, bowing to human perception. The recent endeavors in imminent trends of video compression, e.g. deep learning based coding tools and end-to-end image/video coding, and MPEG-7 compact feature descriptor standards, i.e. Compact Descriptors for Visual Search and Compact Descriptors for Video Analysis, promote the sustainable and fast development in their own directions, respectively. In this paper, thanks to booming AI technology, e.g. prediction and generation models, we carry out exploration in the new area, Video Coding for Machines (VCM), arising from the emerging MPEG standardization efforts1. Towards collaborative compression and intelligent analytics, VCM attempts to bridge the gap between feature coding for machine vision and video coding for human vision. Aligning with the rising Analyze then Compress instance Digital Retina, the definition, formulation, and paradigm of VCM are given first. Meanwhile, we systematically review state-of-the-art techniques in video compression and feature compression from the unique perspective of MPEG standardization, which provides the academic and industrial evidence to realize the collaborative compression of video and feature streams in a broad range of AI applications. Finally, we come up with potential VCM solutions, and the preliminary results have demonstrated the performance and efficiency gains. Further direction is discussed as well. Ling-Yu Duan, Jiaying Liu 0001, Wenhan Yang, Tiejun Huang 0001, Wen Gao 0001 |
IEEE Trans. Image Process. | 2 |
| 2020 | A Comprehensive Benchmark for Single Image Compression Artifact ReductionabstractWe present a comprehensive study and evaluation of existing single image compression artifact removal algorithms using a new 4K resolution benchmark. This benchmark is called the Large-Scale Ideal Ultra high-definition 4K (LIU4K), and it includes including diversified foreground objects and background scenes with rich structures. Compression artifact removal, as a common post-processing technique, aims at alleviating undesirable artifacts, such as blockiness, ringing, and banding caused by quantization and approximation in the compression process. In this work, a systematic listing of the reviewed methods is presented based on their basic models (handcrafted models and deep networks). The main contributions and novelties of these methods are highlighted, and the main development directions are summarized, including architectures, multi-domain sources, signal structures, and new targeted units. Furthermore, based on a unified deep learning configuration (i.e.same training data, loss function, optimization algorithm,etc.), we evaluate recent deep learning-based methods based on diversified evaluation measures. The experimental results show state-of-the-art performance comparisons of existing methods based on both full-reference, non-reference, and task-driven metrics. Our survey gives a comprehensive reference source for future research on single image compression artifact removal and inspires new directions in related fields. Jiaying Liu 0001, Dong Liu 0002, Wenhan Yang, Sifeng Xia, Xiaoshuai Zhang, Yuanying Dai |
IEEE Trans. Image Process. | 1 |
| 2020 | LR3M: Robust Low-Light Enhancement via Low-Rank Regularized Retinex ModelabstractNoise causes unpleasant visual effects in low-light image/video enhancement. In this paper, we aim to make the enhancement model and method aware of noise in the whole process. To deal with heavy noise which is not handled in previous methods, we introduce a robust low-light enhancement approach, aiming at well enhancing low-light images/videos and suppressing intensive noise jointly. Our method is based on the proposed Low-Rank Regularized Retinex Model (LR3M), which is the first to inject low-rank prior into a Retinex decomposition process to suppress noise in the reflectance map. Our method estimates a piece-wise smoothed illumination and a noise-suppressed reflectance sequentially, avoiding remaining noise in the illumination and reflectance maps which are usually presented in alternative decomposition methods. After getting the estimated illumination and reflectance, we adjust the illumination layer and generate our enhancement result. Furthermore, we apply our LR3M to video low-light enhancement. We consider inter-frame coherence of illumination maps and find similar patches through reflectance maps of successive frames to form the low-rank prior to make use of temporal correspondence. Our method performs well for a wide variety of images and videos, and achieves better quality both in enhancing and denoising, compared with the state-of-the-art methods. Xutong Ren, Wenhan Yang, Wen-Huang Cheng, Jiaying Liu 0001 |
IEEE Trans. Image Process. | 4 |
| 2020 | Modality Compensation Network: Cross-Modal Adaptation for Action RecognitionabstractWith the prevalence of RGB-D cameras, multimodal video data have become more available for human action recognition. One main challenge for this task lies in how to effectively leverage their complementary information. In this work, we propose a Modality Compensation Network (MCN) to explore the relationships of different modalities, and boost the representations for human action recognition. We regard RGB/ optical flow videos as source modalities, skeletons as auxiliary modality. Our goal is to extract more discriminative features from source modalities, with the help of auxiliary modality. Built on deep Convolutional Neural Networks (CNN) and Long Short Term Memory (LSTM) networks, our model bridges data from source and auxiliary modalities by a modality adaptation block to achieve adaptive representation learning, that the network learns to compensate for the loss of skeletons at test time and even at training time. We explore multiple adaptation schemes to narrow the distance between source and auxiliary modal distributions from different levels, according to the alignment of source and auxiliary data in training. In addition, skeletons are only required in the training phase. Our model is able to improve the recognition performance with source data when testing. Experimental results reveal that MCN outperforms stateof- the-art approaches on four widely-used action recognition benchmarks. Sijie Song, Jiaying Liu 0001, Yanghao Li, Zongming Guo |
IEEE Trans. Image Process. | 2 |
| 2020 | Consistent Video Style Transfer via Relaxation and RegularizationabstractIn recent years, neural style transfer has attracted more and more attention, especially for image style transfer. However, temporally consistent style transfer for videos is still a challenging problem. Existing methods, either relying on a significant amount of video data with optical flows or using singleframe regularizers, fail to handle strong motions or complex variations, therefore have limited performance on real videos. In this paper, we address the problem by jointly considering the intrinsic properties of stylization and temporal consistency. We first identify the cause of the conflict between style transfer and temporal consistency, and propose to reconcile this contradiction by relaxing the objective function, so as to make the stylization loss term more robust to motions. Through relaxation, style transfer is more robust to inter-frame variation without degrading the subjective effect. Then, we provide a novel formulation and understanding of temporal consistency. Based on the formulation, we analyze the drawbacks of existing training strategies and derive a new regularization. We show by experiments that the proposed regularization can better balance the spatial and temporal performance. Based on relaxation and regularization, we design a zero-shot video style transfer framework. Moreover, for better feature migration, we introduce a new module to dynamically adjust inter-channel distributions. Quantitative and qualitative results demonstrate the superiority of our method over state-of-the-art style transfer methods. Wenjing Wang 0001, Shuai Yang 0001, Jizheng Xu, Jiaying Liu 0001 |
IEEE Trans. Image Process. | 4 |
| 2020 | Removing Arbitrary-Scale Rain Streaks via Fractal Band Learning With Self-SupervisionabstractData-driven rain streak removal methods, most of which rely on synthesized paired data, usually come across the generalization problem when being applied in real scenarios. In this paper, we propose a novel deep-learning based rain streak removal method injected with self-supervision to obtain the capacity of removing more varied-scale rain streaks in practical applications. To this end, in this work, efforts are made from two perspectives. First, considering that rain streak removal is highly correlated with texture characteristics, we create a fractal band learning (FBL) network based on frequency band recovery. It integrates commonly seen band feature operations as neural forms and effectively improves the capacity to capture discriminative features for deraining. Second, to further improve the generalization ability of FBL to remove rain streaks of varied scales, we incorporate scale-robust self-supervision to regularize the network training. The constraint forces the extracted features of an input rain image at different scales to be equivalent after rescaling operations. Therefore, our method can offer similar responses based on solely image content without the interference of scale change and is capable to remove varied-scale rain streaks. Extensive experiments in quantitative and qualitative evaluations demonstrate the superiority of our method for rain streak removal, especially for the real cases where very large rain streaks exist, and prove the effectiveness of each component. Wenhan Yang, Shiqi Wang 0001, Jiaying Liu 0001 |
IEEE Trans. Image Process. | 3 |
| 2020 | Advancing Image Understanding in Poor Visibility Environments: A Collective Benchmark StudyabstractExisting enhancement methods are empirically expected to help the high-level end computer vision task: however, that is observed to not always be the case in practice. We focus on object or face detection in poor visibility enhancements caused by bad weathers (haze, rain) and low light conditions. To provide a more thorough examination and fair comparison, we introduce three benchmark sets collected in real-world hazy, rainy, and low-light conditions, respectively, with annotated objects/faces. We launched the UG2+ challenge Track 2 competition in IEEE CVPR 2019, aiming to evoke a comprehensive discussion and exploration about whether and how low-level vision techniques can benefit the high-level automatic visual recognition in various scenarios. To our best knowledge, this is the first and currently largest effort of its kind. Baseline results by cascading existing enhancement and detection models are reported, indicating the highly challenging nature of our new data as well as the large room for further technical innovations. Thanks to a large participation from the research community, we are able to analyze representative team solutions, striving to better identify the strengths and limitations of existing mindsets as well as the future directions. Wenhan Yang, Ye Yuan 0012, Wenqi Ren, Jiaying Liu 0001, Walter J. Scheirer, Zhangyang Wang, Taiheng Zhang, Qiaoyong Zhong, Di Xie, Shiliang Pu, Yuqiang Zheng, Yanyun Qu, Yuhong Xie, Hao Jiang 0014, Siyuan Yang 0001, Yan Liu 0041, Xiaochao Qu, Pengfei Wan 0001, Shuai Zheng 0005, Minhui Zhong, Taiyi Su, Lingzhi He, Yandong Guo, Yao Zhao 0001, Zhenfeng Zhu, Jinxiu Liang, Jingwen Wang 0003, Yuhui Quan, Yong Xu 0007, Bo Liu 0112, Xin Liu 0012, Tingyu Lin 0003, Xiaochuan Li 0001, Feng Lu 0005, Lin Gu 0003, Shengdi Zhou, Cong Cao 0005, Cheng Chi 0003, Chubin Zhuang, Zhen Lei 0001, Stan Z. Li, Shizheng Wang, Ruizhe Liu, Dong Yi, Zheming Zuo, Jianning Chi, Huan Wang 0014, Kai Wang 0036, Yixiu Liu, Xingyu Gao 0001, Zhenyu Chen 0003, Yongzhou Li, Huicai Zhong, Jing Huang 0017, Heng Guo 0003, Jianfei Yang 0001, Wenjuan Liao, Jiangang Yang, Liguo Zhou, Mingyue Feng, Likun Qin |
IEEE Trans. Image Process. | 4 |
| 2020 | Deep Reference Generation With Multi-Domain Hierarchical Constraints for Inter PredictionabstractInter prediction is an important module in video coding for temporal redundancy removal, where similar reference blocks are searched from previously coded frames and employed to predict the block to be coded. Although existing video codecs can estimate and compensate for block-level motions, their inter prediction performance is still heavily affected by the remaining inconsistent pixel-wise displacement caused by irregular rotation and deformation. In this paper, we address the problem by proposing a deep frame interpolation network to generate additional reference frames in coding scenarios. First, we summarize the previous adaptive convolutions used for frame interpolation and propose a factorized kernel convolutional network to improve the modeling capacity and simultaneously keep its compact form. Second, to better train this network, multi-domain hierarchical constraints are introduced to regularize the training of our factorized kernel convolutional network. For spatial domain, we use a gradually down-sampled and up-sampled auto-encoder to generate the factorized kernels for frame interpolation at different scales. For quality domain, considering the inconsistent quality of the input frames, the factorized kernel convolution is modulated with quality-related features to learn to exploit more information from high quality frames. For frequency domain, a sum of absolute transformed difference loss that performs frequency transformation is utilized to facilitate network optimization from the view of coding performance. With the well-designed frame interpolation network regularized by multi-domain hierarchical constraints, our method surpasses HEVC on average 3.8% BD-rate saving for the luma component under the random access configuration and also obtains on average 0.83% BD-rate saving over the upcoming VVC. Jiaying Liu 0001, Sifeng Xia, Wenhan Yang |
IEEE Trans. Multim. | 1 |
| 2020 | Image/Video Restoration via Multiplanar Autoregressive Model and Low-Rank OptimizationabstractIn this article, we introduce an image/video restoration approach by utilizing the high-dimensional similarity in images/videos. After grouping similar patches from neighboring frames, we propose to build a multiplanar autoregressive (AR) model to exploit the correlation in cross-dimensional planes of the patch group, which has long been neglected by previous AR models. To further utilize the nonlocal self-similarity in images/videos, a joint multiplanar AR and low-rank based approach is proposed (MARLow) to reconstruct patch groups more effectively. Moreover, for video restoration, the temporal smoothness of the restored video is constrained by the Markov random field (MRF), where MRF encodes a priori knowledge about consistency of patches from neighboring frames. Specifically, we treat different restoration results (from different patch groups) of a certain patch as labels of an MRF, and temporal consistency among these restored patches is imposed. The proposed method is also suitable for other restoration applications such as interpolation and text removal. Extensive experimental results demonstrate that the proposed approach obtains encouraging performance comparing with state-of-the-art methods. Mading Li, Jiaying Liu 0001, Xiaoyan Sun 0001, Zhiwei Xiong |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2020 | A Benchmark Dataset and Comparison Study for Multi-modal Human Action AnalyticsabstractLarge-scale benchmarks provide a solid foundation for the development of action analytics. Most of the previous activity benchmarks focus on analyzing actions in RGB videos. There is a lack of large-scale and high-quality benchmarks for multi-modal action analytics. In this article, we introduce PKU Multi-Modal Dataset (PKU-MMD), a new large-scale benchmark for multi-modal human action analytics. It consists of about 28,000 action instances and 6.2 million frames in total and provides high-quality multi-modal data sources, including RGB, depth, infrared radiation (IR), and skeletons. To make PKU-MMD more practical, our dataset comprises two subsets under different settings for action understanding, namely Part I and Part II. Part I contains 1,076 untrimmed video sequences with 51 action classes performed by 66 subjects, while Part II contains 1,009 untrimmed video sequences with 41 action classes performed by 13 subjects. Compared to Part I, Part II is more challenging due to short action intervals, concurrent actions and heavy occlusion. PKU-MMD can be leveraged in two scenarios: action recognition with trimmed video clips and action detection with untrimmed video sequences. For each scenario, we provide benchmark performance on both subsets by conducting different methods with different modalities under two evaluation protocols, respectively. Experimental results show that PKU-MMD is a significant challenge to many state-of-the-art methods. We further illustrate that the features learned on PKU-MMD can be well transferred to other datasets. We believe this large-scale dataset will boost the research in the field of action analytics for the community. Jiaying Liu 0001, Sijie Song, Chunhui Liu 0002, Yanghao Li, Yueyu Hu |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2019 | Temporal Bilinear Networks for Video Action RecognitionabstractTemporal modeling in videos is a fundamental yet challenging problem in computer vision. In this paper, we propose a novel Temporal Bilinear (TB) model to capture the temporal pairwise feature interactions between adjacent frames. Compared with some existing temporal methods which are limited in linear transformations, our TB model considers explicit quadratic bilinear transformations in the temporal domain for motion evolution and sequential relation modeling. We further leverage the factorized bilinear model in linear complexity and a bottleneck network design to build our TB blocks, which also constrains the parameters and computation cost. We consider two schemes in terms of the incorporation of TB blocks and the original 2D spatial convolutions, namely wide and deep Temporal Bilinear Networks (TBN). Finally, we perform experiments on several widely adopted datasets including Kinetics, UCF101 and HMDB51. The effectiveness of our TBNs is validated by comprehensive ablation analyses and comparisons with various state-of-the-art methods. Yanghao Li, Sijie Song, Jiaying Liu 0001 |
AAAI | 4 |
| 2019 | Weakly Supervised Scene Parsing with Point-Based Distance Metric LearningabstractSemantic scene parsing is suffering from the fact that pixellevel annotations are hard to be collected. To tackle this issue, we propose a Point-based Distance Metric Learning (PDML) in this paper. PDML does not require dense annotated masks and only leverages several labeled points that are much easier to obtain to guide the training process. Concretely, we leverage semantic relationship among the annotated points by encouraging the feature representations of the intra- and intercategory points to keep consistent, i.e. points within the same category should have more similar feature representations compared to those from different categories. We formulate such a characteristic into a simple distance metric loss, which collaborates with the point-wise cross-entropy loss to optimize the deep neural networks. Furthermore, to fully exploit the limited annotations, distance metric learning is conducted across different training images instead of simply adopting an image-dependent manner. We conduct extensive experiments on two challenging scene parsing benchmarks of PASCALContext and ADE 20K to validate the effectiveness of our PDML, and competitive mIoU scores are achieved. Rui Qian 0003, Yunchao Wei, Humphrey Shi, Jiachen Li 0003, Jiaying Liu 0001, Thomas S. Huang |
AAAI | 5 |
| 2019 | TET-GAN: Text Effects Transfer via Stylization and DestylizationabstractText effects transfer technology automatically makes the text dramatically more impressive. However, previous style transfer methods either study the model for general style, which cannot handle the highly-structured text effects along the glyph, or require manual design of subtle matching criteria for text effects. In this paper, we focus on the use of the powerful representation abilities of deep neural features for text effects transfer. For this purpose, we propose a novel Texture Effects Transfer GAN (TET-GAN), which consists of a stylization subnetwork and a destylization subnetwork. The key idea is to train our network to accomplish both the objective of style transfer and style removal, so that it can learn to disentangle and recombine the content and style features of text effects images. To support the training of our network, we propose a new text effects dataset with as much as 64 professionally designed styles on 837 characters. We show that the disentangled feature representations enable us to transfer or remove all these styles on arbitrary glyphs using one network. Furthermore, the flexible network design empowers TET-GAN to efficiently extend to a new text style via oneshot learning where only one example is required. We demonstrate the superiority of the proposed method in generating high-quality stylized text over the state-of-the-art methods. Shuai Yang 0001, Jiaying Liu 0001, Wenjing Wang 0001, Zongming Guo |
AAAI | 2 |
| 2019 | Unsupervised Person Image Generation With Semantic Parsing TransformationabstractIn this paper, we address unsupervised pose-guided person image generation, which is known challenging due to non-rigid deformation. Unlike previous methods learning a rock-hard direct mapping between human bodies, we propose a new pathway to decompose the hard mapping into two more accessible subtasks, namely, semantic parsing transformation and appearance generation. Firstly, a semantic generative network is proposed to transform between semantic parsing maps, in order to simplify the non-rigid deformation learning. Secondly, an appearance generative network learns to synthesize semantic-aware textures. Thirdly, we demonstrate that training our framework in an end-to-end manner further refines the semantic maps and final results accordingly. Our method is generalizable to other semantic-aware person image generation tasks, e.g., clothing texture transfer and controlled image manipulation. Experimental results demonstrate the superiority of our method on DeepFashion and Market-1501 datasets, especially in keeping the clothing attributes and better body shapes. Sijie Song, Wei Zhang 0031, Jiaying Liu 0001, Tao Mei 0001 |
CVPR | 3 |
| 2019 | Typography With Decor: Intelligent Text Style TransferabstractText effects transfer can dramatically make the text visually pleasing. In this paper, we present a novel framework to stylize the text with exquisite decor, which is ignored by the previous text stylization methods. Decorative elements pose a challenge to spontaneously handle basal text effects and decor, which are two different styles. To address this issue, our key idea is to learn to separate, transfer and recombine the decors and the basal text effect. A novel text effect transfer network is proposed to infer the styled version of the target text. The stylized text is finally embellished with decor where the placement of the decor is carefully determined by a novel structure-aware strategy. Furthermore, we propose a domain adaptation strategy for decor detection and a one-shot training strategy for text effects transfer, which greatly enhance the robustness of our network to new styles. We base our experiments on our collected topography dataset including 59,000 professionally styled text and demonstrate the superiority of our method over other state-of-the-art style transfer methods. Wenjing Wang 0001, Jiaying Liu 0001, Shuai Yang 0001, Zongming Guo |
CVPR | 2 |
| 2019 | Iterative Reorganization With Weak Spatial Constraints: Solving Arbitrary Jigsaw Puzzles for Unsupervised Representation LearningabstractLearning visual features from unlabeled image data is an important yet challenging task, which is often achieved by training a model on some annotation-free information. We consider spatial contexts, for which we solve so-called jigsaw puzzles, i.e., each image is cut into grids and then disordered, and the goal is to recover the correct configuration. Existing approaches formulated it as a classification task by defining a fixed mapping from a small subset of configurations to a class set, but these approaches ignore the underlying relationship between different configurations and also limit their applications to more complex scenarios. This paper presents a novel approach which applies to jigsaw puzzles with an arbitrary grid size and dimensionality. We provide a fundamental and generalized principle, that weaker cues are easier to be learned in an unsupervised manner and also transfer better. In the context of puzzle recognition, we use an iterative manner which, instead of solving the puzzle all at once, adjusts the order of the patches in each step until convergence. In each step, we combine both unary and binary features of each patch into a cost function judging the correctness of the current configuration. Our approach, by taking similarity between puzzles into consideration, enjoys a more efficient way of learning visual knowledge. We verify the effectiveness of our approach from two aspects. First, it solves arbitrarily complex puzzles, including high-dimensional puzzles, that prior methods are difficult to handle. Second, it serves as a reliable way of network initialization, which leads to better transfer performance in visual recognition tasks including classification, detection and segmentation. Chen Wei 0005, Lingxi Xie, Xutong Ren, Yingda Xia, Chi Su, Jiaying Liu 0001, Qi Tian 0001, Alan L. Yuille |
CVPR | 6 |
| 2019 | Frame-Consistent Recurrent Video Deraining With Dual-Level FlowabstractIn this paper, we address the problem of rain removal from videos by proposing a more comprehensive framework that considers the additional degradation factors in real scenes neglected in previous works. The proposed framework is built upon a two-stage recurrent network with dual-level flow regularizations to perform the inverse recovery process of the rain synthesis model for video deraining. The rain-free frame is estimated from the single rain frame at the first stage. It is then taken as guidance along with previously recovered clean frames to help obtain a more accurate clean frame at the second stage. This two-step architecture is capable of extracting more reliable motion information from the initially estimated rain-free frame at the first stage for better frame alignment and motion modeling at the second stage. Furthermore, to keep the motion consistency between frames that facilitates a frame-consistent deraining model at the second stage, a dual-level flow based regularization is proposed at both coarse flow and fine pixel levels. To better train and evaluate the proposed video deraining network, a novel rain synthesis model is developed to produce more visually authentic paired training and evaluation videos. Extensive experiments on a series of synthetic and real videos verify not only the superiority of the proposed method over state-of-the-art but also the effectiveness of network design and its each component. Wenhan Yang, Jiaying Liu 0001, Jiashi Feng |
CVPR | 2 |
| 2019 | Disentangled Image MattingabstractMost previous image matting methods require a roughly-specificed trimap as input, and estimate fractional alpha values for all pixels that are in the unknown region of the trimap. In this paper, we argue that directly estimating the alpha matte from a coarse trimap is a major limitation of previous methods, as this practice tries to address two difficult and inherently different problems at the same time: identifying true blending pixels inside the trimap region, and estimate accurate alpha values for them. We propose AdaMatting, a new end-to-end matting framework that disentangles this problem into two sub-tasks: trimap adaptation and alpha estimation. Trimap adaptation is a pixel-wise classification problem that infers the global structure of the input image by identifying definite foreground, background, and semi-transparent image regions. Alpha estimation is a regression problem that calculates the opacity value of each blended pixel. Our method separately handles these two sub-tasks within a single deep convolutional neural network (CNN). Extensive experiments show that AdaMatting has additional structure awareness and trimap fault-tolerance. Our method achieves the state-of-the-art performance on Adobe Composition-1k dataset both qualitatively and quantitatively. It is also the current best-performing method on the alphamatting.com online evaluation for all commonly-used metrics. Shaofan Cai, Xiaoshuai Zhang, Haoqiang Fan, Jiangyu Liu, Jiaying Liu 0001, Jue Wang 0001, Jian Sun 0001 |
ICCV | 7 |
| 2019 | Controllable Artistic Text Style Transfer via Shape-Matching GANabstractArtistic text style transfer is the task of migrating the style from a source image to the target text to create artistic typography. Recent style transfer methods have considered texture control to enhance usability. However, controlling the stylistic degree in terms of shape deformation remains an important open challenge. In this paper, we present the first text style transfer network that allows for real-time control of the crucial stylistic degree of the glyph through an adjustable parameter. Our key contribution is a novel bidirectional shape matching framework to establish an effective glyph-style mapping at various deformation levels without paired ground truth. Based on this idea, we propose a scale-controllable module to empower a single network to continuously characterize the multi-scale shape features of the style image and transfer these features to the target text. The proposed method demonstrates its superiority over previous state-of-the-arts in generating diverse, controllable and high-quality stylized text. Shuai Yang 0001, Zhangyang Wang, Ning Xu 0007, Jiaying Liu 0001, Zongming Guo |
ICCV | 5 |
| 2019 | Partition Tree Guided Progressive Rethinking Network for in-Loop Filtering of HEVCabstractIn-Loop filter is a key part in High Efficiency Video Coding (HEVC) which effectively removes the compression artifacts. Recently, many newly proposed methods combine residual learning and dense connection to construct a deeper network for better in-loop filtering performance. However, the long-term dependency between blocks is neglected, and information usually passes between blocks only after dimension compression. To address these issues, we propose the Progressive Rethinking Block (PRB) to deliver long-term memory between the neighboring blocks and allow information to flow without compression, which is similar to human decision mechanism - usually reviewing the complete past memorized experiences to decide in the present, not just based on simple principles summarized before. PRBs further establish the Progressive Rethinking Network (PRN). In addition, we calculate the Multi-scale Mean value of Coding Units (MM-CU) to generate the side information maps which guide the training of the network by novelly telling the network architecture of the entire coding partition tree. Experimental results show that our proposed partition tree guided PRN provides 10.1% BD-rate reduction on average compared to the HEVC baseline. Dezhao Wang, Sifeng Xia, Wenhan Yang, Yueyu Hu, Jiaying Liu 0001 |
ICIP | 5 |
| 2019 | Deep Inter Prediction Via Pixel-Wise Motion Oriented Reference GenerationabstractInter prediction is an important module in video coding for temporal redundancy removal, where the reference blocks are searched from the previously coded frames and employed to predict the block to be coded. However, apart from regular block-wise shift motion, there usually exists inconsistent pixel-wise motion such as rotation and deformation between blocks, which will largely degrade the prediction performance. In this paper, we propose a Multiscale Adaptive Separable Convolutional Neural Network (MASCNN) to generate pixel-wise closer reference frames for inter prediction. A multiscale network is built to interpolate the target frame from coarse to fine. Reconstruction losses are enforced on each scale to make the network infer the main structure at small scales, which improves the interpolation accuracy of the network. Furthermore, a sum of absolute transformed difference (SATD) loss function is proposed to regularize the network training, which further improves the coding performance. Compared with HEVC, our method can obtain on average 5.7% BD-rate saving and up to 9.9% BD-rate saving for the luma component under the random access configuration. Sifeng Xia, Wenhan Yang, Yueyu Hu, Jiaying Liu 0001 |
ICIP | 4 |
| 2019 | Dynamically Unfolding Recurrent Restorer: A Moving Endpoint Control Method for Image Restoration
Xiaoshuai Zhang, Yiping Lu 0001, Jiaying Liu 0001, Bin Dong 0001 |
ICLR (Poster) | 3 |
| 2019 | Deep Pyramid Variation Learning for Image InterpolationabstractPrevious learning-based interpolation methods do not consider multi-scale structural information, which is generally effective for image modeling. In this paper, we design a deep network based on a novel pyramid variation learning approach with multi-scale structure modeling. An image is represented as multi-dimensional features. Besides two spatial dimensions, the features include a neighboring variation dimension where every pixel is encoded as the variation to its nearest low-resolution pixel, and a scale dimension along which the feature maps generated by a gradual down-sampling process are stacked. Thus, these multi-dimensional features are constructed to model local dependency and multi-scale similarity jointly. Inspired by this feature design, we build an end-to-end trainable Recurrent Multi-Path Aggregation Network (RMPAN) for image interpolation, where the scale dimension is unfolded to form a multi-path aggregation network to apply joint filters at different scales recurrently. Location aware sampling layers are used in RMPAN to transform feature maps into different scales with only location changes in each convolution path, which aggregate the context information without resolution loss. Comprehensive experiments demonstrate that our method leads to a superior performance and offers new state-of-the-art benchmark. Wenhan Yang, Jiaying Liu 0001 |
ICME | 4 |
| 2019 | Switch Mode Based Deep Fractional Interpolation in Video CodingabstractFractional interpolation is a significant technology in motion compensation of video coding. It generates sub-pixel level reference samples in inter prediction to facilitate temporal redundancy removal between video frames. Recently, some methods explore to introduce the deep learning technique for fractional interpolation and have obtained better compression results. However, existing deep learning based methods still treat fractional interpolation as a traditional interpolation problem but fail to adjust it to the motion compensation scenario. In this paper, we design a switch mode based deep fractional interpolation method to introduce integer pixels of different positions to the interpolation of sub-pixel position samples. By switching between integer pixels of different positions, our method can infer the sub-pixels with smaller variations and achieve better fractional interpolation results. Consequently the motion compensation performance can be further improved. Experimental results have also verified the efficiency of the switch mode based deep fractional interpolation. Compared with High Efficiency Video Coding, our method achieves 2.8% bit saving on average and up to 6.2% bit saving under low-delay P configuration. Sifeng Xia, Wenhan Yang, Yueyu Hu, Wen-Huang Cheng, Jiaying Liu 0001 |
ISCAS | 5 |
| 2019 | Optimized Skeleton-based Action Recognition via Sparsified Graph RegressionabstractWith the prevalence of accessible depth sensors, dynamic human body skeletons have attracted much attention as a robust modality for action recognition. Previous methods model skeletons based on RNN or CNN, which has limited expressive power for irregular skeleton joints. While graph convolutional networks (GCN) have been proposed to address irregular graph-structured data, the fundamental graph construction remains challenging. In this paper, we represent skeletons naturally on graphs, and propose a graph regression based GCN (GR-GCN) for skeleton-based action recognition, aiming to capture the spatio-temporal variation in the data. As the graph representation is crucial to graph convolution, we first propose graph regression to statistically learn the underlying graph from multiple observations. In particular, we provide spatio-temporal modeling of skeletons and pose an optimization problem on the graph structure over consecutive frames, which enforces the sparsity of the underlying graph for efficient representation. The optimized graph not only connects each joint to its neighboring joints in the same frame strongly or weakly, but also links with relevant joints in the previous and subsequent frames. We then feed the optimized graph into the GCN along with the coordinates of the skeleton sequence for feature learning, where we deploy high-order and fast Chebyshev approximation of spectral graph convolution. Further, we provide analysis of the variation characterization by the Chebyshev approximation. Experimental results validate the effectiveness of the proposed graph regression and show that the proposed GR-GCN achieves the state-of-the-art performance on the widely used NTU RGB+D, UT-Kinect and SYSU 3D datasets. Xiang Gao 0014, Wei Hu 0003, Jiaxiang Tang, Jiaying Liu 0001, Zongming Guo |
ACM Multimedia | 4 |
| 2019 | FashionOn: Semantic-guided Image-based Virtual Try-on with Detailed Human and Clothing InformationabstractThe image-based virtual try-on system has attracted a lot of research attention. The virtual try-on task is challenging since synthesizing try-on images involves the estimation of 3D transformation from 2D images, which is an ill-posed problem. Therefore, most of the previous virtual try-on systems cannot solve difficult cases, e.g., body occlusions, wrinkles of clothes, and details of the hair. Moreover, the existing systems require the users to upload the image for the target pose, which is not user-friendly. In this paper, we aim to resolve the above challenges by proposing a novel FashionOn network to synthesize user images fitting different clothes in arbitrary poses to provide comprehensive information about how suitable the clothes are. Specifically, given a user image, an in-shop clothing image, and a target pose (can be arbitrarily manipulated by joint points), FashionOn learns to synthesize the try-on images by three important stages: pose-guided parsing translation, segmentation region coloring, and salient region refinement. Extensive experiments demonstrate that FashionOn maintains the details of clothing information (e.g., logo, pleat, lace), as well as resolves the body occlusion problem, and thus achieves the state-of-the-art virtual try-on performance both qualitatively and quantitatively. Chia-Wei Hsieh, Chieh-Yun Chen, Chien-Lung Chou, Hong-Han Shuai, Jiaying Liu 0001, Wen-Huang Cheng |
ACM Multimedia | 5 |
| 2019 | A Progressive Background Updating Based Coding Scheme for Surveillance VideosabstractBackground based coding is an effective scheme to improve the coding efficiency of surveillance videos. However, it takes a long time to generate a high quality background picture (BG-picture). And the encoding of the high quality BG-picture will increase the bitrate abruptly. To solve these problems, a progressive background updating based coding scheme is proposed in this paper. In the proposed scheme, the BG-picture is updated block by block. To improve the overall coding efficiency, an importance map is designed to select the valid background blocks (B-blocks) progressively which will be encoded with high quality. It is worth noting that only the valid B-blocks is encoded instead of the entire BG-picture. Compared with the reference software of Versatile Video Coding (VVC), the proposed scheme achieves about 23.3 percent bit-rate saving on average. Yunhui Shi, Wenpeng Ding, Jiaying Liu 0001 |
VCIP | 5 |
| 2019 | Reduced-reference quality assessment of image super-resolution by energy change and texture variation
Yuming Fang 0001, Jiaying Liu 0001, Yabin Zhang 0002, Weisi Lin, Zongming Guo |
J. Vis. Commun. Image Represent. | 2 |
| 2019 | Selfie retoucher: subject-oriented self-portrait enhancement
Sifeng Xia, Shuai Yang 0001, Jiaying Liu 0001 |
Multim. Tools Appl. | 3 |
| 2019 | Multi-Modality Multi-Task Recurrent Neural Network for Online Action DetectionabstractOnline action detection is a brand new challenge and plays a critical role in visual surveillance analytics. It goes one step further than a conventional action recognition task, which recognizes human actions from well-segmented clips. Online action detection is desired to identify the action type and localize action positions on the fly from the untrimmed stream data. In this paper, we propose a multi-modality multi-task recurrent neural network, which incorporates both RGB and Skeleton networks. We design different temporal modeling networks to capture specific characteristics from various modalities. Then, a deep long short-term memory subnetwork is utilized effectively to capture the complex long-range temporal dynamics, naturally avoiding the conventional sliding window design and thus ensuring high computational efficiency. Constrained by a multi-task objective function in the training phase, this network achieves superior detection performance and is capable of automatically localizing the start and end points of actions more accurately. Furthermore, embedding subtask of regression provides the ability to forecast the action prior to its occurrence. We evaluate the proposed method and several other methods in action detection and forecasting on the online action detection data set and gaming action data set datasets. Experimental results demonstrate that our model achieves the state-of-the-art performance on both tasks. Jiaying Liu 0001, Yanghao Li, Sijie Song, Junliang Xing, Cuiling Lan, Wenjun Zeng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | Reference-Guided Deep Super-Resolution via Manifold Localized External CompensationabstractThe rapid development of social network and online multimedia technology makes it possible to address traditional image and video enhancement problems, with the aid of online similar reference data. In this paper, we tackle the problem of super-resolution (SR) in this way, specifically aiming to handle the “one-to-many” problem between the image patches of low resolution (LR) and high resolution (HR). We propose a manifold localized deep external compensation (MALDEC) network to additionally utilize reference images, i. e., retrieved similar images in cloud database and reference HR frame in a video, to provide an accurate localization and mapping to the HR manifold, and compensate the lost high-frequency details. The proposed network employs a three-step recovery: 1) internal structure inference, which uses the LR image itself and the internally inferred high frequency information to preserve main structure of the HR image; 2) manifold localization, which localizes the HR manifold and constructs the correspondence between the internal inferred image and the external images; and 3) external compensation, which introduces the external references of retrieved similar patches based on manifold localization information to reconstruct the high-frequency details. The learnable components of MALDEC, internal structure inference, and external compensation, are trained jointly to make a good tradeoff between these two terms for an optimal SR result. Finally, the proposed method is examined under three tasks: cloud-based image SR, multi-pose face reconstruction, and reference frame-guided video SR. Extensive experiments demonstrate the superiority of our method than the state-of-the-art SR methods in both objective and subjective evaluations, and our method offers new state-of-the-art performance. Wenhan Yang, Sifeng Xia, Jiaying Liu 0001, Zongming Guo |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | Content-Aware Convolutional Neural Network for In-Loop Filtering in High Efficiency Video CodingabstractRecently, convolutional neural network (CNN) has attracted tremendous attention and has achieved great success in many image processing tasks. In this paper, we focus on CNN technology combined with image restoration to facilitate video coding performance and propose the content-aware CNN based in-loop filtering for high-efficiency video coding (HEVC). In particular, we quantitatively analyze the structure of the proposed CNN model from multiple dimensions to make the model interpretable and optimal for CNN-based loop filtering. More specifically, each coding tree unit (CTU) is treated as an independent region for processing, such that the proposed content-aware multimodel filtering mechanism is realized by the restoration of different regions with different CNN models under the guidance of the discriminative network. To adapt the image content, the discriminative neural network is learned to analyze the content characteristics of each region for the adaptive selection of the deep learning model. The CTU level control is also enabled in the sense of rate-distortion optimization. To learn the CNN model, an iterative training method is proposed by simultaneously labeling filter categories at the CTU level and fine-tuning the CNN model parameters. The CNN based in-loop filter is implemented after sample adaptive offset in HEVC, and extensive experiments show that the proposed approach significantly improves the coding performance and achieves up to 10.0% bit-rate reduction. On average, 4.1%, 6.0%, 4.7%, and 6.0% bit-rate reduction can be obtained under all intra, low delay, low delay P, and random access configurations, respectively. Chuanmin Jia, Shiqi Wang 0001, Xinfeng Zhang 0001, Shanshe Wang, Jiaying Liu 0001, Shiliang Pu, Siwei Ma 0001 |
IEEE Trans. Image Process. | 5 |
| 2019 | One-for-All: Grouped Variation Network-Based Fractional Interpolation in Video CodingabstractFractional interpolation is used to provide sub-pixel level references for motion compensation in the interprediction of video coding, which attempts to remove temporal redundancy in video sequences. Traditional handcrafted fractional interpolation filters face the challenge of modeling discontinuous regions in videos, while existing deep learning-based methods are either designed for a single quantization parameter (QP), only generating half-pixel samples, or need to train a model for each sub-pixel position. In this paper, we present a one-for-all fractional interpolation method based on a grouped variation convolutional neural network (GVCNN). Our method can deal with video frames coded using different QPs and is capable of generating all sub-pixel positions at one sub-pixel level. Also, by predicting variations between integer-position pixels and sub-pixels, our network offers more expressive power. Moreover, we perform specific measurements in training data generation to simulate practical situations in video coding, including blurring the down-sampled sub-pixel samples to avoid aliasing effects and coding integer pixels to simulate reconstruction errors. In addition, we analyze the impact of the size of blur kernels theoretically. Experimental results verify the efficiency of GVCNN. Compared with HEVC, our method achieves 2.2% in bit saving on average and up to 5.2% under low-delay P configuration. Jiaying Liu 0001, Sifeng Xia, Wenhan Yang, Mading Li, Dong Liu 0002 |
IEEE Trans. Image Process. | 1 |
| 2019 | D3R-Net: Dynamic Routing Residue Recurrent Network for Video Rain RemovalabstractIn this paper, we address the problem of video rain removal by considering rain occlusion regions, i.e., very low light transmittance for rain streaks. Different from additive rain streaks, in such occlusion regions, the details of backgrounds are completely lost. Therefore, we propose a hybrid rain model to depict both rain streaks and occlusions. Integrating the hybrid model and useful motion segmentation context information, we present a Dynamic Routing Residue Recurrent Network (D3R-Net). D3R-Net first extracts the spatial features by a residual network. Then, the spatial features are aggregated by recurrent units along the temporal axis. In the temporal fusion, the context information is embedded into the network in a "dynamic routing" way. A heap of recurrent units takes responsibility for handling the temporal fusion in given contexts, e.g., rain or non-rain regions. In the certain forward and backward processes, one of these recurrent units is mainly activated. Then, a context selection gate is employed to detect the context and select one of these temporally fused features generated by these recurrent units as the final fused feature. Finally, this last feature plays a role of "residual feature." It is combined with the spatial feature and then used to reconstruct the negative rain streaks. In such a D3R-Net, we incorporate motion segmentation, which denotes whether a pixel belongs to fast moving edges or not, and rain type indicator, indicating whether a pixel belongs to rain streaks, rain occlusions, and non-rain regions, as the context variables. Extensive experiments on a series of synthetic and real videos with rain streaks verify not only the superiority of the proposed method over state of the art but also the effectiveness of our network design and its each component. Jiaying Liu 0001, Wenhan Yang, Shuai Yang 0001, Zongming Guo |
IEEE Trans. Image Process. | 1 |
| 2019 | Context-Aware Text-Based Binary Image Stylization and SynthesisabstractIn this work, we present a new framework for the stylization of text-based binary images. First, our method stylizes the stroke-based geometric shape like text, symbols and icons in the target binary image based on an input style image. Second, the composition of the stylized geometric shape and a background image is explored. To accomplish the task, we propose legibilitypreserving structure and texture transfer algorithms, which progressively narrow the visual differences between the binary image and the style image. The stylization is then followed by a contextaware layout design algorithm, where cues for both seamlessness and aesthetics are employed to determine the optimal layout of the shape in the background. Given the layout, the binary image is seamlessly embedded into the background by texture synthesis under a context-aware boundary constraint. According to the contents of binary images, our method can be applied to many fields.We show that the proposed method is capable of addressing the unsupervised text stylization problem and is superior to stateof- the-art style transfer methods in automatic artistic typography creation. Besides, extensive experiments on various tasks, such as visual-textual presentation synthesis, icon/symbol rendering and structure-guided image inpainting, demonstrate the effectiveness of the proposed method. Shuai Yang 0001, Jiaying Liu 0001, Wenhan Yang, Zongming Guo |
IEEE Trans. Image Process. | 2 |
| 2019 | Scale-Free Single Image Deraining Via Visibility-Enhanced Recurrent Wavelet LearningabstractIn this paper, we address a rain removal problem from a single image, even in the presence of large rain streaks and rain streak accumulation (where individual streaks cannot be seen, and thus visually similar to mist or fog). For rain streak removal, the mismatch problem between different streak sizes in training and testing phases leads to a poor performance, especially when there are large streaks. To mitigate this problem, we embed a hierarchical representation of wavelet transform into a recurrent rain removal process: 1) rain removal on the low-frequency component; 2) recurrent detail recovery on highfrequency components under the guidance of the recovered lowfrequency component. Benefiting from the recurrent multi-scale modeling of wavelet transform-like design, the proposed network trained on streaks with one size can adapt to those with larger sizes, which significantly favors real rain streak removal. The dilated residual dense network is used as the basic model of the recurrent recovery process. The network includes multiple paths with different receptive fields, thus can make full use of multi-scale redundancy and utilize context information in large regions. Furthermore, to handle heavy rain cases where rain streak accumulation is presented, we construct a detail appearing rain accumulation removal to not only improve the visibility but also enhance the details in dark regions. The evaluation on both synthetic and real images, particularly on those containing large rain streaks and heavy accumulation, shows the effectiveness of our novel models, which significantly outperforms the state-ofthe- art methods. Wenhan Yang, Jiaying Liu 0001, Shuai Yang 0001, Zongming Guo |
IEEE Trans. Image Process. | 2 |
| 2019 | Fine-Grained Quality Assessment for Compressed ImagesabstractImage quality assessment (IQA) has attracted more and more attention due to the urgent demand in image services. The perceptual-based image compression is one of the most prominent applications that require IQA metrics to be highly correlated with human vision. To explore IQA algorithms that are more consistent with human vision, several calibrated databases have been constructed. However, the distorted images in the existing databases are usually generated by corrupting the pristine images with various distortions in coarse levels, such that the IQA algorithms validated on them may be inefficient to optimize the perceptual-based image compression with fine-grained quality differences. In this paper, we construct a large-scale image database which can be used for fine-grained quality assessment of compressed images. In the proposed database, reference images are compressed at constant bitrate levels by JPEG encoders with different optimization methods. To distinguish subtle differences, the pair-wise comparison method is utilized to rank them in subjective experiments. We select 100 reference images for the proposed database, and each image is compressed into three target bitrates by four different JPEG optimization methods, such that 1200 distorted images are generated in total. Sixteen well-known IQA algorithms are evaluated and analyzed on the proposed database. With the devised fine-grained IQA database, we expect to further promote image quality assessment by shifting it from a coarse-grained stage to a fine-grained stage. The database is available at: https://sites.google.com/site/zhangxinf07/fg-iqa. Xinfeng Zhang 0001, Weisi Lin, Shiqi Wang 0001, Jiaying Liu 0001, Siwei Ma 0001, Wen Gao 0001 |
IEEE Trans. Image Process. | 4 |
| 2019 | Progressive Spatial Recurrent Neural Network for Intra PredictionabstractIntra prediction is an important component of modern video codecs, which is able to efficiently squeeze out the spatial redundancy in video frames. With preceding pixels as the context, traditional intra prediction schemes generate linear predictions based on several predefined directions (i.e., modes) for blocks to be encoded. However, these modes are relatively simple and their predictions may fail when facing blocks with complex textures, which leads to additional bits encoding the residue. In this paper, we design a progressive spatial recurrent neural network (PS-RNN) that learns to conduct intra prediction. Specifically, our PS-RNN consists of three spatial recurrent units and progressively generates predictions by passing information along from preceding contents to blocks to be encoded. To make our network generate predictions considering both distortion and bit rate, we propose using sum of absolute transformed difference (SATD) as the loss function to train PS-RNN since SATD is able to measure rate-distortion cost of encoding a residue block. Moreover, our method supports variable-block-size for intra prediction, which is more practical in real coding conditions. The proposed intra prediction scheme achieves on average 2.5% bit-rate reduction on variable-block-size settings under the same reconstruction quality compared with HEVC. Yueyu Hu, Wenhan Yang, Mading Li, Jiaying Liu 0001 |
IEEE Trans. Multim. | 4 |
| 2018 | Deep Retinex Decomposition for Low-Light Enhancement
Chen Wei 0005, Wenjing Wang 0001, Wenhan Yang, Jiaying Liu 0001 |
BMVC | 4 |
| 2018 | Erase or Fill? Deep Joint Recurrent Rain Removal and Reconstruction in VideosabstractIn this paper, we address the problem of video rain removal by constructing deep recurrent convolutional networks. We visit the rain removal case by considering rain occlusion regions, i.e. the light transmittance of rain streaks is low. Different from additive rain streaks, in such rain occlusion regions, the details of background images are completely lost. Therefore, we propose a hybrid rain model to depict both rain streaks and occlusions. With the wealth of temporal redundancy, we build a Joint Recurrent Rain Removal and Reconstruction Network (J4R-Net) that seamlessly integrates rain degradation classification, spatial texture appearances based rain removal and temporal coherence based background details reconstruction. The rain degradation classification provides a binary map that reveals whether a location is degraded by linear additive streaks or occlusions. With this side information, the gate of the recurrent unit learns to make a trade-off between rain streak removal and background details reconstruction. Extensive experiments on a series of synthetic and real videos with rain streaks verify the superiority of the proposed method over previous state-of-the-art methods. Jiaying Liu 0001, Wenhan Yang, Shuai Yang 0001, Zongming Guo |
CVPR | 1 |
| 2018 | Attentive Generative Adversarial Network for Raindrop Removal From a Single ImageabstractRaindrops adhered to a glass window or camera lens can severely hamper the visibility of a background scene and degrade an image considerably. In this paper, we address the problem by visually removing raindrops, and thus transforming a raindrop degraded image into a clean one. The problem is intractable, since first the regions occluded by raindrops are not given. Second, the information about the background scene of the occluded regions is completely lost for most part. To resolve the problem, we apply an attentive generative network using adversarial training. Our main idea is to inject visual attention into both the generative and discriminative networks. During the training, our visual attention learns about raindrop regions and their surroundings. Hence, by injecting this information, the generative network will pay more attention to the raindrop regions and the surrounding structures, and the discriminative network will be able to assess the local consistency of the restored regions. This injection of visual attention to both generative and discriminative networks is the main contribution of this paper. Our experiments show the effectiveness of our approach, which outperforms the state of the art methods quantitatively and qualitatively. Rui Qian 0003, Robby T. Tan, Wenhan Yang, Jiajun Su, Jiaying Liu 0001 |
CVPR | 5 |
| 2018 | Enhanced Intra Prediction with Recurrent Neural Network in Video CodingabstractIntra prediction is one of the important parts in video/image codec. With intra prediction mechanism, spatial redundancy can be largely removed for further bit saving. However, current state-of-the-art intra prediction method does not produce satisfactory prediction result due to its limits in reference samples and modeling ability. To enhance the intra prediction in HEVC, in this paper, a deep neural network featuring spatial RNN, which models the spatial dependency of pixels as sequential dynamics, is proposed to generate better prediction signals. Experimental results show improvement in BD-Rate for the proposed method compared with the original HEVC prediction scheme. Yueyu Hu, Wenhan Yang, Sifeng Xia, Wen-Huang Cheng, Jiaying Liu 0001 |
DCC | 5 |
| 2018 | A Group Variational Transformation Neural Network for Fractional Interpolation of Video CodingabstractMotion compensation is an important technology in video coding to remove the temporal redundancy between coded video frames. In motion compensation, fractional interpolation is used to obtain more reference blocks at sub-pixel level. Existing video coding standards commonly use fixed interpolation filters for fractional interpolation, which are not efficient enough to handle diverse video signals well. In this paper, we design a group variational transformation convolutional neural network (GVTCNN) to improve the fractional interpolation performance of the luma component in motion compensation. GVTCNN infers samples at different sub-pixel positions from the input integer-position sample. It first extracts a shared feature map from the integer-position sample to infer various sub-pixel position samples. Then a group variational transformation technique is used to transform a group of copied shared feature maps to samples at different sub-pixel positions. Experimental results have identified the interpolation efficiency of our GVTCNN. Compared with the interpolation method of High Efficiency Video Coding, our method achieves 1.9% bit saving on average and up to 5.6% bit saving under low-delay P configuration. Sifeng Xia, Wenhan Yang, Yueyu Hu, Siwei Ma 0001, Jiaying Liu 0001 |
DCC | 5 |
| 2018 | GLADNet: Low-Light Enhancement Network with Global AwarenessabstractIn this paper, we address the problem of lowlight enhancement. Our key idea is to first calculate a global illumination estimation for the low-light input, then adjust the illumination under the guidance of the estimation and supplement the details using a concatenation with the original input. Considering that, we propose a GLobal illumination Aware and Detail-preserving Network (GLADNet). The input image is rescaled to a certain size and then put into an encoder-decoder network to generate global priori knowledge of the illumination. Based on the global prior and the original input image, a convolutional network is employed for detail reconstruction. For training GLADNet, we use a synthetic dataset generated from RAW images. Extensive experiments demonstrate the superiority of our method over other compared methods on the real low-light images captured in various conditions. Wenjing Wang 0001, Chen Wei 0005, Wenhan Yang, Jiaying Liu 0001 |
FG | 4 |
| 2018 | Soft Decoding of Light Field Images Using Pocs and Fast Graph Spectrayl FiltersabstractLight field data captured by a lenslet-based image sensor is typically demosaicked, aligned and rearranged into a series of sub-aperture (viewpoint) images, before a disparity-compensated coding scheme is employed for compression. In this paper, we focus on the problem of soft decoding of block-based compressed sub-aperture images at the decoder: given quantization bin indices of DCT coefficients of non-overlapping code blocks, we select appropriate coefficient values that are low-pass filtered using graph spectral filters and view-consistent across sub-aperture images via projection on convex sets (POCS). Specifically, after an initial pixel estimate, we low-pass filter each pixel block using accelerated graph filters based on the Lanczos method. We then map filtered pixels to a neighborhood of sub-aperture images based on estimated disparity to enforce indexed quantization bin constraints of multiple images. Experimental results show that our algorithm achieves PSNR gain of 2.34dB over JPEG hard decoding. Shuai Yang 0001, Gene Cheung, Jiaying Liu 0001, Zongming Guo |
ICASSP | 3 |
| 2018 | Restoration of Unevenly Illuminated ImagesabstractIn this paper, we tackle the problem of restoring unevenly illuminated images. Generally, there exist three kinds of exposure conditions in these images: under-, normal-, and over-exposures. Thus, a three-component generalized Gaussian mixture model (3GGMM) is used to fit the histogram of the illuminance image, and probabilistically characterize the three exposure states. Based on the 3GGMM, separate optimal tone mapping functions are designed to enhance under- and overexposed regions by maximizing expected contrast of these regions. The output illumination can be obtained by fusing the restoration results in different exposure states. Experimental results validate the effectiveness of the proposed image restoration approach. Mading Li, Jiaying Liu 0001, Zongming Guo |
ICIP | 3 |
| 2018 | Dmcnn: Dual-Domain Multi-Scale Convolutional Neural Network for Compression Artifacts RemovalabstractJPEG is one of the most commonly used standards among lossy image compression methods. However, JPEG compression inevitably introduces various kinds of artifacts, especially at high compression rates, which could greatly affect the Quality of Experience (QoE). Recently, convolutional neural network (CNN) based methods have shown excellent performance for removing the JPEG artifacts. Lots of efforts have been made to deepen the CNN s and extract deeper features, while relatively few works pay attention to the receptive field of the network. In this paper, we illustrate that the quality of output images can be significantly improved by enlarging the receptive fields in many cases. One step further, we propose a Dual-domain Multi-scale CNN (DMCNN) to take full advantage of redundancies on both the pixel and DCT domains. Experiments show that DMCNN sets a new state-of-the-art for the task of JPEG artifact removal. Xiaoshuai Zhang, Wenhan Yang, Yueyu Hu, Jiaying Liu 0001 |
ICIP | 4 |
| 2018 | Skeleton-Indexed Deep Multi-Modal Feature Learning for High Performance Human Action RecognitionabstractThis paper presents a new framework for action recognition with multi-modal data. A skeleton-indexed feature learning procedure is developed to further exploit the detailed local features from RGB and optical flow videos. In particular, the proposed framework is built based on a deep Convolutional Network (ConvNet) and a Recurrent Neural Network (RNN) with Long Short Term Memory (LSTM). A skeleton-indexed transform layer is designed to automatically extract visual features around key joints, and a part-aggregated pooling is developed to uniformly regulate the visual features from different body parts and actors. Besides, several fusion schemes are explored to take advantage of multi-modal data. The proposed deep architecture is end-to-end trainable and can better incorporate different modalities to learn effective feature representations. Quantitative experiment results on two datasets, the NTU RGB+D dataset and the MSR dataset, demonstrate the excellent performance of our scheme over other state-of-the-arts. To our knowledge, the performance obtained by the proposed framework is currently the best on the challenging NTU RGB+D dataset. Sijie Song, Cuiling Lan, Junliang Xing, Wenjun Zeng 0001, Jiaying Liu 0001 |
ICME | 5 |
| 2018 | Joint Enhancement and Denoising Method via Sequential DecompositionabstractMany low-light enhancement methods ignore intensive noise in original images. As a result, they often simultaneously enhance the noise as well. Furthermore, extra denoising procedures adopted by most methods ruin the details. In this paper, we introduce a joint low-light enhancement and denoising strategy, aimed at obtaining well-enhanced low-light images while getting rid of the inherent noise issue simultaneously. The proposed method performs Retinex model based decomposition in a successive sequence, which sequentially estimates a piece-wise smoothed illumination and a noise-suppressed reflectance. After getting the illumination and reflectance map, we adjust the illumination layer and generate our enhancement result. In this noise-suppressed sequential decomposition process we enforce the spatial smoothness on each component and skillfully make use of weight matrices to suppress the noise and improve the contrast. Results of extensive experiments demonstrate the effectiveness and practicability of our method. It performs well for a wide variety of images, and achieves better or comparable quality compared with the state-of-the-art methods. Xutong Ren, Mading Li, Wen-Huang Cheng, Jiaying Liu 0001 |
ISCAS | 4 |
| 2018 | Dual Recovery Network with Online Compensation for Image Super-ResolutionabstractImage super-resolution (SR) methods essentially lead to a loss of some high-frequency (HF) information when predicting high-resolution (HR) images from low-resolution (LR) images without using external references. To address this issue, we additionally utilize online retrieved data to facilitate image SR in a unified deep framework. A novel dual high-frequency recovery network (DHN) is proposed to predict an HR image with three parts: an LR image, an internal inferred HF (IHF) map (HF missing part inferred solely from the LR image) and an external extracted HF (EHF) map. In particular, we infer the HF information based on both the LR image and similar HR references which are retrieved online. For the EHF map, we align the references with affine transformation and then in the aligned references, part of HF signals are extracted by the proposed DHN to compensate for the HF loss. Extensive experimental results demonstrate that our DHN achieves notably better performance than state-of-the-art SR methods. Sifeng Xia, Wenhan Yang, Jiaying Liu 0001, Zongming Guo |
ISCAS | 3 |
| 2018 | Session details: Panel-2
Jiaying Liu 0001, Wen-Huang Cheng |
ACM Multimedia | 1 |
| 2018 | AI + Multimedia Make Better Life?abstractNo abstract available. Wen-Huang Cheng, Jiaying Liu 0001, Mohan Kankanhalli, Abdulmotaleb El Saddik, Benoit Huet |
ACM Multimedia | 2 |
| 2018 | Context-Aware Unsupervised Text StylizationabstractIn this work, we present a novel algorithm to stylize the text without supervision, which provides a flexible and convenient way to invoke fantastic text expressions. Rather than employing the fixed pair of target text and source style images, our unsupervised framework establishes an implicit mapping for them by using an abstract imagery of the style image as bridges. Based on the mapping, we progressively narrow the visual discrepancy between text and style images by the proposed legibility-preserving structure transfer and texture transfer algorithms, which effectively balance the text legibility and style consistency. Furthermore, we explore a seamless composition of the stylized text and a background image, in which the optimal text layout is determined by a context-aware layout design algorithm utilizing cues for both seamlessness and aesthetics. Given the layout, the text can be seamlessly embedded into the background by texture synthesis under a context-aware boundary constraint. Experimental results demonstrate the effectiveness of the proposed method in automatic artistic typography creation and visual-textual presentation synthesis. Shuai Yang 0001, Jiaying Liu 0001, Wenhan Yang, Zongming Guo |
ACM Multimedia | 2 |
| 2018 | Optimized Spatial Recurrent Network for Intra Prediction in Video CodingabstractIntra prediction in modern video codecs is able to efficiently reduce spatial redundancy in video frames. With preceding pixels as context, traditional intra prediction schemes generate linear predictions based on several predefined directions (i.e. modes) for the current prediction unit (PU). However, these modes are relatively simple and are not able to handle complex textures, which leads to additional bits encoding the residue. In this paper, we design a convolutional neural network (CNN) guided spatial recurrent neural network (RNN) to improve the intra prediction in High-Efficiency Video Coding (HEVC). By exploring the correlations between pixels, the network learns to generate prediction signal in a progressive manner. The progressive model solves the problem of asymmetry in intra prediction naturally. As the model is designed for global context modeling, no flags for intra prediction modes selection need to be encoded. Our proposed intra prediction scheme achieves on average 1.2% bit-rate saving compared with HEVC. Yueyu Hu, Wenhan Yang, Sifeng Xia, Jiaying Liu 0001 |
VCIP | 4 |
| 2018 | Text effects transfer via distribution-aware texture synthesis
Shuai Yang 0001, Jiaying Liu 0001, Zhouhui Lian, Zongming Guo |
Comput. Vis. Image Underst. | 2 |
| 2018 | Video super-resolution based on spatial-temporal recurrent residual networks
Wenhan Yang, Jiashi Feng, Guosen Xie, Jiaying Liu 0001, Zongming Guo, Shuicheng Yan |
Comput. Vis. Image Underst. | 4 |
| 2018 | Kernel Wiener filtering model with low-rank approximation for image denoising
Yongqin Zhang, Jinsheng Xiao, Jiaying Liu 0001, Zongming Guo, Xiaopeng Zong |
Inf. Sci. | 5 |
| 2018 | Blind visual quality assessment for image super-resolution by convolutional neural network
Yuming Fang 0001, Chi Zhang 0027, Wenhan Yang, Jiaying Liu 0001, Zongming Guo |
Multim. Tools Appl. | 4 |
| 2018 | Automatic portrait oil painter: joint domain stylization for portrait images
Saboya Yang, Shuai Yang 0001, Wenhan Yang, Jiaying Liu 0001 |
Multim. Tools Appl. | 4 |
| 2018 | Adaptive Batch Normalization for practical domain adaptation
Yanghao Li, Naiyan Wang, Jianping Shi, Jiaying Liu 0001 |
Pattern Recognit. | 5 |
| 2018 | Isophote-Constrained Autoregressive Model With Adaptive Window Extension for Image InterpolationabstractThe autoregressive (AR) model is widely used in image interpolations. Traditional AR models consider utilizing the dependence between pixels to model the image signal. However, they ignore the valuable patch-level information for image modeling. In this paper, we propose to integrate both the pixel-level and patch-level information to depict the relationship between high-resolution and low-resolution pixels and obtain better image interpolation results. In particular, we propose an isophote-constrained AR (ICAR) model to perform AR-flavored interpolation within an identified joint stable region and further develop an AR interpolation with an adaptive window extension. Considering the smoothness along the isophote curve, the ICAR model searches only several successive similar patches along the isophote curve over a large region to construct an adaptive window. These overlapped patches, representing the patch-level structure similarity, are used to construct a joint AR model. To better characterize the piecewise stationarity and determine whether a pixel is suitable for AR estimation, we further propose pixel-level and patch-level similarity metrics and embed them into the ICAR model, introducing a weighted ICAR model. Comprehensive experiments demonstrate that our method can effectively reconstruct the edge structures and suppress jaggy or ringing artifacts. In the objective quality evaluation, our method achieves the best results in terms of both peak signal-to-noise ratio and structural similarity for both simple size doubling (two times) and for arbitrary scale enlargements. Wenhan Yang, Jiaying Liu 0001, Mading Li, Zongming Guo |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | Joint-Feature Guided Depth Map Super-Resolution With Face PriorsabstractIn this paper, we present a novel method to super-resolve and recover the facial depth map nicely. The key idea is to exploit the exemplar-based method to obtain the reliable face priors from high-quality facial depth map to improve the depth image. Specifically, a new neighbor embedding (NE) framework is designed for face prior learning and depth map reconstruction. First, face components are decomposed to form specialized dictionaries and then reconstructed, respectively. Joint features, i.e., low-level depth, intensity cues and high-level position cues, are put forward for robust patch similarity measurement. The NE results are used to obtain the face priors of facial structures and smooth maps, which are then combined in an uniform optimization framework to recover high-quality facial depth maps. Finally, an edge enhancement process is implemented to estimate the final high resolution depth map. Experimental results demonstrate the superiority of our method compared to state-of-the-art depth map super-resolution techniques on both synthetic data and real-world data from Kinect. Shuai Yang 0001, Jiaying Liu 0001, Yuming Fang 0001, Zongming Guo |
IEEE Trans. Cybern. | 2 |
| 2018 | Structure-Revealing Low-Light Image Enhancement Via Robust Retinex ModelabstractLow-light image enhancement methods based on classic Retinex model attempt to manipulate the estimated illumination and to project it back to the corresponding reflectance. However, the model does not consider the noise, which inevitably exists in images captured in low-light conditions. In this paper, we propose the robust Retinex model, which additionally considers a noise map compared with the conventional Retinex model, to improve the performance of enhancing low-light images accompanied by intensive noise. Based on the robust Retinex model, we present an optimization function that includes novel regularization terms for the illumination and reflectance. Specifically, we use norm to constrain the piece-wise smoothness of the illumination, adopt a fidelity term for gradients of the reflectance to reveal the structure details in low-light images, and make the first attempt to estimate a noise map out of the robust Retinex model. To effectively solve the optimization problem, we provide an augmented Lagrange multiplier based alternating direction minimization algorithm without logarithmic transformation. Experimental results demonstrate the effectiveness of the proposed method in low-light image enhancement. In addition, the proposed method can be generalized to handle a series of similar problems, such as the image enhancement for underwater or remote sensing and in hazy or dusty conditions. Mading Li, Jiaying Liu 0001, Wenhan Yang, Xiaoyan Sun 0001, Zongming Guo |
IEEE Trans. Image Process. | 2 |
| 2018 | Spatio-Temporal Attention-Based LSTM Networks for 3D Action Recognition and DetectionabstractHuman action analytics has attracted a lot of attention for decades in computer vision. It is important to extract discriminative spatio-temporal features to model the spatial and temporal evolutions of different actions. In this paper, we propose a spatial and temporal attention model to explore the spatial and temporal discriminative features for human action recognition and detection from skeleton data. We build our networks based on the recurrent neural networks with long short-term memory units. The learned model is capable of selectively focusing on discriminative joints of skeletons within each input frame and paying different levels of attention to the outputs of different frames. To ensure effective training of the network for action recognition, we propose a regularized cross-entropy loss to drive the learning process and develop a joint training strategy accordingly. Moreover, based on temporal attention, we develop a method to generate the action temporal proposals for action detection. We evaluate the proposed method on the SBU Kinect Interaction data set, the NTU RGB + D data set, and the PKU-MMD data set, respectively. Experiment results demonstrate the effectiveness of our proposed model on both action recognition and action detection. Sijie Song, Cuiling Lan, Junliang Xing, Wenjun Zeng 0001, Jiaying Liu 0001 |
IEEE Trans. Image Process. | 5 |
| 2018 | Structure-Guided Image Inpainting Using Homography TransformationabstractIn this paper, we present a novel structure-guided framework for exemplar-based image inpainting to maintain the neighborhood consistence and structure coherence of an inpainted region. The proposed method consists of a data term for pixel validity and boundary continuity, a smoothness term to depict the compatibility of neighboring pixels for contextual continuity, and a coherence term to investigate image inherent regularities to ensure image self-similarity. To better reconstruct image structures, the method utilizes image regularity statistics to extract dominant linear structures of the target image. Guided by these structures, homography transformations are estimated and combined to globally repair the missing region using the Markov random field model. To reduce computational complexity, a hierarchical process is implemented to utilize the regularity effectively. The experimental results demonstrate that our method yields better results for various real-world scenes than existing state-of-the-art image inpainting techniques. Jiaying Liu 0001, Shuai Yang 0001, Yuming Fang 0001, Zongming Guo |
IEEE Trans. Multim. | 1 |
| 2018 | Photo Stylistic Brush: Robust Style Transfer via Superpixel-Based Bipartite GraphabstractWith the rapid development of social network and multimedia technology, customized image and video stylization have been widely used for various social-media applications. In this paper, we explore the problem of exemplar-based photo style transfer, which provides a flexible and convenient way to invoke fantastic visual impression. Rather than investigating some fixed artistic patterns to represent certain styles as was done in some previous works, our work emphasizes styles related to a series of visual effects in the photograph (e.g., color, tone, and contrast). We propose a photo stylistic brush, an automatic robust style transfer approach based on Super pixel-based BIpartite Graph (SuperBIG). A two-step bipartite graph algorithm with different granularity levels is employed to aggregate pixels into superpixels and find their correspondences. In the first step, with the extracted hierarchical features, a bipartite graph is constructed to describe the content similarity for pixel partition to produce superpixels. In the second step, superpixels in the input/reference image are rematched to form a new superpixel-based bipartite graph, and superpixel-level correspondences are generated by bipartite matching. Finally, the refined correspondence guides SuperBIG to perform the transformation in a decorrelated color space. Extensive experimental results demonstrate the effectiveness and robustness of the proposed method for transferring various styles of exemplar images, even for some challenging cases, such as night images. Jiaying Liu 0001, Wenhan Yang, Xiaoyan Sun 0001, Wenjun Zeng 0001 |
IEEE Trans. Multim. | 1 |
| 2017 | An End-to-End Spatio-Temporal Attention Model for Human Action Recognition from Skeleton DataabstractHuman action recognition is an important task in computer vision. Extracting discriminative spatial and temporal features to model the spatial and temporal evolutions of different actions plays a key role in accomplishing this task. In this work, we propose an end-to-end spatial and temporal attention model for human action recognition from skeleton data. We build our model on top of the Recurrent Neural Networks (RNNs) with Long Short-Term Memory (LSTM), which learns to selectively focus on discriminative joints of skeleton within each frame of the inputs and pays different levels of attention to the outputs of different frames. Furthermore, to ensure effective training of the network, we propose a regularized cross-entropy loss to drive the model learning process and develop a joint training strategy accordingly. Experimental results demonstrate the effectiveness of the proposed model, both on the small human action recognition dataset of SBU and the currently largest NTU dataset. Sijie Song, Cuiling Lan, Junliang Xing, Wenjun Zeng 0001, Jiaying Liu 0001 |
AAAI | 5 |
| 2017 | Temporal Perceptive Network for Skeleton-Based Action Recognition
Yueyu Hu, Chunhui Liu 0002, Yanghao Li, Jiaying Liu 0001 |
BMVC | 4 |
| 2017 | Awesome Typography: Statistics-Based Text Effects TransferabstractIn this work, we explore the problem of generating fantastic special-effects for the typography. It is quite challenging due to the model diversities to illustrate varied text effects for different characters. To address this issue, our key idea is to exploit the analytics on the high regularity of the spatial distribution for text effects to guide the synthesis process. Specifically, we characterize the stylized patches by their normalized positions and the optimal scales to depict their style elements. Our method first estimates these two features and derives their correlation statistically. They are then converted into soft constraints for texture transfer to accomplish adaptive multi-scale texture synthesis and to make style element distribution uniform. It allows our algorithm to produce artistic typography that fits for both local texture patterns and the global spatial distribution in the example. Experimental results demonstrate the superiority of our method for various text effects over conventional style transfer methods. In addition, we validate the effectiveness of our algorithm with extensive artistic typography library generation. Shuai Yang 0001, Jiaying Liu 0001, Zhouhui Lian, Zongming Guo |
CVPR | 2 |
| 2017 | Deep Joint Rain Detection and Removal from a Single ImageabstractIn this paper, we address a rain removal problem from a single image, even in the presence of heavy rain and rain streak accumulation. Our core ideas lie in our new rain image model and new deep learning architecture. We add a binary map that provides rain streak locations to an existing model, which comprises a rain streak layer and a background layer. We create a model consisting of a component representing rain streak accumulation (where individual streaks cannot be seen, and thus visually similar to mist or fog), and another component representing various shapes and directions of overlapping rain streaks, which usually happen in heavy rain. Based on the model, we develop a multi-task deep learning architecture that learns the binary rain streak map, the appearance of rain streaks, and the clean background, which is our ultimate output. The additional binary map is critically beneficial, since its loss function can provide additional strong information to the network. To handle rain streak accumulation (again, a phenomenon visually similar to mist or fog) and various shapes and directions of overlapping rain streaks, we propose a recurrent rain detection and removal network that removes rain streaks and clears up the rain accumulation iteratively and progressively. In each recurrence of our method, a new contextualized dilated network is developed to exploit regional contextual information and to produce better representations for rain detection. The evaluation on real images, particularly on heavy rain, shows the effectiveness of our models and architecture. Wenhan Yang, Robby T. Tan, Jiashi Feng, Jiaying Liu 0001, Zongming Guo, Shuicheng Yan |
CVPR | 4 |
| 2017 | General scale interpolation via context-aware autoregressive model and multiplanar constraintabstractIn this paper, we propose a novel image interpolation algorithm suitable for general scale enlargement. Different from previous AR-based interpolation algorithms which employ predetermined reference configuration to predict pixel values, we consider the context information when building AR models. Optimal references are selected by incorporating nonlocal-based correlation coefficient and the indicator for local edge direction. Furthermore, the multiplanar constraint among similar patches is applied to enhance the correlation within the estimation window and serves as a kind of supplement to data fidelity term in AR model. The experimental results show that our method is effective in several enlargement scales and successfully alleviate the artifacts nearby edges and preserve their sharpness. The comparison experiments demonstrate that the proposed method can obtain desirable performance in terms of both objective and subjective results. Shihong Deng, Jiaying Liu 0001, Mading Li, Wenhan Yang, Zongming Guo |
ICASSP | 2 |
| 2017 | Online action detection and forecast via Multitask deep Recurrent Neural NetworksabstractOnline human action detection and forecast on untrimmed 3D skeleton sequences is a novel task based on traditional action recognition and has not been fully studied. Its aim is to localize and recognize one action in a long sequence while doing forecasting task at the same time. In this paper, we propose an online detection algorithm featuring Multi-Task Recurrent Neural Network to solve this problem. First, a deep Long Short Term Memory (LSTM) network is designed for feature extraction and temporal dynamic modeling. Then we utilize a classification subnetwork to classify one action, and predict the status of it at the same time. To forecast the occurrence of actions and estimate the accurate time of occurrence, we incorporate a regression subnetwork to our model. Then we split the action classes to three stages and train the model by optimizing a joint classification regression objective function. Experimental results show that the proposed model achieves satisfactory results on online action detection and forecast. Chunhui Liu 0002, Yanghao Li, Yueyu Hu, Jiaying Liu 0001 |
ICASSP | 4 |
| 2017 | 1+N fusion: Cascaded self-portrait enhancementabstractIn this paper, we present a novel cascaded framework to solve a self-portrait enhancement problem we call “1+N” problem, in which a self-portrait is enhanced with the help of N supporting photos that share the same scene and similar shooting time. The key idea is to exploit the extra information of these N photos to expand the field of view of the self-portrait and improve its lighting style. We achieve this by alternatingly optimizing two complementary tasks, namely illumination unification and photo registration. Based on the correspondences extracted in the input 1+N photos, our method estimates and updates the illumination and registration coefficients in a cascaded manner. Then a Markov Random Field formulation is proposed to globally fuse the aligned photos. Experimental results demonstrate the proposed method achieves high-quality results in this novel application scenario. Shuai Yang 0001, Jiaying Liu 0001, Sifeng Xia, Zongming Guo |
ICASSP | 2 |
| 2017 | Factorized Bilinear Models for Image RecognitionabstractAlthough Deep Convolutional Neural Networks (CNNs) have liberated their power in various computer vision tasks, the most important components of CNN, convolutional layers and fully connected layers, are still limited to linear transformations. In this paper, we propose a novel Factorized Bilinear (FB) layer to model the pairwise feature interactions by considering the quadratic terms in the transformations. Compared with existing methods that tried to incorporate complex non-linearity structures into CNNs, the factorized parameterization makes our FB layer only require a linear increase of parameters and affordable computational cost. To further reduce the risk of overfitting of the FB layer, a specific remedy called DropFactor is devised during the training process. We also analyze the connection between FB layer and some existing models, and show FB layer is a generalization to them. Finally, we validate the effectiveness of FB layer on several widely adopted datasets including CIFAR-10, CIFAR-100 and ImageNet, and demonstrate superior results compared with various state-of-the-art deep models. Yanghao Li, Naiyan Wang, Jiaying Liu 0001 |
ICCV | 3 |
| 2017 | Deep joint discriminative learning for vehicle re-identification and retrievalabstractIn this paper, we propose a novel vehicle re-identification method based on a Deep Joint Discriminative Learning (DJDL) model, which utilizes a deep convolutional network to effectively extract discriminative representations for vehicle images. To exploit properties and relationship among samples in different views, we design a unified framework to combine several different tasks efficiently, including identification, attribute recognition, verification and triplet tasks. The whole network is optimized jointly via a specific batch composition design. Extensive experiments are conducted on a large-scale VehicleID [1] dataset. Experimental results demonstrate the effectiveness of our method and show that it achieves the state-of-the-art performance on both vehicle re-identification and retrieval. Yanghao Li, Hongfei Yan, Jiaying Liu 0001 |
ICIP | 4 |
| 2017 | Variation learning guided convolutional network for image interpolationabstractIn this paper, we propose a variational learning model that effectively exploits the structural similarities for image representation, and construct a deep network based on this model for image interpolation. Based on the local dependency, our learning model represents an image as the three-dimensional features. Besides two coordinate dimensions, an additional neighboring variation dimension is added to encode every pixel as the variation to its nearest low-resolution pixel by the local similarity. This added dimension lowers the risk of over-fitting for learning approaches and constructs abundant structural correspondences for inferring the missing information lost in image degradation. Then, this three-dimensional features are naturally modeled, extracted and refined by an end-to-end trainable recurrent convolutional network for image interpolation. Comprehensive experiments demonstrate that our method leads to a surprisingly superior performance and offers new state-of-the-art benchmark. Wenhan Yang, Jiaying Liu 0001, Sifeng Xia, Zongming Guo |
ICIP | 2 |
| 2017 | Soft segmentation-guided bipartite graph image stylizationabstractIn this paper, we propose a photo stylistic brush, an automatic robust style transfer approach based on soft segmentation-guided bipartite graph. A two-step bipartite graph algorithm with different granularity levels is employed to aggregate pixels into superpixel and find their correspondences. In the first step, with the extracted hierarchical features, a bipartite graph is constructed to describe the content similarity for pixel partition to produce superpixels. In the second step, superpixels in the input/reference image are rematched to form a new soft segmentation-guided bipartite graph, and superpixel-level correspondences are generated by a bipartite matching. Finally, the refined correspondence guides our approach to perform the transfer in a decorrelated color space. Extensive experimental results demonstrate the effectiveness and robustness of the proposed method for transferring various styles of exemplar images, even for some challenging cases, such as night images. Saboya Yang, Jiaying Liu 0001, Wenhan Yang, Shuai Yang 0001, Chunpeng Li |
ICIP | 2 |
| 2017 | Demystifying Neural Style TransferabstractNeural Style Transfer has recently demonstrated very exciting results which catches eyes in both academia and industry. Despite the amazing results, the principle of neural style transfer, especially why the Gram matrices could represent style remains unclear. In this paper, we propose a novel interpretation of neural style transfer by treating it as a domain adaptation problem. Specifically, we theoretically show that matching the Gram matrices of feature maps is equivalent to minimize the Maximum Mean Discrepancy (MMD) with the second order polynomial kernel. Thus, we argue that the essence of neural style transfer is to match the feature distributions between the style images and the generated images. To further support our standpoint, we experiment with several other distribution alignment methods, and achieve appealing results. We believe this novel interpretation connects these two important research fields, and could enlighten future researches. Yanghao Li, Naiyan Wang, Jiaying Liu 0001 |
IJCAI | 3 |
| 2017 | Joint-domain unsupervised stylization for portraitsabstractPeople wish to own a portrait painting of themselves by Da Vinci. Unfortunately, it is impossible to make this dream come true; nevertheless, it may give us an opportunity by transferring some artistic features from one single reference painting. To address this issue, we propose a joint-domain image stylization approach, particularly for portrait oil paintings. From the view of artistic appreciation, we analyze an amount of oil painting artworks and summarize three critical factors to depict the figure, i.e. color, structure and texture. First, the tone of the input image is recolored based on semantic regions corresponding to the reference. Those semantic regions are segmented automatically via the color swatch, by considering the constraints of colors and positions. Then, we exploit sparse representation to reconstruct the layout by acquiring the structure from the reference. The paired training set for sparse dictionary learning is built with the guidance of edge features. Third, considering that texture is usually locally stochastic but regularly repetitive in global, a coarse-to-fine texture synthesis is used to enhance the detail pattern. Subjective results demonstrate the proposed method achieves desirable results compared with state-of-art methods while keeping consistent with artist's style. Saboya Yang, Jiaying Liu 0001, Shuai Yang 0001, Wenhan Yang, Zongming Guo |
ISCAS | 2 |
| 2017 | Real-Time Deep Video SpaTial Resolution UpConversion SysTem (STRUCT++ Demo)abstractImage and video super-resolution (SR) has been explored for several decades. However, few works are integrated into practical systems for real-time image and video SR. In this work, we present a real-time deep video SpaTial Resolution UpConversion SysTem (STRUCT++). Our demo system achieves real-time performance (50 fps on CPU for CIF sequences and 45 fps on GPU for HDTV videos) and provides several functions: 1) batch processing; 2) full resolution comparison; 3) local region zooming in. These functions are convenient for super-resolution of a batch of videos (at most 10 videos in parallel), comparisons with other approaches and observations of local details of the SR results. The system is built on a Global context aggregation and Local queue jumping Network (GLNet). It has a thinner and deeper network structure to aggregate global context with an additional local queue jumping path to better model local structures of the signal. GLNet achieves state-of-the-art performance for real-time video SR. Wenhan Yang, Shihong Deng, Yueyu Hu, Junliang Xing, Jiaying Liu 0001 |
ACM Multimedia | 5 |
| 2017 | Real-time deep image super-resolution via global context aggregation and local queue jumpingabstractDeep learning-based image super-resolution has provided very impressive reconstruction quality. However, their running time still sets barriers for real-time applications. In this paper, we propose a Global context aggregation and Local queue jumping Network (GLNet) which provides the more effective image SR given a certain number of model parameters. In our GLNet, we reconsider the model design of the real-time image SR paradigm. Then, we construct a deep network with fewer channels but a deeper structure to effectively aggregate the global context. The dilated convolutions are used as parts of basic units of our GLNet, which further enlarges the receptive field. Besides, an additional local queue jumping path is employed to connect the first-layer feature map and the last-layer feature map to better model the local signal structure. Extensive experiments demonstrate the superiority of our GLNet which offers new state-of-the-art performance considering both reconstruction quality and time consumption. Yueyu Hu, Jiaying Liu 0001, Wenhan Yang, Shihong Deng, Luyao Zhang 0007, Zongming Guo |
VCIP | 2 |
| 2017 | Objective Quality Assessment of Screen Content Images by Uncertainty WeightingabstractIn this paper, we propose a novel full-reference objective quality assessment metric for screen content images (SCIs) by structure features and uncertainty weighting (SFUW). The input SCI is first divided into textual and pictorial regions. The visual quality of textual regions is estimated based on perceptual structural similarity, where the gradient information is adopted as the structural feature. To predict the visual quality of pictorial regions in SCIs, we extract the structural features and luminance features for similarity computation between the reference and distorted pictorial patches. To obtain the final visual quality of SCI, we design an uncertainty weighting method by perceptual theories to fuse the visual quality of textual and pictorial regions effectively. Experimental results show that the proposed SFUW can obtain better performance of visual quality prediction for SCIs than other existing ones. Yuming Fang 0001, Jiebin Yan, Jiaying Liu 0001, Shiqi Wang 0001, Qiaohong Li, Zongming Guo |
IEEE Trans. Image Process. | 3 |
| 2017 | Deep Edge Guided Recurrent Residual Learning for Image Super-ResolutionabstractIn this paper, we consider the image super-resolution (SR) problem. The main challenge of image SR is to recover high-frequency details of a low-resolution (LR) image that are important for human perception. To address this essentially ill-posed problem, we introduce a Deep Edge Guided REcurrent rEsidual (DEGREE) network to progressively recover the high-frequency details. Different from most of the existing methods that aim at predicting high-resolution (HR) images directly, the DEGREE investigates an alternative route to recover the difference between a pair of LR and HR images by recurrent residual learning. DEGREE further augments the SR process with edge-preserving capability, namely the LR image and its edge map can jointly infer the sharp edge details of the HR image during the recurrent recovery process. To speed up its training convergence rate, by-pass connections across the multiple layers of DEGREE are constructed. In addition, we offer an understanding on DEGREE from the view-point of sub-band frequency decomposition on image signal and experimentally demonstrate how the DEGREE can recover different frequency bands separately. Extensive experiments on three benchmark data sets clearly demonstrate the superiority of DEGREE over the well-established baselines and DEGREE also provides new state-of-the-arts on these data sets. We also present addition experiments for JPEG artifacts reduction to demonstrate the good generality and flexibility of our proposed DEGREE network to handle other image processing tasks. Wenhan Yang, Jiashi Feng, Jianchao Yang, Fang Zhao 0006, Jiaying Liu 0001, Zongming Guo, Shuicheng Yan |
IEEE Trans. Image Process. | 5 |
| 2017 | Retrieval Compensated Group Structured Sparsity for Image Super-ResolutionabstractSparse representation-based image super-resolution is a well-studied topic; however, a general sparse framework that can utilize both internal and external dependencies remains unexplored. In this paper, we propose a group-structured sparse representation approach to make full use of both internal and external dependencies to facilitate image super-resolution. External compensated correlated information is introduced by a two-stage retrieval and refinement. First, in the global stage, the content-based features are exploited to select correlated external images. Then, in the local stage, the patch similarity, measured by the combination of content and high-frequency patch features, is utilized to refine the selected external data. To better learn priors from the compensated external data based on the distribution of the internal data and further complement their advantages, nonlocal redundancy is incorporated into the sparse representation model to form a group sparsity framework based on an adaptive structured dictionary. Our proposed adaptive structured dictionary consists of two parts: one trained on internal data and the other trained on compensated external data. Both are organized in a cluster-based form. To provide the desired over-completeness property, when sparsely coding a given LR patch, the proposed structured dictionary is generated dynamically by combining several of the nearest internal and external orthogonal subdictionaries to the patch instead of selecting only the nearest one as in previous methods. Extensive experiments on image super-resolution validate the effectiveness and state-of-the-art performance of the proposed method. Additional experiments on contaminated and uncorrelated external data also demonstrate its superior robustness. Jiaying Liu 0001, Wenhan Yang, Xinfeng Zhang 0001, Zongming Guo |
IEEE Trans. Multim. | 1 |
| 2016 | MARLow: A Joint Multiplanar Autoregressive and Low-Rank Approach for Image Completion
Mading Li, Jiaying Liu 0001, Zhiwei Xiong, Xiaoyan Sun 0001, Zongming Guo |
ECCV (7) | 2 |
| 2016 | Online Human Action Detection Using Joint Classification-Regression Recurrent Neural Networks
Yanghao Li, Cuiling Lan, Junliang Xing, Wenjun Zeng 0001, Chunfeng Yuan, Jiaying Liu 0001 |
ECCV (7) | 6 |
| 2016 | Joint sub-band based neighbor embedding for image super-resolutionabstractIn this paper, we propose a novel neighbor embedding method based on joint sub-bands for image super-resolution. Rather than directly reconstructing the total spatial variations of the input image, we restore each frequency component separately. The input LR image is decomposed into sub-bands defined by steerable filters to capture structural details on different directional frequency components. Then the neighbor embedding principle is employed to reconstruct each band, respectively. Moreover, taken the diverse characteristics of each band into account, we adopt adaptive similarity criteri-ons for searching nearest neighbors. Finally, we recombine the generated HR sub-bands by applying the inverting subband decomposition to get the final super-resolved result. Experimental results demonstrate the effectiveness of our method both in objective and subjective qualities comparing with other state-of-the-art methods. Sijie Song, Yanghao Li, Jiaying Liu 0001, Zongming Quo |
ICASSP | 3 |
| 2016 | Structure-guided image completion via regularity statisticsabstractIn this paper, we propose a novel hierarchical image completion approach using regularity statistics, considering structure features. Guided by dominant structures, the target image is used to generate reference images in a self-reproductive way by image data enhancement. The structure-guided image data enhancement allows us to expand the search space for samples. A Markov Random Field model is used to guide the enhanced image data combination to globally reconstruct the target image. For lower computational complexity and more accurate structure estimation, a hierarchical process is implemented. Experiments demonstrate the effectiveness of our method comparing to several state-of-the-art image completion techniques. Shuai Yang 0001, Jiaying Liu 0001, Sijie Song, Mading Li, Zongming Quo |
ICASSP | 2 |
| 2016 | A new reversible data hiding scheme exploiting high-dimensional prediction-error histogramabstractPairwise prediction-error expansion (pairwise PEE) is an improvement of the conventional PEE and it can provide excellent performance for reversible data hiding (RDH). Unlike PEE in which the prediction-errors are modified individually, the correlation among prediction-errors is exploited in pairwise PEE by jointly modifying each prediction-error pair. In this paper, the idea of pairwise PEE is developed and a new RDH scheme is proposed. A three-dimensional prediction-error histogram (3D-PEH) is generated by counting every non-overlapped prediction-error triple. Then, data embedding is conducted by modifying the 3D-PEH with a specifically designed reversible mapping. By using 3D-PEH and the proposed reversible mapping, the inter-correlation of prediction-errors is better exploited, and the performance of PEE is significantly enhanced. Moreover, the superiority of our method over pairwise PEE and some other state-of-the-art RDH methods is also experimentally verified. The proposed method is an effective extension of PEE towards the direction of high-dimensional histogram modification. Siren Cai, Xiaolong Li 0001, Jiaying Liu 0001, Zongming Guo |
ICIP | 3 |
| 2016 | Quality assessment for image super-resolution based on energy change and texture variationabstractIn this paper, we propose a novel reduced-reference quality assessment metric for image super-resolution (RRIQA-SR) based on the low-resolution (LR) image information. First, we use the Markov Random Field (MRF) to model the pixel correspondence between LR and high-resolution (HR) images. Based on the pixel correspondence, we predict the perceptual similarity between image patches of LR and HR images by two components: the energy change and texture variation. The overall quality of HR images is estimated by the perceptual similarity between local image patches of LR and HR images. Experimental results demonstrate that the proposed method can obtain better performance of quality prediction for HR images than other existing ones, even including some full-reference (FR) metrics. Yuming Fang 0001, Jiaying Liu 0001, Yabin Zhang 0002, Weisi Lin, Zongming Guo |
ICIP | 2 |
| 2016 | Local ternary pattern based on path integral for steganalysisabstractThe least significant bit (LSB) matching is a steganographic method which embeds the stego signal into cover images in the spatial domain. However, the stego signal disturbs the correlation of neighboring pixels in cover image and this can be utilized for steganalysis. Local binary pattern (LBP) is an effective image texture descriptor, and it can summarize the correlation of neighboring pixels. In this paper, a LBP-based steganalyzer is proposed to identify the deviations of the correlation violated by the stego noise. Specifically, our paper proposes the local ternary pattern based on path integral (pi-LTP) to enhance the feature discrimination in large-scale pixels. Moreover, a greedy incremental algorithm is utilized in our method to select the optimal subspace of pi-LTP features. Experimental results show our method has a better performance than the state-of-the-art steganalysis methods. Qiuyan Lin, Jiaying Liu 0001, Zongming Guo |
ICIP | 2 |
| 2016 | Robust and automatic video colorization via multiframe reordering refinementabstractIn this paper, we propose a robust video colorization method automatically through limited color references in a video sequence. The proposed method first estimates motion vectors between a monochrome frame and colored reference frames for initial matching by optical flow. Then it transfers color information to matched points in the monochrome frame and further propagates color information of matched points to other parts of the monochrome frame. Furthermore, we design a multiframe reordering refinement to colorize video sequences robustly. Experimental results demonstrate that the proposed method achieves much better performance in video colorization than state-of-the-art methods. Sifeng Xia, Jiaying Liu 0001, Yuming Fang 0001, Wenhan Yang, Zongming Guo |
ICIP | 2 |
| 2016 | Facial depth map enhancement via neighbor embeddingabstractThe simple yet subtle structures of faces make it difficult to capture the fine differences between different facial regions in the depth map, especially for consumer devices like Kinect. To address this issue, we present a novel method to super-solve and recover the facial depth map nicely. The key idea of our approach is to exploit the learning-based method to obtain the reliable face priors from high quality facial depth map to further improve the depth image. Specifically, we utilize the neighbor embedding framework. First, face components are decomposed to train specialized dictionaries and reconstructed, respectively. Joint features, i.e. color, depth and position cues, are put forward for robust patch similarity measurement. The neighbor embedding results form high frequency cues of facial depth details and gradients. Finally, an optimization function is defined to combine these high frequency information to yield depth maps that fit the actual face structures better. Experimental results demonstrate the superiority of our method compared to state-of-the-art techniques in recovering both synthetic data and real world data from Kinect. Shuai Yang 0001, Sijie Song, Qikun Guo, Xiaoqing Lu, Jiaying Liu 0001 |
ICPR | 5 |
| 2016 | Computational modeling of artistic intention: Quantify lighting surprise for painting analysisabstractThe use of strong lighting contrast to accentuate objects and figures in a painting—called Chiaroscuro—is popular among Renaissance painters such as Caravaggio, La Tour and Rembrandt. In this paper, we propose a new metric called LuCo to quantify the extent to which Chiaroscuro is employed by an artist in a painting. This measurement could be used to assess the capability of any system to fulfill the original artistic intention and consequently ensure minimal disruptions of Quality of Experience. We first argue that Chiaroscuro is a device for artists to draw attention to specific spatial regions; thus it can be understood as a restricted notion of visual saliency computed using only luminance features. Operationally, using a set of local luminance patches we first compute a Bayesian surprise value, where the prior and posterior probabilities are computed assuming a Gaussian Markov Random Field (GMRF) model. Inverse covariance matrices of the GMRF model are estimated via sparse graph learning for robustness. We construct a histogram using the computed surprise values from different local patches in a painting. Finally, we compute a skewness parameter for the constructed histogram as our LuCo score: large skewness means luminance surprises are either very small or very large, meaning that the artist accentuated lighting contrast in the painting. Experimental results show that paintings by Chiaroscuro artists have higher LuCo scores than 19th century French Impressionists, and Rembrandt's self-portraits have increasingly higher LuCo scores as he aged except for his late period—both trends are in agreement with art historians' interpretations. Saboya Yang, Gene Cheung, Patrick Le Callet, Jiaying Liu 0001, Zongming Guo |
QoMEX | 4 |
| 2016 | Autoregressive image interpolation via context modeling and multiplanar constraintabstractIn this paper, we propose a novel image interpolation algorithm by context-aware autoregressive (AR) model and multiplanar constraint. Different from existing AR based methods which employ predetermined reference configuration to predict pixel values, the proposed method considers the anisotropic pixel dependencies in natural images and adaptively chooses the optimal prediction context by utilizing the nonlocal redundancy to interpolate pixels. Furthermore, the multiplanar constraint is applied to enhance the correlations within the estimation window by exploiting the self-similarity property of natural images. Similar patches are collected by the combination of patch-wise pixel values and the gradient information. And the inter-patch dependencies are adopted to improve the interpolation. The experimental results show that our method is effective in image interpolation and successfully decreases the artifacts nearby the sharp edges. The comparison experiments demonstrate that the proposed method can obtain better performance than other related ones in terms of both objective and subjective results. Shihong Deng, Jiaying Liu 0001, Mading Li, Wenhan Yang, Zongming Guo |
VCIP | 2 |
| 2015 | Image Restoration Based on 3-D Autoregressive Model via Low-Rank MinimizationabstractDue to all kinds of need of customers and the complicated transmitting environment of digital image and video resources, numerous practical applications emerge, e.g. Image in painting, interpolation, super-resolution and the removal of salt and pepper noise. One thing these cases all have in common is that there are plenty of missing pixels randomly distributed in an image. Existing image restoration methods aiming at solving this problem include kernel regression [1], matrix completion [2] and total variation (TV) model [3]. The 3-D AR model has also been proposed to detect and interpolate the missing data in video sequences. However, the missing rate or missing region in these papers is usually small. With the missing rate increasing, known pixels in a local neighborhood are not going to be enough to form a solvable linear system. Thus, generally speaking, AR model is not suitable for image restoration from high missing rates. Nevertheless, with proper preliminary processing as proposed in this paper, AR models can be well utilized and present good results even in high missing rates. In this paper, we propose a novel method for image restoration. For the first time, the 3-D AR model is utilized in a single image to simultaneously measure correlation within and between similar patches. 2-D AR model combining with a multiscale structure reconstruct the image using its low-resolution versions to preserve important perceptual statistics such as edges. After obtaining the preliminary reconstruction of the reconstructed full size image, similar patches are collected and the 3-D AR model is applied to form a more local-consistent patch set. Then, an iterative singular value thresholding (SVT) method is utilized to solve the low-rank minimization problem. Instead of aggregating all the overlapped patches after each patch set is processed, we perform SVT for each patch set and aggregate all the overlapped patches into an intermediate image, then the iterative regularization is carried out on the image to produce the newly output for next iteration. Experimental results demonstrate that the proposed method achieves higher PSNR and SSIM than state-of-the-art methods [1-3] and the processed images possess a better visual quality especially in edge structures and texture regions. Mading Li, Jiaying Liu 0001, Zongming Guo |
DCC | 2 |
| 2015 | Neighborhood regression for edge-preserving image super-resolutionabstractThere have been many proposed works on image super-resolution via employing different priors or external databases to enhance HR results. However, most of them do not work well on the reconstruction of high-frequency details of images, which are more sensitive for human vision system. Rather than reconstructing the whole components in the image directly, we propose a novel edge-preserving super-resolution algorithm, which reconstructs low- and high-frequency components separately. In this paper, a Neighborhood Regression method is proposed to reconstruct high-frequency details on edge maps, and low-frequency part is reconstructed by the traditional bicubic method. Then, we perform an iterative combination method to obtain the estimated high resolution result, based on an energy minimization function which contains both low-frequency consistency and high-frequency adaptation. Extensive experiments evaluate the effectiveness and performance of our algorithm. It shows that our method is competitive or even better than the state-of-art methods. Yanghao Li, Jiaying Liu 0001, Wenhan Yang, Zongming Guo |
ICASSP | 2 |
| 2015 | Novel autoregressive model based on adaptive window-extension and patch-geodesic distance for image interpolationabstractIn this paper, we propose a novel autoregressive (AR) model based on the adaptive window and the patch-geodesic distance for the image interpolation. The model combines the information of inner/inter-patch correlation. To model the inner-patch correlation, we introduce a patch-geodesic distance similarity metric. The proposed metric shows the desirable capacity to depict the piecewise-stationarity of natural images. For the inter-patch correlation, we introduce the inter-patch structure variation and propose an adaptive window-extension AR model. The model extends the interpolation window according to the local structural variation, increasing the adaptation without violating the consistency. Comprehensive experiments demonstrate that the proposed method is better than or competitive with state-of-the-art interpolation methods in both objective and subjective quality evaluations. Wenhan Yang, Jiaying Liu 0001, Shuai Yang 0001, Zongming Guo |
ICASSP | 2 |
| 2015 | Multi-pose face hallucination via neighbor embedding for facial componentsabstractIn this paper, we propose a novel multi-pose face hallucination method based on Neighbor Embedding for Facial Components (NEFC) to magnify face images with various poses and expressions. To represent the structure of a face, a facial component decomposition is employed on each face image. Then, a neighbor embedding reconstruction method with locality-constraint is performed for each facial component. For the video scenario, we utilize optical flow to locate the position of each patch among the neighboring frames and make use of the Intra and Inter Nonlocal Means method to preserve consistency between neighboring frames. Experimental results evaluate the effectiveness and adaptability of our algorithm. It shows that our method achieves better performance than the state-of-the-art methods, especially on the face images with various poses and expressions. Yanghao Li, Jiaying Liu 0001, Wenhan Yang, Zongming Guo |
ICIP | 2 |
| 2015 | Adaptive autoregressive model with window extension via explicit geometry for image interpolationabstractIn this paper, we propose a novel adaptive autoregressive (AR) model constructed with an explicit geometry based extended window for image interpolation. Geometric features are chosen as criterions to include more useful pixels. These features are estimated explicitly and guide the interpolation window to extend adaptively. To characterize the piecewise stationary of images, the patch-geodesic distance based similarity is proposed and modulated into the adaptive AR model. For increasing the precision of the parameter estimation, a weighted ridge regression based estimation is employed. With the estimation, the multicollinearity between parameters, which occurs in piecewise stationarity conditions, is eliminated. Experimental results demonstrate that the proposed method is better than or competitive with state-of-the-art interpolation methods in both objective and subjective quality evaluations. Qingyun Wang 0007, Jiaying Liu 0001, Wenhan Yang, Zongming Guo |
ICIP | 2 |
| 2015 | Hierarchical oil painting stylization with limited reference via sparse representationabstractTraditional image stylization is enforced by learning the mappings with an external paired training set. But in practice, people usually encounter a specific stylish image and want to transfer its style to their own pictures without the external dataset. Thus, we propose a hierarchical stylization model with limited reference particularly for oil paintings. First, the edge patch based dictionary is trained to build connections between images and limited reference, then reconstruct the structure layer. Due to the highly structured property of saliency regions, the saliency mask is extracted to integrate the structure layer and the texture layer with different weights. Hence, the advantages of both sparse representation based methods and example based methods are integrated. Moreover, the color layer and the surface layer are considered to make the output more consistent with the artist's individual oil painting style. Subjective results demonstrate the proposed method produces desirable results with state-of-art methods while keeping consistent with the artist's oil painting style. Saboya Yang, Jiaying Liu 0001, Shuai Yang 0001, Sifeng Xia, Zongming Guo |
MMSP | 2 |
| 2015 | Image super-resolution via nonlocal similarity and group structured sparse representationabstractSparse prior provides an effective tool for the image reconstruction. However, the sparse coding for independent patches leads to the unstable sparse decomposition. In this paper, we propose a group structured sparse representation model by considering the nonlocal similarity. The nonlocal similar patches are collected and classified into groups. Patches in the same group are reconstructed based the same basis of dictionaries. The dictionary is organized as the combination of many orthogonal sub-dictionaries. To provide the redundancy, the dictionary used for the sparse coding is generated online with several sub-dictionaries, thus it is over-complete. We apply the proposed model into a gradual SR framework. The framework enlarges LR to HR by a patch enhancement and an alternative sparse reconstruction on the patch and group. Objective quality evaluation shows that our proposed SR method achieves highest PSNR results comparing with the state-of-the-art methods. And subjective results demonstrate the proposed method reduces artifacts and preserves more details. Wenhan Yang, Jiaying Liu 0001, Saboya Yang, Zongming Quo |
VCIP | 2 |
| 2015 | Adaptive General Scale Interpolation Based on Weighted Autoregressive ModelsabstractThe autoregressive (AR) model has been widely used in signal processing for its effective estimation, especially in image processing. Many dedicated 2× interpolation algorithms adopt the AR model to describe the strong correlation between low-resolution (LR) pixels and high-resolution (HR) pixels. However, these AR model-based methods closely depend on the fixed relative position between LR pixels and HR pixels that are nonexistent in the general scale interpolation. In this paper, we present an adaptive general scale interpolation algorithm that is capable of arbitrary scaling factors considering the nonstationarity of natural images. Different from other dedicated 2× interpolation methods, the proposed AR terms are modeled by pixels with their adjacent unknown HR neighbors. To compensate for the information loss caused by mismatches of AR models, we consider a weighting scheme suitable for general scale situations based on the pixel similarity to increase accuracy of the estimation. Comprehensive experiments demonstrate the effectiveness of the proposed method on general scaling factors. The maximum gain of peak signal-to-noise ratio is 2.07 dB compared with segment adaptive gradient angle in 1.5× enlargements. To evaluate the performance in resolution adaptive video coding, we have also tested our method on Joint Scalable Video Model codec and obtained better subjective quality and rate-distortion performance. Mading Li, Jiaying Liu 0001, Jie Ren 0012, Zongming Guo |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2015 | Image Super-Resolution Based on Structure-Modulated Sparse RepresentationabstractSparse representation has recently attracted enormous interests in the field of image restoration. The conventional sparsity-based methods enforce sparse coding on small image patches with certain constraints. However, they neglected the characteristics of image structures both within the same scale and across the different scales for the image sparse representation. This drawback limits the modeling capability of sparsity-based super-resolution methods, especially for the recovery of the observed low-resolution images. In this paper, we propose a joint super-resolution framework of structure-modulated sparse representations to improve the performance of sparsity-based image super-resolution. The proposed algorithm formulates the constrained optimization problem for high-resolution image recovery. The multistep magnification scheme with the ridge regression is first used to exploit the multiscale redundancy for the initial estimation of the high-resolution image. Then, the gradient histogram preservation is incorporated as a regularization term in sparse modeling of the image super-resolution problem. Finally, the numerical solution is provided to solve the super-resolution problem of model parameter estimation and sparse representation. Extensive experiments on image super-resolution are carried out to validate the generality, effectiveness, and robustness of the proposed algorithm. Experimental results demonstrate that our proposed algorithm, which can recover more fine structures and details from an input low-resolution image, outperforms the state-of-the-art methods both subjectively and objectively in most cases. Yongqin Zhang, Jiaying Liu 0001, Wenhan Yang, Zongming Guo |
IEEE Trans. Image Process. | 2 |
| 2015 | Video Compression Artifact Reduction via Spatio-Temporal Multi-Hypothesis PredictionabstractAnnoying compression artifacts exist in most of lossy coded videos at low bit rates, which are caused by coarse quantization of transform coefficients or motion compensation from distorted frames. In this paper, we propose a compression artifact reduction approach that utilizes both the spatial and the temporal correlation to form multi-hypothesis predictions from spatio-temporal similar blocks. For each transform block, three predictions with their reliabilities are estimated, respectively. The first prediction is constructed by inversely quantizing transform coefficients directly, and its reliability is determined by the variance of quantization noise. The second prediction is derived by representing each transform block with a temporal auto-regressive (TAR) model along its motion trajectory, and its corresponding reliability is estimated from local prediction errors of the TAR model. The last prediction infers the original coefficients from similar blocks in non-local regions, and its reliability is estimated based on the distribution of coefficients in these similar blocks. Finally, all the predictions are adaptively fused according to their reliabilities to restore high-quality videos. The experimental results show that the proposed method can efficiently reduce most of the compression artifacts and improve both subjective and objective quality of block transform coded videos. Xinfeng Zhang 0001, Ruiqin Xiong, Weisi Lin, Siwei Ma 0001, Jiaying Liu 0001, Wen Gao 0001 |
IEEE Trans. Image Process. | 5 |
| 2014 | BSIK-SVD: A dictionary-learning algorithm for block-sparse representationsabstractSparse dictionary learning has attracted enormous interest in image processing and data representation in recent years. To improve the performance of dictionary learning, we propose an efficient block-structured incoherent K-SVD algorithm for the sparse representation of signals. Without relying on any prior knowledge of the group structure for the input data, we develop a two-stage agglomerative hierarchical clustering method for block sparse representations. This clustering method adaptively identifies the underlying block structure of the dictionary under the restricted conditions of both a maximal block size and a minimal distance between the blocks. Furthermore, to meet the constraints of both the upper bound and the lower bound of the mutual coherence of dictionary atoms, we introduce a regularization term for the objective function to suppress the block coherence of the overcomplete dictionary. The experiments on synthetic data and real images demonstrate that the proposed algorithm has lower representation error, higher visual quality and better reconstructed results than other state-of-the-art methods. Yongqin Zhang, Jiaying Liu 0001, Mading Li, Zongming Guo |
ICASSP | 2 |
| 2014 | General scale interpolation based on fine-grained isophote model with consistency constraintabstractIn this paper, we propose a fine-grained isophote model with consistency constraint to characterize the piecewise-stationarity of image signals. According to this model, we present a novel interpolation algorithm. In this model, the displacement coefficient is used to model the isophote. Then fine-grained pixel intensity information is introduced to correct the displacement calculation and make the isophote estimation more robust. In order to handle the piecewise-stationarity, we force the isophote direction consistent in the local window when an interpolated line is piecewise-stationary. The proposed algorithm can accommodate the general scale enlargement. Experimental results demonstrate that the proposed approach achieves better performances in both objective and subjective quality assessment. Wenhan Yang, Jiaying Liu 0001, Mading Li, Zongming Guo |
ICIP | 2 |
| 2014 | Exploiting multi-scale spatial structures for sparsity based single image super-resolutionabstractTo improve the performance of sparsity-based single image super-resolution (SR), we propose a joint SR framework of structure prior based sparse representation (SPSR). The proposed SPSR algorithm exploits the multi-scale spatial structural self-similarities, the gradient prior and nonlocally centralized sparse representation to formulate a constrained optimization problem for high-resolution image recovery. The high-resolution image is firstly initialized by exploiting cross-scale patch redundancy in an image pyramid from single input low-resolution image. Then the sparse modeling of the image SR problem is proposed to refine it further, where the gradient histogram preservation is incorporated as a regularization term. Finally, an iterative solution is provided to solve the problem of model parameter estimation and sparse representation. Experimental results on image super-resolution validate the generality, effectiveness and robustness of the proposed SPSR algorithm. Yongqin Zhang, Jiaying Liu 0001, Wei Bai 0002, Zongming Guo |
ICIP | 2 |
| 2014 | Segmentation-based scale-invariant nonlocal means super resolutionabstractZooming in/out appears frequently in video shooting, which makes scale vary between frames. And object motion in videos may cause scale change of the object. It leads to the difficulty in finding similar patches and causes the invalidation of nonlocal means super resolution (NLM SR). In this paper, we propose a novel scale-compensated NLM SR algorithm. First, by considering the parameter model, the image is segmented in order to detect regions with different scales. Then, scale variations in different regions are computed based on SIFT descriptor. And patches extracted from different regions are compensated into the same scale to eliminate the effect of scale change. It is shown by experimental results that our proposed algorithm achieves the average PSNR by up to 0.678dB comparing with the state-of-the-art methods. Subjective results demonstrate the proposed method reduces artifacts and preserves more details. Saboya Yang, Jiaying Liu 0001, Qiaochu Li, Zongming Guo |
ISCAS | 2 |
| 2014 | Image transformation using limited reference with application to photo-sketch synthesisabstractImage transformation refers to transforming images from a source image space to a target image space. Contemporary image transformation methods achieve this by learning coupled dictionaries from a set of paired images. However, in practical use, such paired training images are not easy to get especially when the target image style is not fixed. Thus in most cases, the reference is limited. In this paper, we propose a sparse representation based framework of transforming images with limited reference, which can be used for the typical image transformation application, photo-sketch synthesis. In the learning stage, the edge features are utilized to map patches between different style images, thus building the coupled database for dictionary learning. In the reconstruction stage, sparse representation can well preserve the basic structure of image contents. In addition, a texture synthesis strategy is introduced to enhance target-like textures in the output image. Experimental results show that the performance of our method is comparable to state-of-the-art methods even with limited reference, which is very efficient and less restrictive for practical use. Wei Bai 0002, Yanghao Li, Jiaying Liu 0001, Zongming Guo |
VCIP | 3 |
| 2014 | Patch-based image deblocking using geodesic distance weighted low-rank approximationabstractTransform coding based on the discrete cosine transform (DCT) has been widely used in image coding standards. However, the coded images often suffer from severe visual distortions such as blocking artifacts. In this paper, we propose a novel image deblocking method to address the blocking artifacts reduction problem in a patch-based scheme. Image patches are clustered and reconstructed by the low-rank approximation, which is weighted by the geodesic distance. Experimental results show that the proposed method achieves higher PSNR than the state-of-the-art deblocking and denoising methods and the processed images present good visual quality. Mading Li, Jiaying Liu 0001, Jie Ren 0012, Zongming Guo |
VCIP | 2 |
| 2014 | Joint image denoising using adaptive principal component analysis and self-similarity
Yongqin Zhang, Jiaying Liu 0001, Mading Li, Zongming Guo |
Inf. Sci. | 2 |
| 2013 | Single-Pass Dependent Bit Allocation in Temporal Scalability Video CodingabstractSummary form only given. In the scalable video coding, we refer to a group-of-pictures (GOP) structure that is composed of hierarchically aligned B-pictures. It employs generalized B-pictures that can be used as a reference to following inter-coded frames. Although it introduces a structural encoding delay of one GOP size, it provides much higher coding efficiency than the conventional GOP structures [2]. Moreover, due to its natural capability of providing the temporal scalability, it is employed as a GOP structure of H.264/SVC [3]. Because of the complex inter-layer dependence of hierarchical B-pictures, the development of an efficient and effective bit allocation algorithm for H.264/SVC is a challenging task. There are several bit allocation algorithms that considered the inter-layer dependence in the literature before. Schwarz et al. proposed the QP cascading scheme that applies a fixed quantization parameter (QP) difference between adjacent temporal layers. Liu et al. introduced constant weights to temporal layers in their H.264/SVC rate control algorithm. Although these algorithms achieve superior coding efficiency, they are limited in two aspects. First, the inter-layer dependence is heuristically addressed. Second, the input video characteristics are not taken into account. For these reasons, the optimality of these bit allocation algorithms cannot be guaranteed. We propose a single-pass dependent bit allocation algorithm for scalable video coding with hierarchical B-pictures in this work. It is generally perceived that dependent bit allocation algorithms cannot be practically employed due to their extremely high complexity requirement. To develop a practical single-pass bit allocation algorithm, we use the number of skipped blocks and the ratio of the mean absolute difference (MAD) as features to measure the inter-layer signal dependence of input video signals. The proposed algorithm performs bit allocation at the target bit rate with two mechanisms: 1) the GOP based rate control and 2) adaptive temporal layer QP decision. The superior performance of the proposed algorithm is demonstrated by experimental results, which is benchmarked by two other single-pass bit allocation algorithms in the literature. The rate and the PSNR coding performance of the proposed scheme and two benchmarks at various target bit rates for GOP-4 and GOP-8, respectively. We see that the proposed rate control algorithm achieves about 0.2-0.3dB improvement in coding efficiency as compared to JSVM. Furthermore, the proposed rate control algorithm outperforms Liu's Algorithm by a significant margin. Jiaying Liu 0001, Yongjin Cho, Zongming Guo |
DCC | 1 |
| 2013 | Image Blocking Artifacts Reduction via Patch Clustering and Low-Rank MinimizationabstractSummary form only given. Block-based Discrete Cosine Transform (BDCT) has been widely used in image and video compression due to its energy compacting property and relative ease of implementation. However, BDCT has a major drawback, which is usually referred to as blocking artifacts. Blocking artifacts appear as grid noise along the block boundaries because each block is transformed and quantized independently. Image deblocking techniques can reduce these distortions and alleviate the conflict between bit rate reduction and visual quality preservation. Many state-of-the-art image deblocking algorithms treated the blocking artifacts reduction of the compressed image as an inverse restoration problem. Natural image prior models are well utilized into the blocking artifacts reduction processing, such as the local sparsity prior model and non-local similarity property of natural images. These two local and non-local models characterize the image prior information in two complementary perspectives. Therefore, it is necessary to combine these two models in a unified framework. In this paper, we propose a novel method to reduce the blocking artifacts of blockcoded images via patch clustering and low-rank minimization, which simultaneously exploits the local and non-local sparse representations in a unified framework. First, the whole compressed image are divided into small patches. For each patch, we perform patch clustering to collect similar patches into a group. Then the whole group are simultaneously reconstructed by a low-rank minimization approach. Singular value thresholding (SVT) algorithm is employed to solve the low-rank minimization problem. To further improve the performance of the proposed algorithm, we adopt an iterative procedure to utilize the newly output data in each iteration and update the noise and signal variance adaptively. Experimental results show that the proposed method achieves higher PSNR and SSIM than the state-of-the-art methods. Comparing to the state-of-theart algorithms and, the proposed algorithm achieves about 0.37dB and 0.11dB improvement on average. For visual quality assessment, the deblocking images produced by the proposed algorithm reveal much more sharp edge structures and richer textures. Jie Ren 0012, Jiaying Liu 0001, Mading Li, Wei Bai 0002, Zongming Guo |
DCC | 2 |
| 2013 | Postprocessing of block-coded videos for deflicker and deblockingabstractIn this paper, we propose a novel postprocessing method to suppress both the flickering and blocking artifacts in block-coded videos. For reducing the flickering effect between adjacent frames, we propose an adaptive multi-scale motion filtering method to maintain the motion coherence of processed video. For blocking artifacts suppression, we adopt a patch-based scheme in which similar patches are grouped in a spatio-temporal domain and each patch group is recovered by solving a low rank matrix completion problem. Experimental results show that the proposed method can significantly reduce the flickering and blocking artifacts in the decoded videos. Jie Ren 0012, Jiaying Liu 0001, Mading Li, Zongming Guo |
ICASSP | 2 |
| 2013 | Adaptive general scale interpolation based on similar pixels weightingabstractIn this paper, we propose an adaptive general scale interpolation algorithm considering the non-stationarity of natural images in local areas. In image 2× enlargement, there are fixed relative positions between low-resolution (LR) pixels and high-resolution (HR) pixels. Unknown HR pixels can be estimated by their available LR neighbors. However, such relative positions are not fixed in the general-scale enlargement situations. The number and position of available LR pixels are indeterminate, therefore HR pixels can not be estimated by LR pixels. To make our method suitable for general scaling factors, we construct autoregressive (AR) models with pixels' neighbors instead of their available LR neighbors. Simultaneously, we introduce the similarity between pixels within a local window, which improves the method's performance by modeling the non-stationarity of image signals. Experimental results demonstrate the effectiveness of the proposed method on general scaling factors. Mading Li, Jiaying Liu 0001, Jie Ren 0012, Zongming Guo |
ISCAS | 2 |
| 2013 | Illumination-invariance and nonlocal means based super resolutionabstractIn this paper, we propose a novel algorithm for multi-frame super resolution (SR) with illumination-invariance. Traditional multi-frame SR methods fail to handle images with illumination changes, so in our approach, we adjust the contrast between different search windows and select proper candidate patches to take full advantage of intensity information. We simplify Speed Up Robust Features to get local structure information and incorporate the local structure information into similarity measurement, which does not change significantly in complex illumination situation. By combining intensity and structure information in a proper way, our algorithm Illumination-Invariant Nonlocal Means SR could find more potential similar patches in frames where there are illumination changes than Nonlocal Means SR (NLM SR). Experimental results demonstrate that our algorithm has better performance both in objective and subjective perception with complex illumination conditions and is comparable to NLM SR in stable illumination situation. Mengyan Wang, Jiaying Liu 0001, Wei Bai 0002, Zongming Guo |
ISCAS | 2 |
| 2013 | Multi-frame Super Resolution Using Refined Exploration of Extensive Self-examples
Wei Bai 0002, Jiaying Liu 0001, Mading Li, Zongming Guo |
MMM (1) | 2 |
| 2013 | Image super resolution using saliency-modulated context-aware sparse decompositionabstractThis paper presents a novel saliency-modulated sparse representation algorithm for image super resolution. In images, regions salient to human eyes appear to be more organized and structured. This property is utilized in both the dictionary learning and the sparse coding process to capture more structural details for the reconstructed image. Apart from a general dictionary, example patches from the salient regions are extracted to train a salient dictionary. We also incorporate context-aware sparse decomposition to model dependencies between dictionary atoms of adjacent patches, especially in the salient regions. Experiments show the proposed method outperforms state-of-the-art methods with the highest PSNR gain. Subjective results demonstrate the proposed method reduces artifacts and preserves more details. Wei Bai 0002, Saboya Yang, Jiaying Liu 0001, Jie Ren 0012, Zongming Guo |
VCIP | 3 |
| 2013 | Joint image denoising using self-similarity based low-rank approximationsabstractThe observed images are usually noisy due to data acquisition and transmission process. Therefore, image denoising is a necessary procedure prior to post-processing applications. The proposed algorithm exploits the self-similarity based low rank technique to approximate the real-world image in the multivariate analysis sense. It consists of two successive steps: adaptive dimensionality reduction of similar patch groups, and the collaborative filtering. For each target patch, the singular value decomposition (SVD) is used to factorize the similar patch group collected in a local search window by block-matching. Parallel analysis automatically selects the principal signal components by discarding the nonsignificant singular values. After the inverse SVD transform, the denoised image is reconstructed by the weighted averaging approach. Finally, the collaborative Wiener filtering is applied to further remove the noise. Experimental results show that the proposed algorithm surpasses the state-of-the-art methods in most cases. Yongqin Zhang, Jiaying Liu 0001, Saboya Yang, Zongming Guo |
VCIP | 2 |
| 2013 | Guided image filtering using signal subspace projectionabstractThere are various image filtering approaches in computer vision and image processing that are effective for some types of noise, but they invariably make certain assumptions about the properties of the signal and/or noise which lack the generality for diverse image noise reduction. This study describes a novel generalised guided image filtering method with the reference image generated by signal subspace projection (SSP) technique. It adopts refined parallel analysis with Monte Carlo simulations to select the dimensionality of signal subspace in the patch‐based noisy images. The noiseless image is reconstructed from the noisy image projected onto the significant eigenimages by component analysis. Training/test image are utilised to determine the relationship between the optimal parameter value and noise deviation that maximises the output peak signal‐to‐noise ratio (PSNR). The optimal parameters of the proposed algorithm can be automatically selected using noise deviation estimation based on the smallest singular value of the patch‐based image by singular value decomposition (SVD). Finally, we present a quantitative and qualitative comparison of the proposed algorithm with the traditional guided filter and other state‐of‐the‐art methods with respect to the choice of the image patch and neighbourhood window sizes. Yongqin Zhang, Jiaying Liu 0001, Zongming Guo |
IET Image Process. | 3 |
| 2013 | Dependent R/D Modeling Techniques and Joint T-Q Layer Bit Allocation for H.264/SVCabstractWe investigate dependent rate/distortion (R/D) modeling techniques for H.264/SVC videos. We introduce a self-domain (S-domain) analysis method for characterizing the dependent R/D behaviors, where the R/D characteristics of a base layer are employed as the observation domain for those of dependent layers. Based onS-domain observations, we propose empirical dependent R/D models and analyze physical implications of the proposed models. As an application of the proposed R/D models, we examine a joint temporal-quality layer bit allocation algorithm formulated as a Lagrange optimization problem. The proposed R/D models enable us to derive an analytical solution to the joint optimization problem. Finally, it is demonstrated by experimental results that our bit allocation algorithm outperforms JSVM benchmark by a significant margin (10%-20%) at various bit rates. Yongjin Cho, Do-Kyoung Kwon, Jiaying Liu 0001, C.-C. Jay Kuo |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2013 | Context-Aware Sparse Decomposition for Image Denoising and Super-ResolutionabstractImage prior models based on sparse and redundant representations are attracting more and more attention in the field of image restoration. The conventional sparsity-based methods enforce sparsity prior on small image patches independently. Unfortunately, these works neglected the contextual information between sparse representations of neighboring image patches. It limits the modeling capability of sparsity-based image prior, especially when the major structural information of the source image is lost in the following serious degradation process. In this paper, we utilize the contextual information of local patches (denoted as context-aware sparsity prior) to enhance the performance of sparsity-based restoration method. In addition, a unified framework based on the markov random fields model is proposed to tune the local prior into a global one to deal with arbitrary size images. An iterative numerical solution is presented to solve the joint problem of model parameters estimation and sparse recovery. Finally, the experimental results on image denoising and super-resolution demonstrate the effectiveness and robustness of the proposed context-aware method. Jie Ren 0012, Jiaying Liu 0001, Zongming Guo |
IEEE Trans. Image Process. | 2 |
| 2013 | Rate-Distortion Analysis of Dead-Zone Plus Uniform Threshold Scalar Quantization and Its Application - Part I: Fundamental TheoryabstractThis paper provides a systematic rate-distortion (R-D) analysis of the dead-zone plus uniform threshold scalar quantization (DZ+UTSQ) with nearly uniform reconstruction quantization (NURQ) for generalized Gaussian distribution (GGD), which consists of two aspects: R-D performance analysis and R-D modeling. In R-D performance analysis, we first derive the preliminary constraint of optimum entropy-constrained DZ+UTSQ/NURQ for GGD, under which the property of the GGD distortion-rate (D-R) function is elucidated. Then for the GGD source of actual transform coefficients, the refined constraint and precise conditions of optimum DZ+UTSQ/NURQ are rigorously deduced in the real coding bit rate range, and efficient DZ+UTSQ/NURQ design criteria are proposed to reasonably simplify the utilization of effective quantizers in practice. In R-D modeling, inspired by R-D performance analysis, the D-R function is first developed, followed by the novel rate-quantization (R-Q) and distortion-quantization (D-Q) models derived using analytical and heuristic methods. The D-R, R-Q, and D-Q models form the source model describing the relationship between the rate, distortion, and quantization steps. One application of the proposed source model is the effective two-pass VBR coding algorithm design on an encoder of H.264/AVC reference software, which achieves constant video quality and desirable rate control accuracy. Jun Sun 0012, Yizhou Duan, Jiaying Liu 0001, Zongming Guo |
IEEE Trans. Image Process. | 4 |
| 2013 | Rate-Distortion Analysis of Dead-Zone Plus Uniform Threshold Scalar Quantization and Its Application - Part II: Two-Pass VBR Coding for H.264/AVCabstractIn the first part of this paper, we derive a source model describing the relationship between the rate, distortion, and quantization steps of the dead-zone plus uniform threshold scalar quantizers with nearly uniform reconstruction quantizers for generalized Gaussian distribution. This source model consists of rate-quantization, distortion-quantization (D-Q), and distortion-rate (D-R) models. In this part, we first rigorously confirm the accuracy of the proposed source model by comparing the calculated results with the coding data of JM 16.0. Efficient parameter estimation strategies are then developed to better employ this source model in our two-pass rate control method for H.264 variable bit rate coding. Based on our D-Q and D-R models, the proposed method is of high stability, low complexity and is easy to implement. Extensive experiments demonstrate that the proposed method achieves: 1) average peak signal-to-noise ratio variance of only 0.0658 dB, compared to 1.8758 dB of JM 16.0's method, with an average rate control error of 1.95% and 2) significant improvement in smoothing the video quality compared with the latest two-pass rate control method. Jun Sun 0012, Yizhou Duan, Jiaying Liu 0001, Zongming Guo |
IEEE Trans. Image Process. | 4 |
| 2012 | Nonlocal based Super Resolution with rotation invariance and search window relocationabstractIn this paper, we present a novel method for Super Resolution (SR) reconstruction with rotation invariance and search window relocation. To combine complementary information in observed images to generate a higher resolution image, we first relocate search window to involve potential similar patches and then use rotation invariance similarity measure to find accurate similar patches. Comparing with Nonlocal Means SR, our algorithm can find more similar patches for weighted average. Experimental results demonstrate superior performance of the proposed method in terms of both objective measurements and subjective evaluation. Yue Zhuo 0005, Jiaying Liu 0001, Jie Ren 0012, Zongming Guo |
ICASSP | 2 |
| 2012 | Single pass dependent bit allocation for H.264 temporal scalabilityabstractIn this paper, we propose a single-pass dependent bit allocation algorithm for H.264/SVC hierarchical B-pictures. To develop a practical bit allocation algorithm, we use the number of skipped blocks and the ratio of the mean absolute difference (MAD) as features to measure the inter-layer signal dependence of input video signals. The proposed algorithm performs bit allocation at the target bit rate with two steps: the group-of-picture (GOP) based rate control and adaptive temporal layer quantization parameter (QP) decision. The superior performance of the proposed algorithm is demonstrated by experimental results, which is compared with two other one-pass bit allocation algorithms in the literature. Jiaying Liu 0001, Yongjin Cho, Zongming Guo |
ICIP | 1 |
| 2012 | Illumination-invariant non-local means based video denoisingabstractIn this paper, we present a robust illumination-invariant non-local means (NLM) based video denoising algorithm with special illumination handling. Illumination variances pose a challenge to the NLM-based denoising algorithms. To address this issue, we first propose several possible technical improvements, and verify their efficacy of eliminating the influence of illumination changes. Then, by analyzing and comparing these techniques, a histogram processing based technique is integrated into the non-local means denoising framework. Experimental results on synthesis and real video denoising show that the proposed method is able to fully explore the non-local self-similarity property in natural videos under variable illumination conditions. Jie Ren 0012, Yue Zhuo 0005, Jiaying Liu 0001, Zongming Guo |
ICIP | 3 |
| 2012 | Image super-resolution by structural sparse coding
Jie Ren 0012, Jiaying Liu 0001, Mengyan Wang, Zongming Guo |
ICPR | 2 |
| 2012 | Visual-weighted motion compensation frame interpolation with motion vector refinementabstractIn this paper, we propose a novel frame rate up-conversion algorithm based on joint motion vector refinement and visual-weighted motion compensation interpolation (MCI). It utilizes a hierarchical motion vector refinement to correct inaccurate motion vectors (MVs), which is composed of the global level and the local level. In the global level, distinct inaccurate MVs are detected by global controlling and then corrected by neighborhood information. Afterwards, the local level performs the local controlling to pick out local outliers and re-estimate them with the maximum likelihood method. Finally, plausible weights for each block in the interpolated frame, computed by the similarity index(SSIM), are applied for visual compensation. The experimental results demonstrate that compared with the conventional algorithm EBME, the proposed algorithm achieved the average PSNR by up to 2.7dB while the visual quality improvement is also remarkable. Wei Bai 0002, Jiaying Liu 0001, Jie Ren 0012, Zongming Guo |
ISCAS | 2 |
| 2012 | Optimized bit extraction of SVC exploiting linear error modelabstractThe Scalable Video Coding (SVC) extension of the H.264/AVC video coding standard supports fidelity or quality (SNR) scalability. The quality enhancement packets would be discarded in case of limited network capacity, which calls for an optimized bit extraction strategy. In this paper, we first analyze the linear feature in H.264/AVC video coding. A linear error model is also constructed using this feature in case of SVC quality scalability. Then based on the linear error model, the rate and distortion (R-D) impact of each quality enhancement packet over the whole sequence is obtained. Finally a new priority assigning algorithm is designed for a more efficient extraction, giving high rank to those with great R-D impacts. Extensive experiments are presented to demonstrate the accuracy of the linear error model and the validity of the priority assigning algorithm. Tests on the set of eight standard video sequences show the quality promotion under any bitrate constraint, and a fidelity gain up to 0.4 dB PSNR is achieved by the proposed strategy, compared to the JSVM reference software with Quality Layer information. Jun Sun 0012, Jiaying Liu 0001, Zongming Guo |
ISCAS | 3 |
| 2011 | Joint Spatial-Temporal Layer Bit Allocation with S-Domain Dependent R-D ModelingabstractSummary form only given. H.264/SVC, as a scalable extension of H.264/AVC, is finally standardized in 2007. Scalable video stream has achieved great flexibility and adaptability in terms of frame rates, display resolutions and quality levels. With three dimensions of the scalability, each coding unit in an H.264/SVC video is subject to highly complicated inter-dependency, which gives one of major challenges for bit allocating. In this work, we study an optimal solution to joint spatial-temporal (S-T) bit allocation problem with self-domain R-D modeling. The self-domain (ιS-domain) analysis employs the R-D characteristics of the reference layer as the observation domain of those of dependent layers. Jiaying Liu 0001, Yongjin Cho, Zongming Guo |
DCC | 1 |
| 2011 | Practical rate control algorithm for temporal scalability in scalable video codingabstractA rate control algorithm for hierarchical B-pictures in Scalable Video Coding (SVC) is proposed in this work. The complex inter-frame dependency issue is effectively addressed by the Q-distance policy decision rule while the statistical smoothing effect enables the GOP-based precise bit rate control. The simplicity of the decision processes greatly reduces the encoder complexity providing an efficient and effective rate control algorithm with hierarchical B-pictures. Experimental results verify the significant performance gain by the proposed algorithm. Jiaying Liu 0001, Yongjin Cho, Zongming Guo |
ICIP | 1 |
| 2011 | Similarity modulated block estimation for image interpolationabstractModeling the nonstationarity of image signals is one of the challenging issues for image interpolation. In this paper, we propose a similarity probability modeling to faithfully characterize the nonstationarity of image signals, and present a novel image interpolation algorithm based on the proposed model. The missing pixels are estimated in groups by weighted block estimation. The weight of each pixel inside the block is defined as the similarity probability between itself and the centered to-be-interpolated pixel. It is demonstrated by the experimental results that the proposed method preserves the edge structures of the interpolated images better than the state-of-the-art interpolation methods. Annoying artifacts nearby the sharp edges are also greatly reduced. Jie Ren 0012, Jiaying Liu 0001, Wei Bai 0002, Zongming Guo |
ICIP | 2 |
| 2011 | Efficient dead-zone plus uniform threshold scalar quantization of generalized Gaussian random variablesabstractThis paper studies the rate-distortion (R-D) performance of entropy-constrained dead-zone plus uniform threshold scalar quantization (DZ+UTSQ) and nearly-uniform reconstruction quantization (NURQ) for generalized Gaussian distribution (GGD). We first derive the preliminary constraint of R-D optimized DZ+UTSQ/NURQ for GGD. Then for GGD source of actual DCT coefficients, the refined constraint and precise conditions of optimum DZ+UTSQ/NURQ are rigorously deduced in the real coding bit rate range. Based on above analysis, efficient DZ+UTSQ/NURQ design criteria are proposed to reasonably simplify the implementation of effective quantizer in practice. Yizhou Duan, Jun Sun 0012, Jiaying Liu 0001, Zongming Guo |
VCIP | 3 |
| 2011 | A novel parallel encoding framework for scalable video codingabstractIn this paper, we first propose a new parallel video coding framework, considering three important factors: parallel strategy, computational complexity and task scheduling. Then combining the characteristics of scalable video coding (SVC), a novel parallel encoding structure for temporal and quality scalabilities is introduced to obtain a high speedup of parallel SVC. Since the data dependencies in SVC are complex and time variant, the scheduling of parallel SVC is extremely difficult. In order to find the optimal scheduling solution, directed acyclic graph (DAG) is exploited to model the dependencies of encoding tasks, and the complexities of the encoding tasks which are accurately estimated by the Kalman filter to weight the scheduling tasks. Finally, two heuristic scheduling algorithms are also proposed to achieve a high encoding speed of parallel SVC. Experimental results show that the speedup of our parallel method was higher (about 60%) than previous work. Using the proposed method, high definition (HD) SVC videos can be encoded in real time. Jun Sun 0012, Jiaying Liu 0001, Zongming Guo, Longshe Huo |
VCIP | 3 |
| 2010 | Bit Allocation for Spatial Scalability Coding of H.264/SVC With Dependent Rate-Distortion AnalysisabstractWe propose a model-based spatial layer bit allocation algorithm for H.264/scalable video coding (SVC) in this paper. The challenge of this problem lies in the fact that the rate-distortion (R-D) behavior of an enhancement layer is dependent on its preceding layers because of inter-layer prediction. To solve it, we first focus on the case of two spatial layers, derive the distortion and rate models of the dependent layer analytically, and develop a low-complexity bit allocation algorithm. It is shown by experimental results that the proposed two-layer bit allocation algorithm can achieve the coding performance close to the optimal R-D performance based on the full search method. Then, we extend this result to multilayer bit allocation by performing the two-layer allocation scheme recursively. Finally, we compare the performance of group of pictures-based and frame-based spatial layer bit allocation schemes at a fixed temporal resolution. The superior performance of the proposed spatial layer bit allocation algorithm is demonstrated using Joint Scalable Video Model reference software algorithm and two prior H.264/SVC rate control algorithms as the benchmarks. Jiaying Liu 0001, Yongjin Cho, Zongming Guo, C.-C. Jay Kuo |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2009 | H.264/SVC temporal bit allocation with dependent distortion modelabstractThe bit allocation problem for hierarchical B-pictures in H.264/SVC is studied with a GOP-based dependent distortion model in this work. Inter-dependency between temporal layers of H.264/SVC is often neglected because of the complexity involved, which often leads to poorer rate control performance. To address this shortcoming, we propose a distortion model that takes inter-dependency into consideration while preserving the low complexity of the encoding process. It is demonstrated by experimental results that the new distortion model results in a highly efficient bit allocation scheme, which outperforms the rate control algorithm in the JSVM 9.12 reference codec by a significant margin. Yongjin Cho, Jiaying Liu 0001, Do-Kyoung Kwon, C.-C. Jay Kuo |
ICASSP | 2 |
| 2009 | Bit allocation for joint spatial-quality scalability in H.264/SVCabstractIn this work, we propose a model-based layer bit allocation algorithm for joint spatial-quality (S-Q) scalability in H.264/SVC. The complicated inter-layer dependency is decoupled by the proposed spatial and quality rate and distortion (R-D) models. We show that the R-D characteristics of a dependent layer can be represented by a number of independent functions with GOP as a basic coding unit. Then, the joint bit allocation problem is formulated as a two-step optimization problem by the Lagrangian multiplier method, which can be numerically solved using the proposed R-D models. Finally, we develop a low-complexity bit allocation algorithm for the combined spatial and quality scalability in H.264/SVC. It is shown by experimental results that our proposed bit allocation algorithm achieves the coding performance significantly improved from current reference software JSVM. Jiaying Liu 0001, Zongming Guo, Yongjin Cho |
ICIP | 1 |
| 2009 | Frame-based bit allocation for spatial scalability in H.264/SVCabstractThe spatial scalability of H.264/SVC is achieved by a multi-layer approach, where an enhancement layer is dependent on its preceding layers. To address this dependent issue, we propose a model-based spatial layer bit allocation algorithm for H.264/SVC in this work. The inter-layer dependency is decoupled by analyzing the signal flow in the H.264/SVC encoder. We show that the rate and the distortion (R-D) characteristics of a dependent layer with a frame as a basic coding unit. Finally, a low complexity spatial layer bit allocation scheme is developed using the proposed frame-based R-D models. It is shown by experimental results that our proposed bit allocation algorithm can achieve the coding performance close to the optimal R-D performance of full search and is highly improved from current reference codec JSVM. Jiaying Liu 0001, Yongjin Cho, Zongming Guo |
ICME | 1 |
| 2009 | Joint Quality-temporal (Q-T) Bit Allocation for H.264/SVCabstractA joint quality-temporal (Q-T) bit allocation scheme is proposed for H.264/SVC in this work. First, rate and distortion (R-D) models for dependent quality layers are derived, where the complex inter-layer dependency is considered. Then, the joint Q-T bit allocation problem is formulated as an optimization problem using the Lagrange method and solved numerically with the derived R-D models. As a result, we develop a low-complexity bit allocation scheme for the joint Q-T scalability of H.264/SVC. It is demonstrated by experimental results that the new R-D models results in a highly efficient bit allocation scheme, which outperforms the JSVM benchmark by a significant margin. Yongjin Cho, Jiaying Liu 0001, Do-Kyoung Kwon, C.-C. Jay Kuo |
ISCAS | 2 |
| 2008 | Efficient intra-4×4 mode decision based on bit-rate estimation in H.264/AVCabstractRate-distortion optimization (RDO) technique is widely employed by H.264/AVC for the purpose of determining the best mode. However, such technique results in dramatic increase in the computation complexity of the underlying encoder. In this paper, we address this problem by presenting an efficient intra-4×4 mode decision algorithm. The algorithm works by approximating the bit-rate so as to reduce the computational cost of RDO and the main idea is the following: First we have found a quick way to estimate the bit-rate via the number of DCT coefficients to be quantized to 0 and that to be quantized to ±1. The parameters of the estimated function are adaptively obtained by using the Least Squares Fitting method of the above and the left block in the current frame and the co-location one in the previous encoded frame as the feedback. This close loop bit-rate estimation would skip the processes of quantization, inverse transform, entropy coding and reconstruction; we then use the estimated bit-rate and Sum of Absolute Differences (SAD) to simplify the optimization process of R-D cost function. Experimental results show that our scheme decreases the time for intra coding by 50% with negligible loss of PSNR, and that comparing with those fast mode-decision algorithms based-on local edge direction, the optimal prediction mode obtained via our scheme is closer to that obtained via the original RDO in statistic. Jiaying Liu 0001, Zongming Guo |
ISCAS | 1 |
| 2008 | Bit allocation for spatial scalability in H.264/SVCabstractWe propose a model-based spatial layer bit allocation algorithm for H.264/SVC in this work. The spatial scalability of H.264/SVC is achieved by a multi-layer approach, where an enhancement layer is bound by the dependency on its preceding layers. The inter-layer dependency is decoupled in our analysis by a careful examination of the signal flow in the H.264/SVC encoder. We show that the rate and the distortion (R-D) characteristics of a dependent layer can be represented by a number of independent functions with a group of pictures (GOP) as a basic coding unit. Finally, a low complexity spatial layer bit allocation scheme is developed using the proposed GOP-based R-D models. It is shown by experimental results that our proposed bit allocation algorithm can achieve the coding performance close to the optimal R-D performance of full search and is significantly improved from current reference software JSVM. Jiaying Liu 0001, Yongjin Cho, Zongming Guo, C.-C. Jay Kuo |
MMSP | 1 |