Liang Li 0039

dblp:14/1395-39 · DBLP profile ↗
← Back
16ranked-venue papers
3as first author
5since 2021 · last 2022
0000-0002-9792-7069ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author
YearPublicationVenuePosition
2022 Feature-Guided Blind Face Restoration with GAN Prior
abstract
Blind face restoration (BFR) aims to restore high-quality face images from inputs with complex degradation, which is key to extensive applications. Existing methods usually learn a black-box mapping to achieve the goal, which however often produce over-smoothed results. This work proposes a novel feature-guided framework via leveraging prior from a pre-trained Generative Adversarial Network (GAN) model to recover reasonable textures. Furthermore, we design a feature fusion module to guide the generator with low-level spatial content information in the degraded input, for the sake of holding the consistency on both facial structure and background details. The proposed method can be seamlessly integrated with a learned GAN through a simple yet effective principle to recover realistic results under complex degradation circumstances. Extensive comparisons demonstrate the superiority of our strategy over other state-of-the-art methods in terms of restoration quality and training cost.
Zhengzhang Hou, Liang Li 0039, Xiaojie Guo 0001
ICME2
2022 Synthetic-to-Real Generalization for Semantic Segmentation
abstract
The discrepancy between synthetic and real data is crucial to the performance of domain generalization for semantic segmentation. Since real data is not always accessible, a popular line of approaches is to enhance the diversity of synthetic data via either complex adversarial generation or unstable stylization. However, the internal structure of the synthetic image is often neglected. To largely explore useful information in synthetic data, we observe that, although objects of the same category have different texture patterns between domains, their shapes are quite similar. Based on this observation, we argue that focusing on structural information and alleviating texture dependence are effective ways to improve generalization capability. In this work, we propose an end-to-end network, which explicitly constrains the network to learn shapes and spatial knowledge, and implicitly relieves the texture reliance of the network. Extensive experiments verify the effectiveness of our proposed method and demonstrate its clear advantages over other competitors.
Liang Li 0039, Xiaojie Guo 0001
ICME2
2022 BiAttnNet: Bilateral Attention for Improving Real-Time Semantic Segmentation
abstract
Semantic segmentation requires both speed and accuracy. This paper presents a two-branch network BiAttnNet with a unique Bilateral Attention structure that separates all attention modules into the Detail Branch to contribute semantic detail selections for specialized detail exploring. Specifically, the Detail Branch comprises AttnTrans entirely, which provides a better alternate for regular convolution. AttnTrans is a computationally efficient filtration entirely composed of concurrent spatial and channel attention. Meanwhile, a Context Branch is implemented with FCN-ResNet for rough segmentation. By combining two branches’ outputs, BiAttnNet achieves a good balance between speed and accuracy. Evaluations on the Cityscapes testing set conclude that BiAttnNet achieves 74.7% mIoU at 89.2 FPS at a quarter (512 × 1024) resolution with only 2.2 million parameters, running on a single GTX 2080 Ti card.
Genling Li, Liang Li 0039, Jiawan Zhang
IEEE Signal Process. Lett.2
2022 Hierarchical Semantic Broadcasting Network for Real-Time Semantic Segmentation
abstract
Semantic segmentation has been one of the essential tasks in computer vision. More complicated and computationally intensive mechanisms are integrated into segmentation models to get more accurate results, leading to increased processing time and resource usage. Based on the idea that pixels with similar high-level features are more likely to have similar semantic labels, we propose a computationally efficient mechanism named Hierarchical Semantic Broadcasting (HSB), which can infer earlier-stage semantic label maps from lower-level feature maps by referring to semantic label maps of higher-level feature maps. Since lower-level feature maps have higher resolution and richer context, HSB can provide additional details for better semantic segmentation. By integrating HSB into a general-purpose light network, we propose Hierarchical Semantic Broadcasting Network (HSB-Net) for real-time semantic segmentation, which achieves a good trade-off between accuracy and speed. Evaluated by the Cityscapes dataset, HSB-Net can run at 123.7 FPS for a 512$ \boldsymbol{\times }$1024 input on a single GeForce RTX 2080 Ti card while achieving 73.1% mean IoU.
Genling Li, Liang Li 0039, Jiawan Zhang
IEEE Signal Process. Lett.2
2021 Image Demoiréing with a Dual-Domain Distilling Network
abstract
Due to slight discrepancy of spatial frequency between the camera sensor array and sub-pixel layout of LCD monitor, moiré pattern artifacts appear in various shapes and colours which seriously degrade the quality of captured images. It is challenging yet practically crucial to remove moiré artifacts from a single camera-captured screen image. In this paper, we propose a dual-domain distilling network (3DNet for short) to tackle this problem in an end-to-end manner. The 3DNet consists of a dual-branch student network (a.k.a. demoiréing network), and two teacher networks. The two branches of student network exploit knowledge in both spatial-domain and frequency-domain for the sake of removing moiré artifacts, based on the observation that rich image details can be discovered in the frequency-domain while structure information can be well kept in the spatial-domain. The demoiréing process of two branches is supervised by the knowledge distilled from two teacher networks trained for reconstructing clear images in the spatial and frequency domains respectively. Comprehensive experimental results are conducted to demonstrate the efficacy of our design, and reveal its superiority over state-of-the-art alternatives.
Qiaoyu Tian, Liang Li 0039, Xiaojie Guo 0001
ICME3
2020 End-to-end trainable network for superpixel and image segmentation
Liang Li 0039, Jiawan Zhang
Pattern Recognit. Lett.2
2019 Fast Superpixel Segmentation with Deep Features
Mubinun Awaisu, Liang Li 0039, Jiawan Zhang
CGI2
2018 Structure-Texture Decomposition via Joint Structure Discovery and Texture Smoothing
abstract
Structure-texture decomposition from an image (a.k.a. structure-preserving image smoothing) is important for a variety of multimedia, computer vision and graphics tasks. Its performance heavily depends on the precision of indicating where are structural edges to maintain and where are textures to remove. An intuitive thought for constructing indication is to directly execute edge detection on the input image, which however would suffer from rich textures. Feeding inaccurate or erroneous indications into the smoother is at high risk of generating unsatisfactory results. It is almost sure that edge detectors can do a better job on inputs with textures removed. The above two components, say the smoother and the indicator, turn out to be in a chicken-egg situation. To address this issue, we propose a method to jointly detect structural edges and remove textures, by iteratively smoothing the input based on the edges detected from the previous smoothed result and refining the edges based on the newly processed image. Experiments on a number of challenging cases are conducted to show that the edge detection task and the smoothing task can benefit from each other, and reveal the superiority of our method over other state-of-the-art alternatives. Our code is publicly available at https://sites.google.com/view/xjguo/sdts.
Xiaojie Guo 0001, Siyuan Li 0001, Liang Li 0039, Jiawan Zhang
ICME3
2018 Soft Clustering Guided Image Smoothing
abstract
Image smoothing, which aims to remove unwanted textures and preserve desired structures, plays an important role in many multimedia and computer vision tasks. The key to image smoothing, despite different applications, is to distinguish the structures from the textures. This paper presents a novel image smoothing method, following the principle that, for a certain pixel, its neighbors in both space and intensity should contribute more on smoothing, while the distant ones be insulated for avoiding over-smoothing. Intuitively, clustering is a good candidate to achieve the goal. However, due to rich textures and clutters within images, simply performing the clustering on the input likely obtains inaccurate results, and thus leads to unsatisfied smoothing results. In addition, for our task, using traditional hard clustering techniques is at high risk of generating staircase artifacts. For addressing these issues, an algorithm is customized, which on the one hand adopts the soft clustering to more faithfully assign pixels, on the other hand iterates the soft clustering and smoothing, expecting to improve each other. Experiments on several challenging images are provided to show the efficacy of our method, and its superiority over other prevailing approaches.
Liang Li 0039, Xiaojie Guo 0001, Wei Feng 0005, Jiawan Zhang
ICME1
2018 Co-Saliency Detection via Hierarchical Consistency Measure
abstract
Co-saliency detection is a newly emerging research topic in multimedia and computer vision, the goal of which is to extract common salient objects from multiple images. Effectively seeking the global consistency among multiple images is critical to the performance. To achieve the goal, this paper designs a novel model with consideration of a hierarchical consistency measure. Different from most existing co-saliency methods that only exploit common features (such as color and texture), this paper further utilizes the shape of object as another cue to evaluate the consistency among common salient objects. More specifically, for each involved image, an intra-image saliency map is firstly generated via a single image saliency detection algorithm. Having the intra-image map constructed, the consistency metrics at object level and superpixel level are designed to measure the corresponding relationship among multiple images and obtain the inter saliency result by considering multiple visual attention features and multiple constrains. Finally, the intra-image and inter-image saliency maps are fused to produce the final map. Experiments on benchmark datasets are conducted to demonstrate the effectiveness of our method, and reveal its advances over other state-of-the-art alternatives.
Liang Li 0039, Runmin Cong, Xiaojie Guo 0001, Jiawan Zhang
ICME2
2018 Embedded learning for computerized production of movie trailers
Jiachuan Sheng, Yuzhi Li, Liang Li 0039
Multim. Tools Appl.4
2014 Contrast enhancement based single image dehazing VIA TV-l1 minimization
abstract
In this paper, we propose a general algorithm to removing haze from single images using total variation minimization. Our approach stems from two simple yet fundamental observations about haze-free images and the haze itself. First, clear-day images usually have stronger contrast than images plagued by bad weather; and second, the variations in natural atmospheric veil, which highly depends on the depth of objects, always tend to be smooth. Integrating these two criteria together leads to a new effective dehazing model, which encourages the gradient ℓ1sparsity of atmospheric veil and implicitly maximizes the global contrast of haze-free image in the meanwhile. We also show that the proposed dehazing model can be efficiently solved using the TV-ℓ1minimization. Compared to alternative state-of-the-art methods, our approach is physically plausible and works well for all types of hazy situations. Comparative study and quantitative evaluation on both synthetic and natural images validate the superior performance and the generality of our approach.
Liang Li 0039, Wei Feng 0005, Jiawan Zhang
ICME1
2013 Maximum Cohesive Grid of Superpixels for Fast Object Localization
abstract
This paper addresses a challenging problem of regularizing arbitrary super pixels into an optimal grid structure, which may significantly extend current low-level vision algorithms by allowing them to use super pixels (SPs) conveniently as using pixels. For this purpose, we aim at constructing maximum cohesive SP-grid, which is composed of real nodes, i.e SPs, and dummy nodes that are meaningless in the image with only position-taking function in the grid. For a given formation of image SPs and proper number of dummy nodes, we first dynamically align them into a grid based on the centroid localities of SPs. We then define the SP-grid coherence as the sum of edge weights, with SP locality and appearance encoded, along all direct paths connecting any two nearest neighboring real nodes in the grid. We finally maximize the SP-grid coherence via cascade dynamic programming. Our approach can take the regional objectness as an optional constraint to produce more semantically reliable SP-grids. Experiments on object localization show that our approach outperforms state-of-the-art methods in terms of both detection accuracy and speed. We also find that with the same searching strategy and features, object localization at SP-level is about 100-500 times faster than pixel-level, with usually better detection accuracy.
Liang Li 0039, Wei Feng 0005, Jiawan Zhang
CVPR1
2012 StoryWizard: a framework for fast stylized story illustration
Jiawan Zhang, Yukun Hao, Liang Li 0039, Di Sun 0001
Vis. Comput.3
2011 Video dehazing with spatial and temporal coherence
Jiawan Zhang, Liang Li 0039, Yi Zhang 0070, Guoqiang Yang, Xiaochun Cao
Vis. Comput.2
2010 Local albedo-insensitive single image dehazing
Jiawan Zhang, Liang Li 0039, Guoqiang Yang, Yi Zhang 0070
Vis. Comput.2