Lingbo Yang

dblp:197/0031 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
7since 2021 · last 2022
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2022 Semantic Segmentation Based on Temporal Features: Learning of Temporal-Spatial Information From Time-Series SAR Images for Paddy Rice Mapping
abstract
Synthetic aperture radar (SAR) can be used to obtain remote sensing images of different growth stages of crops under all weather conditions. Such time-series SAR images can provide an abundance of temporal and spatial features for use in large-scale crop mapping and analysis. In this study, we propose a temporal feature-based segmentation (TFBS) model for accurate crop mapping using time-series SAR images. This model first extracts deep-seated temporal features and then learns the spatial context of the extracted temporal features for crop mapping. The results indicate that the TFBS model significantly outperforms traditional long short-term memory (LSTM), U-network, and convolutional LSTM models in crop mapping based on time-series SAR images. TFBS demonstrates better generalizability than other models in the study area, which makes it more transferable, and the results show that data augmentation can significantly improve this generalizability. The visualization of the temporal features extracted by the TFBS shows that there is a high degree of intraclass homogeneity among rice fields and interclass heterogeneity between rice fields and other features. TFBS also achieved the highest accuracy of the four deep learning models for multicrop classification in the study area. This study presents a feasible way of producing high-accuracy large-scale crop maps based on the proposed model.
Lingbo Yang, Jingfeng Huang, Tao Lin 0008, Limin Wang 0005, Ruzemaimaiti Mijiti, Pengliang Wei, Jie Shao 0002, Qiangzi Li, Xin Du 0004
IEEE Trans. Geosci. Remote. Sens.1
2022 Conceptual Compression via Deep Structure and Texture Synthesis
abstract
Existing compression methods typically focus on the removal of signal-level redundancies, while the potential and versatility of decomposing visual data into compact conceptual components still lack further study. To this end, we propose a novel conceptual compression framework that encodes visual data into compact structure and texture representations, then decodes in a deep synthesis fashion, aiming to achieve better visual reconstruction quality, flexible content manipulation, and potential support for various vision tasks. In particular, we propose to compress images by a dual-layered model consisting of two complementary visual features: 1) structure layer represented by structural maps and 2) texture layer characterized by low-dimensional deep representations. At the encoder side, the structural maps and texture representations are individually extracted and compressed, generating the compact, interpretable, inter-operable bitstreams. During the decoding stage, a hierarchical fusion GAN (HF-GAN) is proposed to learn the synthesis paradigm where the textures are rendered into the decoded structural maps, leading to high-quality reconstruction with remarkable visual realism. Extensive experiments on diverse images have demonstrated the superiority of our framework with lower bitrates, higher reconstruction quality, and increased versatility towards visual analysis and content manipulation tasks.
Jianhui Chang, Zhenghui Zhao, Chuanmin Jia, Shiqi Wang 0001, Lingbo Yang, Qi Mao 0002, Jian Zhang 0018, Siwei Ma 0001
IEEE Trans. Image Process.5
2021 Progressive Semantic-Aware Style Transformation for Blind Face Restoration
abstract
Face restoration is important in face image processing, and has been widely studied in recent years. However, previous works often fail to generate plausible high quality (HQ) results for real-world low quality (LQ) face images. In this paper, we propose a new progressive semantic-aware style transformation framework, named PSFR-GAN, for face restoration. Specifically, instead of using an encoder-decoder framework as previous methods, we formulate the restoration of LQ face images as a multi-scale progressive restoration procedure through semantic-aware style transformation. Given a pair of LQ face image and its corresponding parsing map, we first generate a multi-scale pyramid of the inputs, and then progressively modulate different scale features from coarse-to-fine in a semantic-aware style transfer way. Compared with previous networks, the proposed PSFR-GAN makes full use of the semantic (parsing maps) and pixel (LQ images) space information from different scales of input pairs. In addition, we further introduce a semantic aware style loss which calculates the feature style loss for each semantic region individually to improve the details of face textures. Finally, we pretrain a face parsing network which can generate decent parsing maps from real-world LQ face images. Experiment results show that our model trained with synthetic data can not only produce more realistic high-resolution results for synthetic LQ inputs but also generalize better to natural LQ face images compared with state-of-the-art methods.
Chaofeng Chen, Xiaoming Li 0002, Lingbo Yang, Xianhui Lin, Lei Zhang 0006, Kwan-Yee Kenneth Wong
CVPR3
2021 Thousand to One: Semantic Prior Modeling for Conceptual Coding
abstract
Conceptual coding has been an emerging research topic recently, which encodes natural images into disentangled conceptual representations for compression. However, the compression performance of the existing methods is still suboptimal due to the lack of comprehensive consideration of rate constraint and reconstruction quality. To this end, we propose a novel end-to-end semantic prior modeling based conceptual coding scheme towards extremely low bitrate image compression, which leverages semantic-wise deep representations as a unified prior for entropy estimation and texture synthesis. Specifically, we employ semantic segmentation maps as structural guidance for extracting deep semantic prior, which provides fine-grained texture distribution modeling for better detail construction and higher flexibility in subsequent high-level vision tasks. Moreover, a cross-channel entropy model is proposed to further exploit the inter-channel correlation of the spatially independent semantic prior, leading to more accurate entropy estimation for rate-constrained training. The proposed scheme achieves an ultra-high 1000× compression ratio, while still enjoying high visual reconstruction quality and versatility towards visual processing and analysis tasks.
Jianhui Chang, Zhenghui Zhao, Lingbo Yang, Chuanmin Jia, Jian Zhang 0018, Siwei Ma 0001
ICME3
2021 Intrinsic Temporal Regularization for High-resolution Human Video Synthesis
abstract
Fashion video synthesis has attracted increasing attention due to its huge potential in immersive media, virtual reality and online retail applications, yet traditional 3D graphic pipelines often require extensive manual labor on data capture and model rigging. In this paper, we investigate an image-based approach to this problem that generates a fashion video clip from a still source image of the desired outfit, which is then rigged in a framewise fashion under the guidance of a driving video. A key challenge for this task lies in the modeling of feature transformation across source and driving frames, where fine-grained transform helps promote visual details at garment regions, but often at the expense of intensified temporal flickering. To resolve this dilemma, we propose a novel framework with 1) a multi-scale transform estimation and feature fusion module to preserve fine-grained garment details, and 2) an intrinsic regularization loss to enforce temporal consistency of learned transform between adjacent frames. Our solution is capable of generating 512\times512 fashion videos with rich garment details and smooth fabric movements beyond existing results. Extensive experiments over the FashionVideo benchmark dataset have demonstrated the superiority of the proposed framework over several competitive baselines.
Lingbo Yang, Zhanning Gao, Siwei Ma 0001, Wen Gao 0001
ACM Multimedia1
2021 Single Scanner BLS System for Forest Plot Mapping
abstract
The 3-D information collected from sample plots is significant for forest inventories. Terrestrial laser scanning (TLS) has been demonstrated to be an effective device in data acquisition of forest plots. Although TLS is able to achieve precise measurements, multiple scans are usually necessary to collect more detailed data, which generally requires more time in scan preparation and field data acquisition. In contrast, mobile laser scanning (MLS) is being increasingly utilized in mapping due to its mobility. However, the geometrical peculiarity of forests introduces challenges. In this article, a test backpack-based MLS system, i.e., backpack laser scanning (BLS), is designed for forest plot mapping without a global navigation satellite system/inertial measurement unit (GNSS-IMU) system. To achieve accurate matching, this article proposes to combine the line and point features for calculating transformation, in which the line feature is derived from trunk skeletons. Then, a scan-to-map matching strategy is proposed for correcting positional drift. Finally, this article evaluates the effectiveness and the mapping accuracy of the proposed method in forest sample plots. The experimental results indicate that the proposed method achieves accurate forest plot mapping using the BLS; meanwhile, compared to the existing methods, the proposed method utilizes the geometric attributes of the trees and reaches a lower mapping error, in which the mean errors and the root square mean errors for the horizontal/vertical direction in plots are less than 3 cm.
Jie Shao 0002, Wuming Zhang, Nicolas Mellado, Shuangna Jin, Shangshu Cai, Lei Luo 0005, Lingbo Yang, Guangjian Yan, Guoqing Zhou 0001
IEEE Trans. Geosci. Remote. Sens.7
2021 Towards Fine-Grained Human Pose Transfer With Detail Replenishing Network
abstract
Human pose transfer (HPT) is an emerging research topic with huge potential in fashion design, media production, online advertising and virtual reality. For these applications, the visual realism of fine-grained appearance details is crucial for production quality and user engagement. However, existing HPT methods often suffer from three fundamental issues: detail deficiency, content ambiguity and style inconsistency, which severely degrade the visual quality and realism of generated images. Aiming towards real-world applications, we develop a more challenging yet practical HPT setting, termed as Fine-grained Human Pose Transfer (FHPT), with a higher focus on semantic fidelity and detail replenishment. Concretely, we analyze the potential design flaws of existing methods via an illustrative example, and establish the core FHPT methodology by combing the idea of content synthesis and feature transfer together in a mutually-guided fashion. Thereafter, we substantiate the proposed methodology with a Detail Replenishing Network (DRN) and a corresponding coarse-to-fine model training scheme. Moreover, we build up a complete suite of fine-grained evaluation protocols to address the challenges of FHPT in a comprehensive manner, including semantic analysis, structural detection and perceptual quality assessment. Extensive experiments on the DeepFashion benchmark dataset have verified the power of proposed benchmark against start-of-the-art works, with 12%-14% gain on top-10 retrieval recall, 5% higher joint localization accuracy, and near 40% gain on face identity preservation. Our codes, models and evaluation tools will be released at https://github.com/Lotayou/RATE.
Lingbo Yang, Pan Wang 0008, Chang Liu 0047, Zhanning Gao, Peiran Ren, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001, Xian-Sheng Hua 0001, Wen Gao 0001
IEEE Trans. Image Process.1
2020 Region-Adaptive Texture Enhancement For Detailed Person Image Synthesis
abstract
The ability to produce convincing textural details is essential for the fidelity of synthesized person images. Existing methods typically follow a “warping-based” strategy that propagates appearance features through the same pathway used for pose transfer. However, most fine-grained features would be lost during down-sampling, leading to over-smoothed clothes and missing details in the output images. In this paper we presents RATE-Net, a novel framework for synthesizing person images with sharp texture details. The proposed framework leverages an additional texture enhancing module to extract appearance information from the source image and estimate a fine-grained residual texture map, which helps to refine the coarse estimation from the pose transfer module. In addition, we design an effective alternate updating strategy to promote mutual guidance between two modules for better shape and appearance consistency. Experiments conducted on DeepFashion benchmark dataset have demonstrated the superiority of our framework compared with existing networks.
Lingbo Yang, Pan Wang 0008, Xinfeng Zhang 0001, Shanshe Wang, Zhanning Gao, Peiran Ren, Xuansong Xie, Siwei Ma 0001, Wen Gao 0001
ICME1
2020 HiFaceGAN: Face Renovation via Collaborative Suppression and Replenishment
abstract
Existing face restoration researches typically rely on either the image degradation prior or explicit guidance labels for training, which often lead to limited generalization ability over real-world images with heterogeneous degradation and rich background contents. In this paper, we investigate a more challenging and practical "dual-blind" version of the problem by lifting the requirements on both types of prior, termed as "Face Renovation"(FR). Specifically, we formulate FR as a semantic-guided generation problem and tackle it with a collaborative suppression and replenishment (CSR) approach. This leads to HiFaceGAN, a multi-stage framework containing several nested CSR units that progressively replenish facial details based on the hierarchical semantic guidance extracted from the front-end content-adaptive suppression modules. Extensive experiments on both synthetic and real face images have verified the superior performance of our HiFaceGAN over a wide range of challenging restoration subtasks, demonstrating its versatility, robustness and generalization ability towards real-world face processing applications. Code is available at https://github.com/Lotayou/Face-Renovation.
Lingbo Yang, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001, Chang Liu 0047, Pan Wang 0008, Peiran Ren
ACM Multimedia1