EDBT 2026 Demo / reviewers in the wild / expert
Lingyu Liang
dblp:01/10400
· DBLP profile ↗
51ranked-venue papers
12as first author
32since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 6 first-author · 14 since 2021Artificial intelligence and machine learning · 21 · 4 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 3 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 7 · 3 first-author · 3 since 2021Computer networks · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FreeMem: Enhancing Consistency in Long Video Generation via Tuning-Free MemoryabstractText-to-Video (T2V) generation has advanced greatly, yet maintaining consistency remains challenging, especially for tuning-free long video generation. We attribute the consistency problem to cumulative deviations for long video generation at three levels: the random noise lacking correlation results initial deviation between frames; discrepancy in semantic feature tokens between denoising network blocks gradually accumulates as the frame count grows, leading to greater deviations; attention mechanisms struggle to capture global relationships across distant frames in long videos. To address these, we propose FreeMem, a tuning-free framework leveraging hierarchical memory update and injection: the noise memory stabilizes consistency by manipulating low and high frequency components in the initial noise space; the token memory combats inconsistency through adaptive fusion of historical and current semantic feature tokens between denoising network blocks; and the attention memory establishes persistent cache to model long-range relationships within self attention layers. Evaluated on VBench, FreeMem improves subject and background consistency matrics across various methods, offering a practical solution for low-cost, high-consistency long video generation. Jibin Peng, Di Lin 0002, Zhecheng Xu, Wuyuan Xie, Miaohui Wang, Lingyu Liang, Qing Guo 0005 |
AAAI | 8 |
| 2025 | Deep Unfolding for Task-Decomposed Image Restoration Under Diverse Degradations
Jiaxuan Cheng, Lingyu Liang, Xinchao Li, Guoxi Sun, Shuangping Huang |
ICIC (3) | 2 |
| 2025 | Personalized Cross-Silo Federated Learning with Adaptive Proximal Relationships and Gradient-Based Aggregation
Lingyu Liang, Jiaxuan Cheng |
ICIC (9) | 2 |
| 2025 | COS-SLAM: Coordinate Attention Semantic SLAM with Pixel-to-Line Transformer
Handong Shen, Lingyu Liang, Xiaohao Liu, Xinchao Li, Guoxi Sun, Shuangping Huang |
ICIG (2) | 2 |
| 2025 | ReplayCAD: Generative Diffusion Replay for Continual Anomaly DetectionabstractContinual Anomaly Detection (CAD) enables anomaly detection models in learning new classes while preserving knowledge of historical classes. CAD faces two key challenges: catastrophic forgetting and segmentation of small anomalous regions. Existing CAD methods store image distributions or patch features to mitigate catastrophic forgetting, but they fail to preserve pixel-level detailed features for accurate segmentation. To overcome this limitation, we propose ReplayCAD, a novel diffusion-driven generative replay framework that replay high-quality historical data, thus effectively preserving pixel-level detailed features. Specifically, we compress historical data by searching for a class semantic embedding in the conditional space of the pre-trained diffusion model, which can guide the model to replay data with fine-grained pixel details, thus improving the segmentation performance. However, relying solely on semantic features results in limited spatial diversity. Hence, we further use spatial features to guide data compression, achieving precise control of sample space, thereby generating more diverse data. Our method achieves state-of-the-art performance in both classification and segmentation, with notable improvements in segmentation: 11.5% on VisA and 8.1% on MVTec. Our source code is available at https://github.com/HULEI7/ReplayCAD. Lei Hu 0012, Zhiyong Gan, Ling Deng, Jinglin Liang 0001, Lingyu Liang, Shuangping Huang, Tianshui Chen |
IJCAI | 5 |
| 2025 | ELFSurf: Explicit Latent Fusion for Implicit Surface ReconstructionabstractImplicit neural networks are widely used to reconstruct 3D surfaces from noisy point clouds. In order to construct accurate implicit fields, existing methods encode input point clouds into either feature vectors for each point (point latents) or regular grid features (grid latents). Point latents can capture higher frequency features and are good for detailed reconstruction. However, the reconstruction results are easily affected by noise and point distribution, resulting in incomplete reconstruction. In contrast, grid latents are coarser and therefore some details are lost, but the reconstruction results are more complete and smoother. In order to make full use of these two types of latents, we propose a new method to explicit fuse them for more accurate 3D surface reconstruction, called ELFSurf. Specifically, ELFSurf obtains the feature vectors for each point by our Point Convolution Module (PCM). The grid features are obtained by Grid Transformer Module (GTM). In addition, we design a Sparse Decoding Module (SDM) to sparsify and fuse the extracted features to improve the decoding performance. The decoder finally maps the features to occupancy probabilities. The experiments demonstrate that our method can achieve more accurate 3D surface reconstruction on object-level and scene-level datasets with better generalization compared to the state-of-the-art approaches. Jijun Zhou, Zhuhua Yang, Lingyu Liang, Yutian Yang |
IJCNN | 3 |
| 2025 | FedCare: towards interactive diagnosis of federated learning systems
Tian-Ye Zhang, Haozhe Feng, Wenqi Huang 0002, Lingyu Liang, Huanming Zhang, Zexian Chen, Anthony K. H. Tung, Wei Chen 0001 |
Frontiers Comput. Sci. | 4 |
| 2025 | Demand-Driven Sparse Mobile Crowdsensing With Neighborhood-Aware ReconstructionabstractSparse mobile crowdsensing (SMCS) is a cost-effective paradigm aimed at recruiting workers to complete sensing tasks and inferring the remaining unobserved data, with broad applications in large-scale, fine-grained monitoring services. In SMCS, spatial coverage of the sensing area or global completion accuracy is typically used as the performance metric. However, in many real-world service scenarios (e.g., temperature, humidity, air quality monitoring), users are generally only interested in data from their specific regions and expect the highest possible data accuracy. In such cases, relying solely on coverage or global completion error fails to adequately assess the quality of the platform’s service. To address this and satisfy users’ sensing demands as much as possible while maintaining low sensing costs, we propose the Demand-Driven Framework with Neighborhood-Aware Data Reconstruction (D2-SMCS), which integrates regional population demand calculation, dynamic clustering, and data reconstruction. Unlike existing approaches, we introduce quality of service (QoS) as a performance metric based on regional population demand. First, we quantify the interest level of sensing tasks in different regions by considering factors such as population demand and data fluctuation. Based on this quantification, the dynamic clustering module selects the regions most beneficial for accurate data completion. Finally, to overcome the limitation of traditional matrix completion methods in capturing short-term variations, we propose an innovative Neighborhood-Aware Latent Matrix Completion (NALMC) approach to infer and complete the unobserved regions. Extensive experiments on real-world datasets demonstrate the effectiveness of our framework. Qihang Zhou, Guoqiang Deng, Lingyu Liang, Xinglin Zhang 0001 |
IEEE Internet Things J. | 5 |
| 2025 | Reads: A Personalized Federated Learning Framework With Fine-Grained Layer Aggregation and Decentralized ClusteringabstractThe heterogeneity of local data and client performance, along with real-world system risks, is driving the evolution of federated learning (FL) towards personalized, model-heterogeneous, and decentralized approaches. However, due to the differing structures of heterogeneous models, it is hard to use them to identify clients with similar data distributions and further enhance the personalization of local models. Therefore, how to deal with data heterogeneity to obtain superior personalized local models for clients, while simultaneously addressing model heterogeneity and system risks is a challenging problem. In this paper, we propose a novel personalized FL framework with fine-gRained layEr aggregAtion andDecentralized cluStering (${\sf Reads}$), which integrates four key components: (1) deep mutual learning with privacy guarantee for model training and privacy preservation, (2) fine-grained layer similarity computation among heterogeneous model layers, (3) fully decentralized clustering for soft clustering of clients based on layer similarities, and (4) personalized layer aggregation for capturing common knowledge from other clients. Through${\sf Reads}$, clients obtain personalized models that accommodate model heterogeneity, while the system ensures robustness against a single point of failure. Extensive experiments demonstrate the efficacy of${\sf Reads}$in achieving these goals. Haoyu Fu, Fengsen Tian, Guoqiang Deng, Lingyu Liang, Xinglin Zhang 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2025 | A Pricing Game for Federated Learning Supporting Lightweight Local Model TrainingabstractThe pervasive distribution of data across clients with privacy concerns and heterogeneous performance in edge networks presents a significant opportunity to enhance AI model performance. Federated learning (FL) enables a model owner (MO) to recruit these clients, offering compensation for their contributions, and to improve model quality by aggregating knowledge from their locally trained models. However, several challenges arise in this process. Clients may decline participation if they do not achieve positive utility. Moreover, due to constraints in memory, computing, and communication resources, some clients can only train lightweight models that represent partial versions of the global model. Importantly, the MO's pricing for client contributions and the proportions of local model training are interdependent, collectively influencing client utilities and participation decisions. To address these challenges, we first model the utility functions of both the MO and the clients, accommodating the support for lightweight local models. We then formulate their interactions as a Stackelberg game and theoretically prove the existence of a Nash equilibrium. Based on this equilibrium, we derive optimal collaboration strategies for both the MO and the clients. Additionally, we design an efficient approximation algorithm to enable the MO to maximize its utility by selecting suitable clients to participate in FL. Finally, extensive experiments validate our theoretical findings, demonstrating the superior performance and effectiveness of the proposed algorithms Fengsen Tian, Mingzi Wang, Guoqiang Deng, Lingyu Liang, Xinglin Zhang 0001 |
IEEE Trans. Mob. Comput. | 5 |
| 2024 | KMTalk: Speech-Driven 3D Facial Animation with Key Motion Embedding
Shengjie Gong, Jiapeng Tang, Lingyu Liang, Yining Huang, Shuangping Huang |
ECCV (56) | 4 |
| 2024 | DEGAN: Discrimination Enhanced GAN for Perceptual-Oriented Super-ResolutionabstractRecent years, generative adversarial networks (GANs) have gained significant prominence in single image super-resolution (SISR) tasks. This can mainly be attributed to their exceptional ability to generate intricate details. However, the instability and lack of realism in the details generated by GANs have been challenges. Existing methods mainly concentrate on improving the generator and designing complex loss functions, often overlooking the important role of discrimination. To this end, we propose our discrimination enhanced GAN (DEGAN) by improving the discriminator and simplify the discrimination task. We introduce an efficient wide activation UNet to enhance the discriminator, enabling a more comprehensive and nuanced analysis of the input image. Additionally, we introduce a texture aware mask that provides more precise guidance and alleviates the difficulty of discrimination. Our DEGAN is simple yet effective. Quantitative and visual comparisons with state-of-the-art methods on benchmark datasets demonstrate the superiority of our method. Xiaoyu Jin, Wenqi Huang 0002, Lingyu Liang, Yang Wu 0001, Qunsheng Zeng, Ruiye Zhou, Zhuojun Cai, Jianing Shang, Wenming Yang |
ICASSP | 3 |
| 2024 | Document Image Dewarping Guided by 3D Geometry and Layout PriorsabstractDocument image dewarping aims to reconstruct the flat document image from distorted inputs. Previous methods often use geometric or text-line priors to guide the dewarping process. However, document images contain diverse contents, including figures, tables, or paragraph structures, image dewarping without considering the layout structure may fail to obtain global optimization. This paper proposes an encoder-decoder neural network, called DocTLNet, to achieve document image dewarping, which uses both 3D geometry and layout as constraints to refine the content details. To further enhance the layout details, a layout-aug loss is also proposed to explicitly guides the network to handle the distorted layout boundaries. Qualitative and quantitative experiments were conducted on DocUNet benchmark, and the results indicate that our DocTLNet is superior to related methods. Lingyu Liang, Shuangping Huang |
ICME | 2 |
| 2024 | Handwriting Trajectory Recovery Via Trajectory Transformer With Global Radical Context-Aware Module
Junxiang Lin, Zhounan Chen, Lingyu Liang, Wenjie Peng, Shuangping Huang |
ICPR (20) | 3 |
| 2024 | Voxel Proposal Network via Multi-Frame Knowledge Distillation for Semantic Scene CompletionabstractSemantic scene completion is a difficult task that involves completing the geometry and semantics of a scene from point clouds in a large-scale environment. Many current methods use 3D/2D convolutions or attention mechanisms, but these have limitations in directly constructing geometry and accurately propagating features from related voxels, the completion likely fails while propagating features in a single pass without considering multiple potential pathways. And they are generally only suitable for static scenes and struggle to handle dynamic aspects. This paper introduces Voxel Proposal Network (VPNet) that completes scenes from 3D and Bird's-Eye-View (BEV) perspectives. It includes Confident Voxel Proposal based on voxel-wise coordinates to propose confident voxels with high reliability for completion. This method reconstructs the scene geometry and implicitly models the uncertainty of voxel-wise semantic labels by presenting multiple possibilities for voxels. VPNet employs Multi-Frame Knowledge Distillation based on the point clouds of multiple adjacent frames to accurately predict the voxel-wise labels by condensing various possibilities of voxel relationships. VPNet has shown superior performance and achieved state-of-the-art results on the SemanticKITTI and SemanticPOSS datasets. Lubo Wang, Di Lin 0002, Kairui Yang, Qing Guo 0005, Wuyuan Xie, Miaohui Wang, Lingyu Liang, Ping Li 0016 |
NeurIPS | 8 |
| 2024 | Multi-view Depth Estimation with Adaptive Feature Extraction and Region-Aware Depth Prediction
Chi Zhang 0098, Lingyu Liang, Jijun Zhou, Yong Xu 0007 |
PRCV (6) | 2 |
| 2024 | An Industrial Scene Text Detection with Spectral Domain Enhancement and Graph Fourier MappingabstractText detection is a task of great significance in different scenarios, which has wide applications for downstream tasks such as text recognition and text retrieval. Varieties of text detection methods have been proposed to solve this problem in natural scenes and have achieved good results. However, these methods cannot get satisfied performance in industrial scenes for various interferences caused by the industrial environment like the image noise, background material texture and low contrast. To deal with these challenges, we first propose a contour modeling algorithm based on graph Fourier transform mapping to represent arbitrary shaped text contours. A refined fast Fourier convolutional network module with ability of spectral-domain sensing is also intruduced to enhance text feature and suppress interference. Based on these two components, we construct a novel network to achieve accurate industrial scene text detection. Quantitative evaluations are conducted on benchmark datasets MPSC and IcText, experimental results show that our method obtains the state-of-the-art detection accuracy. Wocheng Xiao, Lingyu Liang, Shuangping Huang |
SMC | 2 |
| 2024 | GNF-Net: An Adaptive Mesh Denoising Method with GCN-Based Guided Normal FilteringabstractMesh Denoising has become a popular area of research, many traditional and learning-based methods have been proposed to remove noise from meshes. However, most approaches only focus on denoising meshes with low levels of noise. When the noise is high-frequency, it can be difficult to recover the original shape. In this paper, we present a mesh denoising approach that utilizes graph convolution representations to enhance the understanding of the mesh characteristics. It integrates precisely designed graphs to explore the inherent compositional structure of the mesh. When analyzing meshes under the influence of various noises, we extract information about the original features of the mesh by capturing spatial geometric features through graph convolution calculations. Our method is based on Guided Normal Filtering (GNF) to design a Graphical Representation Module (GRM) and a GCN-Based Normal Prediction Module (NPM). It can adaptively obtain the optimal guided normal vectors for noisy meshes. We have compared and analyzed the various methods to produce state-of-the-art results. Lingyu Liang, Yutian Yang, Shuangping Huang |
SMC | 2 |
| 2024 | A Multi-Stream Structure-Enhanced Network for Mesh DenoisingabstractTriangular meshes provide an efficient representation of 3D shapes. Various applications such as 3D simulation suffer from degradation in geometric quality. This paper proposes a novel Multi-stream Structure-Enhanced Network (MSE-Net) based on graph convolutional networks. The network uses multi-scale features besides vertex position to guide face normal filtering, which can better preserve the geometric feature during the denoising process. In contrast to former methods that focus on filtering vertex coordinate and face normal apart, MSE-Net innovatively fuses more structure features like face area, inner product between face normal and vertex normals, and the interior angles of face to guide the face normal and vertex position updating, utilizing the inherent structural characteristic of Mesh. Our method achieves state-of-the-art performance on several publicly available datasets, demonstrating its effectiveness. Yutian Yang, Lingyu Liang, Yong Xu 0007 |
SMC | 2 |
| 2024 | Multi-view depth estimation based on multi-feature aggregation for 3D reconstruction
Chi Zhang 0098, Lingyu Liang, Jijun Zhou, Yong Xu 0007 |
Comput. Graph. | 2 |
| 2024 | Federated Graph Augmentation for Semisupervised Node ClassificationabstractSemisupervised node classification is a prevalent task on graphs, which involves predicting the labels of unlabeled nodes based on limited labeled data available. At present, centralized approaches to training models for this task are unsustainable due to the increasing demand for computational power, storage capacity, and privacy. An approach of potential is federated graph learning (FGL), which allows multiple clients to collaborate on learning a model while maintaining data privacy. However, current methods suffer from the inability to consider the topology of the graph data and inadequate use of unlabeled data. To address these issues, we propose federated graph augmentation (FedGA) by combining graph neural network (GNN) models to utilize similar topologies existing in different client graphs and augment the client data. Furthermore, we develop FedGA-L based on FedGA, which integrates pseudolabeling and label-injection to improve the utilization of unlabeled data. FedGA-L allows pseudolabels to be used as additional information to enhance data augmentation and further improve the accuracy of node classification. We evaluate the effectiveness of FedGA and FedGA-L through experiments on multiple datasets. The results demonstrate improved accuracy in solving typical classification tasks and their compatibility with a variety of federated learning (FL) frameworks. On widely recognized datasets for graph learning, we achieve an accuracy improvement of 5%–7% compared to vanilla federated learning algorithms. Zhichang Xia, Xinglin Zhang 0001, Lingyu Liang, Yun Li 0002, Yue-Jiao Gong |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2023 | An Entity Alignment Method Based on Graph Attention Network with Pre-classification
Wenqi Huang 0002, Lingyu Liang, Yongjie Liang, Jiaxuan Hou, Xuanang Li |
WISA | 2 |
| 2023 | CVSformer: Cross-View Synthesis Transformer for Semantic Scene CompletionabstractSemantic scene completion (SSC) requires an accurate understanding of the geometric and semantic relationships between the objects in the 3D scene for reasoning the occluded objects. The popular SSC methods voxelize the 3D objects, allowing the deep 3D convolutional network (3D CNN) to learn the object relationships from the complex scenes. However, the current networks lack the controllable kernels to model the object relationship across multiple views, where appropriate views provide the relevant information for suggesting the existence of the occluded objects. In this paper, we propose Cross-View Synthesis Transformer (CVSformer), which consists of Multi-View Feature Synthesis and Cross-View Transformer for learning cross-view object relationships. In the multi-view feature synthesis, we use a set of 3D convolutional kernels rotated differently to compute the multi-view features for each voxel. In the cross-view transformer, we employ the cross-view fusion to comprehensively learn the cross-view relationships, which form useful information for enhancing the features of individual views. We use the enhanced features to predict the geometric occupancies and semantic labels of all voxels. We evaluate CVSformer on public datasets, where CVS-former yields state-of-the-art results. Our code is available at https://github.com/donghaotian123/CVSformer. Haotian Dong, Enhui Ma, Lubo Wang, Miaohui Wang, Wuyuan Xie, Qing Guo 0005, Ping Li 0016, Lingyu Liang, Kairui Yang, Di Lin 0002 |
ICCV | 8 |
| 2023 | HisDoc R-CNN: Robust Chinese Historical Document Text Line Detection with Dynamic Rotational Proposal Network and Iterative Attention Head
Cheng Jian, Lingyu Liang, Chongyu Liu |
ICDAR (1) | 3 |
| 2023 | Adaptive Low-Light Image Enhancement Optimization Framework with Algorithm Unrolling
Qichang He, Lingyu Liang, Wocheng Xiao, Mingju Liang |
PRCV (11) | 2 |
| 2023 | Constituent Attention for Vision Transformers
Haoling Li, Mengqi Xue, Jie Song 0011, Haofei Zhang, Wenqi Huang 0002, Lingyu Liang, Mingli Song |
Comput. Vis. Image Underst. | 6 |
| 2022 | Identification of Bird's Nest Hazard Level of Transmission Line Based on Improved Yolov5 and Location Constraints
Yang Wu 0001, Qunsheng Zeng, Wenqi Huang 0002, Lingyu Liang |
PRCV (4) | 5 |
| 2022 | Regression Guided by Relative Ranking Using Convolutional Neural Network (R$^3$3CNN) for Facial Beauty PredictionabstractFacial beauty prediction (FBP) aims to automatically assess facial attractiveness consistently with judgements based on human perception. Most of previous methods formulate FBP as a classification, regression or ranking problem of machine learning. However, humans not only represent facial attractiveness as a score, but also perceive the relative aesthetics of faces. Inspired by this observation, we formulate FBP as a specific regression problem guided by ranking information. Specifically, we propose a general CNN architecture, called R$^3$CNN, to integrate the relative ranking of faces in terms of aesthetics to improve performance of FBP. As R$^3$CNN consists of both regression and ranking components, it is challenging to train and fine-tune it by existing techniques. To tackle this problem, we propose the following learning schemes for R$^3$CNN: 1) a hard pair sampling strategy that generates challenging-to-predicted image pairs and pseudo ranking labels from true rating scores; 2) an assemble loss function that combines regression loss and pairwise ranking loss (PR-Loss), where PR-Loss can be a hinge-form loss or a log-sum-exp pairwise loss; 3) a cascaded fine-tuning method that further improves prediction. Moreover, we build a benchmark dataset, called SCUT-FBP5500, containing 5,500 facial images with diverse properties (male/female, Asian/Caucasian, ages) and labels (face landmarks, rating scores within [1, 5], rating score distribution). Experiments were performed on both the SCUT-FBP and the SCUT-FBP5500 benchmark datasets, where our method achieves state-of-the-art performance on different evaluation settings. Comparisons with related CNN models highlight the effectiveness of the R$^3$CNN architecture for FBP. Luojun Lin, Lingyu Liang |
IEEE Trans. Affect. Comput. | 2 |
| 2022 | PDE Learning of Filtering and Propagation for Task-Aware Facial Intrinsic Image AnalysisabstractFiltering and propagation are two basic operations in image analysis and rendering, and they are also widely used in computer graphics and machine learning. However, the models of filtering and propagation were based on diverse mathematical formulations, which have not been fully understood. This article aims to explore the properties of both filtering and propagation models from a partial differential equation (PDE) learning perspective. We propose a unified PDE learning framework based on nonlinear reaction-diffusion with a guided map, graph Laplacian, and reaction weight. It reveals that: 1) the guided map and reaction weight determines whether the PDE produces filtering or propagation diffusion and 2) the kernel of graph Laplacian controls the diffusion pattern. Based on the proposed PDE framework, we derive the mathematical relations between different models, including learning to diffusion (LTD) model, label propagation, edit propagation, and edge-aware filter. In practical verification, we apply the PDE framework to design diffusion operations with the adaptive kernel to tackle the ill-posed problem of facial intrinsic image analysis (FIIA). A flexible task-aware FIIA system is built to achieve various facial rendering effects, such as face image relighting and delighting, artistic illumination transfer, illumination-aware face swapping, or transfiguring. Qualitative and quantitative experiments show the effectiveness and flexibility of task-aware FIIA and provide new insights on PDE learning for visual analysis and rendering. Lingyu Liang, Yong Xu 0007 |
IEEE Trans. Cybern. | 1 |
| 2022 | Deep LSAC for Fine-Grained RecognitionabstractFine-grained recognition emphasizes the identification of subtle differences among object categories given objects that appear in different shapes and poses. These variances should be reduced for reliable recognition. We propose a fine-grained recognition system that incorporates localization, segmentation, alignment, and classification in a unified deep neural network. The input to the classification module includes functions that enable backward-propagation (BP) in constructing the solver. Our major contribution is to propose a valve linkage function (VLF) for BP chaining and form our deep localization, segmentation, alignment, and classification (LSAC) system. The VLF can adaptively compromise errors of classification and alignment when training the LSAC model. It in turn helps to update the localization and segmentation. We evaluate our framework on two widely used fine-grained object data sets. The performance confirms the effectiveness of our LSAC system. Di Lin 0002, Yi Wang 0031, Lingyu Liang, Ping Li 0016, C. L. Philip Chen |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | Fourier Contour Embedding for Arbitrary-Shaped Text DetectionabstractOne of the main challenges for arbitrary-shaped text detection is to design a good text instance representation that allows networks to learn diverse text geometry variances. Most of existing methods model text instances in image spatial domain via masks or contour point sequences in the Cartesian or the polar coordinate system. However, the mask representation might lead to expensive post-processing, while the point sequence one may have limited capability to model texts with highly-curved shapes. To tackle these problems, we model text instances in the Fourier domain and propose one novel Fourier Contour Embedding (FCE) method to represent arbitrary shaped text contours as compact signatures. We further construct FCENet with a backbone, feature pyramid networks (FP-N) and a simple post-processing with the Inverse Fourier Transformation (IFT) and Non-Maximum Suppression (N-MS). Different from previous methods, FCENet first pre-dicts compact Fourier signatures of text instances, and then reconstructs text contours via IFT and NMS during test. Extensive experiments demonstrate that FCE is accurate and robust to fit contours of scene texts even with highly-curved shapes, and also validate the effectiveness and the good generalization of FCENet for arbitrary-shaped text detection. Furthermore, experimental results show that our FCENet is superior to the state-of-the-art (SOTA) meth-ods on CTW1500 and Total-Text, especially on challenging highly-curved text subset. Yiqin Zhu, Jianyong Chen, Lingyu Liang, Zhanghui Kuang, Wayne Zhang 0001 |
CVPR | 3 |
| 2021 | Illumination-Aware Image Quality Assessment for Enhanced Low-Light Image
Sigan Yao, Yiqin Zhu, Lingyu Liang, Tao Wang 0047 |
PRCV (3) | 3 |
| 2020 | SCUT-HCCDoc: A new benchmark dataset of handwritten Chinese text in unconstrained camera-captured documents
Hesuo Zhang, Lingyu Liang |
Pattern Recognit. | 2 |
| 2019 | Attribute-Aware Convolutional Neural Networks for Facial Beauty PredictionabstractFacial beauty prediction (FBP) aims to develop a machine that automatically makes facial attractiveness assessment. To a large extent, the perception of facial beauty for a human is involved with the attributes of facial appearance, which provides some significant visual cues for FBP. Deep convolution neural networks (CNNs) have shown its power for FBP, but convolution filters with fixed parameters cannot take full advantage of the facial attributes for FBP. To address this problem, we propose an Attribute-aware Convolutional Neural Network (AaNet) that modulates the filters of the main network, adaptively, using parameter generators that take beauty-related attributes as extra inputs. The parameter generators update the filters in the main network in two different manners: filter tuning or filter rebirth. However, AaNet takes attributes information as prior knowledge, that is ill-suited to those datasets merely with task-oriented labels. Therefore, imitating the design of AaNet, we further propose a Pseudo Attribute-aware Convolutional Neural Network (P-AaNet) that modulates filters conditioned on global context embeddings (pseudo attributes) of input faces learnt by a lightweight pseudo attribute distiller. Extensive ablation studies show that the AaNet and P-AaNet improve the performance of FBP when compared to conventional convolution and attention scheme, which validates the effectiveness of our method. Luojun Lin, Lingyu Liang, Weijie Chen 0006 |
IJCAI | 2 |
| 2019 | Adaptive GNN for Image Analysis and EditingabstractGraph neural network (GNN) has powerful representation ability, but optimal configurations of GNN are non-trivial to obtain due to diversity of graph structure and cascaded nonlinearities. This paper aims to understand some properties of GNN from a computer vision (CV) perspective. In mathematical analysis, we propose an adaptive GNN model by recursive definition, and derive its relation with two basic operations in CV: filtering and propagation operations. The proposed GNN model is formulated as a label propagation system with guided map, graph Laplacian and node weight. It reveals that 1) the guided map and node weight determine whether a GNN leads to filtering or propagation diffusion, and 2) the kernel of graph Laplacian controls diffusion pattern. In practical verification, we design a new regularization structure with guided feature to produce GNN-based filtering and propagation diffusion to tackle the ill-posed inverse problems of quotient image analysis (QIA), which recovers the reflectance ratio as a signature for image analysis or adjustment. A flexible QIA-GNN framework is constructed to achieve various image-based editing tasks, like face illumination synthesis and low-light image enhancement. Experiments show the effectiveness of the QIA-GNN, and provide new insights of GNN for image analysis and editing. Lingyu Liang, Yong Xu 0007 |
NeurIPS | 1 |
| 2019 | Adaptive Label Propagation for Facial Appearance TransferabstractFacial appearance transfer (FAT) is a critical component of various facial editing tasks. It aims to transfer the facial appearance of a reference into a target with good visual consistency. When there are considerable visual differences between a reference and a target, however, it may introduce visual artifacts into the results. To tackle this problem, we propose a facial appearance map with illumination-aware and region-aware properties that allows seamless FAT. We formulate the appearance-map generation as label propagation (LP) on a similarity graph, and propose a new regularization structure to facilitate the adaptive appearance-map diffusion. Solving the original LP model of appearance map in general requires on the order$O(kn^2)$time for an$n$-nodes graph where each node has$k$neighbors. It may be computationally prohibitive for an image with a large spatial resolution. To tackle this problem, we mathematically analyze the graph-based LP model and propose a fast algorithm with smart subset sampling. It selects a subset with$m$nodes of the graph with$n$nodes ($m\ll n$) to approximate the solution to the original system, which significantly reduces its computational requirements from$O(kn^2)$to$O(m^2n)$. Based on the adaptive LP-based appearance map, we construct a framework to achieve various editing effects with FAT, including face replacement, face dubbing, face swapping, and transfiguring. Comparisons with related methods show the effectiveness of the adaptive LP model for FAT. Qualitative and quantitative evaluations verify the computational improvements of the approximation algorithm. Lingyu Liang, Xinglin Zhang 0001 |
IEEE Trans. Multim. | 1 |
| 2018 | SCUT-FBP5500: A Diverse Benchmark Dataset for Multi-Paradigm Facial Beauty PredictionabstractFacial beauty prediction (FBP) is a significant visual recognition problem to make assessment of facial attractiveness that is consistent to human perception. To tackle this problem, various data-driven models, especially state-of-the-art deep learning techniques, were introduced, and benchmark dataset become one of the essential elements to achieve FBP. Previous works have formulated the recognition of facial beauty as a specific supervised learning problem of classification, regression or ranking, which indicates that FBP is intrinsically a computation problem with multiple paradigms. However, most of FBP benchmark datasets were built under specific computation constrains, which limits the performance and flexibility of the computational model trained on the dataset. In this paper, we argue that FBP is a multi-paradigm computation problem, and propose a new diverse benchmark dataset, called SCUT-FBP5500, to achieve multi-paradigm facial beauty prediction. The SCUT-FBP5500 dataset has totally 5500 frontal faces with diverse properties (male/female, Asian/Caucasian, ages) and diverse labels (face landmarks, beauty scores within [1], [5], beauty score distribution), which allows different computational models with different FBP paradigms, such as appearance-based/shape-based facial beauty classification/regression model for male/female of Asian/Caucasian. We evaluated the SCUT-FBP5500 dataset for FBP using different combinations of feature and predictor, and various deep learning methods. The results indicates the improvement of FBP and the potential applications based on the SCUT-FBP5500. Lingyu Liang, Luojun Lin, Duorui Xie, Mengru Li |
ICPR | 1 |
| 2018 | R2-ResNeXt: A ResNeXt-Based Regression Model with Relative Ranking for Facial Beauty PredictionabstractThe purpose of facial beauty prediction (FBP) is to develop a machine that automatically evaluates facial attractiveness in a human perceptual manner. One of the essential problem of facial beauty prediction is the discriminative facial representation of the prediction model. Previous methods formulate FBP as a specific supervised learning of classification, regression, or ranking. We find that the relative ranking information is useful to improve the regression model of FBP. Based on this observation, this paper proposes a regression model guided by the relative ranking with the state-of-the-art Res NeXt structure to achieve FBP, and we call the model as R2-ResNeXt. The R2-ResNeXt facilitates to learn the representation and predictor guided by relative ranking for facial attractiveness assessment in an end-to-end manner. To train the R2-ResNeXt, we develop an aggregated loss that combines regression loss and pairwise ranking loss linearly. We also design a method to construct a dataset containing relatively -labelled image pairs whose individual images are sampled from the SCUT-FBP benchmark database. The experimental results on the SCUT-FBP benchmark show that our R2-ResNeXt achieves the state-of-the-art performance compared with related literatures, and further indicates the effectiveness of the deep residual architecture and relative beauty ranking into regression task for facial beauty prediction. Luojun Lin, Lingyu Liang |
ICPR | 2 |
| 2018 | Content-Aware Face Blending by Label Propagation
Lingyu Liang, Xinglin Zhang 0001 |
PRCV (3) | 1 |
| 2018 | Learning to Generate Realistic Scene Chinese Character Images by Multitask Coupled GAN
Qingxiang Lin, Lingyu Liang, Yaoxiong Huang |
PRCV (3) | 2 |
| 2017 | Facial attractiveness prediction using psychologically inspired convolutional neural network (PI-CNN)abstractThis paper proposes a psychologically inspired convolutional neural network (PI-CNN) to achieve automatic facial beauty prediction. Different from the previous methods, the PI-CNN is a hierarchical model that facilitates both the facial beauty representation learning and predictor training. Inspired by the recent psychological studies, significant appearance features of facial detail, lighting and color were used to optimize the PI-CNN facial beauty predictor using a new cascaded fine-tuning method. Experiments indicate that the cascaded fine-tuned PI-CNN predictor is robust to facial appearance variances, and obtains the highest correlation of 0.87 in the SCUT-FBP benchmark database, which is superior to the related hand-designed feature and related deep learning methods. Jie Xu 0041, Lingyu Liang, Ziyong Feng, Duorui Xie, Huiyun Mao |
ICASSP | 3 |
| 2017 | Region-aware scattering convolution networks for facial beauty predictionabstractThis paper proposes a scattering convolutional network with region-aware facial attributes to obtain a mid-level representation for facial beauty prediction (FBP). Different from the previous works that only focus on the discriminative representation for prediction, this paper also considers the invariant properties of the facial representation that reduces the variances caused by the image transformations, such as rotations and translation. The proposed region-aware scattering convolution network (RegionScatNet) is based on a deep convolution network of scattering transforms (ScatNet) integrated with facial texture and shape features. It consists of three components, including: 1) Region Extraction to obtain the significant facial perception region with implicit shape features using region-aware mask; 2) Attributes Decomposition to separate the extracted region into detail and structure facial layers by a guided filter; 3) Scattering Convolution that computes the roto-translation invariant representation of facial detail and structure for FBP by cascading three-layer wavelets filters and non-linear modulus pooling. The comparisons with related deep learning-based methods illustrate the effectiveness of RegionScatNet for FBP. The evaluations with various prediction model (like SVR and Gaussian process) and with different facial variances (like rotaion) indicate the robustness of the RegionScatNet-based features. Lingyu Liang, Duorui Xie, Jie Xu 0041, Mengru Li, Luojun Lin |
ICIP | 1 |
| 2017 | Edge-Aware Label Propagation for Mobile Facial Enhancement on the CloudabstractThis paper proposes a facial enhancement framework with mask generation for cloud-based mobile applications. We mathematically analyze and unify the mask generation, as well as the state-of-the-art region-aware mask and edit propagation techniques, from a graph-based semi supervised learning perspective. Then we propose a label propagation model with a new edge-aware structure and guided feature for mask generation. The limit analysis of the model leads to a fast algorithm, which reduces the intrinsic computation cost. Then we develop a flexible and efficient cloud-based PaaS system, called FaceMore, for intelligent mobile face enhancement. The flexibility and extendibility of the cloud-based architecture facilitates intelligent facial enhancement applications and the parallel processing effectively improves the efficiency of the algorithm. Qualitative and quantitative evaluations were performed for mask propagation and face enhancement. Comparisons with the previous methods and five representative commercial systems, including PicTreat, Portraiture, Portrait+, Meitu, and Baidu Motu, illustrate the robustness and effectiveness of our method. Lingyu Liang, Deng Liu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2015 | FaceMore: A Face Beautification Platform on the CloudabstractFace More, a cloud-based face beautification platform for intelligent face manipulation, is developed in this work. It provides flexible and efficient cloud API to develop automatic or interactive face retouching applications. A web-site, www.facemore.net, is built on Face More, where user can upload images and obtain various online face beautification services. To obtain automatic inhomogeneous editing effects in a natural and efficient manner, we analyze the novel tool, called region-aware mask, from semi-supervised learning perspective. We reformulate the optimized-based model of region-aware mask using label propagation, and propose a fast approximate algorithm for mask generation, which leads to about 40% speed improvement while maintain high visual quality for face beatification. Qualitative and quantitative evaluations were performed for mask generation and face beatification. Comparisons with five representative commercial systems, including PicTreat, Portraiture, Portrait+, Meitu and Baidu Motu, illustrate the effectiveness of our system for image enhancement of facial lighting, smoothness and color. Lingyu Liang, Deng Liu |
SMC | 1 |
| 2015 | SCUT-FBP: A Benchmark Dataset for Facial Beauty PerceptionabstractIn this paper, a novel face dataset with attractiveness ratings, namely the SCUT-FBP dataset, is developed for automatic facial beauty perception. This dataset provides a benchmark to evaluate the performance of different methods for facial attractiveness prediction, including the state-of-the-art deep learning method. The SCUT-FBP dataset contains face portraits of 500 Asian female subjects with attractiveness ratings, all of which have been verified in terms of rating distribution, standard deviation, consistency, and self-consistency. Benchmark evaluations for facial attractiveness prediction were performed with different combinations of facial geometrical features and texture features using classical statistical learning methods and the deep learning method. The best Pearson correlation 0.8187 was achieved by the CNN model. The results of the experiments indicate that the SCUT-FBP dataset provides a reliable benchmark for facial beauty perception. Duorui Xie, Lingyu Liang, Jie Xu 0041, Mengru Li |
SMC | 2 |
| 2015 | Multiple Facial Image Editing Using Edge-Aware PDE LearningabstractThis paper introduces a novel facial editing tool, called edge-aware mask, to achieve multiple photo-realistic rendering effects in a unified framework. The edge-aware masks facilitate three basic operations for adaptive facial editing, including region selection, edit setting and region blending. Inspired by the state-of-the-art edit propagation and partial differential equation (PDE) learning method, we propose an adaptive PDE model with facial priors for masks generation through edge-aware diffusion. The edge-aware masks can automatically fit the complex region boundary with great accuracy and produce smooth transition between different regions, which significantly improves the visual consistence of face editing and reduce the human intervention. Then, a unified and flexible facial editing framework is constructed, which consists of layer decomposition, edge-aware masks generation, and layer/mask composition. The combinations of multiple facial layers and edge-aware masks can achieve various facial effects simultaneously, including face enhancement, relighting, makeup and face blending etc. Qualitative and quantitative evaluations were performed using different datasets for different facial editing tasks. Experiments demonstrate the effectiveness and flexibility of our methods, and the comparisons with the previous methods indicate that improved results are obtained using the combination of multiple edge-aware masks. Lingyu Liang, Xin Zhang 0013, Yong Xu 0007 |
Comput. Graph. Forum | 1 |
| 2014 | Similar handwritten Chinese character recognition by kernel discriminative locality alignment
Dapeng Tao, Lingyu Liang, Yan Gao 0011 |
Pattern Recognit. Lett. | 2 |
| 2014 | Facial Skin Beautification Using Adaptive Region-Aware MasksabstractIn this paper, we propose a unified facial beautification framework with respect to skin homogeneity, lighting, and color. A novel region-aware mask is constructed for skin manipulation, which can automatically select the edited regions with great precision. Inspired by the state-of-the-art edit propagation techniques, we present an adaptive edge-preserving energy minimization model with a spatially variant parameter and a high-dimensional guided feature space for mask generation. Using region-aware masks, our method facilitates more flexible and accurate facial skin enhancement while the complex manipulations are simplified considerably. In our beautification framework, a portrait is decomposed into smoothness, lighting, and color layers by an edge-preserving operator. Next, facial landmarks and significant features are extracted as input constraints for mask generation. After three region-aware masks have been obtained, a user can perform facial beautification simply by adjusting the skin parameters. Furthermore, the combinations of parameters can be optimized automatically, depending on the data priors and psychological knowledge. We performed both qualitative and quantitative evaluation for our method using faces with different genders, races, ages, poses, and backgrounds from various databases. The experimental results demonstrate that our technique is superior to previous methods and comparable to commercial systems, for example, PicTreat, Portrait+ , and Portraiture. Lingyu Liang, Xuelong Li 0001 |
IEEE Trans. Cybern. | 1 |
| 2013 | Facial Skin Beautification Using Region-Aware MaskabstractIn this paper, we present a facial attractiveness enhancement system for three major skin attributes: homogeneity, lighting, and color. To implement specific facial manipulation, we propose a new region-aware mask generation approach based on the state-of-the-art edit propagation techniques. The region-aware masks allow the user to beautify faces without the tedious and time-consuming region selection. In our system, an input portrait is decomposed into smoothness, lighting and color layers by edge-preserving filter. Then, taking the facial landmarks and significant edges as the constraint, layer masks are generated based on the image-guided energy minimization framework. The system could beautify faces automatically using the attractiveness priors of example-sets and psychological knowledge. Experiments illustrate the effectiveness of our method on various face databases, and the comparison with the previous methods indicates the reliability of our region-aware mask. Lingyu Liang |
SMC | 1 |
| 2013 | Image-Based Rendering for Ink PaintingabstractInk painting is one of the traditional forms of expression in oriental art and it is fascinating to generate the distinctive ink painting effect using modern non-photo realistic rendering (NPR) technique. Since most stroke-based rendering (SBR) methods require much user interaction and professional knowledge about painting, it should be more convenient to employ image-based rendering (IBR) methods to directly produce the ink wash effect from photos. According to our knowledge, however, most current IBR methods are based on ad-hoc algorithm, and they fail to perform well in generical scenario. In this paper, we propose a new IBR framework for ink painting, which could effectively simulate ink diffusion in absorbent paper, and produce various types of black or color ink painting effects. First, significant edges of the original image are detected as the constraint regions. Second, the pixel values of the edges are propagated to the blank regions out of them using an edge-preserving energy minimization model in edit propagation technique. Third, absorbent paper appearance is simulated through texture synthesis and detail manipulation. Finally, we compose the ink diffusion results with the absorbent paper background to generate the whole ink painting. Experiments illustrate that various types of ink diffusion effects and absorbent paper appearance could be effectively produced by our method. Lingyu Liang |
SMC | 1 |
| 2011 | Similar Handwritten Chinese Character Recognition Using Discriminative Locality Alignment Manifold LearningabstractThe discriminant analysis for Similar Handwritten Chinese Character Recognition (SHCR) is essential for the improvement of handwritten Chinese character recognition performance. In this paper, a new manifold based subspace learning algorithm, Discriminative Locality Alignment (DLA), is introduced into SHCR. Experimental results demonstrate that DLA is consistently superior to LDA (Linear Discriminant Analysis) in terms of discriminate information extraction, dimension reduction and recognition accuracy. In addition, DLA reveals some attractive intrinsic properties for numeric calculation, e.g. it can overcome the matrix singular problem and small sample size problem in SHCR. Dapeng Tao, Lingyu Liang, Yan Gao 0011 |
ICDAR | 2 |