EDBT 2026 Demo / reviewers in the wild / expert
Fayao Liu
dblp:91/9687
· DBLP profile ↗
66ranked-venue papers
8as first author
49since 2021 · last 2026
0000-0001-6649-7660ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 50 · 5 first-author · 35 since 2021Graphics, computer vision, multimedia, augmented reality and games · 38 · 3 first-author · 30 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SinColor: Uncertainty-Guided Single-Step Diffusion for Image ColorizationabstractImage colorization is a fundamental yet challenging task in computer vision, aiming to recover plausible and spatially coherent colors from grayscale images. Recent advancements in diffusion models have enabled significant progress in this field, yet existing methods predominantly rely on multi-step diffusion processes. While effective for generating high-frequency details, these approaches are suboptimal for colorization, as color information is inherently low-frequency, spatially smooth, and globally consistent. This mismatch leads to two critical limitations: 1) color artifacts and inconsistency due to excessive noise in the color space, and 2) high computational cost that hinders practical application. In this work, we propose a novel single-step diffusion framework for efficient and high-quality image colorization. We introduce a color uncertainty estimation (CUE) module to identify reliable and uncertain regions in the image, allowing the model to prioritize local certainty while reasoning about confused regions. To focus the model on low-frequency color generation, we directly encode the grayscale image into a latent representation, remove structural components in the output, and reconstruct the final image via efficient decoding. Extensive experiments on ImageNet, COCO-Stuff, and Extended COCO-Stuff demonstrate that our approach achieves state-of-the-art performance while reducing inference time by 98% and trainable parameters by 97% compared to leading multi-step diffusion methods. Our contributions include a systematic analysis of diffusion-based colorization, a lightweight yet effective uncertainty-aware framework, and comprehensive validation of its efficiency and effectiveness. Yutong Gao 0001, Congyan Lang, Yidian Liu, Fayao Liu, Guoshun Nan, Yunchao Wei |
IEEE Trans. Image Process. | 6 |
| 2026 | Investigate Interactive Semantic Segmentation via an Uncertainty Mining ViewabstractWith the rapid development of intelligence media, traditional semantic segmentation has shown excellent potential in application scenarios like autonomous driving. However, due to limited performance, traditional segmentation models usually lead to poor user experiences in applications that require high segmentation precision. Therefore, interactive semantic segmentation (ISS) is gaining the attention is gaining attention due to its capability to generate high-precision semantic segmentation results through a few user-provided clicks for experience improvement, which thus has a promising development prospect in fine-grained application scenarios,e.g., virtual reality, smart medical, data annotation,etc.. For good interaction efficiency, most existing interactive methods make efforts to conduct suitable click simulation strategies and reasonable click encoding methods, aiming at the robust understanding of diverse user clicks and translating comprehensible user intent,i.e., assign the correct category to the clicked area, for the neural network. Though proved effective, their designs ignore the uncertainty hiding in the extracted interaction features, which reflects the interaction difficulty and the user clicking intents. This can lead to inappropriate click simulation and click encoding, limiting the interaction efficiency. Hence we focus on exploring a reasonable ISS scheme via an uncertainty mining view. Specifically, we propose an uncertainty-based class-balanced click sampling (UCCS) simulation strategy by considering both the uncertainty of the click simulation region and its semantic imbalance, to form a reasonable click distribution. Furthermore, we propose a semantic uncertainty residual encoding (SURE) method to better embed the user's intention into the localization maps, by mining semantic confusion between the click and misprediction classes. We prove the effectiveness of our design through extensive experiments and initially analyze the importance of uncertainty mining for the ISS. Our model can achieve state-of-the-art performance on three semantic segmentation benchmarks. Yutong Gao 0001, Congyan Lang, Fayao Liu, Xun Xu 0002, Yuanzhouhan Cao, Yunchao Wei |
IEEE Trans. Multim. | 3 |
| 2025 | CADCrafter: Generating Computer-Aided Design Models from Unconstrained ImagesabstractCreating CAD digital twins from the physical world is crucial for manufacturing, design, and simulation. However, current methods typically rely on costly 3D scanning with labor-intensive post-processing. To provide a user-friendly design process, we explore the problem of reverse engineering from unconstrained real-world CAD images that can be easily captured by users of all experiences. However, the scarcity of real-world CAD data poses challenges in directly training such models. To tackle these challenges, we propose CADCrafter, an image-to-parametric CAD model generation framework that trains solely on synthetic textureless CAD data while testing on real-world images. To bridge the significant representation disparity between images and parametric CAD models, we introduce a geometry encoder to accurately capture diverse geometric features. Moreover, the texture-invariant properties of the geometric features can also facilitate the generalization to real-world scenarios. Since compiling CAD parameter sequences into explicit CAD models is a non-differentiable process, the network training inherently lacks explicit geometric supervision. To impose geometric validity constraints, we employ direct preference optimization (DPO) to fine-tune our model with the automatic code checker feedback on CAD sequence quality. Furthermore, we collected a real-world dataset, comprised of multi-view images and corresponding CAD command sequence pairs, to evaluate our method. Experimental results demonstrate that our approach can robustly handle real unconstrained CAD images, and even generalize to unseen general objects. Jiacheng Wei, Tianrun Chen, Chi Zhang 0007, Shangzhan Zhang, Bingchen Yang, Chuan-Sheng Foo, Guosheng Lin, Qixing Huang, Fayao Liu |
CVPR | 11 |
| 2025 | MagicArticulate: Make Your 3D Models Articulation-ReadyabstractWith the explosive growth of 3D content creation, there is an increasing demand for automatically converting static 3D models into articulation-ready versions that support realistic animation. Traditional approaches rely heavily on manual annotation, which is both time-consuming and labor-intensive. Moreover, the lack of large-scale benchmarks has hindered the development of learning-based solutions. In this work, we present MagicArticulate, an effective framework that automatically transforms static 3D models into articulation-ready assets. Our key contributions are threefold. First, we introduce Articulation-XL, a large-scale benchmark containing over 33k 3D models with high-quality articulation annotations, carefully curated from Objaverse-XL. Second, we propose a novel skeleton generation method that formulates the task as a sequence modeling problem, leveraging an autoregressive transformer to naturally handle varying numbers of bones or joints within skeletons and their inherent dependencies across different 3D models. Third, we predict skinning weights using a functional diffusion process that incorporates volumetric geodesic distance priors between vertices and joints. Extensive experiments demonstrate that MagicArticulate significantly outperforms existing methods across diverse object categories, achieving high-quality articulation that enables realistic animation. Project page: https://chaoyuesong.github.io/MagicArticulate. Chaoyue Song, Xiu Li 0001, Fan Yang 0103, Zhongcong Xu, Jun Hao Liew, Fayao Liu, Jiashi Feng, Guosheng Lin |
CVPR | 9 |
| 2025 | FIND: Few-Shot Anomaly Inspection with Normal-Only Multi-Modal Data
Fayao Liu, Jingyi Liao, Sichao Tian, Chuan-Sheng Foo, Xulei Yang |
ICCV | 2 |
| 2025 | Leveraging Large-Scale Pretrained Vision Foundation Models for Label-Efficient 3D Point Cloud Segmentation
Fayao Liu, Rui Yao 0006, Guosheng Lin |
ICIG (2) | 2 |
| 2025 | Robust-PIFu: Robust Pixel-aligned Implicit Function for 3D Human Digitalization from a Single ImageabstractExisting methods for 3D clothed human digitalization perform well when the input image is captured in ideal conditions that assume the lack of any occlusion. However, in reality, images may often have occlusion problems such as incomplete observation of the human subject's full body, self-occlusion by the human subject, and non-frontal body pose. When given such input images, these existing methods fail to perform adequately. Thus, we propose Robust-PIFu, a pixel-aligned implicit model that capitalized on large-scale, pretrained latent diffusion models to address the challenge of digitalizing human subjects from non-ideal images that suffer from occlusions.
Robust-PIfu offers four new contributions. Firstly, we propose a 'disentangling' latent diffusion model. This diffusion model, pretrained on billions of images, takes in any input image and removes external occlusions, such as inter-person occlusions, from that image. Secondly, Robust-PIFu addresses internal occlusions like self-occlusion by introducing a `penetrating' latent diffusion model. This diffusion model outputs multi-layered normal maps that by-pass occlusions caused by the human subject's own limbs or other body parts (i.e. self-occlusion). Thirdly, in order to incorporate such multi-layered normal maps into a pixel-aligned implicit model, we introduce our Layered-Normals Pixel-aligned Implicit Model, which improves the structural accuracy of predicted clothed human meshes. Lastly, Robust-PIFu proposes an optional super-resolution mechanism for the multi-layered normal maps. This addresses scenarios where the input image is of low or inadequate resolution. Though not strictly related to occlusion, this is still an important subproblem. Our experiments show that Robust-PIFu outperforms current SOTA methods both qualitatively and quantitatively. Our code will be released to the public. Kennard Yanting Chan, Fayao Liu, Guosheng Lin, Chuan-Sheng Foo, Weisi Lin |
ICLR | 2 |
| 2025 | Text-to-Image Rectified Flow as Plug-and-Play PriorsabstractLarge-scale diffusion models have achieved remarkable performance in generative tasks. Beyond their initial training applications, these models have proven their ability to function as versatile plug-and-play priors. For instance, 2D diffusion models can serve as loss functions to optimize 3D implicit models. Rectified Flow, a novel class of generative models, has demonstrated superior performance across various domains. Compared to diffusion-based methods, rectified flow approaches surpass them in terms of generation quality and efficiency. In this work, we present theoretical and experimental evidence demonstrating that rectified flow based methods offer similar functionalities to diffusion models — they can also serve as effective priors. Besides the generative capabilities of diffusion priors, motivated by the unique time-symmetry properties of rectified flow models, a variant of our method can additionally perform image inversion. Experimentally, our rectified flow based priors outperform their diffusion counterparts — the SDS and VSD losses — in text-to-3D generation. Our method also displays competitive performance in image inversion and editing. Code is available at: https://github.com/yangxiaofeng/rectified_flow_prior. Xulei Yang, Fayao Liu, Guosheng Lin |
ICLR | 4 |
| 2025 | Exploring Active Learning for Label-Efficient Training of Semantic Neural Radiance FieldabstractNeural Radiance Field (NeRF) models are implicit neural scene representation methods that offer unprecedented capabilities in novel view synthesis. Semantically-aware NeRFs not only capture the shape and radiance of a scene, but also encode semantic information of the scene. The training of semantically-aware NeRFs typically requires pixel-level class labels, which can be prohibitively expensive to collect. In this work, we explore active learning as a potential solution to alleviate the annotation burden. We investigate various design choices for active learning of semantically-aware NeRF, including selection granularity and selection strategies. We further propose a novel active learning strategy that takes into account 3D geometric constraints in sample selection. Our experiments demonstrate that active learning can effectively reduce the annotation cost of training semantically-aware NeRF, achieving more than 2× reduction in annotation cost compared to random sampling. Yuzhe Zhu, Lile Cai, Kangkang Lu 0001, Fayao Liu, Xulei Yang |
ICME | 4 |
| 2025 | Puppeteer: Rig and Animate Your 3D ModelsabstractModern interactive applications increasingly demand dynamic 3D content, yet the transformation of static 3D models into animated assets constitutes a significant bottleneck in content creation pipelines. While recent advances in generative AI have revolutionized static 3D model creation, rigging and animation continue to depend heavily on expert intervention. We present \textbf{Puppeteer}, a comprehensive framework that addresses both automatic rigging and animation for diverse 3D objects.
Our system first predicts plausible skeletal structures via an auto-regressive transformer that introduces a joint-based tokenization strategy for compact representation and a hierarchical ordering methodology with stochastic perturbation that enhances bidirectional learning capabilities. It then infers skinning weights via an attention-based architecture incorporating topology-aware joint attention that explicitly encodes inter-joint relationships based on skeletal graph distances.
Finally, we complement these rigging advances with a differentiable optimization-based animation pipeline that generates stable, high-fidelity animations while being computationally more efficient than existing approaches.
Extensive evaluations across multiple benchmarks demonstrate that our method significantly outperforms state-of-the-art techniques in both skeletal prediction accuracy and skinning quality. The system robustly processes diverse 3D content, ranging from professionally designed game assets to AI-generated shapes, producing temporally coherent animations that eliminate the jittering issues common in existing methods. Chaoyue Song, Xiu Li 0001, Fan Yang 0103, Zhongcong Xu, Jiacheng Wei, Fayao Liu, Jiashi Feng, Guosheng Lin |
NeurIPS | 6 |
| 2025 | MoDA: Modeling Deformable 3D Objects from Casual Videos
Chaoyue Song, Jiacheng Wei, Chuan-Sheng Foo, Fayao Liu, Guosheng Lin |
Int. J. Comput. Vis. | 6 |
| 2025 | Weakly Supervised Segmentation on Outdoor 4D Point Clouds With Progressive 4D GroupingabstractRecently, some weakly supervised 3D point cloud segmentation methods have been proposed to develop effective models with minimum annotation efforts. Our previous work, W4DTS, proposes a challenging task that utilizes only 0.001% points in outdoor point cloud datasets to achieve an effective segmentation model. However, under an extremely limited annotation budget, the quality of pseudo labels generated by W4DTS is unsatisfactory, which limits the segmentation performance in such scenarios. To solve this issue, we propose a progressive 4D grouping approach to group the annotated and unannotated points across space and time, which can generate high-quality pseudo labels with very sparse annotated points. Moreover, to further improve our progressive 4D grouping approach, we design a cross-frame contrastive learning and a local consistency learning to improve the quality of our 4D grouping. Experimental results reveal that with only 0.001% annotations, our solution significantly outperforms the previous best approach on SemanticKITTI. We also evaluate our framework on the SemanticPOSS dataset and ScribbleKITTI dataset, and achieve performances close to our fully supervised backbone models. Hanyu Shi 0002, Fayao Liu, Yi Xu 0002, Guosheng Lin |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Mining Semantic Correlations Between Mispredictions and Corrections for Interactive Semantic SegmentationabstractInteractive semantic segmentation pursues high-quality segmentation results at the cost of a small number of user clicks. It is attracting more and more research attention for its convenience in labeling semantic pixel-level data. Existing interactive segmentation methods often pursue higher interaction efficiency by mining the latent information of user clicks or exploring efficient interaction manners. However, these works neglect to explicitly exploit the semantic correlations between user corrections and model mispredictions, thus suffering from two flaws. First, similar prediction errors frequently occur in actual use, causing users to repeatedly correct them. Second, the interaction difficulty of different semantic classes varies across images, but existing models use monotonic parameters for all images which lack semantic pertinence. Therefore, in this article, we explore the semantic correlations existing in corrections and mispredictions by proposing a simple yet effective online learning solution to the above problems, named correction-misprediction correlation mining (CM2). Specifically, we leverage the correction-misprediction similarities to design a confusion memory module (CMM) for automatic correction when similar prediction errors reappear. Furthermore, we measure the semantic interaction difficulty by counting the correction-misprediction pairs and design a challenge adaptive convolutional layer (CACL), which can adaptively switch different parameters according to interaction difficulties to better segment the challenging classes. Our method requires no extra training besides the online learning process and can effectively improve interaction efficiency. Our proposed CM2 achieves state-of-the-art results on three public semantic segmentation benchmarks. Yutong Gao 0001, Congyan Lang, Fayao Liu, Chuan-Sheng Foo, Yuanzhouhan Cao, Yunchao Wei |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Few-Shot Image Generation via Style Adaptation and Content PreservationabstractTraining a generative model with limited data (e.g., 10) is a very challenging task. Many works propose to fine-tune a pretrained GAN model. However, this can easily result in overfitting. In other words, they manage to adapt the style but fail to preserve the content, where style denotes the specific properties that define a domain while content denotes the domain-irrelevant information that represents diversity. Recent works try to maintain a predefined correspondence to preserve the content, however, the diversity is still not enough and it may affect style adaptation. In this work, we propose a paired image reconstruction approach for content preservation. We propose to introduce an image translation module to GAN transferring, where the module teaches the generator to separate style and content, and the generator provides training data to the translation module in return. Qualitative and quantitative experiments show that our method consistently surpasses the state-of-the-art methods in a few-shot setting. Xiaosheng He, Fan Yang 0103, Fayao Liu, Guosheng Lin |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Temporal and Heterogeneous Graph Neural Network for Remaining Useful Life PredictionabstractPredicting remaining useful life (RUL) plays a crucial role in the prognostics and health management of industrial systems that involve a variety of interrelated sensors. Given a constant stream of time-series sensory data from such systems, deep learning (DL) models have risen to prominence at identifying complex, nonlinear temporal dependencies in these data. In addition to the temporal dependencies of individual sensors, spatial dependencies emerge as important correlations among these sensors, which can be naturally modeled by a temporal graph that describes time-varying spatial relationships. However, the majority of existing studies have relied on capturing discrete snapshots of this temporal graph, a coarse-grained approach that leads to a loss of temporal information. Moreover, given the variety of heterogeneous sensors, it becomes vital that such inherent heterogeneity is leveraged for RUL prediction in temporal sensor graphs. To capture the nuances of the temporal and spatial relationships and heterogeneous characteristics in an interconnected graph of sensors, we introduce a novel model named temporal and heterogeneous graph neural networks (THGNNs). Specifically, THGNN aggregates historical data from neighboring nodes to accurately capture the temporal dynamics and spatial correlations within the stream of sensor data in a fine-grained manner. Moreover, the model leverages feature-wise linear modulation (FiLM) to address the diversity of sensor types, significantly improving the model's capacity to learn the heterogeneity in the data sources. Finally, we have validated the effectiveness of our approach through comprehensive experiments. Our empirical findings demonstrate significant advancements on the N-CMAPSS dataset, achieving improvements of up to 19.2% and 31.6% in terms of two different evaluation metrics over state-of-the-art methods. Zhihao Wen, Yuan Fang 0001, PengCheng Wei, Fayao Liu, Zhenghua Chen, Min Wu 0008 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | CCFL: Customized Client Federated Learning for Unsupervised Person Re-identificationabstractFederated learning-based person re-identification (Re-ID) aims to address the issue of data silos in surveillance systems caused by increasingly stringent regulations on sensitive data. However, due to differences in data collection locations, times, and scales, severe non-independent and identically distributed (non-IID) characteristics exist across different Re-ID datasets. Existing federated learning-based Re-ID methods often adopt a unified model structure, which prevents the model from adapting well to diverse data environments, thereby significantly degrading the overall Re-ID performance. To address the challenges of training neural networks on non-IID data across different datasets, we propose a customizable federated learning framework. First, customizable clients allow each organization to freely select suitable neural network training methods and model architectures based on local data scales and prior knowledge, thus improving training outcomes. Second, since traditional federated learning frameworks cannot achieve knowledge fusion through parameter exchange between models with different architectures, we introduce an independent model, referred to as the interaction model, specifically designed for knowledge exchange among clients. The interaction model learns parameters (knowledge) from local models on each client through distillation learning. Subsequently, the interaction model is uploaded to the server, where it undergoes parameter fusion (knowledge exchange) with interaction models from other clients. Finally, the interaction model, enriched with knowledge from other clients, guides local model training through knowledge distillation. It is worth noting that selecting a lightweight interaction model, while potentially impacting Re-ID performance, can significantly reduce communication costs between the server and clients. Yong Zhou 0003, Fayao Liu, Jiaqi Zhao 0001, Hancheng Zhu, Wen-Liang Du 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Fine Structure-Aware Sampling: A New Sampling Training Scheme for Pixel-Aligned Implicit Models in Single-View Human ReconstructionabstractPixel-aligned implicit models, such as PIFu, PIFuHD, and ICON, are used for single-view clothed human reconstruction. These models need to be trained using a sampling training scheme. Existing sampling training schemes either fail to capture thin surfaces (e.g. ears, fingers) or cause noisy artefacts in reconstructed meshes. To address these problems, we introduce Fine Structured-Aware Sampling (FSS), a new sampling training scheme to train pixel-aligned implicit models for single-view human reconstruction. FSS resolves the aforementioned problems by proactively adapting to the thickness and complexity of surfaces. In addition, unlike existing sampling training schemes, FSS shows how normals of sample points can be capitalized in the training process to improve results. Lastly, to further improve the training process, FSS proposes a mesh thickness loss signal for pixel-aligned implicit models. It becomes computationally feasible to introduce this loss once a slight reworking of the pixel-aligned implicit function framework is carried out. Our results show that our methods significantly outperform SOTA methods qualitatively and quantitatively. Our code is publicly available at https://github.com/kcyt/FSS. Kennard Yanting Chan, Fayao Liu, Guosheng Lin, Chuan-Sheng Foo, Weisi Lin |
AAAI | 2 |
| 2024 | Diverse and Stable 2D Diffusion Guided Text to 3D Generation with Noise RecalibrationabstractIn recent years, following the success of text guided image generation, text guided 3D generation has gained increasing attention among researchers. Dreamfusion is a notable approach that enhances generation quality by utilizing 2D text guided diffusion models and introducing SDS loss, a technique for distilling 2D diffusion model information to train 3D models. However, the SDS loss has two major limitations that hinder its effectiveness. Firstly, when given a text prompt, the SDS loss struggles to produce diverse content. Secondly, during training, SDS loss may cause the generated content to overfit and collapse, limiting the model's ability to learn intricate texture details. To overcome these challenges, we propose a novel approach called Noise Recalibration algorithm. By incorporating this technique, we can generate 3D content with significantly greater diversity and stunning details. Our approach offers a promising solution to the limitations of SDS loss. Fayao Liu, Yi Xu 0002, Hanjing Su, Qingyao Wu, Guosheng Lin |
AAAI | 2 |
| 2024 | Rethinking Few-shot 3D Point Cloud Semantic SegmentationabstractThis paper revisits few-shot 3D point cloud semantic segmentation (FS-PCS), with a focus on two significant is-sues in the state-of-the-art: foreground leakage and sparse point distribution. The former arises from non-uniform point sampling, allowing models to distinguish the density disparities between foreground and background for easier segmentation. The latter results from sampling only 2,048 points, limiting semantic information and deviating from the real-world practice. To address these issues, we in-troduce a standardized FS-PCS setting, upon which a new benchmark is built. Moreover, we propose a novel FS-PCS model. While previous methods are based on feature op-timization by mainly refining support features to enhance prototypes, our method is based on correlation optimization, referred to as Correlation Optimization Segmentation (COSeg). Specifically, we compute Class-specific Multi-prototypical Correlation (CMC) for each query point, rep-resenting its correlations to category prototypes. Then, we propose the Hyper Correlation Augmentation (HCA) mod-ule to enhance CMC. Furthermore, tackling the inherent property of few-shot training to incur base susceptibility for models, we propose to learn non-parametric prototypes for the base classes during training. The learned base proto-types are used to calibrate correlations for the background class through a Base Prototypes Calibration (BPC) module. Experiments on popular datasets demonstrate the superior-ity of COSeg over existing methods. The code is available at github.com/ZhaochongAnICOSeg. Zhaochong An, Guolei Sun, Yun Liu 0011, Fayao Liu, Zongwei Wu, Dan Wang 0011, Luc Van Gool, Serge J. Belongie |
CVPR | 4 |
| 2024 | R-Cyclic Diffuser: Reductive and Cyclic Latent Diffusion for 3D Clothed Human DigitalizationabstractRecently, the authors of Zero-1-to-3 demonstrated that a latent diffusion model, pretrained with Internet-scale data, can not only address the single-view 3D object reconstruction task but can even attain SOTA results in it. However, when applied to the task of single-view 3D clothed human reconstruction, Zero-1-to-3 (and related models) are unable to compete with the corresponding SOTA methods in this field despite being trained on clothed human data. In this work, we aim to tailor Zero-1-to-3's approach to the single-view 3D clothed human reconstruction task in a much more principled and structured manner. To this end, we propose R-Cyclic Diffuser, a framework that adapts Zero-1-to-3's novel approach to clothed human data by fusing it with a pixel-aligned implicit model. R-Cyclic Diffuser offers a total of three new contributions. The first and primary contribution is R-Cyclic Diffuser's cyclical conditioning mechanism for novel view synthesis. This mechanism directly addresses the view inconsistency problem faced by Zero-1-to-3 and related models. Secondly, we further enhance this mechanism with two key features - Lateral Inversion Constraint and Cyclic Noise Selection. Both features are designed to regularize and restrict the randomness of outputs generated by a latent diffusion model. Thirdly, we show how SMPL-X body priors can be incorporated in a latent diffusion model such that novel views of clothed human bodies can be generated much more accurately. Our experiments show that R-Cyclic Diffuser is able to outperform current SOTA methods in singleview 3D clothed human reconstruction both qualitatively and quantitatively. Our code is made publicly available at https://github.com/kcyt/r-cyclic-diffuser. Kennard Yanting Chan, Fayao Liu, Guosheng Lin, Chuan-Sheng Foo, Weisi Lin |
CVPR | 2 |
| 2024 | Sculpt3D: Multi-View Consistent Text-to-3D Generation with Sparse 3D PriorabstractRecent works on text-to-3d generation show that using only 2D diffusion supervision for 3D generation tends to produce results with inconsistent appearances (e.g., faces on the back view) and inaccurate shapes (e.g., animals with extra legs). Existing methods mainly address this issue by retraining diffusion models with images rendered from 3D data to ensure multi-view consistency while struggling to balance 2D generation quality with 3D consistency. In this paper, we present a new framework Sculpt3D that equips the current pipeline with explicit injection of 3D priors from retrieved reference objects without re-training the 2D diffusion model. Specifically, we demonstrate that high-quality and diverse 3D geometry can be guaranteed by keypoints supervision through a sparse ray sampling approach. Moreover, to ensure accurate appearances of different views, we further modulate the output of the 2D diffusion model to the correct patterns of the template views without altering the generated object's style. These two decoupled designs effectively harness 3D information from reference objects to generate 3D objects while preserving the generation quality of the 2D diffusion model. Extensive experiments show our method can largely improve the multi-view consistency while retaining fidelity and diversity. Our project page is available at: https://stellarcheng.github.io/Sculpt3D/. Fan Yang 0103, Chengzeng Feng, Zhoujie Fu, Chuan-Sheng Foo, Guosheng Lin, Fayao Liu |
CVPR | 8 |
| 2024 | REACTO: Reconstructing Articulated Objects from a Single VideoabstractIn this paper, we address the challenge of reconstructing general articulated 3D objects from a single video. Existing works employing dynamic neural radiance fields have advanced the modeling of articulated objects like humans and animals from videos, but face challenges with piece-wise rigid general articulated objects due to limitations in their deformation models. To tackle this, we propose Quasi-Rigid Blend Skinning, a novel deformation model that enhances the rigidity of each part while maintaining flexible deformation of the joints. Our primary insight combines three distinct approaches: 1) an enhanced bone rigging system for improved component modeling, 2) the use of quasi-sparse skinning weights to boost part rigidity and reconstruction fidelity, and 3) the application of geodesic point assignment for precise motion and seamless deformation. Our method outperforms previous works in producing higher-fidelity 3D reconstructions of general articulated objects, as demonstrated on both real and synthetic datasets. Project page: https://chaoyuesong.github.io/REACTO. Chaoyue Song, Jiacheng Wei, Chuan-Sheng Foo, Guosheng Lin, Fayao Liu |
CVPR | 5 |
| 2024 | 3DFG-PIFu: 3D Feature Grids for Human Digitization from Sparse Views
Kennard Yanting Chan, Fayao Liu, Guosheng Lin, Chuan-Sheng Foo, Weisi Lin |
ECCV (24) | 2 |
| 2024 | Learn to Optimize Denoising Scores: A Unified and Improved Diffusion Prior for 3D Generation
Chi Zhang 0007, Yi Xu 0002, Xulei Yang, Fayao Liu, Guosheng Lin |
ECCV (44) | 7 |
| 2024 | PromptAD: Zero-shot Anomaly Detection using Text PromptsabstractWe consider the problem of zero-shot anomaly detection in which a model is pre-trained to detect anomalies in images belonging to seen classes, and expected to detect anomalies from unseen classes at test time. State-of-the-art anomaly detection (AD) methods can often achieve exceptional results when training images are abundant, but they catastrophically fail in zero-shot scenarios with a lack of real examples. However, with the emergence of multi-modal models such as CLIP, it is possible to use knowledge from other modalities (e.g. text) to compensate for the lack of visual information and improve AD performance. In this work, we propose PromptAD, a dual-branch framework which uses prior knowledge about both normal and abnormal behaviours in the form of text prompts to detect anomalies even in unseen classes. More specifically, it uses CLIP as a backbone encoder network and an additional dual-branch vision-language decoding network for both normality and abnormality information. The normality branch establishes a profile of normality, while the abnormality branch models anomalous behaviors, guided by natural language text prompts. As the two branches capture complementary information or ‘views’, we propose a ‘cross-view contrastive learning’ (CCL) component which regularizes each view with additional reference information from the other view. We further propose a cross-view mutual interaction (CMI) strategy to promote the mutual exploration of useful knowledge from each branch. We show that PromptAD outperforms existing baselines in zero-shot anomaly detection on key benchmark datasets and analyse the role of each component in ablation studies. Adam Goodge, Fayao Liu, Chuan-Sheng Foo |
WACV | 3 |
| 2024 | Learning Temporal Variations for 4D Point Cloud Segmentation
Hanyu Shi 0002, Jiacheng Wei, Hao Wang 0094, Fayao Liu, Guosheng Lin |
Int. J. Comput. Vis. | 4 |
| 2024 | Neural Radiance Selector: Find the best 2D representations of 3D data for CLIP based 3D tasks
Fayao Liu, Guosheng Lin |
Knowl. Based Syst. | 2 |
| 2024 | Training neural networks with classification rules for incorporating domain knowledge
Wenyu Zhang 0003, Fayao Liu, Cuong Manh Nguyen, Zhong Liang Ou Yang, Savitha Ramasamy, Chuan-Sheng Foo |
Knowl. Based Syst. | 2 |
| 2024 | Multi-level self attention for unsupervised learning person re-identification
Jiaqi Zhao 0001, Yong Zhou 0003, Fayao Liu, Rui Yao 0006, Hancheng Zhu, Abdulmotaleb El Saddik |
Multim. Tools Appl. | 4 |
| 2024 | LCReg: Long-tailed image classification with Latent Categories based Recognition
Weide Liu, Henghui Ding, Fayao Liu, Jie Lin 0001, Guosheng Lin |
Pattern Recognit. | 5 |
| 2024 | Dense Supervision Propagation for Weakly Supervised Semantic Segmentation on 3D Point CloudsabstractSemantic segmentation on 3D point clouds is an important task for 3D scene understanding. While dense labeling on 3D data is expensive and time-consuming, only a few works address weakly supervised semantic point cloud segmentation methods to relieve the labeling cost by learning from simpler and cheaper labels. Meanwhile, there are still huge performance gaps between existing weakly supervised methods and state-of-the-art fully supervised methods. In this paper, we propose Dense Supervision Propagation (DSP) to train a semantic point cloud segmentation network with only a small portion of points being labeled. We argue that we can better utilize the limited supervision information as we densely propagate the supervision signal from the labeled points to other points within and across the input samples. Specifically, we propose a cross-sample feature reallocating module to transfer similar features and therefore re-route the gradients across two samples with common classes and an intra-sample feature redistribution module to propagate supervision signals on unlabeled points across and within point cloud samples. We conduct extensive experiments on public datasets S3DIS and ScanNet. Our weakly supervised method with only 10% and 1% of labels can produce competitive results with the fully supervised counterpart. Jiacheng Wei, Guosheng Lin, Kim-Hui Yap, Fayao Liu, Tzu-Yi Hung |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Dynamic Interaction Dilation for Interactive Human ParsingabstractInteractive segmentation pursues generating high-quality pixel-level predictions with a few user-provided clicks, which is gaining attention for its convenience in segmentation data annotation. Users are allowed to iteratively refine the prediction by adding clicks until the result is satisfactory. Existing interactive methods usually transform the clicks into a set of localization maps by Euclidian distance computation or RGB texture extraction to guide the segmentation, which makes the click transformation a core module in interactive segmentation networks. However, when adopted in human images where large poses, occlusions, and bad illuminations are prevailing, prior transformation methods tend to cause uncorrectable overlapping across localization maps which are difficult to form a good match among human parts. Furthermore, the inappropriately transformed information is hard to be refined with the static transformation manner which is out of tune with the dynamically refined interaction process. Hence, we design a dynamic transformation scheme for interactive human parsing (IHP) named Dynamic Interaction Dilation Net (DID-Net), which serves as an initial attempt to break the limitations of static transformation while capturing long-range dependencies of clicks within each human part. Specifically, we construct a Dynamic Dilation Module (DD-Module) to dilate clicks radially in several directions assisted by human body edge detection to refine the dilation quality in each interaction iteration. Furthermore, we propose an Adaptive Interaction Excitation Block (AIE-Block) to exploit potential semantic clues buried in the dilated clicks. Our DID-Net achieves state-of-the-art performance on 3 public human parsing benchmarks. Yutong Gao 0001, Congyan Lang, Fayao Liu, Yuanzhouhan Cao, Yunchao Wei |
IEEE Trans. Multim. | 3 |
| 2024 | Neural Logic Vision Language ExplainerabstractIf we compare how humans reason and how deep models reason, humans reason in a symbolic manner with a formal language called logic, while most deep models reason in black-box. A natural question to ask is “Do the trained deep models reason similar as humans?” or “Can we explain the reasoning of deep models in the language of logic?”. In this work, we presentNeurLogXto explain the reasoning process of deep vision language models in the language of logic. Given a trained vision language model, our method starts by generating reasoning facts through augmenting the input data. We then develop a differentiable inductive logic programming framework to learn interpretable logic rules from the facts. We show our results on various popular vision language models. Interestingly, we observe that almost all of the tested models can reason logically. Fayao Liu, Guosheng Lin |
IEEE Trans. Multim. | 2 |
| 2024 | On Representation Knowledge Distillation for Graph Neural NetworksabstractKnowledge distillation (KD) is a learning paradigm for boosting resource-efficient graph neural networks (GNNs) using more expressive yet cumbersome teacher models. Past work on distillation for GNNs proposed the local structure preserving (LSP) loss, which matches local structural relationships defined over edges across the student and teacher's node embeddings. This article studies whether preserving the global topology of how the teacher embeds graph data can be a more effective distillation objective for GNNs, as real-world graphs often contain latent interactions and noisy edges. We propose graph contrastive representation distillation (G-CRD), which uses contrastive learning to implicitly preserve global topology by aligning the student node embeddings to those of the teacher in a shared representation space. Additionally, we introduce an expanded set of benchmarks on large-scale real-world datasets where the performance gap between teacher and student GNNs is non-negligible. Experiments across four datasets and 14 heterogeneous GNN architectures show that G-CRD consistently boosts the performance and robustness of lightweight GNNs, outperforming LSP (and a global structure preserving (GSP) variant of LSP) as well as baselines from 2-D computer vision. An analysis of the representational similarity among teacher and student embedding spaces reveals that G-CRD balances preserving local and global relationships, while structure preserving approaches are best at preserving one or the other. Chaitanya K. Joshi, Fayao Liu, Xu Xun, Jie Lin 0001, Chuan-Sheng Foo |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Collaborative Propagation on Multiple Instance Graphs for 3D Instance Segmentation with Single-point SupervisionabstractInstance segmentation on 3D point clouds has been attracting increasing attention due to its wide applications, especially in scene understanding areas. However, most existing methods operate on fully annotated data while manually preparing ground-truth labels at point-level is very cumbersome and labor-intensive. To address this issue, we propose a novel weakly supervised method RWSeg that only requires labeling one object with one point. With these sparse weak labels, we introduce a unified framework with two branches to propagate semantic and instance information respectively to unknown regions using self-attention and a cross-graph random walk method. Specifically, we propose a Cross-graph Competing Random Walks (CRW) algorithm that encourages competition among different instance graphs to resolve ambiguities in closely placed objects, improving instance assignment accuracy. RWSeg generates high-quality instance-level pseudo labels. Experimental results on ScanNet-v2 and S3DIS datasets show that our approach achieves comparable performance with fully-supervised methods and outperforms previous weakly-supervised methods by a substantial margin. Ruibo Li, Jiacheng Wei, Fayao Liu, Guosheng Lin |
ICCV | 4 |
| 2023 | ELFNet: Evidential Local-global Fusion for Stereo MatchingabstractAlthough existing stereo matching models have achieved continuous improvement, they often face issues related to trustworthiness due to the absence of uncertainty estimation. Additionally, effectively leveraging multi-scale and multi-view knowledge of stereo pairs remains unexplored. In this paper, we introduce the Evidential Local-global Fusion (ELF) framework for stereo matching, which endows both uncertainty estimation and confidence-aware fusion with trustworthy heads. Instead of predicting the disparity map alone, our model estimates an evidential-based disparity considering both aleatoric and epistemic uncertainties. With the normal inverse-Gamma distribution as a bridge, the proposed framework realizes intra evidential fusion of multi-level predictions and inter evidential fusion between cost-volume-based and transformer-based stereo matching. Extensive experimental results show that the proposed framework exploits multi-view information effectively and achieves state-of-the-art overall performance both on accuracy and cross-domain generalization. The codes are available at https://github.com/jimmy19991222/ELFNet. Jieming Lou, Weide Liu, Fayao Liu, Jun Cheng 0003 |
ICCV | 4 |
| 2023 | Depth and Video Segmentation Based Visual Attention for Embodied Question AnsweringabstractEmbodied Question Answering (EQA) is a newly defined research area where an agent is required to answer the user's questions by exploring the real-world environment. It has attracted increasing research interests due to its broad applications in personal assistants and in-home robots. Most of the existing methods perform poorly in terms of answering and navigation accuracy due to the absence of fine-level semantic information, stability to the ambiguity, and 3D spatial information of the virtual environment. To tackle these problems, we propose a depth and segmentation based visual attention mechanism for Embodied Question Answering. First, we extract local semantic features by introducing a novel high-speed video segmentation framework. Then guided by the extracted semantic features, a depth and segmentation based visual attention mechanism is proposed for the Visual Question Answering (VQA) sub-task. Further, a feature fusion strategy is designed to guide the navigator's training process without much additional computational cost. The ablation experiments show that our method effectively boosts the performance of the VQA module and navigation module, leading to 4.9 % and 5.6 % overall improvement in EQA accuracy on House3D and Matterport3D datasets respectively. Haonan Luo 0002, Guosheng Lin, Yazhou Yao, Fayao Liu, Zichuan Liu, Zhenmin Tang |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Unsupervised 3D Pose Transfer With Cross Consistency and Dual ReconstructionabstractThe goal of 3D pose transfer is to transfer the pose from the source mesh to the target mesh while preserving the identity information (e.g., face, body shape) of the target mesh. Deep learning-based methods improved the efficiency and performance of 3D pose transfer. However, most of them are trained under the supervision of the ground truth, whose availability is limited in real-world scenarios. In this work, we present X-DualNet, a simple yet effective approach that enables unsupervised 3D pose transfer. In X-DualNet, we introduce a generator G which contains correspondence learning and pose transfer modules to achieve 3D pose transfer. We learn the shape correspondence by solving an optimal transport problem without any key point annotations and generate high-quality meshes with our elastic instance normalization (ElaIN) in the pose transfer module. With G as the basic component, we propose a cross consistency learning scheme and a dual reconstruction objective to learn the pose transfer without supervision. Besides that, we also adopt an as-rigid-as-possible deformer in the training process to fine-tune the body shape of the generated results. Extensive experiments on human and animal data demonstrate that our framework can successfully achieve comparable performance as the state-of-the-art supervised approaches. Chaoyue Song, Jiacheng Wei, Ruibo Li, Fayao Liu, Guosheng Lin |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Integrating topology beyond descriptions for zero-shot learning
Yutong Gao 0001, Congyan Lang, Yidong Li, Hongzhe Liu 0001, Fayao Liu |
Pattern Recognit. | 7 |
| 2023 | Temporal Feature Matching and Propagation for Semantic Segmentation on 3D Point Cloud SequencesabstractIn real-world LiDAR-based applications, data is generated in the form of 3D point cloud sequences or 4D point clouds. However, the topic of semantic segmentation on 4D point clouds is under-investigated and existing methods are still not able to achieve satisfactory performance to meet the requirement for real-world applications. The temporal information across different point clouds plays an important role in dynamic scene understanding, which is not well explored in existing work. In this paper, we focus on exploring effective temporal information across two consecutive point clouds for semantic segmentation on point cloud sequences. To this end, we design three novel modules to enhance the features of target frames by extracting different temporal information in the local regions and global regions. Experimental results on SemanticKITTI and SemanticPOSS demonstrate that our method achieves superior performance in 4D semantic segmentation by utilizing temporal information. Hanyu Shi 0002, Ruibo Li, Fayao Liu, Guosheng Lin |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Self-Training Vision Language BERTs With a Unified Conditional ModelabstractNatural language BERTs are trained with language corpus in a self-supervised manner. Unlike natural language BERTs, vision language BERTs need paired data to train, which restricts the scale of VL-BERT pretraining. We propose a self-training approach that allows training VL-BERTs from unlabeled image data. The proposed method starts with our unified conditional model– a vision language BERT model that can perform zero-shot conditional generation. Given different conditions, the unified conditional model can generate captions, dense captions, and even questions. We use the labeled image data to train a teacher model and use the trained model to generate pseudo captions on unlabeled image data. We then combine the labeled data and pseudo labeled data to train a student model. The process is iterated by putting the student model as a new teacher. By using the proposed self-training approach and only 300k unlabeled extra data, we are able to get competitive or even better performances compared to the models of similar model size trained with 3 million extra image data. Fengmao Lv, Fayao Liu, Guosheng Lin |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Effective End-to-End Vision Language Pretraining With Semantic Visual LossabstractCurrent vision language pretraining models are dominated by methods using region visual features extracted from object detectors. Given their good performance, the extract-then-process pipeline significantly restricts the inference speed and therefore limits their real-world use cases. However, training vision language models from raw image pixels is difficult, as the raw image pixels give much less prior knowledge than region features. In this paper, we systematically study how to leverage auxiliary visual pretraining tasks to help training end-to-end vision language models. We introduce three types of visual losses that enable much faster convergence and better finetuning accuracy. Compared with region feature models, our end-to-end models could achieve similar or better performance on down-stream tasks and run more than 10 times faster during inference. Compared with other end-to-end models, our proposed method could achieve similar or better performance when pretrained for only 10% of the pretraining GPU hours. Fayao Liu, Guosheng Lin |
IEEE Trans. Multim. | 2 |
| 2022 | Point Discriminative Learning for Data-efficient 3D Point Cloud Analysisabstract3D point cloud analysis has drawn a lot of research attention due to its wide applications. However, collecting massive labelled 3D point cloud data is both time-consuming and labor-intensive. This calls for data-efficient learning methods. In this work we propose PointDisc, a point discriminative learning method to leverage self-supervisions for data-efficient 3D point cloud classification and segmentation. PointDisc imposes a novel point discrimination loss on the middle and global level features produced by the backbone network. This point discrimination loss enforces learned features to be consistent with points belonging to the corresponding local shape region and inconsistent with randomly sampled noisy points. We conduct extensive experiments on 3D object classification, 3D semantic and part segmentation, showing the benefits of PointDisc for data-efficient learning. Detailed analysis demonstrate that PointDisc learns unsupervised features that well capture local and global geometry. Fayao Liu, Guosheng Lin, Chuan-Sheng Foo, Chaitanya K. Joshi, Jie Lin 0001 |
3DV | 1 |
| 2022 | Weakly Supervised Segmentation on Outdoor 4D point clouds with Temporal Matching and Spatial Graph PropagationabstractExisting point cloud segmentation methods require a large amount of annotated data, especially for the outdoor point cloud scene. Due to the complexity of the outdoor 3D scenes, manual annotations on the outdoor point cloud scene are time-consuming and expensive. In this paper, we study how to achieve scene understanding with limited annotated data. Treating 100 consecutive frames as a sequence, we divide the whole dataset into a series of sequences and annotate only 0.1% points in the first frame of each sequence to reduce the annotation requirements. This leads to a total annotation budget of 0.001%. We propose a novel temporal-spatial framework for effective weakly supervised learning to generate high-quality pseudo labels from these limited annotated data. Specifically, the frame-work contains two modules: an matching module in temporal dimension to propagate pseudo labels across different frames, and a graph propagation module in spatial dimension to propagate the information of pseudo labels to the entire point clouds in each frame. With only 0.001% annotations for training, experimental results on both SemanticKITTI and SemanticPOSS shows our weakly supervised two-stage framework is comparable to some existing fully supervised methods. We also evaluate our framework with 0.005% initial annotations on SemanticKITTI, and achieve a result close to fully supervised backbone model. Hanyu Shi 0002, Jiacheng Wei, Ruibo Li, Fayao Liu, Guosheng Lin |
CVPR | 4 |
| 2022 | CRCNet: Few-Shot Segmentation with Cross-Reference and Region-Global Conditional Networks
Weide Liu, Chi Zhang 0007, Guosheng Lin, Fayao Liu |
Int. J. Comput. Vis. | 4 |
| 2022 | Feature flow: In-network feature flow estimation for video object detection
Ruibing Jin, Guosheng Lin, Changyun Wen, Fayao Liu |
Pattern Recognit. | 5 |
| 2021 | On Automatic Data Augmentation for 3D Point Cloud Classification
Wanyue Zhang, Xun Xu 0002, Fayao Liu, Le Zhang 0001, Chuan-Sheng Foo |
BMVC | 3 |
| 2021 | HCRF-Flow: Scene Flow From Point Clouds With Continuous High-Order CRFs and Position-Aware Flow EmbeddingabstractScene flow in 3D point clouds plays an important role in understanding dynamic environments. Although significant advances have been made by deep neural networks, the performance is far from satisfactory as only per-point translational motion is considered, neglecting the constraints of the rigid motion in local regions. To address the issue, we propose to introduce the motion consistency to force the smoothness among neighboring points. In addition, constraints on the rigidity of the local transformation are also added by sharing unique rigid motion parameters for all points within each local region. To this end, a high-order CRFs based relation module (Con-HCRFs) is deployed to explore both point-wise smoothness and region-wise rigidity. To empower the CRFs to have a discriminative unary term, we also introduce a position-aware flow estimation module to be incorporated into the Con-HCRFs. Comprehensive experiments on FlyingThings3D and KITTI show that our proposed framework (HCRF-Flow) achieves state-of-the-art performance and significantly outperforms previous approaches substantially. Ruibo Li, Guosheng Lin, Tong He 0001, Fayao Liu, Chunhua Shen |
CVPR | 4 |
| 2021 | 3D Pose Transfer with Correspondence Learning and Mesh Refinementabstract3D pose transfer is one of the most challenging 3D generation tasks. It aims to transfer the pose of a source mesh to a target mesh and keep the identity (e.g., body shape) of the target mesh. Some previous works require key point annotations to build reliable correspondence between the source and target meshes, while other methods do not consider any shape correspondence between sources and targets, which leads to limited generation quality. In this work, we propose a correspondence-refinement network to achieve the 3D pose transfer for both human and animal meshes. The correspondence between source and target meshes is first established by solving an optimal transport problem. Then, we warp the source mesh according to the dense correspondence and obtain a coarse warped mesh. The warped mesh will be better refined with our proposed Elastic Instance Normalization, which is a conditional normalization layer and can help to generate high-quality meshes. Extensive experimental results show that the proposed architecture can effectively transfer the poses from source to target meshes and produce better results with satisfied visual performance than state-of-the-art methods. Chaoyue Song, Jiacheng Wei, Ruibo Li, Fayao Liu, Guosheng Lin |
NeurIPS | 4 |
| 2020 | CRNet: Cross-Reference Networks for Few-Shot SegmentationabstractOver the past few years, state-of-the-art image segmentation algorithms are based on deep convolutional neural networks. To render a deep network with the ability to understand a concept, humans need to collect a large amount of pixel-level annotated data to train the models, which is time-consuming and tedious. Recently, few-shot segmentation is proposed to solve this problem. Few-shot segmentation aims to learn a segmentation model that can be generalized to novel classes with only a few training images. In this paper, we propose a cross-reference network (CRNet) for few-shot segmentation. Unlike previous works which only predict the mask in the query image, our proposed model concurrently makes predictions for both the support image and the query image. With a cross-reference mechanism, our network can better find the co-occurrent objects in two images, thus helping the few-shot segmentation task. We also develop a mask refinement module to recurrently refine the prediction of the foreground regions. For the k-shot learning, we propose to finetune parts of networks to take advantage of multiple labeled support images. Experiments on the PASCAL VOC 2012 dataset show that our network achieves state-of-the-art performance. Weide Liu, Chi Zhang 0007, Guosheng Lin, Fayao Liu |
CVPR | 4 |
| 2020 | TRRNet: Tiered Relation Reasoning for Compositional Visual Question Answering
Guosheng Lin, Fengmao Lv, Fayao Liu |
ECCV (21) | 4 |
| 2020 | RefineNet: Multi-Path Refinement Networks for Dense PredictionabstractRecently, very deep convolutional neural networks (CNNs) have shown outstanding performance in object recognition and have also been the first choice for dense prediction problems such as semantic segmentation and depth estimation. However, repeated subsampling operations like pooling or convolution striding in deep CNNs lead to a significant decrease in the initial image resolution. Here, we present RefineNet, a generic multi-path refinement network that explicitly exploits all the information available along the down-sampling process to enable high-resolution prediction using long-range residual connections. In this way, the deeper layers that capture high-level semantic features can be directly refined using fine-grained features from earlier convolutions. The individual components of RefineNet employ residual connections following the identity mapping mindset, which allows for effective end-to-end training. Further, we introduce chained residual pooling, which captures rich background context in an efficient manner. We carry out comprehensive experiments on semantic segmentation which is a dense classification problem and achieve good performance on seven public datasets. We further apply our method for depth estimation and demonstrate the effectiveness of our method on dense regression problems. Guosheng Lin, Fayao Liu, Anton Milan, Chunhua Shen, Ian D. Reid 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | Towards Robust Curve Text Detection With Conditional Spatial ExpansionabstractIt is challenging to detect curve texts due to their irregular shapes and varying sizes. In this paper, we first investigate the deficiency of the existing curve detection methods and then propose a novel Conditional Spatial Expansion (CSE) mechanism to improve the performance of curve text detection. Instead of regarding the curve text detection as a polygon regression or a segmentation problem, we treat it as a region expansion process. Our CSE starts with a seed arbitrarily initialized within a text region and progressively merges neighborhood regions based on the extracted local features by a CNN and contextual information of merged regions. The CSE is highly parameterized and can be seamlessly integrated into existing object detection frameworks. Enhanced by the data-dependent CSE mechanism, our curve text detection system provides robust instance-level text region extraction with minimal post-processing. The analysis experiment shows that our CSE can handle texts with various shapes, sizes, and orientations, and can effectively suppress the false-positives coming from text-like textures or unexpected texts included in the same RoI. Compared with the existing curve text detection algorithms, our method is more robust and enjoys a simpler processing flow. It also creates a new state-of-art performance on curve text benchmarks with Fscore of up to 78.4%. Zichuan Liu, Guosheng Lin, Sheng Yang 0006, Fayao Liu, Weisi Lin, Wang Ling Goh |
CVPR | 4 |
| 2019 | CANet: Class-Agnostic Segmentation Networks With Iterative Refinement and Attentive Few-Shot LearningabstractRecent progress in semantic segmentation is driven by deep Convolutional Neural Networks and large-scale labeled image datasets. However, data labeling for pixel-wise segmentation is tedious and costly. Moreover, a trained model can only make predictions within a set of pre-defined classes. In this paper, we present CANet, a class-agnostic segmentation network that performs few-shot segmentation on new classes with only a few annotated images available. Our network consists of a two-branch dense comparison module which performs multi-level feature comparison between the support image and the query image, and an iterative optimization module which iteratively refines the predicted results. Furthermore, we introduce an attention mechanism to effectively fuse information from multiple support examples under the setting of k-shot learning. Experiments on PASCAL VOC 2012 show that our method achieves a mean Intersection-over-Union score of 55.4% for 1-shot segmentation and 57.1% for 5-shot segmentation, outperforming state-of-the-art methods by a large margin of 14.6% and 13.2%, respectively. Chi Zhang 0007, Guosheng Lin, Fayao Liu, Rui Yao 0006, Chunhua Shen |
CVPR | 3 |
| 2019 | SegEQA: Video Segmentation Based Visual Attention for Embodied Question AnsweringabstractEmbodied Question Answering (EQA) is a newly defined research area where an agent is required to answer the user's questions by exploring the real world environment. It has attracted increasing research interests due to its broad applications in automatic driving system, in-home robots, and personal assistants. Most of the existing methods perform poorly in terms of answering and navigation accuracy due to the absence of local details and vulnerability to the ambiguity caused by complicated vision conditions. To tackle these problems, we propose a segmentation based visual attention mechanism for Embodied Question Answering. Firstly, We extract the local semantic features by introducing a novel high-speed video segmentation framework. Then by the guide of extracted semantic features, a bottom-up visual attention mechanism is proposed for the Visual Question Answering (VQA) sub-task. Further, a feature fusion strategy is proposed to guide the training of the navigator without much additional computational cost. The ablation experiments show that our method boosts the performance of VQA module by 4.2% (68.99% vs 64.73%) and leads to 3.6% (48.59% vs 44.98%) overall improvement in EQA accuracy. Haonan Luo 0002, Guosheng Lin, Zichuan Liu, Fayao Liu, Zhenmin Tang, Yazhou Yao |
ICCV | 4 |
| 2019 | Pyramid Graph Networks With Connection Attentions for Region-Based One-Shot Semantic SegmentationabstractOne-shot image segmentation aims to undertake the segmentation task of a novel class with only one training image available. The difficulty lies in that image segmentation has structured data representations, which yields a many-to-many message passing problem. Previous methods often simplify it to a one-to-many problem by squeezing support data to a global descriptor. However, a mixed global representation drops the data structure and information of individual elements. In this paper, we propose to model structured segmentation data with graphs and apply attentive graph reasoning to propagate label information from support data to query data. The graph attention mechanism could establish the element-to-element correspondence across structured data by learning attention weights between connected graph nodes. To capture correspondence at different semantic levels, we further propose a pyramid-like structure that models different sizes of image regions as graph nodes and undertakes graph reasoning at different levels. Experiments on PASCAL VOC 2012 dataset demonstrate that our proposed network significantly outperforms the baseline method and leads to new state-of-the-art performance on 1-shot and 5-shot segmentation benchmarks. Chi Zhang 0007, Guosheng Lin, Fayao Liu, Jiushuang Guo, Qingyao Wu, Rui Yao 0006 |
ICCV | 3 |
| 2019 | Local fusion networks with chained residual pooling for video action recognitionabstractAction recognition is an important yet challenging problem. We here present a novel method, multistage local fusion networks with residual connections, to boost the performance of video action recognition . In realistic videos, an action instance may have a long time span and some frames may suffer from deteriorated object appearance due to motion blur or video defocus. Our method enhances the per-frame representation by capturing information from neighboring frames. We propose a local fusion block which considers neighboring frames to capture appearance and local motion information for generating per-frame representation. Our local fusion is performed in a multistage manner allowing feature fusion from varying neighborhood sizes in the temporal dimension. We employ residual connections in the fusion blocks to enable effective gradient propagation through the whole network allowing effective end-to-end training. We achieve competitive results on two challenging and public available datasets, namely HMDB51 and UCF101, which shows the effectiveness of the proposed method. Feixiang He, Fayao Liu, Rui Yao 0006, Guosheng Lin |
Image Vis. Comput. | 2 |
| 2018 | Structured Learning of Tree Potentials in CRF for Image SegmentationabstractWe propose a new approach to image segmentation, which exploits the advantages of both conditional random fields (CRFs) and decision trees. In the literature, the potential functions of CRFs are mostly defined as a linear combination of some predefined parametric models, and then, methods, such as structured support vector machines, are applied to learn those linear coefficients. We instead formulate the unary and pairwise potentials as nonparametric forests-ensembles of decision trees, and learn the ensemble parameters and the trees in a unified optimization problem within the large-margin framework. In this fashion, we easily achieve nonlinear learning of potential functions on both unary and pairwise terms in CRFs. Moreover, we learn classwise decision trees for each object that appears in the image. Experimental results on several public segmentation data sets demonstrate the power of the learned nonlinear nonparametric potentials. Fayao Liu, Guosheng Lin, Ruizhi Qiao, Chunhua Shen |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2017 | Structured Learning of Binary Codes with Column Generation for Optimizing Ranking Measures
Guosheng Lin, Fayao Liu, Chunhua Shen, Jianxin Wu 0001, Heng Tao Shen |
Int. J. Comput. Vis. | 2 |
| 2017 | Discriminative Training of Deep Fully Connected Continuous CRFs With Task-Specific LossabstractRecent works on deep conditional random fields (CRFs) have set new records on many vision tasks involving structured predictions. Here, we propose a fully connected deep continuous CRF model with task-specific losses for both discrete and continuous labeling problems. We exemplify the usefulness of the proposed model on multi-class semantic labeling (discrete) and the robust depth estimation (continuous) problems. In our framework, we model both the unary and the pairwise potential functions as deep convolutional neural networks (CNNs), which are jointly learned in an end-to-end fashion. The proposed method possesses the main advantage of continuously valued CRFs, which is a closed-form solution for the maximum a posteriori (MAP) inference. To better take into account the quality of the predicted estimates during the cause of learning, instead of using the commonly employed maximum likelihood CRF parameter learning protocol, we propose task-specific loss functions for learning the CRF parameters. It enables direct optimization of the quality of the MAP estimates during the learning process. Specifically, we optimize the multi-class classification loss for the semantic labeling task and the Tukey's biweight loss for the robust depth estimation problem. Experimental results on the semantic labeling and robust depth estimation tasks demonstrate that the proposed method compare favorably against both baseline and state-of-the-art methods. In particular, we show that although the proposed deep CRF model is continuously valued, with the equipment of task-specific loss, it achieves impressive results even on discrete labeling tasks. Fayao Liu, Guosheng Lin, Chunhua Shen |
IEEE Trans. Image Process. | 1 |
| 2016 | Online unsupervised feature learning for visual tracking
Fayao Liu, Chunhua Shen, Ian D. Reid 0001, Anton van den Hengel |
Image Vis. Comput. | 1 |
| 2016 | Learning Depth from Single Monocular Images Using Deep Convolutional Neural FieldsabstractIn this article, we tackle the problem of depth estimation from single monocular images. Compared with depth estimation using multiple images such as stereo depth perception, depth from monocular images is much more challenging. Prior work typically focuses on exploiting geometric priors or additional sources of information, most using hand-crafted features. Recently, there is mounting evidence that features from deep convolutional neural networks (CNN) set new records for various vision applications. On the other hand, considering the continuous characteristic of the depth values, depth estimation can be naturally formulated as a continuous conditional random field (CRF) learning problem. Therefore, here we present a deep convolutional neural field model for estimating depths from single monocular images, aiming to jointly explore the capacity of deep CNN and continuous CRF. In particular, we propose a deep structured learning scheme which learns the unary and pairwise potentials of continuous CRF in a unified deep CNN framework. We then further propose an equally effective model based on fully convolutional networks and a novel superpixel pooling method, which is about 10 times faster, to speedup the patch-wise convolutions in the deep model. With this more efficient model, we are able to design deeper networks to pursue better performance. Our proposed method can be used for depth estimation of general scenes with no geometric priors nor any extra information injected. In our case, the integral of the partition function can be calculated in a closed form such that we can exactly solve the log-likelihood maximization. Moreover, solving the inference problem for predicting depths of a test image is highly efficient as closed-form solutions exist. Experiments on both indoor and outdoor scene datasets demonstrate that the proposed method outperforms state-of-the-art depth estimation approaches. Fayao Liu, Chunhua Shen, Guosheng Lin, Ian D. Reid 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2015 | Deep convolutional neural fields for depth estimation from a single imageabstractWe consider the problem of depth estimation from a single monocular image in this work. It is a challenging task as no reliable depth cues are available, e.g., stereo correspondences, motions etc. Previous efforts have been focusing on exploiting geometric priors or additional sources of information, with all using hand-crafted features. Recently, there is mounting evidence that features from deep convolutional neural networks (CNN) are setting new records for various vision applications. On the other hand, considering the continuous characteristic of the depth values, depth estimations can be naturally formulated into a continuous conditional random field (CRF) learning problem. Therefore, we in this paper present a deep convolutional neural field model for estimating depths from a single image, aiming to jointly explore the capacity of deep CNN and continuous CRF. Specifically, we propose a deep structured learning scheme which learns the unary and pairwise potentials of continuous CRF in a unified deep CNN framework. The proposed method can be used for depth estimations of general scenes with no geometric priors nor any extra information injected. In our case, the integral of the partition function can be analytically calculated, thus we can exactly solve the log-likelihood optimization. Moreover, solving the MAP problem for predicting depths of a new image is highly efficient as closed-form solutions exist. We experimentally demonstrate that the proposed method outperforms state-of-the-art depth estimation methods on both indoor and outdoor scene datasets. Fayao Liu, Chunhua Shen, Guosheng Lin |
CVPR | 1 |
| 2015 | CRF learning with CNN features for image segmentation
Fayao Liu, Guosheng Lin, Chunhua Shen |
Pattern Recognit. | 1 |
| 2014 | Multiple Kernel Learning in the Primal for Multimodal Alzheimer's Disease ClassificationabstractTo achieve effective and efficient detection of Alzheimer's disease (AD), many machine learning methods have been introduced into this realm. However, the general case of limited training samples, as well as different feature representations typically makes this problem challenging. In this paper, we propose a novel multiple kernel-learning framework to combine multimodal features for AD classification, which is scalable and easy to implement. Contrary to the usual way of solving the problem in the dual, we look at the optimization from a new perspective. By conducting Fourier transform on the Gaussian kernel, we explicitly compute the mapping function, which leads to a more straightforward solution of the problem in the primal. Furthermore, we impose the mixed L21 norm constraint on the kernel weights, known as the group lasso regularization, to enforce group sparsity among different feature modalities. This actually acts as a role of feature modality selection, while at the same time exploiting complementary information among different kernels. Therefore, it is able to extract the most discriminative features for classification. Experiments on the ADNI dataset demonstrate the effectiveness of the proposed method. Fayao Liu, Luping Zhou, Chunhua Shen, Jianping Yin |
IEEE J. Biomed. Health Informatics | 1 |
| 2014 | Efficient Dual Approach to Distance Metric LearningabstractDistance metric learning is of fundamental interest in machine learning because the employed distance metric can significantly affect the performance of many learning methods. Quadratic Mahalanobis metric learning is a popular approach to the problem, but typically requires solving a semidefinite programming (SDP) problem, which is computationally expensive. The worst case complexity of solving an SDP problem involving a matrix variable of size D×D with O(D) linear constraints is about O(D(6.5)) using interior-point methods, where D is the dimension of the input data. Thus, the interior-point methods only practically solve problems exhibiting less than a few thousand variables. Because the number of variables is D(D+1)/2, this implies a limit upon the size of problem that can practically be solved around a few hundred dimensions. The complexity of the popular quadratic Mahalanobis metric learning approach thus limits the size of problem to which metric learning can be applied. Here, we propose a significantly more efficient and scalable approach to the metric learning problem based on the Lagrange dual formulation of the problem. The proposed formulation is much simpler to implement, and therefore allows much larger Mahalanobis metric learning problems to be solved. The time complexity of the proposed method is roughly O(D(3)), which is significantly lower than that of the SDP approach. Experiments on a variety of data sets demonstrate that the proposed method achieves an accuracy comparable with the state of the art, but is applicable to significantly larger problems. We also show that the proposed method can be applied to solve more general Frobenius norm regularized SDP problems approximately. Chunhua Shen, Junae Kim, Fayao Liu, Lei Wang 0001, Anton van den Hengel |
IEEE Trans. Neural Networks Learn. Syst. | 3 |