EDBT 2026 Demo / reviewers in the wild / expert
Weide Liu
dblp:261/9166
· DBLP profile ↗
48ranked-venue papers
14as first author
45since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 28 · 5 first-author · 25 since 2021Artificial intelligence and machine learning · 22 · 10 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Temporal Inconsistency Guidance for Super-resolution Video Quality AssessmentabstractAs super-resolution (SR) techniques introduce unique distortions that fundamentally differ from those caused by traditional degradation processes (e.g., compression), there is an increasing demand for specialized video quality assessment (VQA) methods tailored to SR-generated content. One critical factor affecting perceived quality is temporal inconsistency, which refers to irregularities between consecutive frames. However, existing VQA approaches rarely quantify this phenomenon or explicitly investigate its relationship with human perception. Moreover, SR videos exhibit amplified inconsistency levels as a result of enhancement processes. In this paper, we propose Temporal Inconsistency Guidance for Super-resolution Video Quality Assessment (TIG-SVQA) that underscores the critical role of temporal inconsistency in guiding the quality assessment of SR videos. We first design a perception-oriented approach to quantify frame-wise temporal inconsistency. Based on this, we introduce the Inconsistency Highlighted Spatial Module, which localizes inconsistent regions at both coarse and fine scales. Inspired by the human visual system, we further develop an Inconsistency Guided Temporal Module that performs progressive temporal feature aggregation: (1) a consistency-aware fusion stage in which a visual memory capacity block adaptively determines the information load of each temporal segment based on inconsistency levels, and (2) an informative filtering stage for emphasizing quality-related features. Extensive experiments on both single-frame and multi-frame SR video scenarios demonstrate that our method significantly outperforms state-of-the-art VQA approaches. Xiaoyuan Yang 0003, Weide Liu, Xin Jin 0014, Xu Jia 0012, Yukun Lai, Paul L. Rosin, Hantao Liu, Wei Zhou 0021 |
AAAI | 3 |
| 2026 | Evidential Robust Feature Learning for Generalized Few-Shot Segmentation
Weide Liu, Xiaoyang Zhong, Lu Wang 0001, Chunbo Lang, Yuming Fang 0001, Jun Cheng 0003, Xulei Yang, Gong Cheng 0003 |
Int. J. Comput. Vis. | 1 |
| 2026 | SRNet: Self-supervised structure regularization for stereo matching
Jun Cheng 0003, Zaiwang Gu, Weide Liu, Jiayuan Fan 0001, Zhengguo Li, Chuan-Sheng Foo |
Neurocomputing | 3 |
| 2026 | EHIN: Early-aware hierarchical interaction network for weakly-supervised referring image segmentation
Anqing Chen, Wanli Ma 0001, Weide Liu, Yakun Ju, Paul L. Rosin, Hantao Liu, Wei Zhou 0021 |
Neurocomputing | 6 |
| 2026 | Novel adaptive dissipative synchronization algorithm for memristive fractional-order complex-valued neural networks with mixed time delays
Weide Liu, Haofeng Song, Juxia Xiong |
Neurocomputing | 1 |
| 2026 | Cross-modal attention fusion of RGB and skeleton for multimodal-driven video anomaly detection
Boan Chen, Weide Liu, Jinmei Liu, Baoquan Zhao, Yang Liu 0246 |
Pattern Recognit. | 3 |
| 2026 | Condition-Dependent Causal Discovery: A Polynomial Chaos Framework for Systems With Parametric UncertaintyabstractIdentifying causal structures within industrial processes is vital for effective optimization and control, yet it faces significant challenges due to fluctuating operational conditions. Conventional causal discovery techniques are constrained by their reliance on static relationship assumptions, rendering them inadequate for real-world scenarios where causal dynamics shift with parameters such as temperature and pressure. Furthermore, current uncertainty-aware methods typically focus solely on epistemic uncertainty arising from limited data, neglecting the functional dependency of causal strengths on measurable system parameters. To address this, we propose physics-informed polynomial chaos theory for causal discovery (PIPCT-CD), a novel framework that explicitly models causal edges as polynomial functions of operating parameters. By introducing a novel polynomial chaos expansion-based conditional independence test with a robust score-based learning strategy, PIPCT-CD accurately detects these dynamic interactions and quantifies associated uncertainties. Comprehensive validation on industrial-scale process networks demonstrates that PIPCT-CD achieves an F1-score of 0.800 on an electrical distribution system and 0.632 on a chemical refinery process, outperforming established baseline methods. Weide Liu, Yang Liu 0246 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2026 | Causal Discovery in Dynamic Industrial Systems Under Parametric Uncertainty: A Polynomial Chaos ApproachabstractIndustrial processes are governed by causal relationships whose understanding is essential for root cause analysis and predictive maintenance. However, existing causal discovery methods operate in batch mode and assume static causal structures, conflicting with the inherent nonstationarity caused by equipment degradation and regime changes. Current approaches cannot track how causal effects vary with operating parameters in real-time, nor detect structural change points. This article introduces Online-PCT-CD, a framework that extends parametric causal discovery to streaming data settings. The proposed method integrates incremental Polynomial Chaos Expansion estimation via recursive least squares, dual-layer cumulative sum change point detection, and an adaptive sliding window strategy. Experiments on an industrial chemical process benchmark demonstrate that Online-PCT-CD achieves 93.7% average F1-score, outperforming 17 baseline methods while maintaining 450 samples per second throughput suitable for real-time deployment. Weide Liu |
IEEE Trans. Ind. Informatics | 2 |
| 2026 | Integrating SAM Supervision for 3D Weakly Supervised Point Cloud SegmentationabstractCurrent methods for 3D semantic segmentation propose training models with limited annotations to address the difficulty of annotating large, irregular, and unordered 3D point cloud data. They usually focus on the 3D domain only, without leveraging the complementary nature of 2D and 3D data. Besides, some methods extend original labels or generate pseudo labels to guide the training, but they often fail to fully use these labels or address the noise within them. Meanwhile, the emergence of comprehensive and adaptable foundation models has offered effective solutions for segmenting 2D data. Leveraging this advancement, we present a novel approach that maximizes the utility of sparsely available 3D annotations by incorporating segmentation masks generated by 2D foundation models. We further propagate the 2D segmentation masks into the 3D space by establishing geometric correspondences between 3D scenes and 2D views. We extend the highly sparse annotations to encompass the areas delineated by 3D masks, thereby substantially augmenting the pool of available labels. Furthermore, we apply confidence- and uncertainty-based consistency regularization on augmentations of the 3D point cloud and select the reliable pseudo labels, which are further spread on the 3D masks to generate more labels. This innovative strategy bridges the gap between limited 3D annotations and the powerful capabilities of 2D foundation models, ultimately improving the performance of 3D weakly supervised segmentation. Lechun You, Weide Liu, Xulei Yang, Jun Cheng 0003, Wei Zhou 0021, Bharadwaj Veeravalli, Guosheng Lin |
IEEE Trans. Image Process. | 3 |
| 2026 | Rethinking the Effect of Unimodal Labels in Multimodal Sentiment AnalysisabstractMultimodal sentiment analysis aims to comprehensively understand human sentiment by integrating diverse modalities, such as text, audio, and vision. To improve modality complementarity, the recent Multimodal Multi-task Learning (MML) framework employs joint training of unimodal and multimodal sentiment analysis tasks using sub-annotations of modality. In this work, we further draw attention to the observation that integrating unimodal tasks may introduce conflicting task information, negatively affecting the multimodal task performance. Motivated by this issue, we propose the Multimodal Task Correlation-aware Learning (MTCL) framework to leverage beneficial task correlations and suppress harmful ones. Specifically, MTCL introduces a Correlation-Adaptive Training (CAT) strategy to learn a task-relation aware unimodal encoder for each modality. First, in order to distinguish whether a sample contains conflicting information, CAT strategy incorporates a Dual-Branch Contrast (DBC) module which divides the training set into a beneficial subset and a harmful subset. Based on this division, CAT strategy proposes an adaptive training loss to guide the model in understanding nuanced multitask correlations. The adaptive training loss has two components: (1) For the beneficial subset, a contrastive loss is utilized to improve the model’s ability to extract complementary representations. (2) For the harmful subset, we apply a task-correction loss to mitigate the negative interference caused by harmful task associations. With the CAT strategy, our framework can effectively distinguish beneficial and harmful task correlations to extract distinctive and robust unimodal representations. The superiority of MTCL is verified via extensive experiments on several multimodal video sentiment analysis benchmarks. Our work is publicly available at https://github.com/tiggers23/MTCL . Tianrui Li 0001, Baiyu Lu, Junlin Fang, Desheng Zheng, Wei Zhou 0021, Weide Liu, Fengmao Lv |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2025 | Rectification-specific Supervision and Constrained Estimator for Online Stereo RectificationabstractOnline stereo rectification is critical for autonomous vehicles and robots in dynamic environments, where factors such as vibration, temperature fluctuations, and mechanical stress can affect rectification accuracy and severely degrade downstream stereo depth estimation. Current dominant approaches for online stereo rectification involve estimating relative camera poses in real time to derive rectification homographies. However, they do not directly optimize for rectification constraints. Additionally, the general-purpose correspondence matchers used in these methods are not trained for rectification, while training of these matchers typically requires ground-truth correspondences which are not available in stereo rectification datasets. To address these limitations, we propose a matching-based stereo rectification framework that is directly optimized for rectification and does not require ground-truth correspondence annotations for training. We assume intrinsics are known as they are generally available on modern devices and are relatively stable. Our framework incorporates a rectification-constrained estimator and applies multi-level, rectification-specific supervision that trains the matcher network for rectification without relying on ground-truth correspondences. Additionally, we create a new rectification dataset with ground-truth optical flow annotations, eliminating bias from evaluation metrics used in prior work that relied on pretrained keypoint matching or optical flow models. Extensive experiments show that our approach outperforms both state-of-the-art matching-based and matching-free methods in vertical flow metric by 10.7% on the Carla-Flowguided dataset and 21.3% on the Semi-Truck Highway dataset, offering superior rectification accuracy. Kim-Hui Yap, Weide Liu, Xulei Yang, Jun Cheng 0003 |
CVPR | 3 |
| 2025 | Attribute-formed Class-specific Concept Space: Endowing Language Bottleneck Model with Better Interpretability and ScalabilityabstractLanguage Bottleneck Models (LBMs) are proposed to achieve interpretable image recognition by classifying images based on textual concept bottlenecks. However, current LBMs simply list all concepts together as the bottleneck layer, leading to the spurious cue inference problem and cannot generalized to unseen classes. To address these limitations, we propose the Attribute-formed Language Bottleneck Model (ALBM). ALBM organizes concepts in the attribute-formed class-specific space, where concepts are descriptions of specific attributes for specific classes. In this way, ALBM can avoid the spurious cue inference problem by classifying solely based on the essential concepts of each class. In addition, the cross-class unified attribute set also ensures that the concept spaces of different classes have strong correlations, as a result, the learned concept classifier can be easily generalized to unseen classes. Moreover, to further improve interpretability, we propose Visual Attribute Prompt Learning (VAPL) to extract visual features on fine-grained attributes. Furthermore, to avoid labor-intensive concept annotation, we propose the Description, Summary, and Supplement (DSS) strategy to automatically generate high-quality concept sets with a complete and precise attribute. Extensive experiments on 9 widely used few-shot benchmarks demonstrate the interpretability, transferability, and performance of our approach. The code and collected concept sets are available at https://github.com/tiggers23/ALBM. Jianyang Zhang, Qianli Luo, Guowu Yang, Wenjing Yang 0003, Weide Liu, Guosheng Lin, Fengmao Lv |
CVPR | 5 |
| 2025 | Frequency-Aware Native Resolution Assessment of 8K Omnidirectional ImagesabstractOmnidirectional images (ODIs) serve as fundamental visual medium for presenting virtual reality (VR) contents, supporting fully immersive experiences through 360-degree scene representation. Typically, a high pixel density is essential for visual quality in VR environments, which in turn requires sufficiently high-resolution imagery to achieve. However, capturing native high-resolution ODIs requires expensive omnidirectional cameras with large sensors (e.g., Insta360 TITAN). An alternative approach is to use low-resolution cameras to acquire original images and then enhance their resolution via super-resolution algorithms. In this work, we explore whether super-resolution ODIs can be easily distinguished from native high-resolution ODIs at 8K scale. To this end, we firstly construct the Native Resolution Assessment of 8K Omnidirectional Images (NRA- 8KODI) dataset, whose native 8K ODIs are collected with an Insta360 TITAN camera and 8K super-resolution images are generated from SOTA open-sourced algorithms. Recognizing high-frequency signals are essential for differentiating non-native 8K ODIs, a frequency-aware model is designed to capture high-frequency details. Specially, to maintain high-frequency details kept in high-resolutions while reduce computational costs brought by high-resolutions, we propose a frequency-aware compressor module to suppress feature channels dominated by low-frequency details. Finally, our model achieves 97.2% accuracy in detecting non-native 8K ODIs, implying that super-resolution for ODIs can still be improved for visual experience in VR applications. Jingwen Hou, Zengliang Li, Jiebin Yan, Weide Liu, Yuming Fang 0001, Wei Zhou 0021 |
VCIP | 4 |
| 2025 | Uncertainty Aware Interest Point Detection and DescriptionabstractInterest point detection and description play an important role in many visual tasks, including image registration, pose estimation, 3D reconstruction, and more. State-of-the-art interest point detection techniques are based on deep neural networks (NNs), which are prone to produce overconfident predictions. However, calibrated and ro-bust uncertainty measurement is crucial when deploying deep NN models in safety critical applications. In this work, we propose a novel Uncertainty-Aware interest Point (UAPoint) detection method to address this problem. Our method leverages evidential learning to learn both aleatoric and epistemic uncertainty. We further propose a constrained sampling scheme to construct more efficient training pairs for the descriptor decoder. We evaluate our method on a wide range of benchmarks and show that our method achieves state-of-the-art performance. Code will be released in https://github.com/JingboZeng/UAPoint. Jingbo Zeng, Zaiwang Gu, Weide Liu, Lile Cai, Jun Cheng 0003 |
WACV | 3 |
| 2025 | Improving multi-modal brain tumor segmentation via pre-training and knowledge distillation based post-training
Weide Liu, Jingwen Hou, Xiaoyang Zhong, Huijing Zhan, Jun Cheng 0003, Yuming Fang 0001, Guanghui Yue 0001 |
Neurocomputing | 1 |
| 2025 | Physically-guided open vocabulary segmentation with weighted patched alignment loss
Weide Liu, Jieming Lou, Wei Zhou 0021, Jun Cheng 0003, Xulei Yang |
Neurocomputing | 1 |
| 2025 | Integrating large foundation models into multimodal named entity recognition with evidential fusion
Weide Liu, Xiaoyang Zhong, Jingwen Hou, Haozhe Huang, Wei Zhou 0021, Yuming Fang 0001 |
Neurocomputing | 1 |
| 2025 | Multimodal multitask similarity learning for vision language model on radiological images and reports
Yang Yu 0079, Weide Liu, Ivan Ho Mien, Pavitra Krishnaswamy, Xulei Yang, Jun Cheng 0003 |
Neurocomputing | 3 |
| 2025 | Multitask Auxiliary Network for Perceptual Quality Assessment of Non-Uniformly Distorted Omnidirectional ImagesabstractOmnidirectional image quality assessment (OIQA) has been widely investigated in the past few years and achieved much success. However, most of existing studies are dedicated to solve the uniform distortion problem in OIQA, which has a natural gap with the non-uniform distortion problem, and their ability in capturing non-uniform distortion is far from satisfactory. To narrow this gap, in this paper, we propose a multitask auxiliary network for non-uniformly distorted omnidirectional images, where the parameters are optimized by jointly training the main task and other auxiliary tasks. The proposed network mainly consists of three parts: a backbone for extracting multiscale features from the viewport sequence, a multitask feature selection module for dynamically allocating specific features to different tasks, and auxiliary sub-networks for guiding the proposed model to capture local distortion and global quality change. Extensive experiments conducted on two large-scale OIQA databases demonstrate that the proposed model outperforms other state-of-the-art OIQA metrics, and these auxiliary sub-networks contribute to improve the performance of the proposed model. The source code is available athttps://github.com/RJL2000/MTAOIQA. Jiebin Yan, Jiale Rao, Junjie Chen 0008, Ziwen Tan, Weide Liu, Yuming Fang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Toward Transparent Deep Image Aesthetics Assessment With Tag-Based Content DescriptorsabstractDeep learning approaches for Image Aesthetics Assessment (IAA) have shown promising results in recent years, but the internal mechanisms of these models remain unclear. Previous studies have demonstrated that image aesthetics can be predicted using semantic features, such as pre-trained object classification features. However, these semantic features are learned implicitly, and therefore, previous works have not elucidated what the semantic features are representing. In this work, we aim to create a more transparent deep learning framework for IAA by introducing explainable semantic features. To achieve this, we propose Tag-based Content Descriptors (TCDs), where each value in a TCD describes the relevance of an image to a human-readable tag that refers to a specific type of image content. This allows us to build IAA models from explicit descriptions of image contents. We first propose the explicit matching process to produce TCDs that adopt predefined tags to describe image contents. We show that a simple MLP-based IAA model with TCDs only based on predefined tags can achieve an SRCC of 0.767, which is comparable to most state-of-the-art methods. However, predefined tags may not be sufficient to describe all possible image contents that the model may encounter. Therefore, we further propose the implicit matching process to describe image contents that cannot be described by predefined tags. By integrating components obtained from the implicit matching process into TCDs, the IAA model further achieves an SRCC of 0.817, which significantly outperforms existing IAA methods. Both the explicit matching process and the implicit matching process are realized by the proposed TCD generator. To evaluate the performance of the proposed TCD generator in matching images with predefined tags, we also labeled 5101 images with photography-related tags to form a validation set. And experimental results show that the proposed TCD generator can meaningfully assign photography-related tags to images. Jingwen Hou, Weisi Lin, Yuming Fang 0001, Haoning Wu 0001, Chaofeng Chen, Weide Liu |
IEEE Trans. Image Process. | 7 |
| 2025 | Diffusion-Based Facial Aesthetics Enhancement With 3D Structure GuidanceabstractFacial Aesthetics Enhancement (FAE) aims to improve facial attractiveness by adjusting the structure and appearance of a facial image while preserving its identity as much as possible. Most existing methods adopted deep feature-based or score-based guidance for generation models to conduct FAE. Although these methods achieved promising results, they potentially produced excessively beautified results with lower identity consistency or insufficiently improved facial attractiveness. To enhance facial aesthetics with less loss of identity, we propose the Nearest Neighbor Structure Guidance based on Diffusion (NNSG-Diffusion), a diffusion-based FAE method that beautifies a 2D facial image with 3D structure guidance. Specifically, we propose to extract FAE guidance from a nearest neighbor reference face. To allow for less change of facial structures in the FAE process, a 3D face model is recovered by referring to both the matched 2D reference face and the 2D input face, so that the depth and contour guidance can be extracted from the 3D face model. Then the depth and contour clues can provide effective guidance to Stable Diffusion with ControlNet for FAE. Extensive experiments demonstrate that our method is superior to previous relevant methods in enhancing facial aesthetics while preserving facial identity. Lisha Li, Jingwen Hou, Weide Liu, Yuming Fang 0001, Jiebin Yan |
IEEE Trans. Image Process. | 3 |
| 2025 | Adaptive Cross-Feature Fusion Network With Inconsistency Guidance for Multi-Modal Brain Tumor SegmentationabstractIn the context of contemporary artificial intelligence, increasing deep learning (DL) based segmentation methods have been recently proposed for brain tumor segmentation (BraTS) via analysis of multi-modal MRI. However, known DL-based works usually directly fuse the information of different modalities at multiple stages without considering the gap between modalities, leaving much room for performance improvement. In this paper, we introduce a novel deep neural network, termed ACFNet, for accurately segmenting brain tumor in multi-modal MRI. Specifically, ACFNet has a parallel structure with three encoder-decoder streams. The upper and lower streams generate coarse predictions from individual modality, while the middle stream integrates the complementary knowledge of different modalities and bridges the gap between them to yield fine prediction. To effectively integrate the complementary information, we propose an adaptive cross-feature fusion (ACF) module at the encoder that first explores the correlation information between the feature representations from upper and lower streams and then refines the fused correlation information. To bridge the gap between the information from multi-modal data, we propose a prediction inconsistency guidance (PIG) module at the decoder that helps the network focus more on error-prone regions through a guidance strategy when incorporating the features from the encoder. The guidance is obtained by calculating the prediction inconsistency between upper and lower streams and highlights the gap between multi-modal data. Extensive experiments on the BraTS 2020 dataset show that ACFNet is competent for the BraTS task with promising results and outperforms six mainstream competing methods. Guanghui Yue 0001, Guibin Zhuo, Tianwei Zhou, Weide Liu, Tianfu Wang 0001, Qiuping Jiang |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | Pyramid Network With Quality-Aware Contrastive Loss for Retinal Image Quality AssessmentabstractCaptured retinal images vary greatly in quality. Low-quality images increase the risk of misdiagnosis. This motivates to design effective retinal image quality assessment (RIQA) methods. Current deep learning-based methods usually classify the image into three levels of "Good", "Usable", and "Reject", while ignoring the quantitative feedback for more detailed quality scores. This study proposes a unified RIQA framework, named QAC-Net, that can evaluate the quality of retinal images in both qualitative and quantitative manners. To improve the prediction accuracy, QAC-Net focuses on extracting discriminative features by using two strategies. On the one hand, it adopts a pyramid network structure that simultaneously inputs the scaled images to learn quality-aware features at different scales and purify the feature representation through a consistency loss. On the other hand, to improve feature representation, it utilizes a quality-aware contrastive (QAC) loss that considers quality relationships between different images. The QAC losses for qualitative and quantitative evaluation tasks have different forms in view of the task differences. Considering the shortage of datasets for the quantitative evaluation task, we construct a dataset with 2,300 authentically distorted retinal images, each of which is annotated with a numerical quality score through subjective experiments. Experimental results on public and our constructed datasets show that our QAC-Net is competent for the RIQA tasks with considerable performance. Guanghui Yue 0001, Shaoping Zhang, Tianwei Zhou, Bin Jiang 0003, Weide Liu, Tianfu Wang 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2025 | Subjective and Objective Quality Assessment of Non-Uniformly Distorted Omnidirectional ImagesabstractOmnidirectional image quality assessment (OIQA) has been one of the hot topics in IQA with the continuous development of VR techniques, and achieved much success in the past few years. However, most studies devote themselves to the uniform distortion issue, i.e., all regions of an omnidirectional image are perturbed by the “same amount” of noise, while ignoring the non-uniform distortion issue, i.e., partial regions undergo “different amount” of perturbation with the other regions in the same omnidirectional image. Additionally, nearly all OIQA models are verified on the platforms containing a limited number of samples, which largely increases the over-fitting risk and therefore impedes the development of OIQA. To alleviate these issues, we elaborately explore this topic from both subjective and objective perspectives. Specifically, we construct a large OIQA database containing 10,320 non-uniformly distorted omnidirectional images, each of which is generated by considering quality impairments on one or two camera len(s). Then we meticulously conduct psychophysical experiments and delve into the influence of both holistic and individual factors (i.e., distortion range and viewing condition) on omnidirectional image quality. Furthermore, we propose a perception-guided OIQA model for non-uniform distortion by adaptively simulating users' viewing behavior. Experimental results demonstrate that the proposed model outperforms state-of-the-art methods. Jiebin Yan, Jiale Rao, Xuelin Liu, Yuming Fang 0001, Yifan Zuo 0001, Weide Liu |
IEEE Trans. Multim. | 6 |
| 2025 | Viewport-Unaware Blind Omnidirectional Image Quality Assessment: A Flexible and Effective ParadigmabstractMost of the existing blind omnidirectional image quality assessment (BOIQA) models rely on viewport generation by modeling user viewing behavior or transforming omnidirectional images (OIs) into varying formats; however, these methods are either computationally expensive or less scalable. To solve these issues, in this article, we present a flexible and effective paradigm, which is viewport-unaware and can be easily adapted to 2D plane image quality assessment (2D-IQA). Specifically, the proposed BOIQA model includes an adaptive prior-equator sampling module for extracting a patch sequence from the equirectangular projection (ERP) image in a resolution-agnostic manner, a progressive deformation-unaware feature fusion module which is able to capture patch-wise quality degradation in a deformation-immune way, and a local-to-global quality aggregation module to adaptively map local perception to global quality. Extensive experiments across four OIQA databases (including uniformly distorted OIs and non-uniformly distorted OIs) demonstrate that the proposed model achieves competitive performance with low complexity against other state-of-the-art models, and we also verify its adaptive capacity to 2D-IQA. The source code is available at https://github.com/KangchengWu/OIQA . Jiebin Yan, Kangcheng Wu, Junjie Chen 0008, Ziwen Tan, Yuming Fang 0001, Weide Liu |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2024 | SuperJunction: Learning-Based Junction Detection for Retinal Image RegistrationabstractKeypoints-based approaches have shown to be promising for retinal image registration, which superimpose two or more images from different views based on keypoint detection and description. However, existing approaches suffer from ineffective keypoint detector and descriptor training. Meanwhile, the non-linear mapping from 3D retinal structure to 2D images is often neglected. In this paper, we propose a novel learning-based junction detection approach for retinal image registration, which enhances both the keypoint detector and descriptor training. To improve the keypoint detection, it uses a multi-task vessel detection to regularize the model training, which helps to learn more representative features and reduce the risk of over-fitting. To achieve effective training for keypoints description, a new constrained negative sampling approach is proposed to compute the descriptor loss. Moreover, we also consider the non-linearity between retinal images from different views during matching. Experimental results on FIRE dataset show that our method achieves mean area under curve of 0.850, which is 12.6% higher than 0.755 by the state-of-the-art method. All the codes are available at https://github.com/samjcheng/SuperJunction. Zaiwang Gu, Weide Liu, Wee Siong Ng, Weimin Huang 0002, Jun Cheng 0003 |
AAAI | 4 |
| 2024 | Learning Intra-View and Cross-View Geometric Knowledge for Stereo MatchingabstractGeometric knowledge has been shown to be beneficial for the stereo matching task. However, prior attempts to in-tegrate geometric insights into stereo matching algorithms have largely focused on geometric knowledge from single images while crucial cross-view factors such as occlusion and matching uniqueness have been overlooked. To address this gap, we propose a novel Intra-view and Cross-view Geometric knowledge learning Network (ICGNet), specifically crafted to assimilate both intra-view and cross-view geo-metric knowledge. ICGNet harnesses the power of interest points to serve as a channel for intra-view geometric understanding. Simultaneously, it employs the correspon-dences among these points to capture cross-view geometric relationships. This dual incorporation empowers the proposed ICGNet to leverage both intra-view and cross-view geometric knowledge in its learning process, substantially improving its ability to estimate disparities. Our extensive experiments demonstrate the superiority of the ICGNet over contemporary leading models. The code will be available at https://github.com/DFSDDDDDl199/ICGNet. Weide Liu, Zaiwang Gu, Xulei Yang, Jun Cheng 0003 |
CVPR | 2 |
| 2024 | Clip-Medfake: Synthetic Data Augmentation With AI-Generated Content for Improved Medical Image ClassificationabstractData augmentation is serving as a critical and fundamental technology to improve model generalization and performance in a wide spectrum of machine learning tasks. Despite the increasing interest in developing various pathways to artificially generate new data to reduce the overfitting issue during model training, enriching the diversity of training data in the field of medicine remains facing enormous challenges. By virtue of recent advancements in generative artificial intelligence, we present a novel data augmentation framework, CLIP-MedFake, to address the shortage of training data used in medical image classification. The proposed method first employs the Stable Diffusion model to generate new fake data based on a small amount of training data, and then adopts the paradigm of few-shot learning and uses the CLIP architecture as the backbone to pre-train the model with synthetic data and then fine-tune it with real medical images. Extensive experiment results on two publicly available datasets demonstrate the effectiveness of the proposed method in promoting medical image classification. Honghui Chen, Baoquan Zhao, Guanghui Yue 0001, Weide Liu, Chenlei Lv, Ruomei Wang 0001, Fan Zhou 0001 |
ICIP | 4 |
| 2024 | Predicting Plain Text Imageability for Faithful Prompt-Conditional Image Generation
Guanghui Yue 0001, Weide Liu, Chenlei Lv, Ruomei Wang 0001, Fan Zhou 0001, Baoquan Zhao |
PRICAI (3) | 3 |
| 2024 | Harmonizing Base and Novel Classes: A Class-Contrastive Approach for Generalized Few-Shot Segmentation
Weide Liu, Yuming Fang 0001, Chuan-Sheng Foo, Jun Cheng 0003, Guosheng Lin |
Int. J. Comput. Vis. | 1 |
| 2024 | LCReg: Long-tailed image classification with Latent Categories based Recognition
Weide Liu, Henghui Ding, Fayao Liu, Jie Lin 0001, Guosheng Lin |
Pattern Recognit. | 1 |
| 2024 | Video Quality Assessment for Online Processing: From Spatial to Temporal SamplingabstractWith the rapid development of multimedia processing and deep learning technologies, especially in the field of video understanding, video quality assessment (VQA) has achieved significant progress. Although researchers have moved from designing efficient video quality mapping models to various research directions, in-depth exploration of the effectiveness-efficiency trade-offs of spatio-temporal modeling in VQA models is still less sufficient. Considering the fact that videos have highly redundant information, this paper investigates this problem from the perspective of joint spatial and temporal sampling, aiming to seek the answer to how little information we should keep at least when feeding videos into the VQA models while with acceptable performance sacrifice. To this end, we drastically sample the video’s information from both spatial and temporal dimensions, and the heavily squeezed video is then fed into a stable VQA model. Comprehensive experiments regarding joint spatial and temporal sampling are conducted on six public video quality databases, and the results demonstrate the acceptable performance of the VQA model when throwing away most of the video information. Furthermore, with the proposed joint spatial and temporal sampling strategy, we make an initial attempt to design an online VQA model, which is instantiated by as simple as possible a spatial feature extractor, a temporal feature fusion module, and a global quality regression module. Through quantitative and qualitative experiments, we verify the feasibility of online VQA model by simplifying itself and reducing input. Jiebin Yan, Yuming Fang 0001, Xuelin Liu, Xue Xia 0005, Weide Liu |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | FedCov: Enhanced Trustworthy Federated Learning for Machine RUL Prediction With Continuous-to-Discrete ConversionabstractNumerous approaches have been proposed for predicting machine remaining useful life (RUL), which helps prevent unnecessary downtime and reduces the maintenance cost in industrial systems. Most existing RUL methods rely on centralized learning and require large-scale datasets with manual labels, which are infeasible to collect. As a decentralized learning paradigm, federated learning (FL) has recently been integrated into these approaches, which aims to utilize the distributed data from local users for model training, while preserving their data privacy. However, data heterogeneity in the industry poses a critical challenge for FL, leading to model drifting issue and degraded global model performance. A straightforward method to tackle this problem is to estimate the data distribution of clients. However, it is difficult to apply this method to a regression task, since the prediction space is continuous, which increases the difficulty in estimating the data distribution. Motivated by the digital-to-analog converter in electronics, we propose a novel approach called FedCov, which involves a converter module that transforms continuous RUL values into discrete categories. Subsequently, a generator is trained to aggregate user information based on the discrete label distribution, and it is broadcasted to users as a data enhancement tool to address the data heterogeneity problem. Furthermore, to improve the performance and reliability of our FedCov, an uncertainty estimation module is proposed, which utilizes the confidence level of model predictions to adjust the training direction. Extensive experiments are conducted on theC-MAPSSbenchmark, which demonstrates that our proposed FedCov effectively solves the model drift issues and improves the performances on the RUL task, achieving state-of-the-arts performances. Yuming Fang 0001, Weide Liu, Ruibing Jin, Jun Cheng 0003, Zhenghua Chen |
IEEE Trans. Ind. Informatics | 3 |
| 2024 | Bayesian Uncertainty Calibration for Federated Time Series AnalysisabstractDeep learning models for time series analysis often require large-scale labeled datasets for training. However, acquiring such datasets is cost-intensive and challenging, particularly for individual institutions. To overcome this challenge and concern about data confidentiality among different institutions, federated learning (FL) servers as a viable solution to this dilemma by offering a decentralized learning framework. However, the datasets collected by each institution often suffer from imbalance and may not adhere to uniform protocols, leading to diverse data distributions. To address this problem, we design a global model to approximate the global data distribution of all participant clients, then transfer it to local clients as an induction in the training phase. While discrepancies between the approximate distribution and the actual distribution result in uncertainty in the predicted results. Moreover, the diverse data distributions among various clients within the FL framework, combined with the inherent lack of reliability and interpretability in deep learning models, further amplify the uncertainty of the prediction results. To address these issues, we propose an uncertainty calibration method based on Bayesian deep learning techniques, which captures uncertainty by learning a fidelity transformation to reconstruct the output of time series regression and classification tasks, utilizing deterministic pre-trained models. Extensive experiments on the regression dataset (C-MAPSS) and classification datasets (ESR, Sleep-EDF, HAR, and FD) in the Independent and Identically Distributed (IID) and non-IID settings show that our approach effectively calibrates uncertainty within the FL framework and facilitates better generalization performance in both the regression and classification tasks, achieving state-of-the-art performance. Weide Liu, Xue Xia 0005, Zhenghua Chen, Yuming Fang 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | Towards Explainable Recommendation Via Bert-Guided Explanation GeneratorabstractExplainable recommender system has recently drawn increasing attention due to its capability of providing justification to recommendation. Rather than focusing on certain topics or specific item features, the explanation generated by existing works are too general without the guidance of aspects. However, such information is not given in the practical scenario. To address this issue, we propose a novel Explainable recommender system with BERT-guided explanation generator, named ExBERT to generate reliable explanation with finer granularity. More specifically, a multi-head self-attention based encoder is employed to incorporate pseudo user and item profiles into semantic representation. Moreover, we propose a novel matched explanation prediction task with discriminative ability to enable personalization of the generated sentence. Extensive experiments conducted on two real-world explainable recommendation datasets significantly outperform the state-of-the-art in generation. Huijing Zhan, Ling Li 0012, Weide Liu, Manas Gupta, Alex Chichung Kot |
ICASSP | 4 |
| 2023 | ELFNet: Evidential Local-global Fusion for Stereo MatchingabstractAlthough existing stereo matching models have achieved continuous improvement, they often face issues related to trustworthiness due to the absence of uncertainty estimation. Additionally, effectively leveraging multi-scale and multi-view knowledge of stereo pairs remains unexplored. In this paper, we introduce the Evidential Local-global Fusion (ELF) framework for stereo matching, which endows both uncertainty estimation and confidence-aware fusion with trustworthy heads. Instead of predicting the disparity map alone, our model estimates an evidential-based disparity considering both aleatoric and epistemic uncertainties. With the normal inverse-Gamma distribution as a bridge, the proposed framework realizes intra evidential fusion of multi-level predictions and inter evidential fusion between cost-volume-based and transformer-based stereo matching. Extensive experimental results show that the proposed framework exploits multi-view information effectively and achieves state-of-the-art overall performance both on accuracy and cross-domain generalization. The codes are available at https://github.com/jimmy19991222/ELFNet. Jieming Lou, Weide Liu, Fayao Liu, Jun Cheng 0003 |
ICCV | 2 |
| 2023 | Perceptual Quality Assessment of Enhanced Colonoscopy Images: A Benchmark Dataset and an Objective MethodabstractIn colonoscopy, the captured images are usually with low-quality appearance, such as non-uniform illumination, low contrast, etc., due to the specialized imaging environment, which may provide poor visual feedback and bring challenges to subsequent disease analysis. Many low-light image enhancement (LIE) algorithms have recently proposed to improve the perceptual quality. However, how to fairly evaluate the quality of enhanced colonoscopy images (ECIs) generated by different LIE algorithms remains a rarely-mentioned and challenging problem. In this study, we carry out a pioneering investigation on perceptual quality assessment of ECIs. Firstly, considering the lack of specific datasets, we collect 300 low-light images with diverse contents during the real-world colonoscopy and conduct rigorous subjective studies to compare the performance of 8 popular LIE methods, resulting in a benchmark dataset (named ECIQAD) for ECIs. Secondly, in view of the distinctive distortion characteristics of ECIs, we propose an effective no-reference Enhanced Colonoscopy Image Quality (ECIQ) method to automatically evaluate the perceptual quality of ECIs via analysis of brightness, contrast, colorfulness, naturalness, and noise. Extensive experiments on ECIQAD demonstrate the superiority of our proposed ECIQ method over 14 mainstream no-reference image quality assessment methods. Guanghui Yue 0001, Tianwei Zhou, Jingwen Hou, Weide Liu, Long Xu 0001, Tianfu Wang 0001, Jun Cheng 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Benchmarking Polyp Segmentation Methods in Narrow-Band Imaging Colonoscopy ImagesabstractIn recent years, there has been significant progress in polyp segmentation in white-light imaging (WLI) colonoscopy images, particularly with methods based on deep learning (DL). However, little attention has been paid to the reliability of these methods in narrow-band imaging (NBI) data. NBI improves visibility of blood vessels and helps physicians observe complex polyps more easily than WLI, but NBI images often include polyps with small/flat appearances, background interference, and camouflage properties, making polyp segmentation a challenging task. This paper proposes a new polyp segmentation dataset (PS-NBI2K) consisting of 2,000 NBI colonoscopy images with pixel-wise annotations, and presents benchmarking results and analyses for 24 recently reported DL-based polyp segmentation methods on PS-NBI2K. The results show that existing methods struggle to locate polyps with smaller sizes and stronger interference, and that extracting both local and global features improves performance. There is also a trade-off between effectiveness and efficiency, and most methods cannot achieve the best results in both areas simultaneously. This work highlights potential directions for designing DL-based polyp segmentation methods in NBI colonoscopy images, and the release of PS-NBI2K aims to drive further development in this field. Guanghui Yue 0001, Guibin Zhuo, Tianwei Zhou, Jingfeng Du, Weiqing Yan, Jingwen Hou, Weide Liu, Tianfu Wang 0001 |
IEEE J. Biomed. Health Informatics | 8 |
| 2023 | Interaction-Matrix Based Personalized Image Aesthetics AssessmentabstractPersonalized image aesthetics assessment (IAA) aims to estimate aesthetic experiences subject to the preferences of individual users, contrary to generic IAA that estimates aesthetic experiences subject to average preferences. Most existing personalized IAA methods treat personalized aesthetic experiences as deviations from a generic aesthetic experience, and therefore, personalized IAA models are designed to build upon the prior knowledge on generic IAA. However, we propose that acquiring knowledge on generic IAA is not necessary for building a personalized IAA model. Instead of modeling personalized IAA on the basis of generic IAA, this work proposes to directly estimate personalized aesthetic experiences from the interactions between image contents and user preferences (i.e., preference-content interaction), where interaction-matrices representing preference-content interactions are constructed without needs for prior generic IAA knowledge. To this end, we construct interaction-matrices from content features constructed from pre-trained image classification features and latent preference features. To realize a robust interaction-matrix based personalized IAA model, we discuss in detail on different strategies for constructing interaction-matrices and estimating personalized aesthetic scores from the interaction-matrices. Besides the personalized IAA scenario, we further propose strategies to adapt the proposed personalized IAA model to different scenarios of generic IAA. Extensive experiments show that: 1) our method significantly outperforms 5 previous relevant personalized IAA methods on FLICKR-AES dataset, especially the methods that require generic IAA knowledge as the basis; 2) in terms of generic IAA, the proposed approach also outperforms 13 generic IAA methods on AVA dataset. Jingwen Hou, Weisi Lin, Guanghui Yue 0001, Weide Liu, Baoquan Zhao |
IEEE Trans. Multim. | 4 |
| 2023 | Cross-Image Region Mining With Region Prototypical Network for Weakly Supervised SegmentationabstractWeakly supervised image segmentation trained with image-level labels usually suffers from inaccurate coverage of object areas during the generation of the pseudo groundtruth. This is because the object activation maps are trained with the classification objective and lack the ability to generalize. To improve the generality of the object activation maps, we propose a region prototypical network (RPNet) to explore the cross-image object diversity of the training set. Similar object parts across images are identified via region feature comparison. Object confidence is propagated between regions to discover new object areas while background regions are suppressed. Experiments show that the proposed method generates more complete and accurate pseudo object masks while achieving state-of-the-art performance on PASCAL VOC 2012 and MS COCO. In addition, we investigate the robustness of the proposed method on reduced training sets. The code is available athttps://github.com/liuweide01/RPNet-Weakly-Supervised-Segmentation. Weide Liu, Xiangfei Kong, Tzu-Yi Hung, Guosheng Lin |
IEEE Trans. Multim. | 1 |
| 2023 | Few-Shot Segmentation With Optimal Transport Matching and Message FlowabstractWe tackle the challenging task of few-shot segmentation in this work. It is essential for few-shot semantic segmentation to fully utilize the support information. Previous methods typically adopt masked average pooling over the support feature to extract the support clues as a global vector, usually dominated by the salient part and lost certain essential clues. In this work, we argue that every support pixel’s information is desired to be transferred to all query pixels and propose a Correspondence Matching Network (CMNet) with an Optimal Transport Matching module to mine out the correspondence between the query and support images. Besides, it is critical to fully utilize both local and global information from the annotated support images. To this end, we propose a Message Flow module to propagate the message along the inner-flow inside the same image and cross-flow between support and query images, which greatly helps enhance the local feature representations. Experiments on PASCAL VOC 2012, MS COCO, and FSS-1000 datasets show that our network achieves new state-of-the-art few-shot segmentation performance. Weide Liu, Chi Zhang 0007, Henghui Ding, Tzu-Yi Hung, Guosheng Lin |
IEEE Trans. Multim. | 1 |
| 2022 | CRCNet: Few-Shot Segmentation with Cross-Reference and Region-Global Conditional Networks
Weide Liu, Chi Zhang 0007, Guosheng Lin, Fayao Liu |
Int. J. Comput. Vis. | 1 |
| 2022 | Distilling Knowledge From Object Classification to Aesthetics AssessmentabstractIn this work, we point out that the major dilemma of image aesthetics assessment (IAA) comes from the abstract nature of aesthetic labels. That is, a vast variety of distinct contents can correspond to the same aesthetic label. On the one hand, during inference, the IAA model is required to relate various distinct contents to the same aesthetic label. On the other hand, when training, it would be hard for the IAA model to learn to distinguish different contents merely with the supervision from aesthetic labels, since aesthetic labels are not directly related to any specific content. To deal with this dilemma, we propose to distill knowledge on semantic patterns for a vast variety of image contents from multiple pre-trained object classification (POC) models to an IAA model. Expecting the combination of multiple POC models can provide sufficient knowledge on various image contents, the IAA model can easier learn to relate various distinct contents to a limited number of aesthetic labels. By supervising an end-to-end single-backbone IAA model with the distilled knowledge, the performance of the IAA model is significantly improved by 4.8% in SRCC compared to the version trained only with ground-truth aesthetic labels. On specific categories of images, the SRCC improvement brought by the proposed method can achieve up to 7.2%. Peer comparison also shows that our method outperforms 10 previous IAA methods. Jingwen Hou, Henghui Ding, Weisi Lin, Weide Liu, Yuming Fang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Stability analysis for quaternion-valued inertial memristor-based neural networks with time delays
Weide Liu, Jianliang Huang, Qinghe Yao |
Neurocomputing | 1 |
| 2021 | Guided Co-Segmentation Network for Fast Video Object SegmentationabstractSemi-supervised video object segmentation is a task of propagating instance masks given in the first frame to the entire video. It is a challenging task since it usually suffers from heavy occlusions, large deformation, and large variations of objects. To alleviate these problems, many existing works apply time-consuming techniques such as fine-tuning, post-processing, or extracting optical flow, which makes them intractable for online segmentation. In our work, we focus on online semi-supervised video object segmentation. We propose a GCSeg (Guided Co-Segmentation) Network which is mainly composed of a Reference Module and a Co-segmentation Module, to simultaneously incorporate the short-term, middle-term, and long-term temporal inter-frame relationships. Moreover, we propose an Adaptive Search Strategy to reduce the risk of propagating inaccurate segmentation results in subsequent frames. Our GCSeg network achieves state-of-the-art performance on online semi-supervised video object segmentation on Davis 2016 and Davis 2017 datasets. Weide Liu, Guosheng Lin, Tianyi Zhang 0004, Zichuan Liu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | CRNet: Cross-Reference Networks for Few-Shot SegmentationabstractOver the past few years, state-of-the-art image segmentation algorithms are based on deep convolutional neural networks. To render a deep network with the ability to understand a concept, humans need to collect a large amount of pixel-level annotated data to train the models, which is time-consuming and tedious. Recently, few-shot segmentation is proposed to solve this problem. Few-shot segmentation aims to learn a segmentation model that can be generalized to novel classes with only a few training images. In this paper, we propose a cross-reference network (CRNet) for few-shot segmentation. Unlike previous works which only predict the mask in the query image, our proposed model concurrently makes predictions for both the support image and the query image. With a cross-reference mechanism, our network can better find the co-occurrent objects in two images, thus helping the few-shot segmentation task. We also develop a mask refinement module to recurrently refine the prediction of the foreground regions. For the k-shot learning, we propose to finetune parts of networks to take advantage of multiple labeled support images. Experiments on the PASCAL VOC 2012 dataset show that our network achieves state-of-the-art performance. Weide Liu, Chi Zhang 0007, Guosheng Lin, Fayao Liu |
CVPR | 1 |
| 2020 | Splitting Vs. Merging: Mining Object Regions with Discrepancy and Intersection Loss for Weakly Supervised Semantic Segmentation
Tianyi Zhang 0004, Guosheng Lin, Weide Liu, Jianfei Cai 0001, Alex Chichung Kot |
ECCV (22) | 3 |
| 2020 | Weakly Supervised Segmentation with Maximum Bipartite Graph MatchingabstractIn the weakly supervised segmentation task with only image-level labels, a common step in many existing algorithms is first to locate the image regions corresponding to each existing class with the Class Activation Maps (CAMs), and then generate the pseudo ground truth masks based on the CAMs to train a segmentation network in the fully supervised manner. The quality of the CAMs has a crucial impact on the performance of the segmentation model. We propose to improve the CAMs from a novel graph perspective. We model paired images containing common classes with a bipartite graph and use the maximum matching algorithm to locate corresponding areas in two images. The matching areas are then used to refine the predicted object regions in the CAMs. The experiments on Pascal VOC 2012 dataset show that our network can effectively boost the performance of the baseline model and achieves new state-of-the-art performance. Weide Liu, Chi Zhang 0007, Guosheng Lin, Tzu-Yi Hung, Chunyan Miao |
ACM Multimedia | 1 |