EDBT 2026 Demo / reviewers in the wild / expert
Xulei Yang
dblp:91/10215 · also XuLei Yang
· DBLP profile ↗
93ranked-venue papers
19as first author
64since 2021 · last 2026
0000-0002-7002-4564ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 56 · 16 first-author · 32 since 2021Graphics, computer vision, multimedia, augmented reality and games · 44 · 3 first-author · 39 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-author · 4 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Illumination-Aware Restoration of Metalens-Captured Images: A New Dataset and a Strong BaselineabstractMetalenses offer compelling advantages such as lightweight and ultra-thin design, making them promising alternatives to conventional lenses. However, their widespread adoption is hindered by image quality degradation caused by chromatic and angular aberrations. To mitigate this, restoration processes are often necessary to recover high-quality RGB images from metalens-captured inputs. While recent deep learning-based restoration methods show promise, they typically (1) blur or distort peripheral regions, or (2) fail entirely under unseen illumination conditions. To advance metalens image restoration, we introduce IlluMeta---the first and largest real-world, illumination-aware metalens image dataset—captured across diverse lighting environments. In addition, we propose a novel end-to-end restoration framework that directs attention to challenging regions and adaptively adjusts to varying illuminations via reinforcement learning. Experiments show that our method can be applied in a plug-and-play manner to enhance existing models, significantly improving image restoration quality, especially under unseen lighting conditions, paving the way for broader real-world deployment of metalens technologies. Fen Fang, Xinan Liang, Muli Yang, Jinghong Zheng 0001, Tobias Wilhelm W. Mass, Ying Sun 0001, Xulei Yang, Xuewu Xu, Zhengguo Li |
AAAI | 7 |
| 2026 | Next-Generation Metalens Vision System: Powered by AI and Applied to AIabstractMetalenses have been widely recognized as a key building block of next-generation optical systems, offering unprecedented advantages in compactness, lightweight design, and scalable manufacturing compared to traditional refractive optics. Despite this promise, practical use is limited by optical aberrations, blur, and illumination sensitivity, which degrade both visual quality and machine perception. In this demonstration, we present an end-to-end metalens vision system—from hardware sensing with a custom-built RGB metalens camera, to physics-informed imaging and real-time restoration, and finally to downstream vision applications such as object detection and depth estimation. By integrating spatially-aware attention enhancement and reinforcement learning-based illumination control into a real-time system, our solution transforms degraded raw captures into high-fidelity images that are both visually interpretable and functionally reliable for machine vision. This AI-powered pipeline highlights metalenses as a cornerstone for next-generation imaging, where advances in optics and machine intelligence jointly drive the future of visual perception. Fen Fang, Muli Yang, Henan Wang, Xinan Liang, Tobias Wilhelm W. Mass, Xuewu Xu, Xulei Yang, Zhengguo Li |
AAAI | 7 |
| 2026 | AD-FM: Multimodal LLMs for Anomaly Detection via Multi-Stage Reasoning and Fine-Grained Reward OptimizationabstractWhile Multimodal Large Language Models (MLLMs) demonstrate remarkable capabilities across diverse domains, their application to specialized anomaly detection (AD) remains constrained by domain adaptation challenges. Existing Group Relative Policy Optimization (GRPO) based approaches suffer from two critical limitations: inadequate training data utilization when models produce uniform responses, and insufficient supervision over reasoning processes that encourage immediate binary decisions without deliberative analysis. We propose a comprehensive framework addressing these limitations through two synergistic innovations. First, we introduce a multi-stage deliberative reasoning process that guides models from region identification to focused examination, generating diverse response patterns essential for GRPO optimization while enabling structured supervision over analytical workflows. Second, we develop a fine-grained reward mechanism incorporating classification accuracy and localization supervision, transforming binary feedback into continuous signals that distinguish genuine analytical insight from spurious correctness. Comprehensive evaluation across multiple industrial datasets shows that our method achieves superior accuracy by enabling general-purpose MLLMs to acquire fine-grained visual discrimination for detecting subtle manufacturing defects. Jingyi Liao, Yongyi Su, Rongcheng Tu, Xun Xu 0002, Dacheng Tao, Xulei Yang |
AAAI | 9 |
| 2026 | From Language to Driving: A Dual-Loop SLM-Enhanced Framework for Multi-Planner Scheduling via a Domain-Specific LanguageabstractJiawei Liu, Xun Gong, Muli Yang, Xingrui Yu, Fen Fang, Xulei Yang, Ivor Tsang, Yunfeng hu, Hong Chen, Qing Guo. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xun Gong 0007, Muli Yang, Xingrui Yu, Fen Fang, Xulei Yang, Ivor W. Tsang, Yunfeng Hu 0003, Hong Chen 0003, Qing Guo 0005 |
ACL (1) | 6 |
| 2026 | BL-UDA: Towards Unsupervised Domain-Adaptive Surgical Instrument Segmentation with Source Box LabelsabstractRecent advances in unsupervised domain adaptation (UDA) by adapting the model from one domain to another unseen domain have shown considerable promise in improving surgical instrument segmentation performance across domains. However, existing UDA methods primarily rely on pixel-wise labels, which are always difficult to collect due to the labor-intensive annotation process. In this work, we aim to relax the dependence on pixel-level supervision and investigate a challenging UDA setting - source box annotations, where weak supervision and domain shifts coexist. To achieve this, we introduce a novel unsupervised domain adaptation framework, BL-UDA, which leverages bounding box annotations for surgical instrument segmentation across domains. By utilizing the Segment Anything Model (SAM) for pseudo label generation from box annotations, our method effectively bridges object-level and pixel-level domain adaptation. The proposed BL-UDA framework comprises a teacher-student network with entropy minimization for object detection and an entropy-based label selection strategy for generating box prompts to SAM, facilitating pixel-level domain adaptation. Extensive experiments on the EndoVis 2017 and 2018 datasets demonstrate the superiority of BL-UDA over existing UDA methods, significantly mitigating domain shifts and addressing weak supervision challenges with minimal annotation requirements. Ziyuan Zhao, Yifang Yin, Yichen Zhang 0002, Xulei Yang, Jun Cheng 0003, Roger Zimmermann, Cuntai Guan, Shaohua Kevin Zhou |
ICMR | 5 |
| 2026 | Evidential Robust Feature Learning for Generalized Few-Shot Segmentation
Weide Liu, Xiaoyang Zhong, Lu Wang 0001, Chunbo Lang, Yuming Fang 0001, Jun Cheng 0003, Xulei Yang, Gong Cheng 0003 |
Int. J. Comput. Vis. | 7 |
| 2026 | Calibrating distributions, not just networks: The PR-Q method for low-bit quantization
Xue He, Xue Geng, Tiancheng Zhang 0001, Minghe Yu 0001, Yuhai Zhao, Xulei Yang, Min Wu 0008, Ge Yu 0001 |
Neurocomputing | 6 |
| 2026 | MeLoRA : Probability measures-based low-rank adaptation with Gaussian variational inference
Xue He, Xue Geng, Tiancheng Zhang 0001, Minghe Yu 0001, Yuhai Zhao, Xulei Yang, Min Wu 0008, Ge Yu 0001 |
Knowl. Based Syst. | 6 |
| 2026 | FOCUS: Frequency-Optimized Conditioning of diffUSion models for mitigating catastrophic forgetting during test-time adaptation
Gabriel Tjio, Jie Zhang 0002, Xulei Yang, Nhat Chung, Xiaofeng Cao 0002, Ivor W. Tsang, Chee Keong Kwoh 0001, Qing Guo 0005 |
Mach. Vis. Appl. | 3 |
| 2026 | Toward Accurate Procedure Planning in Instructional Videos: Visual State Generation Helps Task-Selective DiffusionabstractProcedure planning in instructional videos entails predicting an action sequence that transitions a given start state to a desired goal state. This task is particularly challenging due to two key sources of uncertainty: limited visual observations and an enormous decision space. The former results in multiple plausible plan variations due to missing intermediate visual states, while the latter complicates prediction by requiring selection from a large set of potential actions. Unlike prior work that addresses these issues implicitly, we propose an explicit solution. To mitigate the first challenge, we employ image generation models to synthesize diverse intermediate visual states using various text prompts, followed by a prompt selection module integrated within a diffusion model. To tackle the second challenge, we introduce a task-selective diffusion model that applies a task-specific mask to constrain the action space. As the effectiveness of this mask depends on accurate task classification, we further enhance visual representation by leveraging pre-trained vision-language models to generate action-aware, text-enriched multimodal embeddings. Extensive experiments on three benchmark datasets validate the superior performance of our proposed approach. Fen Fang, Muli Yang, Min Wu 0008, Yanhua Yang, Qianli Xu, Joo-Hwee Lim, Xulei Yang, Hongyuan Zhu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2026 | Velocity Space Representation Learning for GPR Keypoint Detection and MatchingabstractReliable localization under Global Positioning System-denied or visually degraded conditions remains a fundamental challenge for autonomous systems. Vision- and Light Detection and Ranging (LiDAR)-based approaches often degrade in low illumination, adverse weather, or appearance-changing environments, as they rely on stable surface texture or geometry. In contrast, ground-penetrating radar (GPR) captures subsurface electromagnetic reflections that remain relatively stable across lighting, seasonal, and weather variations, making it a promising complementary sensing modality for long-term localization. However, spatial variability in subsurface dielectric properties induces fluctuations in electromagnetic wave velocity, leading to geometric distortions in GPR echoes and unstable feature extraction. To address this challenge, we propose the Velocity-Invariant Feature Transform (VIFT), a physics-guided self-supervised learning framework for GPR keypoint detection and description. VIFT explicitly models wave-velocity-induced distortions through a continuous velocity space parameterized by a Beta distribution, and leverages velocity-conditioned wavefield migration as physically consistent data augmentation. A Siamese network is trained with velocity-consistency supervision to jointly learn repeatable keypoint score maps and discriminative local descriptors from unlabeled real GPR scans. To further enhance robustness, sparsity-aware, dispersion, distinctiveness, and orthogonality losses are incorporated to improve repeatability, spatial coverage, and descriptor discriminability. Extensive experiments on public benchmarks and large-scale real-world GPR datasets demonstrate that VIFT consistently outperforms traditional handcrafted methods and recent learning-based Vison and GPR methods, achieving a 5–10% improvement in keypoint repeatability over state-of-the-art methods, particularly under extremely sparse keypoint sampling regimes, while also improving matching accuracy and registration robustness under diverse subsurface conditions. Xieyuanli Chen, Liang Shen 0003, Xulei Yang, Bharadwaj Veeravalli, Shijie Li 0006, Tian Jin 0001, Xiaotao Huang 0001 |
IEEE Trans. Ind. Informatics | 4 |
| 2026 | Integrating SAM Supervision for 3D Weakly Supervised Point Cloud SegmentationabstractCurrent methods for 3D semantic segmentation propose training models with limited annotations to address the difficulty of annotating large, irregular, and unordered 3D point cloud data. They usually focus on the 3D domain only, without leveraging the complementary nature of 2D and 3D data. Besides, some methods extend original labels or generate pseudo labels to guide the training, but they often fail to fully use these labels or address the noise within them. Meanwhile, the emergence of comprehensive and adaptable foundation models has offered effective solutions for segmenting 2D data. Leveraging this advancement, we present a novel approach that maximizes the utility of sparsely available 3D annotations by incorporating segmentation masks generated by 2D foundation models. We further propagate the 2D segmentation masks into the 3D space by establishing geometric correspondences between 3D scenes and 2D views. We extend the highly sparse annotations to encompass the areas delineated by 3D masks, thereby substantially augmenting the pool of available labels. Furthermore, we apply confidence- and uncertainty-based consistency regularization on augmentations of the 3D point cloud and select the reliable pseudo labels, which are further spread on the 3D masks to generate more labels. This innovative strategy bridges the gap between limited 3D annotations and the powerful capabilities of 2D foundation models, ultimately improving the performance of 3D weakly supervised segmentation. Lechun You, Weide Liu, Xulei Yang, Jun Cheng 0003, Wei Zhou 0021, Bharadwaj Veeravalli, Guosheng Lin |
IEEE Trans. Image Process. | 4 |
| 2025 | Trustworthy Disentangled Framework for Multi-Label Medical Image Classification with Multimodal RefinementabstractClinical practice reveals that patients frequently suffer from multiple co-occurring diseases, making multi-label classification (MLC) essential for accurate diagnosis. However, current MLC methods face two major challenges: (1) Disease-specific feature entanglement arising from the complex interdisease correlations among comorbidities; and (2) Untrustworthy results due to single-point estimates that lack confidence measurement. In this paper, we attempt to address these challenges at both the model and optimization levels. Specifically, at the model level, we introduce an improved transformer architecture with multi-CLS tokens for feature disentanglement. This architecture effectively captures the relationships among different diseases, while each CLS token integrates class-wise features, further refined by a multimodal method using a vision language model (VLM). At the optimization level, we propose a novel trustworthy MLC loss that aggregates positive/negative evidence for each class, modeling a multi-Beta distribution based on the Theory of Evidence, to generate reliable predictions with uncertainty estimations. Extensive experiments are conducted on publicly available clinical datasets, and the results demonstrate the effectiveness of our proposed method11The code is available at: https://github.com/CYYukio/Trustworthy-Disentangled-Framework.. Ziyuan Yang 0001, Yongqiang Huang 0003, Xulei Yang, Siyong Yeo, Yi Zhang 0018 |
BIBM | 4 |
| 2025 | Rectification-specific Supervision and Constrained Estimator for Online Stereo RectificationabstractOnline stereo rectification is critical for autonomous vehicles and robots in dynamic environments, where factors such as vibration, temperature fluctuations, and mechanical stress can affect rectification accuracy and severely degrade downstream stereo depth estimation. Current dominant approaches for online stereo rectification involve estimating relative camera poses in real time to derive rectification homographies. However, they do not directly optimize for rectification constraints. Additionally, the general-purpose correspondence matchers used in these methods are not trained for rectification, while training of these matchers typically requires ground-truth correspondences which are not available in stereo rectification datasets. To address these limitations, we propose a matching-based stereo rectification framework that is directly optimized for rectification and does not require ground-truth correspondence annotations for training. We assume intrinsics are known as they are generally available on modern devices and are relatively stable. Our framework incorporates a rectification-constrained estimator and applies multi-level, rectification-specific supervision that trains the matcher network for rectification without relying on ground-truth correspondences. Additionally, we create a new rectification dataset with ground-truth optical flow annotations, eliminating bias from evaluation metrics used in prior work that relied on pretrained keypoint matching or optical flow models. Extensive experiments show that our approach outperforms both state-of-the-art matching-based and matching-free methods in vertical flow metric by 10.7% on the Carla-Flowguided dataset and 21.3% on the Semi-Truck Highway dataset, offering superior rectification accuracy. Kim-Hui Yap, Weide Liu, Xulei Yang, Jun Cheng 0003 |
CVPR | 4 |
| 2025 | SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Groundingabstract3D Visual Grounding (3DVG) aims to locate objects in 3D scenes based on textual descriptions, essential for applications like augmented reality and robotics. Traditional 3DVG approaches rely on annotated 3D datasets and predefined object categories, limiting scalability and adaptability. To overcome these limitations, we introduce SeeGround, a zero-shot 3DVG framework leveraging 2D Vision-Language Models (VLMs) trained on large-scale 2D data. SeeGround represents 3D scenes as a hybrid of query-aligned rendered images and spatially enriched text descriptions, bridging the gap between 3D data and 2D-VLMs input formats. We propose two modules: the Perspective Adaptation Module, which dynamically selects viewpoints for query-relevant image rendering, and the Fusion Alignment Module, which integrates 2D images with 3D spatial descriptions to enhance object localization. Extensive experiments on ScanRefer and Nr3D demonstrate that our approach outperforms existing zero-shot methods by large margins. Notably, we exceed weakly supervised methods and rival some fully supervised ones, outperforming previous SOTA by 7.7% on ScanRefer and 7.1% on Nr3D, showcasing its effectiveness in complex 3DVG tasks. Project website (with demo and code): https://seeground.github.io. Shijie Li 0006, Lingdong Kong, Xulei Yang, Junwei Liang 0001 |
CVPR | 4 |
| 2025 | MedUnifier: Unifying Vision-and-Language Pre-training on Medical Data with Vision Generation Task using Discrete Visual RepresentationsabstractDespite significant progress in Vision-Language Pre-training (VLP), current approaches predominantly emphasize feature extraction and cross-modal comprehension, with limited attention to generating or transforming visual content. This gap hinders the model’s ability to synthesize coherent and novel visual representations from textual prompts, thereby reducing the effectiveness of multi-modal learning. In this work, we propose MedUnifier, a unified VLP framework tailored for medical data. MedUnifier seamlessly integrates text-grounded image generation capabilities with multi-modal learning strategies, including image-text contrastive alignment, image-text matching and image-grounded text generation. Unlike traditional methods that reply on continuous visual representations, our approach employs visual vector quantization, which not only facilitates a more cohesive learning strategy for cross-modal understanding but also enhances multi-modal generation quality by effectively leveraging discrete representations. Our framework’s effectiveness is evidenced by the experiments on established benchmarks, including uni-modal tasks, cross-modal tasks, and multi-modal tasks, where it achieves state-of-the-art performance across various tasks. MedUnifier also offers a highly adaptable tool for a wide range of language and vision tasks in healthcare, marking advancement toward the development of a generalizable AI model for medical applications. Yang Yu 0079, Yucheng Chen 0002, Xulei Yang, Si Yong Yeo |
CVPR | 4 |
| 2025 | Distribution Alignment Informed Thresholding for Semi-Supervised Curvilinear Structure SegmentationabstractCurvilinear structure segmentation using deep neural networks is often limited by the high cost of annotation. Semi-supervised learning (SSL) helps mitigate this dependency on extensive annotated data. State-of-the-art SSL approaches generate pseudo-labels for unlabeled data, which are then used for further model training. These methods primarily focus on calibrating thresholds to binarize the predictions. In this work, we assume that when labeled and unlabeled data are similar, the foreground-to-background ratio should be consistent between them. To leverage this assumption, we calibrate the threshold by minimizing the distribution gap between labeled ground truth and pseudo-labels on unlabeled data. Our proposed threshold calibration can be integrated with existing SSL methods. We evaluate its effectiveness on four datasets, demonstrating that our method outperforms current state-of-the-art SSL techniques, especially in scenarios with very low labeled data. Yuhao Mo, Bihan Wen, Xulei Yang, Ce Zhu, Xun Xu 0002 |
ICASSP | 4 |
| 2025 | Exploiting Vision Language Model for Training-Free 3D Point Cloud OOD Detection via Graph Score PropagationabstractOut-of-distribution (OOD) detection in 3D point cloud data remains a challenge, particularly in applications where safe and robust perception is critical. While existing OOD detection methods have shown progress for 2D image data, extending these to 3D environments involves unique obstacles. This paper introduces a training-free framework that leverages Vision-Language Models (VLMs) for effective OOD detection in 3D point clouds. By constructing a graph based on class prototypes and testing data, we exploit the data manifold structure to enhancing the effectiveness of VLMs for 3D OOD detection. We propose a novel Graph Score Propagation (GSP) method that incorporates prompt clustering and self-training negative prompting to improve OOD scoring with VLM. Our method is also adaptable to few-shot scenarios, providing options for practical applications. We demonstrate that GSP consistently outperforms state-of-the-art methods across synthetic and real-world datasets 3D point cloud OOD detection. Tiankai Chen, Yushu Li, Adam Goodge, Fei Teng 0001, Xulei Yang, Tianrui Li 0001, Xun Xu 0002 |
ICCV | 5 |
| 2025 | Global-Aware Monocular Semantic Scene Completion with State Space Models
Shijie Li 0006, Zhongyao Cheng, Juergen Gall, Xun Xu 0002, Xulei Yang |
ICCV | 7 |
| 2025 | FIND: Few-Shot Anomaly Inspection with Normal-Only Multi-Modal Data
Fayao Liu, Jingyi Liao, Sichao Tian, Chuan-Sheng Foo, Xulei Yang |
ICCV | 6 |
| 2025 | Future-Aware Interaction Network for Motion Forecasting
Shijie Li 0006, Xun Xu 0002, Si Yong Yeo, Xulei Yang |
ICCV | 5 |
| 2025 | Evidential Learning-based Certainty Estimation for Robust Dense Feature MatchingabstractDense feature matching methods aim to estimate a dense correspondence field between images. Inaccurate correspondence can occur due to the presence of unmatchable region, necessitating the need for certainty measurement. This is typically addressed by training a binary classifier to decide whether each predicted correspondence is reliable. However, deep neural network-based classifiers can be vulnerable to image corruptions or perturbations, making it difficult to obtain reliable matching pairs in corrupted scenario. In this work, we propose an evidential deep learning framework to enhance the robustness of dense matching against corruptions. We modify the certainty prediction branch in dense matching models to generate appropriate belief masses and compute the certainty score by taking expectation over the resulting Dirichlet distribution. We evaluate our method on a wide range of benchmarks and show that our method leads to improved robustness against common corruptions and adversarial attacks, achieving up to 10.1\% improvement under severe corruptions. Lile Cai, Chuan-Sheng Foo, Xun Xu 0002, Zaiwang Gu, Jun Cheng 0003, Xulei Yang |
ICLR | 6 |
| 2025 | On the Adversarial Risk of Test Time Adaptation: An Investigation into Realistic Test-Time Data PoisoningabstractTest-time adaptation (TTA) updates the model weights during the inference stage using testing data to enhance generalization. However, this practice exposes TTA to adversarial risks. Existing studies have shown that when TTA is updated with crafted adversarial test samples, also known as test-time poisoned data, the performance on benign samples can deteriorate. Nonetheless, the perceived adversarial risk may be overstated if the poisoned data is generated under overly strong assumptions. In this work, we first review realistic assumptions for test-time data poisoning, including white-box versus grey-box attacks, access to benign data, attack order, and more. We then propose an effective and realistic attack method that better produces poisoned samples without access to benign samples, and derive an effective in-distribution attack objective. We also design two TTA-aware attack objectives. Our benchmarks of existing attack methods reveal that the TTA methods are more robust than previously believed. In addition, we analyze effective defense strategies to help develop adversarially robust TTA methods. The source code is available at https://github.com/Gorilla-Lab-SCUT/RTTDP. Yongyi Su, Yushu Li, Nanqing Liu, Kui Jia, Xulei Yang, Chuan-Sheng Foo, Xun Xu 0002 |
ICLR | 5 |
| 2025 | Text-to-Image Rectified Flow as Plug-and-Play PriorsabstractLarge-scale diffusion models have achieved remarkable performance in generative tasks. Beyond their initial training applications, these models have proven their ability to function as versatile plug-and-play priors. For instance, 2D diffusion models can serve as loss functions to optimize 3D implicit models. Rectified Flow, a novel class of generative models, has demonstrated superior performance across various domains. Compared to diffusion-based methods, rectified flow approaches surpass them in terms of generation quality and efficiency. In this work, we present theoretical and experimental evidence demonstrating that rectified flow based methods offer similar functionalities to diffusion models — they can also serve as effective priors. Besides the generative capabilities of diffusion priors, motivated by the unique time-symmetry properties of rectified flow models, a variant of our method can additionally perform image inversion. Experimentally, our rectified flow based priors outperform their diffusion counterparts — the SDS and VSD losses — in text-to-3D generation. Our method also displays competitive performance in image inversion and editing. Code is available at: https://github.com/yangxiaofeng/rectified_flow_prior. Xulei Yang, Fayao Liu, Guosheng Lin |
ICLR | 3 |
| 2025 | OcSplats: Rendering Occluded Humans with Prior KnowledgeabstractThe task of reconstructing and rendering moving humans from monocular videos, particularly when occlusions are present, is fraught with difficulty due to insufficient visual information. Existing approaches struggle with two primary issues in delivering complete and high-fidelity rendering: the reliance on precise geometry constraints, often failing to account for occluded body parts, and the insufficient observation of unseen body parts, resulting in inconsistent reconstructions. To address these limitations, we introduce OcSplats, a deformable 3D gaussian splatting based method tailored for rendering humans in highly occluded scenarios using prior knowledge. OcSplats designs a human body prior-based geometry completion module to recover the occluded human geometry, ensuring complete reconstruction and rendering. Additionally, OcSplats employs a multi-view diffusion prior to regularize human reconstruction pipeline at novel camera poses beyond those in the occluded monocular video. We evaluate OcSplats on the ZJU-MoCap dataset and challenging OcMotion sequences, and experimental results demonstrate that OcSplats significantly outperforms existing state-of-the-art methods in rendering occluded humans. Jie Zhang 0002, Qiongjie Cui, Xulei Yang, Na Zhao 0004 |
ICME | 3 |
| 2025 | Exploring Active Learning for Label-Efficient Training of Semantic Neural Radiance FieldabstractNeural Radiance Field (NeRF) models are implicit neural scene representation methods that offer unprecedented capabilities in novel view synthesis. Semantically-aware NeRFs not only capture the shape and radiance of a scene, but also encode semantic information of the scene. The training of semantically-aware NeRFs typically requires pixel-level class labels, which can be prohibitively expensive to collect. In this work, we explore active learning as a potential solution to alleviate the annotation burden. We investigate various design choices for active learning of semantically-aware NeRF, including selection granularity and selection strategies. We further propose a novel active learning strategy that takes into account 3D geometric constraints in sample selection. Our experiments demonstrate that active learning can effectively reduce the annotation cost of training semantically-aware NeRF, achieving more than 2× reduction in annotation cost compared to random sampling. Yuzhe Zhu, Lile Cai, Kangkang Lu 0001, Fayao Liu, Xulei Yang |
ICME | 5 |
| 2025 | How Do Images Align and Complement LiDAR? Towards a Harmonized Multi-modal 3D Panoptic SegmentationabstractLiDAR-based 3D panoptic segmentation often struggles with the inherent sparsity of data from LiDAR sensors, which makes it challenging to accurately recognize distant or small objects. Recently, a few studies have sought to overcome this challenge by integrating LiDAR inputs with camera images, leveraging the rich and dense texture information provided by the latter. While these approaches have shown promising results, they still face challenges, such as misalignment during data augmentation and the reliance on post-processing steps. To address these issues, we propose Image-Assists-LiDAR (IAL), a novel multi-modal 3D panoptic segmentation framework. In IAL, we first introduce a modality-synchronized data augmentation strategy, PieAug, to ensure alignment between LiDAR and image inputs from the start. Next, we adopt a transformer decoder to directly predict panoptic segmentation results. To effectively fuse LiDAR and image features into tokens for the decoder, we design a Geometric-guided Token Fusion (GTF) module. Additionally, we leverage the complementary strengths of each modality as priors for query initialization through a Prior-based Query Generation (PQG) module, enhancing the decoder’s ability to generate accurate instance masks. Our IAL framework achieves state-of-the-art performance compared to previous multi-modal 3D panoptic segmentation methods on two widely used benchmarks. Code and models are publicly available at https://github.com/IMPL-Lab/IAL.git. Yining Pan, Qiongjie Cui, Xulei Yang, Na Zhao 0004 |
ICML | 3 |
| 2025 | EFFDNet: A Scribble-Supervised Medical Image Segmentation Method with Enhanced Foreground Feature Discrimination
Jinhua Liu 0003, Shu Yun Tan, Xulei Yang, Yanwu Xu 0004, Si Yong Yeo |
MICCAI (16) | 3 |
| 2025 | Dual Correlation-Aware Mamba for Microvascular Obstruction Identification in Non-contrast Cine Cardiac Magnetic Resonance
Yige Yan, Jun Cheng 0003, Xulei Yang, Shuang Leng, Ru-San Tan, Liang Zhong 0001, Jagath C. Rajapakse |
MICCAI (1) | 3 |
| 2025 | Spatiotemporal-Sensitive Network for Microvascular Obstruction Segmentation from Cine Cardiac Magnetic Resonance
Yang Yu 0079, Christopher Kok 0001, Jun Cheng 0003, Shuang Leng, Ru-San Tan, Liang Zhong 0001, Xulei Yang |
MICCAI (16) | 8 |
| 2025 | SODA: Out-of-Distribution Detection in Domain-Shifted Point Clouds via Neighborhood Propagation
Adam Goodge, Bryan Hooi, Jingyi Liao, Yongyi Su, Wee Siong Ng, Xun Xu 0002, Xulei Yang |
ECML/PKDD (1) | 7 |
| 2025 | AIC3DOD: Advancing Indoor Class-Incremental 3D Object Detection with Point Transformer Architecture and Room Layout ConstraintsabstractOver the recent years, there has been a growing interest in class-incremental 3D object detection based on point clouds. However, the current state-of-the-art (SOTA) methods still fall short of practical adoption, mainly due to two key observations. Firstly, existing SOTA methods are limited by the capability of feature representation from the object detection model. Secondly, these methods overlook the importance of incorporating prior information or geometry constraints, which are crucial elements for 3D point cloud tasks. In this study, we strive to enhance the performance of class-incremental 3D object detection for indoor scenes by proposing AIC3DOD - Advancing Indoor Classincremental 3D Object Detection using the point transformer architecture with room layout constraints. Our approach employs a transformer architecture in our detection model and optimizes the class incremental step in the transformer architecture. Besides, AIC3DOD incorporates additional prior information, namely room layout, to impose physical constraints on detected objects, thereby enhancing overall object detection performance. Extensive experimental results on the ScanNet dataset demonstrate the effectiveness of our approach, showcasing our superior performance compared to other SOTA methods in the class-incremental 3D object detection task. Zhongyao Cheng, Fang Wu 0009, Peisheng Qian, Ziyuan Zhao, Xulei Yang |
WACV | 5 |
| 2025 | Physically-guided open vocabulary segmentation with weighted patched alignment loss
Weide Liu, Jieming Lou, Wei Zhou 0021, Jun Cheng 0003, Xulei Yang |
Neurocomputing | 6 |
| 2025 | Multimodal multitask similarity learning for vision language model on radiological images and reports
Yang Yu 0079, Weide Liu, Ivan Ho Mien, Pavitra Krishnaswamy, Xulei Yang, Jun Cheng 0003 |
Neurocomputing | 6 |
| 2025 | Efficient Distortion-Minimized Layerwise PruningabstractIn this paper, we propose a post-training pruning framework that jointly optimizes layerwise pruning to minimize model output distortion. Through theoretical and empirical analysis, we discover an important additivity property of output distortion from pruning weights/channels in DNNs. Leveraging this property, we reformulate pruning optimization as a combinatorial problem and solve it with dynamic programming, achieving linear time complexity and making the algorithm very fast on CPUs. Furthermore, we optimize additivity-derived distortions using Hessian-based Taylor approximation to enhance pruning efficiency, accompanied by fine-grained complexity reduction techniques. Our method is evaluated on various DNN architectures, including CNNs, ViTs, and object detectors, and on vision tasks such as image classification on CIFAR-10 and ImageNet, and 3D object detection and various datasets. We achieve SoTA with significant FLOPs reductions without accuracy loss. Specifically, on CIFAR-10, we achieve up to $27.9\times$27.9×, $29.2\times$29.2×, and $14.9\times$14.9× FLOPs reductions on ResNet-32, VGG-16, and DenseNet-121, respectively. On ImageNet, we observe no accuracy loss with $1.69\times$1.69× and $2\times$2× FLOPs reductions on ResNet-50 and DeiT-Base, respectively. For 3D object detection, we achieve $\mathbf {3.89}\times, \mathbf {3.72}\times$3.89×,3.72× FLOPs reductions on CenterPoint and PVRCNN models. These results demonstrate the effectiveness and practicality of our approach for improving model performance through layer-adaptive weight pruning. Kaixin Xu, Zhe Wang 0019, Runtao Huang, Xue Geng, Jie Lin 0001, Xulei Yang, Min Wu 0008, Xiaoli Li 0001, Weisi Lin |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Mitigating Missing Feature Channels at Inference Stage: Test-Time Adaptation Through Self-Training With Data ImputationabstractThe robustness of deep learning model can be compromised by out-of-distribution (OOD) testing data. Test-time adaptation (TTA) emerges as an efficient method to mitigate the distribution gap by tuning model weights at inference stage. TTA are mainly demonstrated on robustifying model on additive visual corruptions or adversarial attacks. In this work, we specify an overlooked type of OOD where feature channels could be missing in testing data, potentially due to sensor fault. We reveal that self-training and data imputation can improve the model’s generalization to data with missing feature channel. To address the uncertainty associated with imputed samples, we fuse predictions from imputed and weakly-augmented samples for more reliable pseudo labels. We evaluate the effectiveness on multiple image classification benchmarks with synthesized and realistic missing feature channels, and our proposed method outperforms state-of-the-art TTA methods on all benchmarks. Yongyi Su, Xulei Yang, Xun Xu 0002 |
IEEE Signal Process. Lett. | 3 |
| 2025 | SDCoT++: Improved Static-Dynamic Co-Teaching for Class-Incremental 3D Object DetectionabstractDeep learning approaches have demonstrated high effectiveness in 3D object detection tasks. However, they often suffer from a notable drop in performance on the previously trained classes when learning new classes incrementally without revisiting the old data. This is the "catastrophic forgetting" phenomenon which impedes 3D object detection in real-world scenarios, where intelligent machines must continuously learn to detect previously unseen categories. Furthermore, frequent co-occurrences of old and new classes in scenes exacerbate catastrophic forgetting and cause model confusion. To address these challenges, we propose a novel static-dynamic co-teaching approach. Our framework involves a student model and two teacher models: a static teacher with fixed weights which imparts preserved old knowledge to the student, and a dynamic teacher with continuously updated weights which transfers underlying knowledge from new data to the student. To mitigate the issue of co-occurrence, we generate pseudo labels for base (i.e. old) classes from both static and dynamic sources during incremental learning. Additionally, to mitigate the negative impact of varying occurrence frequencies of classes on fixed thresholding during the selection of pseudo labels, we calibrate the probabilities of base classes to attain more balanced class probabilities. Moreover, our static-dynamic co-teaching framework is backbone-agnostic, making it compatible with different detection architectures. We demonstrate its backbone-agnostic nature by adapting three representative 3D object detectors: VoteNet, 3DETR and CAGroup3D. Extensive experiments showcase the superior performance of our proposed method compared to baseline approaches across indoor and outdoor benchmark datasets and applicability with different backbone models. Na Zhao 0004, Peisheng Qian, Fang Wu 0009, Xun Xu 0002, Xulei Yang, Gim Hee Lee |
IEEE Trans. Image Process. | 5 |
| 2025 | DPPNet: A Depth Pixel-Wise Potential-Aware Network for RGB-D Salient Object DetectionabstractDepth cues are essential for visual perception tasks like Salient Object Detection (SOD). Due to varying depth reliability across scenes, some researchers propose evaluating the overall quality of the depth maps and discarding the less reliable ones to avoid contamination. However, these methods often fail to fully utilize valuable information in depth maps, leading to sub-optimal performance particularly when depth quality is unreliable. Since low-quality depth maps still contain useful information that potentially improves model performance, we propose a Depth Pixel-wise Potential-aware Network to leverage these depth cues effectively. This network includes two novel components designed: 1) A learning strategy for explicitly modeling the confidence of each depth pixel to assist the model in locating valid information in the depth map. 2) A cross-modal adaptive multiple fusion module that fuses features from both RGB and depth modalities. It aims to mitigate the contamination effect of unreliable depth maps and fully exploit the benefits of multiple fusion strategies. Experimental results show that on four publicly available datasets, our method outperforms 17 mainstream methods on various evaluation metrics. Junbin Yuan, Zhoutao Wang, Qingzhen Xu, Bharadwaj Veeravalli, Xulei Yang |
IEEE Trans. Multim. | 6 |
| 2025 | From Algorithm to Hardware: A Survey on Efficient and Safe Deployment of Deep Neural NetworksabstractDeep neural networks (DNNs) have been widely used in many artificial intelligence (AI) tasks. However, deploying them brings significant challenges due to the huge cost of memory, energy, and computation. To address these challenges, researchers have developed various model compression techniques such as model quantization and model pruning. Recently, there has been a surge in research on compression methods to achieve model efficiency while retaining performance. Furthermore, more and more works focus on customizing the DNN hardware accelerators to better leverage the model compression techniques. In addition to efficiency, preserving security and privacy is critical for deploying DNNs. However, the vast and diverse body of related works can be overwhelming. This inspires us to conduct a comprehensive survey on recent research toward the goal of high-performance, cost-efficient, and safe deployment of DNNs. Our survey first covers the mainstream model compression techniques, such as model quantization, model pruning, knowledge distillation, and optimizations of nonlinear operations. We then introduce recent advances in designing hardware accelerators that can adapt to efficient model compression approaches. In addition, we discuss how homomorphic encryption can be integrated to secure DNN deployment. Finally, we discuss several issues, such as hardware evaluation, generalization, and integration of various compression approaches. Overall, we aim to provide a big picture of efficient DNNs from algorithm to hardware accelerators and security perspectives. Xue Geng, Zhe Wang 0019, Chunyun Chen, Qing Xu 0015, Kaixin Xu, Jin Chao, Manas Gupta, Xulei Yang, Zhenghua Chen, Mohamed M. Sabry, Jie Lin 0001, Min Wu 0008, Xiaoli Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2024 | Learning Intra-View and Cross-View Geometric Knowledge for Stereo MatchingabstractGeometric knowledge has been shown to be beneficial for the stereo matching task. However, prior attempts to in-tegrate geometric insights into stereo matching algorithms have largely focused on geometric knowledge from single images while crucial cross-view factors such as occlusion and matching uniqueness have been overlooked. To address this gap, we propose a novel Intra-view and Cross-view Geometric knowledge learning Network (ICGNet), specifically crafted to assimilate both intra-view and cross-view geo-metric knowledge. ICGNet harnesses the power of interest points to serve as a channel for intra-view geometric understanding. Simultaneously, it employs the correspon-dences among these points to capture cross-view geometric relationships. This dual incorporation empowers the proposed ICGNet to leverage both intra-view and cross-view geometric knowledge in its learning process, substantially improving its ability to estimate disparities. Our extensive experiments demonstrate the superiority of the ICGNet over contemporary leading models. The code will be available at https://github.com/DFSDDDDDl199/ICGNet. Weide Liu, Zaiwang Gu, Xulei Yang, Jun Cheng 0003 |
CVPR | 4 |
| 2024 | LPViT: Low-Power Semi-structured Pruning for Vision Transformers
Kaixin Xu, Zhe Wang 0019, Chunyun Chen, Xue Geng, Jie Lin 0001, Xulei Yang, Min Wu 0008, Xiaoli Li 0001, Weisi Lin |
ECCV (71) | 6 |
| 2024 | Learn to Optimize Denoising Scores: A Unified and Improved Diffusion Prior for 3D Generation
Chi Zhang 0007, Yi Xu 0002, Xulei Yang, Fayao Liu, Guosheng Lin |
ECCV (44) | 6 |
| 2024 | Coarse-Grained Mask Regularization for Microvascular Obstruction Identification from Non-contrast Cardiac Magnetic Resonance
Yige Yan, Jun Cheng 0003, Xulei Yang, Zaiwang Gu, Shuang Leng, Ru-San Tan, Liang Zhong 0001, Jagath C. Rajapakse |
MICCAI (1) | 3 |
| 2024 | See, Predict, Plan: Diffusion for Procedure Planning in Robotic Surgical Videos
Ziyuan Zhao, Fen Fang, Xulei Yang, Qianli Xu, Cuntai Guan, Shaohua Kevin Zhou |
MICCAI (6) | 3 |
| 2024 | On-the-fly Point Feature Representation for Point Clouds AnalysisabstractPoint cloud analysis is challenging due to its unique characteristics of unorderness, sparsity and irregularity. Prior works attempt to capture local relationships by convolution operations or attention mechanisms, exploiting geometric information from coordinates implicitly. These methods, however, are insufficient to describe the explicit local geometry, e.g., curvature and orientation. In this paper, we propose On-the-fly Point Feature Representation (OPFR), which captures abundant geometric information explicitly through Curve Feature Generator module. This is inspired by Point Feature Histogram (PFH) from computer vision community. However, the utilization of vanilla PFH encounters great difficulties when applied to large datasets and dense point clouds, as it demands considerable time for feature generation. In contrast, we introduce the Local Reference Constructor module, which approximates the local coordinate systems based on triangle sets. Owing to this, our OPFR only requires extra 1.56ms for inference (65X faster than vanilla PFH) and 0.012M more parameters, and it can serve as a versatile plug-and-play module for various backbones, particularly MLP-based and Transformer-based backbones examined in this study. Additionally, we introduce the novel Hierarchical Sampling module aimed at enhancing the quality of triangle sets, thereby ensuring robustness of the obtained geometric features. Our proposed method improves overall accuracy (OA) on ModelNet40 from 90.7% to 94.5% (+3.8%) for classification, and OA on S3DIS Area-5 from 86.4% to 90.0% (+3.6%) for semantic segmentation, respectively, building upon PointNet++ backbone. When integrated with Point Transformer backbone, we achieve state-of-the-art results on both tasks: 94.8% OA on ModelNet40 and 91.7% OA on S3DIS Area-5. Jiangyi Wang, Zhongyao Cheng, Na Zhao 0004, Jun Cheng 0003, Xulei Yang |
ACM Multimedia | 5 |
| 2024 | Training Binary Neural Networks via Gaussian Variational Inference and Low-Rank Semidefinite ProgrammingabstractCurrent methods for training Binarized Neural Networks (BNNs) heavily rely on the heuristic straight-through estimator (STE), which crucially enables the application of SGD-based optimizers to the combinatorial training problem. Although the STE heuristics and their variants have led to significant improvements in BNN performance, their theoretical underpinnings remain unclear and relatively understudied. In this paper, we propose a theoretically motivated optimization framework for BNN training based on Gaussian variational inference. In its simplest form, our approach yields a non-convex linear programming formulation whose variables and associated gradients motivate the use of latent weights and STE gradients. More importantly, our framework allows us to formulate semidefinite programming (SDP) relaxations to the BNN training task. Such formulations are able to explicitly models pairwise correlations between weights during training, leading to a more accurate optimization characterization of the training problem. As the size of such formulations grows quadratically in the number of weights, quickly becoming intractable for large networks, we apply the Burer-Monteiro approach and only optimize over linear-size low-rank SDP solutions. Our empirical evaluation on CIFAR-10, CIFAR-100, Tiny-ImageNet and ImageNet datasets shows our method consistently outperforming all state-of-the-art algorithms for training BNNs. Lorenzo Orecchia, Xue He, Wang Mark, Xulei Yang, Min Wu 0008, Xue Geng |
NeurIPS | 5 |
| 2023 | Efficient Practices for Profile-to-Frontal Face Synthesis and RecognitionabstractDespite the great progress of deep learning and generative adversarial networks, face frontalization (i.e., profile-to-frontal synthesis) and profile (i.e., non-frontal) face recognition still remain challenging tasks under uncontrolled environments. In this study, we propose three efficient practices to improve the performance of profile-to-frontal face synthesis and recognition. Firstly, the identity preserving module is embedded to constrain synthesized frontal images similar to true frontal faces in feature space. Secondly, facial consistency loss is employed to reduce the artifact of the generated frontal face in pixel space. Lastly, the multi-model embedded scheme enhances the representation learning through diverse facial features extracted by multiple facial feature extractors. The proposed practices are general, though specifically deployed to CR-GAN in this study for performance verification. Experimental results on Multi-PIE and VGGFace2 demonstrate that the proposed practices qualitatively generate more realistic photography frontal faces and quantitatively obtain better face recognition accuracy. Huijiao Wang, Xulei Yang |
ICASSP | 2 |
| 2023 | COCO-TEACH: A Contrastive Co-Teaching Network For Incremental 3D Object DetectionabstractDeep learning (DL) models for 3D object detection from point clouds have shown remarkable progress in various autonomous perception scenarios. However, the issue of catastrophic forgetting seriously hinders the deployment of these models in real-world applications where new classes are encountered over time. In order to address this issue, we present the Contrastive Co-Teaching Network (COCO-TEACH) framework for class-incremental 3D object detection. Our proposed framework consists of two teacher networks: a primary teacher network that detects old class objects in new data and provides them with pseudo-labels and an auxiliary teacher network that leverages the unlabelled objects in new data. The two teacher models transfer their learned knowledge to the target student model through a class-aware consistency loss. To enhance this transfer, a supervised contrastive loss is further incorporated into the loss function. We evaluate the performance of our proposed method against baseline methods through extensive experiments on two benchmark datasets. The results show that our proposed framework achieves state-of-the-art performance on incremental 3D object detection. Zhongyao Cheng, Cen Chen 0002, Ziyuan Zhao, Peisheng Qian, Xiaoli Li 0001, Xulei Yang |
ICIP | 6 |
| 2023 | An Efficient Deep Video Model For Deepfake DetectionabstractThe use of deep learning technology to manipulate images and videos of people in ways that are difficult to distinguish from the real ones, known as deepfake, has become a matter of national security concern in recent years. As a result, many studies have been carried out to detect deepfake and manipulated media. Among these studies, deep video models based on convolutional neural networks have been the preferred method for detecting deepfake in videos. This study presents a novel deep video model called Sequential-Parallel Networks (SPNet) that provides efficient deepfake detection. The SPNet model consists of a simple yet innovative sequential-parallel block that first extracts spatial and temporal features sequentially, then concatenates them together in parallel. As a result, the presented SPNet possesses comparable spatiotemporal modeling abilities as most state-of-the-art deep video methods but with lower computation complexity and fewer parameters. The efficiency of the presented SPNet is demonstrated on a large-scale deepfake benchmark in terms of high recognition accuracy and low computational cost. Ruipeng Sun, Ziyuan Zhao, Zeng Zeng, Bharadwaj Veeravalli, Xulei Yang |
ICIP | 7 |
| 2023 | Controlling Facial Attribute Synthesis by Disentangling Attribute Feature Axes in Latent SpaceabstractIn this study, we propose a novel approach to synthesize high-resolution and hyper-realistic face images with controlled attributes. Firstly, by training an attribute classifier to assign attribute labels to given synthesized face images, we build the links between latent vectors and face attributes. Secondly, we adapt the regression method to match the distributions of latent vectors with the corresponding face attributes, to control the attribute synthesis in the face images. Finally, we use the Gram-Schmidt orthogonalization algorithm to disentangle the attribute feature axes in latent space, such that a change in one attribute will not cause any changes in other attributes. Extensive experiments demonstrate the effectiveness of the proposed approach for high-quality face image synthesis with controlled attributes. Qiyu Wei, Zhongyao Cheng, Zeng Zeng, Xulei Yang |
ICIP | 6 |
| 2023 | SemiGNN-PPI: Self-Ensembling Multi-Graph Neural Network for Efficient and Generalizable Protein-Protein Interaction PredictionabstractProtein-protein interactions (PPIs) are crucial in various biological processes and their study has significant implications for drug development and disease diagnosis. Existing deep learning methods suffer from significant performance degradation under complex real-world scenarios due to various factors, e.g., label scarcity and domain shift. In this paper, we propose a self-ensembling multi-graph neural network (SemiGNN-PPI) that can effectively predict PPIs while being both efficient and generalizable. In SemiGNN-PPI, we not only model the protein correlations but explore the label dependencies by constructing and processing multiple graphs from the perspectives of both features and labels in the graph learning process. We further marry GNN with Mean Teacher to effectively leverage unlabeled graph-structured PPI data for self-ensemble graph learning. We also design multiple graph consistency constraints to align the student and teacher graphs in the feature embedding space, enabling the student model to better learn from the teacher model by incorporating more relationships. Extensive experiments on PPI datasets of different scales with different evaluation settings demonstrate that SemiGNN-PPI outperforms state-of-the-art PPI prediction methods, particularly in challenging scenarios such as training with limited annotations and testing on unseen data. Ziyuan Zhao, Peisheng Qian, Xulei Yang, Zeng Zeng, Cuntai Guan, Tam Wai Leong, Xiaoli Li 0001 |
IJCAI | 3 |
| 2023 | DGSLN: Differentiable graph structure learning neural network for robust graph representations
Xiaofeng Zou, Kenli Li 0001, Cen Chen 0002, Xulei Yang, Wei Wei 0006, Keqin Li 0001 |
Inf. Sci. | 4 |
| 2023 | Non-cooperative game algorithms for computation offloading in mobile edge computing environments
Jianguo Chen 0001, Qingying Deng, Xulei Yang |
J. Parallel Distributed Comput. | 3 |
| 2023 | Toward Communication-Efficient Digital Twin via AI-Powered Transmission and ReconstructionabstractDigital twin technology has recently gathered pace in engineering communities as it allows for the convergence of the real structure and its digital counterpart. 3D point cloud data is a more effective way to describe the real world and to reconstruct the digital counterpart than the conventional 2D images or 360-degree images. Large-scale, e.g., city-scale digital twins, typically collect point cloud data via internet-of-things (IoT) devices and transmit it over wireless networks. However, the existing wireless transmission technology can not carry real-time point cloud transmission for digital twin reconstruction due to mass data volume, high processing overheads, and low delay-tolerance. We propose a novel artificial intelligence (AI) powered end-to-end framework, termed AIRec, for efficient digital twin communication from point cloud compression, wireless channel coding, and digital twin reconstruction. AIRec adopts the encoder-decoder architecture. In the encoder, a novel importance-aware pooling scheme is designed to adaptively select important points with learnable thresholds to reduce the transmission volume. We also design a novel noise-aware joint source and channel coding is proposed to adaptively adjust the transmission strategy based on SNR and map the features to error-resilient channel symbols for wireless transmission to achieve a good tradeoff between the transmission rate and reconstruction quality. The decoder can accurately reconstruct the digital twins from the received symbols. Extensive experiments of typical datasets and comparison with baselines show that we achieve a good reconstruction quality under$24\times $compression ratio. Cen Chen 0002, Xulei Yang, Joey Tianyi Zhou, Tao Zhang 0019, Yangfan Li 0001 |
IEEE J. Sel. Areas Commun. | 3 |
| 2023 | GCM: Efficient video recognition with glance and combine module
Ziyuan Huang 0003, Xulei Yang, Marcelo H. Ang, Teck Khim Ng |
Pattern Recognit. | 3 |
| 2022 | MANET: Mitral Annulus Point Tracking Network in Cardiac Magnetic ResonanceabstractCardiac magnetic resonance (CMR) imaging is frequently recommended for patients at intermediate risk of cardiovascular disease to triage them for medication or invasive aggressive treatment. Mitral annulus (MA) motion and velocities represent the cardiac contraction and relaxation, and hold potential to improve the detection of subtle cardiac dysfunction. However, conventional interpretation of CMR images requires expert manipulation and is often operator-dependent. In this paper, we propose an end-to-end MA Point Tracking Network (MANet) to automatically detect and track MA motion during cardiac cycle. The MANet model consists of MA point detection module and motion tracking module. In MA point detection, we design the convolutional-based feature extraction and elastic regression to detect MA points frame by frame of each CMR video. Then, in MA tracking, we adopt the Deep SORT model to capture spatio-temporal continuity between frames and fine-tune the coordinate position of MA points. 171 CMR videos with 4275 frames are used in comparison experiments, and the results demonstrate that our MANet model achieves promising performance in reference to clinical ground truth (r=0.71, P<0.001). This work provides an important preamble for cardiac motion tracking and cardiac function evaluation. Jianguo Chen 0001, Xulei Yang, Shuang Leng, Ru-San Tan, Zeng Zeng, Liang Zhong 0001 |
ICIP | 2 |
| 2022 | Latent Vector Prototypes Guided Conditional Face SynthesisabstractRecent advances in deep neural networks, especially in generative adversarial networks (GAN), have shown remarkable progress in face image generations. However, most of the existing face image generators can only synthesize random face images, but are not able to control the attributes of the generated face images. Though conditional GAN based methods can manipulate the attributes to some extent, but can only generate low-resolution face images up to 256 × 256. In this study, based on StyleGAN, one of the state-of-the-art image generators for synthesizing high-quality face images, we propose a simple but efficient approach to generate high-resolution and hyper-realistic face images with any desired attribute. By training an attribute classifier to assign attribute labels to given synthesized face images, we build the links between latent vectors and face attributes. In such a way, the latent vectors can be grouped into different clusters, one cluster corresponding to one face attribute, respectively. We then extract the prototypes for the clusters, which are used to control the attribute of the generated face image. Extensive experiments demonstrate the effectiveness of the proposed approach for high-quality face image generation with predefined attributes. Qiyu Wei, Xulei Yang, Tong Sang, Huijiao Wang, Xiaofeng Zou, Zhongyao Cheng, Ziyuan Zhao, Zeng Zeng |
ICIP | 2 |
| 2022 | Iterative Contrastive Learning for Single Image Raindrop RemovalabstractDeep learning has achieved remarkable progress in computer vision and image analysis. However, raindrop removal from single image still remains challenging, due to a wide range of raindrop diversities and surface reflections. In this paper, we propose an iterative neural network with feedback strategy and contrastive learning for single image raindrop removal. First, we design an iterative feedback neural network to refine low-level representations with high-level information, i.e., the output of the previous iteration is used as input for the next iteration, together with the input image with raindrops. As a result, raindrops could be gradually removed through this feedback manner. Then, we deploy contrastive regularization to push the restored image from each iteration close to the clean images without raindrops, but away from rainy images with raindrops. Extensive experiments on two raindrop benchmark datasets demonstrate the effectiveness of the proposed approach in comparison with the state-of-the-art methods. The methodology in this work could be further extended to self-supervised contrastive learning to obtain robust feature representations with less labelled data. Xulei Yang, Peisheng Qian, Li Wang 0057, Cen Chen 0001, Xiaoli Li 0001, Zeng Zeng |
ICIP | 1 |
| 2022 | MMGL: Multi-Scale Multi-View Global-Local Contrastive Learning for Semi-Supervised Cardiac Image SegmentationabstractWith large-scale well-labeled datasets, deep learning has shown significant success in medical image segmentation. However, it is challenging to acquire abundant annotations in clinical practice due to extensive expertise requirements and costly labeling efforts. Recently, contrastive learning has shown a strong capacity for visual representation learning on unlabeled data, achieving impressive performance rivaling supervised learning in many domains. In this work, we propose a novel multi-scale multi-view global-local contrastive learning (MMGL) framework to thoroughly explore global and local features from different scales and views for robust contrastive learning performance, thereby improving segmentation performance with limited annotations. Extensive experiments on the MM-WHS dataset demonstrate the effectiveness of MMGL framework on semi-supervised cardiac image segmentation, outperforming the state-of-the-art contrastive learning methods by a large margin. Ziyuan Zhao, Jinxuan Hu, Zeng Zeng, Xulei Yang, Peisheng Qian, Bharadwaj Veeravalli, Cuntai Guan |
ICIP | 4 |
| 2022 | Algorithms and architecture support of degree-based quantization for graph neural networks
Yilong Guo, Xiaofeng Zou, Xulei Yang, Yuandong Gu |
J. Syst. Archit. | 4 |
| 2022 | A configurable deep learning framework for medical image analysis
Jianguo Chen 0001, Mimi Zhou, Zhaolei Zhang, Xulei Yang |
Neural Comput. Appl. | 5 |
| 2021 | Joint Anomaly Detection and Inpainting for Microscopy Images Via Deep Self-Supervised LearningabstractWhile microscopy enables material scientists to view and analyze microstructures, the imaging results often include defects and anomalies with varied shapes and locations. The presence of such anomalies significantly degrades the quality of microscopy images and the subsequent analytical tasks. Comparing to classic feature-based methods, recent advancements in deep learning provide a more efficient, accurate, and scalable approach to detect and remove anomalies in microscopy images. However, most of the deep inpainting and anomaly detection schemes require a certain level of supervision, i.e., either annotation of the anomalies, or a corpus of purely normal data, which are limited in practice for supervision-starving microscopy applications. In this work, we propose a self-supervised deep learning scheme for joint anomaly detection and inpainting of microscopy images. The proposed anomaly detection model can be trained over a mixture of normal and abnormal microscopy images without any labeling. Instead of a two-stage scheme, our multi-task model can simultaneously detect abnormal regions and remove the defects via jointly training. To benchmark such microscopy application under the real-world setup, we propose a novel dataset of real microscopic images of integrated circuits, dubbed MIIC. The proposed dataset contains tens of thousands of normal microscopic images, while we labeled hundreds of them containing various imaging and manufacturing anomalies and defects for testing. Experiments show that the proposed model outperforms various popular or state-of-the-art competing methods for both microscopy image anomaly detection and inpainting. Deruo Cheng, Xulei Yang, Tong Lin 0001, Yiqiong Shi, Kaiyi Yang, Bah-Hwee Gwee, Bihan Wen |
ICIP | 3 |
| 2021 | Systematic Analysis of Circular Artifacts for StyleganabstractRecent research works have pointed out that the synthesized images by StyleGAN contain prominent circular artifacts which severely degrade the quality of generated images. In this work, we provide a systematic investigation on how those circular artifacts are formed by studying the functionalities of different modules that are used in the Style-GAN architecture. We present both analysis of the StyleGAN mechanism and extensive experiments to verify our claims. The key modules of StyleGAN that promote such undesired artifacts are highlighted based on the analysis. Besides, we propose a simple yet effective solution to remove the prominent circular artifacts for StyleGAN, by applying a simple but efficient pixel-instance normalization layer. The improved StyleGAN model trained via our proposed approach successfully prevents the appearance of circular artifacts in the generated images. Way Tan, Bihan Wen, Cen Chen 0001, Zeng Zeng, Xulei Yang |
ICIP | 5 |
| 2021 | Multiple local 3D CNNs for region-based prediction in smart cities
Yibi Chen, Xiaofeng Zou, Kenli Li 0001, Keqin Li 0001, Xulei Yang, Cen Chen 0002 |
Inf. Sci. | 5 |
| 2020 | Facial Feature Embedded Cyclegan For Vis-Nir TranslationabstractVisible and near-infrared (VIS-NIR) face recognition remains a challenging task due to distinctions between spectral components of two modalities. Inspired by the CycleGAN, this paper presents a method aiming to translate between VIS and NIR face images. To achieve this, we propose a new facial feature embedded CycleGAN. Firstly, to learn the particular feature while preserving common facial representation between VIS and NIR domains, we employ a general facial feature extractor (FFE) to extract effective features. Herein the MobileFaceNet is pre-trained on a VIS face database and serves as the FFE. Secondly, the domain-invariant feature learning is enhanced by proposing a new pixel consistency loss. Lastly, we establish a new WHU VIS-NIR database including varies in face rotation and expressions to enrich the training data. Experimental results on the Oulu-CASIA and our WHU VIS-NIR databases show that the proposed FFE-based CycleGAN (FFE-CycleGAN) outperforms some state-of-the-art methods and achieves 96.5% accuracy. Huijiao Wang, Lei Yu 0006, Li Wang 0057, Xulei Yang |
ICASSP | 5 |
| 2020 | Object Tracking Via ImageNet Classification ScoresabstractObject tracking is a challenging task in computer vision. The correlation filter based trackers are widely used for visual tracking due to their efficiencies. However, they cannot handle occlusion very well. In this paper, an effective method is proposed for occlusion detection based on high-level classification scores from the Convolutional Neural Network (CNN) trained on the ImageNet dataset. Also, we propose a novel tracking method by holistically considering multiple tracking models trained previously. In each frame, multiple correlation filters are first trained using hierarchical convolutional features, and then progressively selected according to the so-called tracking quality (status). Finally, a linear motion model is adopted to effectively re-detect the lost target. Experimental results have demonstrated that our method achieved good performance for handling occlusion. Li Wang 0057, Ting Liu 0009, Bing Wang 0003, Jie Lin 0001, Xulei Yang, Gang Wang 0012 |
ICIP | 5 |
| 2020 | Opencc - an open Benchmark data set for Corpus Callosum Segmentation and EvaluationabstractNeuroimaging studies have revealed that the structural changes of the corpus callosum (CC) are evident in a variety of neurological diseases, such as epilepsy and autism. Segmentation of the CC from magnetic resonance images (MRI) of the brain is a crucial step in the diagnosis of various brain disorders. However, the lack of open benchmark CC datasets has hindered development of CC segmentation techniques. In this work, we present an open benchmark dataset - OpenCC - for CC segmentation and evaluation. The dataset was built through alternative application of automatic segmentation and manual refinement. The automatic segmentation is based on recent advances in deep learning - fully convolutional networks, specifically U-Net, while the manual refinement is done by domain radiologists. The resulting dataset consists of 4643 mid-sagittal (or near mid-sagittal) slices and their corresponding CC masks. Furthermore, we provided some baseline segmentation results on the OpenCC dataset by using two latest deep learning segmentation approaches. The OpenCC dataset can be used for comparison and evaluation of newly developed CC segmentation algorithms. We endeavor that, through the publishing of the OpenCC dataset and baseline segmentation results, we could promote further development of CC segmentation techniques. Xulei Yang, Gabriel Tjio, Cen Chen 0001, Li Wang 0057, Bihan Wen, Yi Su 0001 |
ICIP | 1 |
| 2020 | Automatic detection of anatomical landmarks in brain MR scanning using multi-task deep neural networks
Xulei Yang, Wai Teng Tang, Gabriel Tjio, Si Yong Yeo, Yi Su 0001 |
Neurocomputing | 1 |
| 2019 | Learning Hierarchical Features for Visual Object Tracking With Recursive Neural NetworksabstractRecently, deep learning has achieved very promising results in visual object tracking. Deep neural networks in existing tracking methods require a lot of training data to learn a large number of parameters. However, training data is not sufficient for visual object tracking as annotations of a target object are only available in the first frame of a test sequence. In this paper, we propose to learn hierarchical features for visual object tracking by using tree structure based Recursive Neural Networks (RNN), which have a relatively small number of parameters compared to other deep neural networks (e.g. Convolutional Neural Networks (CNN)) due to all basic modules in RNN share only one set of parameters. Experimental results demonstrate that our feature learning algorithm can significantly improve tracking performance on benchmark datasets. Li Wang 0057, Ting Liu 0009, Bing Wang 0003, Jie Lin 0001, Xulei Yang, Gang Wang 0012 |
ICIP | 5 |
| 2019 | SeSe-Net: Self-Supervised deep learning for segmentation
Zeng Zeng, Xulei Yang, Yu Qiyun, Le Zhang 0001 |
Pattern Recognit. Lett. | 2 |
| 2018 | Exploiting Spatio-Temporal Correlations with Multiple 3D Convolutional Neural Networks for Citywide Vehicle Flow PredictionabstractPredicting vehicle flows is of great importance to traffic management and public safety in smart cities, and very challenging as it is affected by many complex factors, such as spatio-temporal dependencies with external factors (e.g., holidays, events and weather). Recently, deep learning has shown remarkable performance on traditional challenging tasks, such as image classification, due to its powerful feature learning capabilities. Some works have utilized LSTMs to connect the high-level layers of 2D convolutional neural networks (CNNs) to learn the spatio-temporal features, and have shown better performance as compared to many classical methods in traffic prediction. However, these works only build temporal connections on the high-level features at the top layer while leaving the spatio-temporal correlations in the low-level layers not fully exploited. In this paper, we propose to apply 3D CNNs to learn the spatio-temporal correlation features jointly from low-level to high-level layers for traffic data. We also design an end-to-end structure, named as MST3D, especially for vehicle flow prediction. MST3D can learn spatial and multiple temporal dependencies jointly by multiple 3D CNNs, combine the learned features with external factors and assign different weights to different branches dynamically. To the best of our knowledge, it is the first framework that utilizes 3D CNNs for traffic prediction. Experiments on two vehicle flow datasets Beijing and New York City have demonstrated that the proposed framework, MST3D, outperforms the state-of-the-art methods. Cen Chen 0002, Kenli Li 0001, Sin G. Teo, Guizi Chen, Xiaofeng Zou, Xulei Yang, Ramaseshan C. Vijay, Jiashi Feng, Zeng Zeng |
ICDM | 6 |
| 2018 | Deep Learning for Practical Image Recognition: Case Study on Kaggle CompetitionsabstractIn past years, deep convolutional neural networks (DCNN) have achieved big successes in image classification and object detection, as demonstrated on ImageNet in academic field. However, There are some unique practical challenges remain for real-world image recognition applications, e.g., small size of the objects, imbalanced data distributions, limited labeled data samples, etc. In this work, we are making efforts to deal with these challenges through a computational framework by incorporating latest developments in deep learning. In terms of two-stage detection scheme, pseudo labeling, data augmentation, cross-validation and ensemble learning, the proposed framework aims to achieve better performances for practical image recognition applications as compared to using standard deep learning methods. The proposed framework has recently been deployed as the key kernel for several image recognition competitions organized by Kaggle. The performance is promising as our final private scores were ranked 4 out of 2293 teams for fish recognition on the challenge "The Nature Conservancy Fisheries Monitoring" and 3 out of 834 teams for cervix recognition on the challenge "Intel &MobileODT Cervical Cancer Screening", and several others. We believe that by sharing the solutions, we can further promote the applications of deep learning techniques. Xulei Yang, Zeng Zeng, Sin G. Teo, Li Wang 0057, Vijay Chandrasekhar 0001, Steven C. H. Hoi |
KDD | 1 |
| 2018 | Multi-target deep neural networks: Theoretical analysis and implementation
Zeng Zeng, Nanying Liang, Xulei Yang, Steven C. H. Hoi |
Neurocomputing | 3 |
| 2017 | Deep convolutional neural networks for automatic segmentation of left ventricle cavity from cardiac magnetic resonance imagesabstractThis work conducts a feasibility study of deep learning approaches for automatic segmentation of left ventricle (LV) cavity from cardiac magnetic resonance (CMR) images. Automatic LV cavity segmentation is a challenging task, partially due to the small size of the object as compared to the large CMR image background, especially at the apex. To cater for small object segmentation, the authors present a localisation‐segmentation framework, to first locate the object in the large full image, then segment the object within the small cropped region of interest. The localisation is performed by a deep regression model based on convolutional neural networks, while the segmentation is done by the deep neural networks based on U‐Net architecture. They also employ the Dice loss function for the training process of the segmentation models, to investigate its effects on the segmentation performance. The deep learning models are trained and evaluated by using public endocardium‐annotated CMR datasets from York University and MICCAI 2009 LV Challenge websites. The average dice metric values of the authors’ proposed framework are 0.91 and 0.93, respectively, on these two databases. These results are promising as compared to the best results achieved by the current state‐of‐art, which shows the potentials of deep learning approaches for this particular application. Xulei Yang, Zeng Zeng, Yi Su 0001 |
IET Comput. Vis. | 1 |
| 2016 | Cardiac image segmentation by random walks with dynamic shape constraintabstractThe quantitative analysis of the left ventricle (LV) contractile function is one of the key steps in the assessment of cardiovascular disease. Such analysis greatly depends on the accurate delineation of LV boundary from cardiac sequences. However, segmentation of the LV still remains a challenging problem due to its subtle boundary, occlusion, and image inhomogeneity. To overcome such difficulties, the authors propose a novel segmentation method by incorporating a dynamic shape constraint into the weighting function of the random walks segmentation algorithm. This approach involves iterative updates on the intermediate result to achieve the desired solution. The inclusion of a shape constraint restricts the solution space of the segmentation result to handle misleading information that may come from noise, weak boundaries and clutter, leading to increased robustness of the algorithm. The authors describe the details of the proposed method and demonstrate its effectiveness in segmenting the LV from real cardiac magnetic resonance (CMR) image sets. The experimental results demonstrate that the proposed method obtains better segmentation performance than the standard method. Xulei Yang, Yi Su 0001, Rubing Duan, Haijin Fan, Si Yong Yeo, Calvin Chi-Wan Lim, Liang Zhong 0001, Ru-San Tan |
IET Comput. Vis. | 1 |
| 2015 | Kernel online learning algorithm with state feedbacks
Haijin Fan, Qing Song 0001, Xulei Yang |
Knowl. Based Syst. | 3 |
| 2014 | Vicinal support vector classifier using supervised kernel-based clustering
Xulei Yang, Aize Cao, Qing Song 0001, Gerald Schaefer, Yi Su 0001 |
Artif. Intell. Medicine | 1 |
| 2013 | Right Ventricle Segmentation by Temporal Information Constrained Gradient Vector FlowabstractEvaluation of right ventricular (RV) structure and function is of importance in the management of most cardiac disorders. But the segmentation of RV has always been considered challenging due to low contrast of the myocardium with surrounding and high shape variability of the RV. In this paper, we present a 2D + T active contour model for segmentation and tracking of RV endocardium on cardiac magnetic resonance (MR) images. To take into account the temporal information between adjacent frames, we propose to integrate the time-dependent constraints into the energy functional of the classical gradient vector flow (GVF). As a result, the prior motion knowledge of RV is introduced in the deformation process through the time-dependent constraints in the proposed GVF-T model. A weighting parameter is introduced to adjust the weight of the temporal information against the image data itself. The additional external edge forces retrieved from the temporal constraints may be useful for the RV segmentation, such that lead to a better segmentation performance. The effectiveness of the proposed approach is supported by experimental results on synthetic and cardiac MR images. Xulei Yang, Si Yong Yeo, Yi Su 0001, Calvin Chi-Wan Lim, Liang Zhong 0001, Ru-San Tan |
SMC | 1 |
| 2010 | An information-theoretic fuzzy C-spherical shells clustering algorithm
Qing Song 0001, Xulei Yang, Yeng Chai Soh |
Fuzzy Sets Syst. | 2 |
| 2009 | A novel pruning approach for robust data clustering
Xulei Yang, Qing Song 0001, Yi-Lei Wu, Aize Cao |
Neural Comput. Appl. | 1 |
| 2008 | Robust information clustering incorporating spatial information for breast mass detection in digitized mammograms
Aize Cao, Qing Song 0001, Xulei Yang |
Comput. Vis. Image Underst. | 3 |
| 2007 | A robust deterministic annealing algorithm for data clustering
Xulei Yang, Qing Song 0001, Yi-Lei Wu |
Data Knowl. Eng. | 1 |
| 2007 | A Weighted Support Vector Machine for Data ClassificationabstractThis paper presents a weighted support vector machine (WSVM) to improve the outlier sensitivity problem of standard support vector machine (SVM) for two-class data classification. The basic idea is to assign different weights to different data points such that the WSVM training algorithm learns the decision surface according to the relative importance of data points in the training data set. The weights used in WSVM are generated by a robust fuzzy clustering algorithm, kernel-based possibilistic c-means (KPCM) algorithm, whose partition generates relative high values for important data points but low values for outliers. Experimental results indicate that the proposed method reduces the effect of outliers and yields higher classification rate than standard SVM does when outliers exist in the training data set. Xulei Yang, Qing Song 0001, Yue Wang 0005 |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2006 | Clustering Spherical Shells by a Mini-Max Information Algorithm
Xulei Yang, Qing Song 0001 |
ACCV (2) | 1 |
| 2006 | Adaptive Spatial Information Clustering for Image SegmentationabstractThis paper presents a novel image segmentation algorithm that has a new dissimilarity measure which incorporates the spatial information. Our method uses a fully automatic technique to obtain the segmentation result and cluster number, and the new clustering objective function incorporates the spatial information and can compensate for the misclassification errors due to noise shifting. The capacity maximization and structure risk minimization are utilized to evaluate the quality of the clustering result via a trade-off between the number of unreliable data points and model complexity (i.e. cluster number). The weighting factor for neighborhood effect is adaptive to the image content. It enhances the smoothness towards piecewise-homogeneous region and reduces the edge-blurring effect. The experimental results with synthetic and real images demonstrate that the proposed method is effective in determining the optimal cluster number and eliminating the noise artifact. Qing Song 0001, Yeng Chai Soh, Xulei Yang, Kang Sim |
IJCNN | 4 |
| 2006 | Image Segmentation by Deterministic Annealing Algorithm with Adaptive Spatial Constraints
Xulei Yang, Aize Cao, Qing Song 0001 |
ISNN (2) | 1 |
| 2006 | Robust Data Clustering in Mercer Kernel-Induced Feature Space
Xulei Yang, Qing Song 0001, Meng Joo Er |
ISNN (1) | 1 |
| 2006 | A New Cluster Validity for Data ClusteringabstractCluster validity has been widely used to evaluate the fitness of partitions produced by clustering algorithms. This paper presents a new validity, which is called the Vapnik–Chervonenkis-bound (VB) index, for data clustering. It is estimated based on the structural risk minimization (SRM) principle, which optimizes the bound simultaneously over both the distortion function (empirical risk) and the VC-dimension (model complexity). The smallest bound of the guaranteed risk achieved on some appropriate cluster number validates the best description of the data structure. We use the deterministic annealing (DA) algorithm as the underlying clustering technique to produce the partitions. Five numerical examples and two real data sets are used to illustrate the use of VB as a validity index. Its effectiveness is compared to several popular cluster-validity indexes. The results of comparative study show that the proposed VB index has high ability in producing a good cluster number estimate and in addition, it provides a new approach for cluster validity from the view of statistical learning theory. Xulei Yang, Aize Cao, Qing Song 0001 |
Neural Process. Lett. | 1 |
| 2005 | Weighted support vector machine for data classificationabstractThis paper presents a weighted support vector machine (WSVM) to improve the outlier sensitivity problem of standard support vector machine (SVM) for two-class data classification. The basic idea is to assign different weights to different data points such that the WSVM training algorithm learns the decision surface according to the relative importance of data points in the training data set. The weights used in WSVM are generated by kernel-based possibilistic c-means (KPCM) algorithm, whose partition generates relative high values for important data points but low values for outliers. Experimental results indicate that the proposed method reduces the affect of outliers and yields higher classification rate than standard SVM does when outliers exist in the training data set. Xulei Yang, Qing Song 0001, Aizo Cao |
IJCNN | 1 |
| 2005 | Pre-selection of working set for SVM decomposition algorithmabstractThe decomposition algorithm is currently one of the major methods for solving support vector machines (SVM) training problems. The most important issue of this method is the selection of working set, which greatly affects the speed of the decomposition algorithm. In this paper, we propose a novel method for pre-selection of the working set for bound-constrained SVM formulation, which aims to make the training process more efficient. The pre-selection method is implemented based on fuzzy clustering technique in the high dimensional feature space using kernel methods. The effectiveness of the proposed method is supported by experimental results. Xulei Yang, Qing Song 0001 |
IJCNN | 1 |
| 2005 | A robust deterministic annealing algorithm for data clusteringabstractIn this paper, a new robust deterministic annealing (RDA) clustering algorithm is proposed. This method takes advantages of conventional noise clustering (NC) and deterministic annealing (DA) algorithms in terms of independence of data initialization, ability to avoid poor local optima, better performance for unbalanced data, and robustness against noise. The superiority of the proposed RDA clustering algorithm is supported by simulation results. Xulei Yang, Qing Song 0001 |
IJCNN | 1 |
| 2004 | Robust c-shells based deterministic annealing clustering algorithmabstractA new clustering method, robust c-shells based deterministic annealing (RCSDA) algorithm is developed. This development recasts the concept of fuzzy c-shells algorithm into the probability framework and offers several improved features over existing clustering algorithms. First, it is a global or close-to-global minimization algorithm through deterministic annealing rather than a local minimization method in the original fuzzy c-shells approach. Second, it is more effective in boundary detection with compact or hollow spherical shells compared to the original deterministic annealing approach. Finally, the basic idea of Dave's "noise clustering" is introduced into the algorithm which makes it robust against noise. The superiority of the proposed clustering method is supported by experimental results. Xulei Yang, Qing Song 0001, Aize Cao, Chengyi Guo |
FUZZ-IEEE | 1 |
| 2004 | Mammographic mass detection by vicinal support vector machineabstractWe proposed a Vicinal Support Vector Machine (VSVM) as an enhancement learning algorithm for mammographic mass detection on digital mammograms. The detection scheme includes two steps. First, one-class Support Vector Machine (SVM) is applied for the abnormal cases detection, where only normal cases are served as training samples. Then VSVM is investigated for the malignant cases detection. The aim of this step is to decide whether a detected abnormal case is benign or malignant. For the proposed VSVM algorithm, the whole training data are clustered into different soft vicinal areas in feature space by kernel based deterministic annealing (KBDA) method. The choice of different number of clusters makes VSVM be adaptive to different data structures in the input space. We tested the proposed scheme by using 90 clinical mammograms from MIAS database. The corresponding accuracy was observed to be 84%, with an area of A/sub z/=0.89 under the receiver operating characteristics (ROC) curve. The experimental results show that the two-step detection scheme works effective and the proposed VSVM is a promising classifier for breast mass detection. Aize Cao, Qing Song 0001, Xulei Yang, Chengyi Guo |
IJCNN | 3 |