Wuyang Li

dblp:170/0777 · DBLP profile ↗
← Back
43ranked-venue papers
12as first author
43since 2021 · last 2026
0000-0002-7338-9251ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 9 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 9 first-author · 24 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 2 first-author · 16 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 MRM++: Enhanced Masked Relation Modeling for Multi-Modal Medical Pre-training
abstract
Abstract Recent progress in deep learning for automated multi-modal medical diagnosis heavily depend on extensive expert annotations, which is time-intensive and impractical. To mitigate this, masked image modeling (MIM)-based pre-training strategies have emerged, effectively learning generalized representations from unlabelled data for various downstream tasks. Nevertheless, these approaches are tailored for natural images while neglect the distinct characteristics of medical data, resulting in suboptimal generalization in medical diagnosis applications. In this work, we attempt to harness the complementary information of multi-modal medical data to perform self-supervised pre-training and propose MRM++, an enhanced masked relation modeling paradigm. Different from the previous MIM methods that randomly mask input data, causing potentially missing of disease-relevant semantics, we devise prior-guided relation masking to break token-wise feature relation guided by anatomy-aware prior in both self- and cross-modal aspects. This can preserve complete input semantics and enable the model to learn abundant disease-related knowledge. Furthermore, to boost semantic relation modeling, the relation matching is introduced, which aligns sample-wise relations among unmasked and masked features. By exploiting inter-sample relations, the relation matching imposes the global constraints in the feature space, ensuring ample semantic relation for robust feature representation. Additionally, considering that the model may overfit to the pre-training dataset and lead to inherent gap between pre-training and downstream fine-tuning, we conceive task-oriented adapting as a pre-stage before fine-tuning to simultaneously perform self-supervised and task-supervised learning on downstream dataset. It can adaptively transform knowledge from the pre-trained model to be compatible with downstream tasks while maintaining transferable information. Extensive experiments on medical image-text and image-genome benchmarks validate the effectiveness and transfer ability of the proposed framework, outperforming state-of-the-art methods across various downstream diagnostic tasks. Source codes are made publicly available on https://github.com/CUHK-AIM-Group/MRM_plus .
Qiushi Yang, Wuyang Li, Zhe Peng, Fangxiao Cheng, Yixuan Yuan
Int. J. Comput. Vis.2
2026 GAGM: Geometry-aware graph matching framework for weakly supervised gyral hinge correspondence
Wuyang Li, Tianming Liu 0001, Xiang Li 0001, Junwei Han 0001, Yixuan Yuan
Medical Image Anal.2
2026 Harnessing Lightweight Transformer With Contextual Synergic Enhancement for Efficient 3D Medical Image Segmentation
abstract
Transformers have shown remarkable performance in 3D medical image segmentation, but their high computational requirements and need for large amounts of labeled data limit their applicability. To address these challenges, we consider two crucial aspects: model efficiency and data efficiency. Specifically, we propose Light-UNETR, a lightweight transformer designed to achieve model efficiency. Light-UNETR features a Lightweight Dimension Reductive Attention (LIDR) module, which reduces spatial and channel dimensions while capturing both global and local features via multi-branch attention. Additionally, we introduce a Compact Gated Linear Unit (CGLU) to selectively control channel interaction with minimal parameters. Furthermore, we introduce a Contextual Synergic Enhancement (CSE) learning strategy, which aims to boost the data efficiency of Transformers. It first leverages the extrinsic contextual information to support the learning of unlabeled data with Attention-Guided Replacement, then applies Spatial Masking Consistency that utilizes intrinsic contextual information to enhance the spatial context reasoning for unlabeled data. Extensive experiments on various benchmarks demonstrate the superiority of our approach in both performance and efficiency. For example, with only 10% labeled data on the Left Atrial Segmentation dataset, our method surpasses BCP by 1.43% Jaccard while drastically reducing the FLOPs by 90.8% and parameters by 85.8%.
Xinyu Liu 0001, Zhen Chen 0013, Wuyang Li, Chenxin Li, Yixuan Yuan
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 Fiber HGNN: Heterogeneous Graph Neural Network for Fiber Tract Segmentation
abstract
Fiber tract segmentation is crucial for clinical applications such as brain function interpretation and surgical planning. Existing methods typically adopt either a cortical-parcellation-based or fiber clustering approach, but fail to simultaneously integrate heterogeneous information (e.g., streamline shape, point position, anatomical priors). In this work, we propose Fiber HGNN, a novel heterogeneous graph neural network that explicitly models and integrates heterogeneous information of fibers for accurate fiber tract segmentation. We construct a heterogeneous graph comprising three types of nodes: streamline, fiber keypoint and anatomical region. Specifically, fiber keypoints are representative points sampled along each streamline to characterize local geometric features, while anatomical regions provide contextual priors derived from brain atlas. This design enables the network to jointly capture the complementary information of streamline shape, local geometry, and anatomical priors, thus facilitating the learning of more discriminative feature representations. To further leverage implicit anatomical connectivity, we design a Metapath-guided Heterogeneous Information Aggregation (MHIA) network. By analyzing the spatial relationships between streamline keypoints and anatomical regions, the heterogeneous graph is decomposed into anatomical subgraphs for each streamline. In each subgraph, heterogeneous information from metapath-linked nodes is aggregated to obtain the final fiber representation. We evaluate the effectiveness of our framework on the HCP105 and TractoInferno datasets. The experimental results demonstrate that our method significantly outperforms previous state-of-the-art methods. The source code is available at https://github.com/CUHK-AIM-Group/Fiber-HGNN.
Cheng Wang 0043, Wuyang Li, Xinyu Liu 0001, Yifan Liu 0010, Jian Cheng 0002, Yixuan Yuan
IEEE Trans. Medical Imaging2
2025 U-KAN Makes Strong Backbone for Medical Image Segmentation and Generation
abstract
U-Net has become a cornerstone in various visual applications such as image segmentation and diffusion probability models. While numerous innovative designs and improvements have been introduced by incorporating transformers or MLPs, the networks are still limited to linearly modeling patterns as well as the deficient interpretability. To address these challenges, our intuition is inspired by the impressive results of the Kolmogorov-Arnold Networks (KANs) in terms of accuracy and interpretability, which reshape the neural network learning via the stack of non-linear learnable activation functions derived from the Kolmogorov-Anold representation theorem. Specifically, in this paper, we explore the untapped potential of KANs in improving backbones for vision tasks. We investigate, modify and re-design the established U-Net pipeline by integrating the dedicated KAN layers on the tokenized intermediate representation, termed U-KAN. Rigorous medical image segmentation benchmarks verify the superiority of UKAN by higher accuracy even with less computation cost. We further delved into the potential of U-KAN as an alternative U-Net noise predictor in diffusion models, demonstrating its applicability in generating task-oriented model architectures.
Chenxin Li, Xinyu Liu 0001, Wuyang Li, Cheng Wang 0043, Hengyu Liu 0007, Yifan Liu 0010, Zhen Chen 0013, Yixuan Yuan
AAAI3
2025 Universal Domain Adaptive Object Detection via Dual Probabilistic Alignment
abstract
Domain Adaptive Object Detection (DAOD) transfers knowledge from a labeled source domain to an unannotated target domain under closed-set assumption. Universal DAOD (UniDAOD) extends DAOD to handle open-set, partial-set, and closed-set domain adaptation. In this paper, we first unveil two issues: domain-private category alignment is crucial for global-level features, and the domain probability heterogeneity of features across different levels. To address these issues, we propose a novel Dual Probabilistic Alignment (DPA) framework to model domain probability as Gaussian distribution, enabling the heterogeneity domain distribution sampling and measurement. The DPA consists of three tailored modules: the Global-level Domain Private Alignment (GDPA), the Instance-level Domain Shared Alignment (IDSA), and the Private Class Constraint (PCC). GDPA utilizes the global-level sampling to mine domain-private category samples and calculate alignment weight through a cumulative distribution function to address the global-level private category alignment. IDSA utilizes instance-level sampling to mine domain-shared category samples and calculates alignment weight through Gaussian distribution to conduct the domain-shared category domain alignment to address the feature heterogeneity. The PCC aggregates domain-private category centroids between feature and probability spaces to mitigate negative transfer. Extensive experiments demonstrate that our DPA outperforms state-of-the-art UniDAOD and DAOD methods across various datasets and scenarios, including open, partial, and closed sets.
Yuanfan Zheng, Jinlin Wu, Wuyang Li, Zhen Chen 0018
AAAI3
2025 Track Any Anomalous Object: A Granular Video Anomaly Detection Pipeline
abstract
Video anomaly detection (VAD) is crucial in scenarios such as surveillance and autonomous driving, where timely detection of unexpected activities is essential. Albeit existing methods have primarily focused on detecting anomalous objects in videos—either by identifying anomalous frames or objects—they often neglect finer-grained analysis, such as anomalous pixels, which limits their ability to capture a broader range of anomalies. To address this challenge, we propose an innovative VAD framework called Track Any Anomalous Object (TAO), which introduces a Granular Video Anomaly Detection Framework that, for the first time, integrates the detection of multiple fine-grained anomalous objects into a unified framework. Unlike methods that assign anomaly scores to every pixel at each moment, our approach transforms the problem into pixel-level tracking of anomalous objects. By linking anomaly scores to subsequent tasks such as image segmentation and video tracking, our method eliminates the need for threshold selection and achieves more precise anomaly localization, even in long and challenging video sequences. Experiments on extensive datasets demonstrate that TAO achieves state-of-the-art performance, setting a new progress for VAD by providing a practical, granular, and holistic solution. For more information, visit the project page at: https://tao-25.github.io/
Yuzhi Huang, Chenxin Li, Zixu Lin, Yunlong Lin, Hengyu Liu 0007, Wuyang Li, Xinyu Liu 0001, Jiechao Gao, Yue Huang 0001, Xinghao Ding, Yixuan Yuan
CVPR7
2025 FlexGS: Train Once, Deploy Everywhere with Many-in-One Flexible 3D Gaussian Splatting
abstract
3D Gaussian splatting (3DGS) has enabled various applications in 3D scene representation and novel view synthesis due to its efficient rendering capabilities. However, 3DGS demands relatively significant GPU memory, limiting its use on devices with restricted computational resources. Previous approaches have focused on pruning less important Gaussians, effectively compressing 3DGS but often requiring a fine-tuning stage and lacking adaptability for the specific memory needs of different devices. In this work, we present an elastic inference method for 3DGS. Given an input for the desired model size, our method selects and transforms a subset of Gaussians, achieving substantial rendering performance without additional fine-tuning. We introduce a tiny learnable module that controls Gaussian selection based on the input percentage, along with a transformation module that adjusts the selected Gaussians to complement the performance of the reduced model. Comprehensive experiments on ZipNeRF, MipNeRF and Tanks&Temples scenes demonstrate the effectiveness of our approach. Code is available at https://flexgs.github.io/.
Hengyu Liu 0007, Yuehao Wang, Chenxin Li, Ruisi Cai, Wuyang Li, Pavlo Molchanov 0001, Peihao Wang, Zhangyang Wang
CVPR6
2025 ConcealGS: Concealing Invisible Copyright Information in 3D Gaussian Splatting
abstract
As 3D Gaussian Splatting (3D-GS) emerges as a promising technique for 3D reconstruction and novel view synthesis, offering superior rendering quality and efficiency, it becomes crucial to ensure secure transmission and copyright protection of 3D assets in anticipation of widespread distribution. While steganography has advanced significantly in common 3D media like meshes and Neural Radiance Fields (NeRF), research into steganography for 3D- GS representations remains largely unexplored. To address this gap, we propose ConcealGS, a novel 3D steganography method that embeds implicit information into the explicit 3D representation of Gaussian Splatting. By introducing a consistency strategy for the decoder and a gradient optimization approach, ConcealGS overcomes limitations of NeRF-based models, enhancing both the robustness of implicit information and the quality of 3D reconstruction. Extensive evaluations across various potential application scenarios demonstrate that ConcealGS successfully recovers implicit information with negligible impact on rendering quality, offering a groundbreaking approach for embedding invisible yet recoverable information into 3D models. This work paves the way for advanced copyright protection and secure data transmission in the evolving landscape of 3D content creation and distribution. Code is available at https://github.com/zxk1212/ConcealGS.
Hengyu Liu 0007, Chenxin Li, Yining Sun, Wuyang Li, Yifan Liu 0010, Yiyang Lin, Yixuan Yuan, Nanyang Ye 0001
ICASSP5
2025 InfoBridge: Balanced Multimodal Integration through Conditional Dependency Modeling
Chenxin Li, Yifan Liu 0010, Panwang Pan, Hengyu Liu 0007, Xinyu Liu 0001, Wuyang Li, Cheng Wang 0043, Weihao Yu 0004, Yiyang Lin, Yixuan Yuan
ICCV6
2025 Metascope: Optics-Driven Neural Network for Ultra-Micro Metalens Endoscopy
Wuyang Li, Wentao Pan 0001, Zhendong Luo, Chenxin Li, Hengyu Liu 0007, Din Ping Tsai, Mu Ku Chen, Yixuan Yuan
ICCV1
2025 Synthesizing Realistic fMRI: A Physiological Dynamics-Driven Hierarchical Diffusion Model for Efficient fMRI Acquisition
abstract
Functional magnetic resonance imaging (fMRI) is essential for mapping brain activity but faces challenges like lengthy acquisition time and sensitivity to patient movement, limiting its clinical and machine learning applications. While generative models such as diffusion models can synthesize fMRI signals to alleviate these issues, they often underperform due to neglecting the brain's complex structural and dynamic properties. To address these limitations, we propose the Physiological Dynamics-Driven Hierarchical Diffusion Model, a novel framework integrating two key brain physiological properties into the diffusion process: brain hierarchical regional interactions and multifractal dynamics. To model complex interactions among brain regions, we construct hypergraphs based on the prior knowledge of brain functional parcellation reflected by resting-state functional connectivity (rsFC). This enables the aggregation of fMRI signals across multiple scales and generates hierarchical signals. Additionally, by incorporating the prediction of two key dynamics properties of fMRI—the multifractal spectrum and generalized Hurst exponent—our framework effectively guides the diffusion process, ensuring the preservation of the scale-invariant characteristics inherent in real fMRI data. Our framework employs progressive diffusion generation, with signals representing broader brain region information conditioning those that capture localized details, and unifies multiple inputs during denoising for balanced integration. Experiments demonstrate that our model generates physiologically realistic fMRI signals, potentially reducing acquisition time and enhancing data quality, benefiting clinical diagnostics and machine learning in neuroscience.
Yufan Hu, Yu Jiang 0013, Wuyang Li, Yixuan Yuan
ICLR3
2025 InstantSplamp: Fast and Generalizable Stenography Framework for Generative Gaussian Splatting
abstract
With the rapid development of large generative models for 3D, especially the evolution from NeRF representations to more efficient Gaussian Splatting, the synthesis of 3D assets has become increasingly fast and efficient, enabling the large-scale publication and sharing of generated 3D objects. However, while existing methods can add watermarks or steganographic information to individual 3D assets, they often require time-consuming per-scene training and optimization, leading to watermarking overheads that can far exceed the time required for asset generation itself, making deployment impractical for generating large collections of 3D objects. To address this, we propose InstantSplamp a framework that seamlessly integrates the 3D steganography pipeline into large 3D generative models without introducing explicit additional time costs. Guided by visual foundation models,InstantSplamp subtly injects hidden information like copyright tags during asset generation, enabling effective embedding and recovery of watermarks within generated 3D assets while preserving original visual quality. Experiments across various potential deployment scenarios demonstrate that \model~strikes an optimal balance between rendering quality and hiding fidelity, as well as between hiding performance and speed. Compared to existing per-scene optimization techniques for 3D assets, InstantSplamp reduces their watermarking training overheads that are multiples of generation time to nearly zero, paving the way for real-world deployment at scale. Project page: https://gaussian-stego.github.io/.
Chenxin Li, Hengyu Liu 0007, Zhiwen Fan, Wuyang Li, Yifan Liu 0010, Panwang Pan, Yixuan Yuan
ICLR4
2025 Hide-in-Motion: Embedding Steganographic Copyright Information into 4D Gaussian Splatting Assets
abstract
As 4D extensions of 3D Gaussian Splatting (4D-GS) emerge as groundbreaking techniques for dynamic scene reconstruction and novel view synthesis in robotics and computer vision, ensuring the security and trustworthiness of these assets becomes crucial. While steganography has advanced significantly in 2D and 3D media, existing methods are inadequate for the complex, dynamic nature of 4D-GS representations. To address this gap, we propose Hide-in-Motion, a novel 4D steganography method for hiding information through deformation in Gaussian splatting. Our approach introduces a composite attribute and a Decouple Feature Field for coarse-to-fine deformation modeling and embedding implicit information, along with an Opacity-Guided Adaptive strategy. Hide-in-Motion overcomes the limitations of previous techniques, enhancing both the robustness of embedded information and the quality of 4D reconstruction. Extensive evaluations demonstrate that our method successfully embeds and recovers implicit information across various modalities while maintaining high rendering quality in dynamic scenes. This work not only advances copyright protection and secure data transmission for 4D assets but also paves the way for enhancing the security and integrity of 4D digital assets. Code is available at https://github.com/CUHK-AIM-Group/Hide-in-Motion.
Hengyu Liu 0007, Chenxin Li, Wentao Pan 0001, Zhiqin Yang, Yifan Liu 0010, Wuyang Li, Yixuan Yuan
ICRA7
2025 See&Trek: Training-Free Spatial Prompting for Multimodal Large Language Model
abstract
We introduce See&Trek, the first training-free prompting framework tailored to enhance the spatial understanding of Multimodal Large Language Models (MLLMs) under vision-only constraints. While prior efforts have incorporated modalities like depth or point clouds to improve spatial reasoning, purely visual-spatial understanding remains underexplored. See&Trek addresses this gap by focusing on two core principles: increasing visual diversity and motion reconstruction. For visual diversity, we conduct Maximum Semantic Richness Sampling, which employs an off-the-shell perception model to extract semantically rich keyframes that capture scene structure. For motion reconstruction, we simulate visual trajectories and encode relative spatial positions into keyframes to preserve both spatial relations and temporal coherence. Our method is training&GPU-free, requiring only a single forward pass, and can be seamlessly integrated into existing MLLMs. Extensive experiments on the VSI-Bench and STI-Bench show that See&Trek consistently boosts various MLLMs performance across diverse spatial reasoning tasks with the most +3.5% improvement, offering a promising path toward stronger spatial intelligence.
Pengteng Li, Pinhao Song, Wuyang Li, Huizai Yao, Weiyu Guo, Yijie Xu, Dugang Liu, Hui Xiong 0001
NeurIPS3
2025 VoxDet: Rethinking 3D Semantic Scene Completion as Dense Object Detection
abstract
Semantic Scene Completion (SSC) aims to reconstruct the 3D geometry and semantics of the surrounding environment. With dense voxel labels, prior works typically formulate SSC as a *dense segmentation task*, independently classifying each voxel. However, this paradigm neglects critical instance-centric discriminability, leading to instance-level incompleteness and adjacent ambiguities. To address this, we highlight a "free lunch" of SSC labels: the voxel-level class label has implicitly told the instance-level insight, which is ever-overlooked by the community. Motivated by this observation, we first introduce a training-free **Voxel-to-Instance (VoxNT) trick**: a simple yet effective method that freely converts voxel-level class labels into instance-level offset labels. Building on this, we further propose **VoxDet**, an instance-centric framework that reformulates the voxel-level SSC as *dense object detection* by decoupling it into two sub-tasks: offset regression and semantic prediction. Specifically, based on the lifted 3D volume, VoxDet first uses (a) Spatially-decoupled Voxel Encoder to generate disentangled feature volumes for the two sub-tasks, which learn task-specific spatial deformation in the densely projected tri-perceptive space. Then, we deploy (b) Task-decoupled Dense Predictor to address SSC via dense detection. Here, we first regress a 4D offset field to estimate distances (6 directions) between voxels and the corresponding object boundaries in the voxel space. The regressed offsets are then used to guide the instance-level aggregation in the classification branch, achieving instance-aware scene completion. VoxDet can be deployed on both camera and LiDAR input and jointly achieves state-of-the-art results on both benchmarks, which gives 63.0 IoU on the SemanticKITTI test set, **ranking 1$^{st}$** on the online leaderboard.
Wuyang Li, Alexandre Alahi
NeurIPS1
2025 IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering
abstract
Vision-language models (VLMs) excel at descriptive tasks, but whether they truly understand scenes from visual observations remains uncertain. We introduce IR3D-Bench, a benchmark challenging VLMs to demonstrate understanding through active creation rather than passive recognition. Grounded in the analysis-by-synthesis paradigm, IR3D-Bench tasks Vision-Language Agents (VLAs) with actively using programming and rendering tools to recreate the underlying 3D structure of an input image, achieving agentic inverse rendering through tool use. This ''understanding-by-creating'' approach probes the tool-using generative capacity of VLAs, moving beyond the descriptive or conversational capacity measured by traditional scene understanding benchmarks. We provide a comprehensive suite of metrics to evaluate geometric accuracy, spatial relations, appearance attributes, and overall plausibility. Initial experiments on agentic inverse rendering powered by various state-of-the-art VLMs highlight current limitations, particularly in visual precision rather than basic tool usage. IR3D-Bench, including data and evaluation protocols, is released to facilitate systematic study and development of tool-using VLAs towards genuine scene understanding by creating.
Hengyu Liu 0007, Chenxin Li, Yipeng Wu, Wuyang Li, Zhiqin Yang, Zhenyuan Zhang 0001, Yunlong Lin, Sirui Han, Brandon Yushan Feng
NeurIPS5
2025 FM-APP: Foundation Model for Any Phenotype Prediction via fMRI to sMRI Knowledge Transfer
abstract
Predicting individual-level non-neuroimaging phenotypes (e.g., fluid intelligence) using brain imaging data is a fundamental goal of neuroscience. Recent research has focused on utilizing high-cost functional magnetic resonance imaging (fMRI) to predict phenotypes seen during training. However, these methods 1) only consider predicting seen phenotypes, failing to achieve zero-shot inference for unseen phenotypes; 2) overlook the knowledge transfer from fMRI to structural MRI (sMRI), missing out on utilizing cost-effective sMRI for accurate predictions. To address these challenges, we propose a Foundational Model for Any Phenotype Prediction via fMRI to sMRI knowledge transfer (FM-APP), consisting of a Phenotypes Text Memory Bank (PTMB) module, Any Phenotype Prediction (APP) module, and fMRI to sMRI Knowledge Transfer (F2SKT) module. Our proposed FM-APP adapts to downstream tasks by generating regressor parameters instead of fine-tuning the model itself. Specifically, to retain important clues from seen phenotype descriptions, PTMB utilizes the BiomedCLIP model to store semantic features of seen phenotypes. To achieve any phenotype prediction, the APP introduces a regressor synthesizer for zero-shot inference. Additionally, to improve sMRI prediction accuracy while preserving its cost advantage, the F2SKT uses the PTMB to construct phenotype active maps, guiding adaptive knowledge transfer from fMRI to sMRI. Experiments on the Human Connectome Project (HCP) and HCP Aging datasets demonstrate our approach outperforms state-of-the-art methods, showcasing strong zero-shot inference capabilities and providing a novel framework for analyzing brain structure and phenotypes. Our code: https://github.com/ZhibinHe/FM-APP.
Wuyang Li, Yifan Liu 0010, Xinyu Liu 0001, Junwei Han 0001, Yixuan Yuan
IEEE Trans. Medical Imaging2
2025 LLM-Guided Decoupled Probabilistic Prompt for Continual Learning in Medical Image Diagnosis
abstract
Deep learning-based traditional diagnostic models typically exhibit limitations when applied to dynamic clinical environments that require handling the emergence of new diseases. Continual learning (CL) offers a promising solution, aiming to learn new knowledge while preserving previously learned knowledge. Though recent rehearsal-free CL methods employing prompt tuning (PT) have shown promise, they rely on deterministic prompts that struggle to handle diverse fine-grained knowledge. Moreover, existing PT methods utilize randomly initialized prompts that are trained under standard classification constraints, impeding expert knowledge integration and optimal performance acquisition. In this paper, we propose an LLM-guided Decoupled Probabilistic Prompt (LDPP) for Continual Learning in medical image diagnosis. Specifically, we develop an Expert Knowledge Generation (EKG) module that leverages LLM to acquire decoupled expert knowledge and comprehensive category descriptions. Then, we introduce a Decoupled Probabilistic Prompt pool (DePP) to construct a shared decoupled probabilistic prompt pool, which constructs a shared prompt pool with probabilistic prompts derived from the expert knowledge set. These prompts dynamically provide diverse and flexible descriptions for input images. Finally, We design a Steering Prompt Pool (SPP) to enhance intra-class compactness and promote model performance by learning non-shared prompts. With extensive experimental validation, LDPP consistently sets state-of-the-art performance under the challenging class-incremental setting in CL. Code is available at: https://github.com/CUHK-AIM-Group/LDPP.
Yiwen Luo, Wuyang Li, Xiang Li 0001, Tianming Liu 0001, Tianye Niu, Yixuan Yuan
IEEE Trans. Medical Imaging2
2025 ToothMaker: Realistic Panoramic Dental Radiograph Generation via Disentangled Control
abstract
Generating high-fidelity dental radiographs is essential for training diagnostic models. Despite the development of numerous methods for other medical data, generative approaches in dental radiology remain unexplored. Due to the intricate tooth structures and specialized terminology, these methods often yield ambiguous tooth regions and incorrect dental concepts when applied to dentistry. In this paper, we take the first attempt to investigate diffusion-based teeth X-ray image generation and propose ToothMaker, a novel framework specifically designed for the dental domain. Firstly, to synthesize X-ray images that possess accurate tooth structures and realistic radiological styles simultaneously, we design control-disentangled fine-tuning (CDFT) strategy. Specifically, we present two separate controllers to handle style and layout control respectively, and introduce a gradient-based decoupling method that optimizes each using their corresponding disentangled gradients. Secondly, to enhance model's understanding of dental terminology, we propose prior-disentangled guidance module (PDGM), enabling precise synthesis of dental concepts. It utilizes large language model to decompose dental terminology into a series of meta-knowledge elements and performs interactions and refinements through hypergraph neural network. These elements are then fed into the network to guide the generation of dental concepts. Extensive experiments demonstrate the high fidelity and diversity of the images synthesized by our approach. By incorporating the generated data, we achieve substantial performance improvements on downstream segmentation and visual question answering tasks, indicating that our method can greatly reduce the reliance on manually annotated data. Code will be public available at https://github.com/CUHK-AIM-Group/ToothMaker.
Weihao Yu 0004, Xiaoqing Guo, Wuyang Li, Xinyu Liu 0001, Hui Chen 0032, Yixuan Yuan
IEEE Trans. Medical Imaging3
2024 SVP: Enhancing Security and Scalability for Metaverse Blockchain Through Integrating Stake in Voting-based Consensus Protocol
abstract
Blockchain has now become a critical infrastructure in Metaverse for storing and managing the digital resources of users, bridging the real and virtual worlds. However, consensus protocols in blockchains constrain the performance of their applications. While existing voting-based consensus protocols such as HotStuff and other Byzantine Fault Tolerance (BFT) protocols have optimized efficiency and scalability, they simply adopt a one-person-one-vote rule that is not aligned with the human-centric values of most blockchain applications, including Metaverse. Therefore, we propose Stake Voting Protocol (SVP), a secure and scalable consensus protocol, whose design philosophy is to consider validators’ stakes in the BFT protocol and introduce flexibility through a sliding window. We also propose a certification rule within the pipelined two-chain consensus process to enhance security. Furthermore, our epoch change and incentive mechanisms ensure dynamics and liveness, respectively. Finally, our analytical and experimental results demonstrate that the proposed SVP satisfies correctness and can resist specific attacks with low latency and high throughput.
Wuyang Li, Hui Li 0022, Qiufan Wu, Han Wang 0022, Weimin Zeng, Yanping Zhang 0008, Ping Lu 0008, Runhuai Huang
IEEE Big Data1
2024 CLIFF: Continual Latent Diffusion for Open-Vocabulary Object Detection
Wuyang Li, Xinyu Liu 0001, Jiayi Ma 0001, Yixuan Yuan
ECCV (55)1
2024 F2TNet: FMRI to T1w MRI Knowledge Transfer Network for Brain Multi-phenotype Prediction
Wuyang Li, Yu Jiang 0013, Zhihao Peng 0002, Pengyu Wang 0005, Xiang Li 0001, Tianming Liu 0001, Junwei Han 0001, Yixuan Yuan
MICCAI (11)2
2024 👦 Endora: Video Generation Models as Endoscopy Simulators
Chenxin Li, Hengyu Liu 0007, Yifan Liu 0010, Brandon Yushan Feng, Wuyang Li, Xinyu Liu 0001, Zhen Chen 0013, Yixuan Yuan
MICCAI (6)5
2024 From Static to Dynamic Diagnostics: Boosting Medical Image Analysis via Motion-Informed Generative Videos
Wuyang Li, Xinyu Liu 0001, Qiushi Yang, Yixuan Yuan
MICCAI (3)1
2024 LGS: A Light-Weight 4D Gaussian Splatting for Efficient Surgical Scene Reconstruction
Hengyu Liu 0007, Yifan Liu 0010, Chenxin Li, Wuyang Li, Yixuan Yuan
MICCAI (3)4
2024 When 3D Partial Points Meets SAM: Tooth Point Cloud Segmentation with Sparse Labels
Yifan Liu 0010, Wuyang Li, Cheng Wang 0043, Hui Chen 0032, Yixuan Yuan
MICCAI (11)2
2024 DiffRect: Latent Diffusion Label Rectification for Semi-supervised Medical Image Segmentation
Xinyu Liu 0001, Wuyang Li, Yixuan Yuan
MICCAI (12)2
2024 Flaws can be Applause: Unleashing Potential of Segmenting Ambiguous Objects in SAM
abstract
As the vision foundation models like the Segment Anything Model (SAM) demonstrate potent universality, they also present challenges in giving ambiguous and uncertain predictions. Significant variations in the model output and granularity can occur with simply subtle changes in the prompt, contradicting the consensus requirement for the robustness of a model. While some established works have been dedicated to stabilizing and fortifying the prediction of SAM, this paper takes a unique path to explore how this flaw can be inverted into an advantage when modeling inherently ambiguous data distributions. We introduce an optimization framework based on a conditional variational autoencoder, which jointly models the prompt and the granularity of the object with a latent probability distribution. This approach enables the model to adaptively perceive and represent the real ambiguous label distribution, taming SAM to produce a series of diverse, convincing, and reasonable segmentation outputs controllably. Extensive experiments on several practical deployment scenarios involving ambiguity demonstrates the exceptional performance of our framework. Project page: \url{https://a-sa-m.github.io/}.
Chenxin Li, Yuzhi Huang, Wuyang Li, Hengyu Liu 0007, Xinyu Liu 0001, Qing Xu 0014, Zhen Chen 0013, Yue Huang 0001, Yixuan Yuan
NeurIPS3
2024 Decoupled Unbiased Teacher for Source-Free Domain Adaptive Medical Object Detection
abstract
Source-free domain adaptation (SFDA) aims to adapt a lightweight pretrained source model to unlabeled new domains without the original labeled source data. Due to the privacy of patients and storage consumption concerns, SFDA is a more practical setting for building a generalized model in medical object detection. Existing methods usually apply the vanilla pseudo-labeling technique, while neglecting the bias issues in SFDA, leading to limited adaptation performance. To this end, we systematically analyze the biases in SFDA medical object detection by constructing a structural causal model (SCM) and propose an unbiased SFDA framework dubbed decoupled unbiased teacher (DUT). Based on the SCM, we derive that the confounding effect causes biases in the SFDA medical object detection task at the sample level, feature level, and prediction level. To prevent the model from emphasizing easy object patterns in the biased dataset, a dual invariance assessment (DIA) strategy is devised to generate counterfactual synthetics. The synthetics are based on unbiased invariant samples in both discrimination and semantic perspectives. To alleviate overfitting to domain-specific features in SFDA, we design a cross-domain feature intervention (CFI) module to explicitly deconfound the domain-specific prior with feature intervention and obtain unbiased features. Besides, we establish a correspondence supervision prioritization (CSP) strategy for addressing the prediction bias caused by coarse pseudo-labels by sample prioritizing and robust box supervision. Through extensive experiments on multiple SFDA medical object detection scenarios, DUT yields superior performance over previous state-of-the-art unsupervised domain adaptation (UDA) and SFDA counterparts, demonstrating the significance of addressing the bias issues in this challenging task. The code is available at https://github.com/CUHK-AIM-Group/Decoupled-Unbiased-Teacher.
Xinyu Liu 0001, Wuyang Li, Yixuan Yuan
IEEE Trans. Neural Networks Learn. Syst.2
2023 Adjustment and Alignment for Unbiased Open Set Domain Adaptation
abstract
Open Set Domain Adaptation (OSDA) transfers the model from a label-rich domain to a label-free one containing novel-class samples. Existing OSDA works overlook abundant novel-class semantics hidden in the source domain, leading to a biased model learning and transfer. Although the causality has been studied to remove the semantic-level bias, the non-available novel-class samples result in the failure of existing causal solutions in OSDA. To break through this barrier, we propose a novel causality-driven solution with the unexplored front-door adjustment theory, and then implement it with a theoretically grounded framework, coined Adjustment and Alignment (ANNA), to achieve an unbiased OSDA. In a nutshell, ANNA consists of Front-Door Adjustment (FDA) to correct the biased learning in the source domain and Decoupled Causal Alignment (DCA) to transfer the model unbiasedly. On the one hand, FDA delves into fine-grained visual blocks to discover novel-class regions hidden in the base-class image. Then, it corrects the biased model optimization by implementing causal debiasing. On the other hand, DCA disentangles the base-class and novel-class regions with orthogonal masks, and then adapts the decoupled distribution for an unbiased model transfer. Extensive experiments show that ANNA achieves state-of-the-art results. The code is available at https://github.com/CityU-AIM-Group/Anna.
Wuyang Li, Jie Liu 0044, Bo Han 0003, Yixuan Yuan
CVPR1
2023 Novel Scenes & Classes: Towards Adaptive Open-set Object Detection
abstract
Domain Adaptive Object Detection (DAOD) transfers an object detector to a novel domain free of labels. However, in the real world, besides encountering novel scenes, novel domains always contain novel-class objects de facto, which are ignored in existing research. Thus, we formulate and study a more practical setting, Adaptive Open-set Object Detection (AOOD), considering both novel scenes and classes. Directly combing off-the-shelled cross-domain and open-set approaches is sub-optimal since their low-order dependence, e.g., the confidence score, is insufficient for the AOOD with two dimensions of novel information. To address this, we propose a novel Structured Motif Matching (SOMA) framework for AOOD, which models the high-order relation with motifs, i.e., statistically significant subgraphs, and formulates AOOD solution as motif matching to learn with high-order patterns. In a nutshell, SOMA consists of Structure-aware Novel-class Learning (SNL) and Structure-aware Transfer Learning (STL). As for SNL, we establish an instance-oriented graph to capture the class-independent object feature hidden in different base classes. Then, a high-order metric is proposed to match the most significant motif as high-order patterns, serving for motif-guided novel-class learning. In STL, we set up a semantic-oriented graph to model the class-dependent relation across domains, and match unlabelled objects with high-order motifs to align the crossdomain distribution with structural awareness. Extensive experiments demonstrate that the proposed SOMA achieves state-of-the-art performance. Codes are available at https://github.com/CityU-AIM-Group/SOMA.
Wuyang Li, Xiaoqing Guo, Yixuan Yuan
ICCV1
2023 MRM: Masked Relation Modeling for Medical Image Pre-Training with Genetics
abstract
Modern deep learning techniques on automatic multi-modal medical diagnosis rely on massive expert annotations, which is time-consuming and prohibitive. Recent masked image modeling (MIM)-based pre-training methods have witnessed impressive advances for learning meaningful representations from unlabeled data and transferring to downstream tasks. However, these methods focus on natural images and ignore the specific properties of medical data, yielding unsatisfying generalization performance on downstream medical diagnosis. In this paper, we aim to leverage genetics to boost image pre-training and present a masked relation modeling (MRM) framework. Instead of explicitly masking input data in previous MIM methods leading to loss of disease-related semantics, we design relation masking to mask out token-wise feature relation in both self- and cross-modality levels, which preserves intact semantics within the input and allows the model to learn rich disease-related information. Moreover, to enhance semantic relation modeling, we propose relation matching to align the sample-wise relation between the intact and masked features. The relation matching exploits inter-sample relation by encouraging global constraints in the feature space to render sufficient semantic relation for feature representation. Extensive experiments demonstrate that the proposed framework is simple yet powerful, achieving state-of-the-art transfer performance on various downstream diagnosis tasks. Codes are available at https://github.com/CityU-AIM-Group/MRM.
Qiushi Yang, Wuyang Li, Baopu Li, Yixuan Yuan
ICCV2
2023 $\mathrm {H^{2}}$GM: A Hierarchical Hypergraph Matching Framework for Brain Landmark Alignment
Wuyang Li, Yixuan Yuan
MICCAI (10)2
2023 Medical federated learning with joint graph purification for noisy label learning
Zhen Chen 0013, Wuyang Li, Xiaohan Xing, Yixuan Yuan
Medical Image Anal.2
2023 SIGMA++: Improved Semantic-Complete Graph Matching for Domain Adaptive Object Detection
abstract
Domain Adaptive Object Detection (DAOD) generalizes the object detector from an annotated domain to a label-free novel one. Recent works estimate prototypes (class centers) and minimize the corresponding distances to adapt the cross-domain class conditional distribution. However, this prototype-based paradigm 1) fails to capture the class variance with agnostic structural dependencies, and 2) ignores the domain-mismatched classes with a sub-optimal adaptation. To address these two challenges, we propose an improved SemantIc-complete Graph MAtching framework, dubbed SIGMA++, for DAOD, completing mismatched semantics and reformulating adaptation with hypergraph matching. Specifically, we propose a Hypergraphical Semantic Completion (HSC) module to generate hallucination graph nodes in mismatched classes. HSC builds a cross-image hypergraph to model class conditional distribution with high-order dependencies and learns a graph-guided memory bank to generate missing semantics. After representing the source and target batch with hypergraphs, we reformulate domain adaptation with a hypergraph matching problem, i.e., discovering well-matched nodes with homogeneous semantics to reduce the domain gap, which is solved with a Bipartite Hypergraph Matching (BHM) module. Graph nodes are used to estimate semantic-aware affinity, while edges serve as high-order structural constraints in a structure-aware matching loss, achieving fine-grained adaptation with hypergraph matching. The applicability of various object detectors verifies the generalization of SIGMA++, and extensive experiments on nine benchmarks show its state-of-the-art performance on both AP$_{50}$and adaptation gains.
Wuyang Li, Xinyu Liu 0001, Yixuan Yuan
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 GRAB-Net: Graph-Based Boundary-Aware Network for Medical Point Cloud Segmentation
abstract
Point cloud segmentation is fundamental in many medical applications, such as aneurysm clipping and orthodontic planning. Recent methods mainly focus on designing powerful local feature extractors and generally overlook the segmentation around the boundaries between objects, which is extremely harmful to the clinical practice and degenerates the overall segmentation performance. To remedy this problem, we propose a GRAph-based Boundary-aware Network (GRAB-Net) with three paradigms, Graph-based Boundary-perception Module (GBM), Outer-boundary Context-assignment Module (OCM), and Inner-boundary Feature-rectification Module (IFM), for medical point cloud segmentation. Aiming to improve the segmentation performance around boundaries, GBM is designed to detect boundaries and interchange complementary information inside semantic and boundary features in the graph domain, where semantics-boundary correlations are modelled globally and informative clues are exchanged by graph reasoning. Furthermore, to reduce the context confusion that degenerates the segmentation performance outside the boundaries, OCM is proposed to construct the contextual graph, where dissimilar contexts are assigned to points of different categories guided by geometrical landmarks. In addition, we advance IFM to distinguish ambiguous features inside boundaries in a contrastive manner, where boundary-aware contrast strategies are proposed to facilitate the discriminative representation learning. Extensive experiments on two public datasets, IntrA and 3DTeethSeg, demonstrate the superiority of our method over state-of-the-art methods.
Yifan Liu 0010, Wuyang Li, Jie Liu 0044, Hui Chen 0032, Yixuan Yuan
IEEE Trans. Medical Imaging2
2023 SCAN++: Enhanced Semantic Conditioned Adaptation for Domain Adaptive Object Detection
abstract
Domain Adaptive Object Detection (DAOD) transfers an object detector from the labeled source domain to a novel unlabelled target domain. Recent advances bridge the domain gap by aligning category-agnostic feature distribution and minimizing the domain discrepancy for adapting semantic distribution. Though great success, these methods model domain discrepancy with prototypes within a batch, yielding a biased estimation of domain-level statistics. Moreover, the category-agnostic alignment leads to the disagreement of the cross-domain semantic distribution with inevitable classification errors. To address these two issues, we propose an enhanced Semantic Conditioned AdaptatioN (SCAN++) framework, which leverages unbiased semantics for DAOD. Specifically, in the source domain, we design the conditional kernel to sample Pixel of Interests (PoIs), and aggregate PoIs with a cross-image graph to estimate an unbiased semantic sequence. Conditioned on the semantic sequence, we further update the parameter of the conditional kernel in a semantic conditioned manifestation module, and establish a novel conditional graph in the target domain to model unlabeled semantics. After modeling the semantic distribution in both domains, we integrate the conditional kernel into adversarial alignment to achieve semantic-aware adaptation in a Conditional Kernel guided Alignment (CKA) module. Meanwhile, the Semantic Sequence guided Transport (SST) module is proposed to transfer reliable semantic knowledge to the target domain through solving the cross-domain Optimal Transport (OT) assignment, achieving unbiased adaptation at the semantic level. Comprehensive experiments on four adaptation scenarios demonstrate that SCAN++ achieves state-of-the-art results. The code is available athttps://github.com/CityU-AIM-Group/SCAN/tree/SCAN++.
Wuyang Li, Xinyu Liu 0001, Yixuan Yuan
IEEE Trans. Multim.1
2022 SCAN: Cross Domain Object Detection with Semantic Conditioned Adaptation
abstract
The domain gap severely limits the transferability and scalability of object detectors trained in a specific domain when applied to a novel one. Most existing works bridge the domain gap by minimizing the domain discrepancy in the category space and aligning category-agnostic global features. Though great success, these methods model domain discrepancy with prototypes within a batch, yielding a biased estimation of domain-level distribution. Besides, the category-agnostic alignment leads to the disagreement of class-specific distributions in the two domains, further causing inevitable classification errors. To overcome these two challenges, we propose a novel Semantic Conditioned AdaptatioN (SCAN) framework such that well-modeled unbiased semantics can support semantic conditioned adaptation for precise domain adaptive object detection. Specifically, class-specific semantics crossing different images in the source domain are graphically aggregated as the input to learn an unbiased semantic paradigm incrementally. The paradigm is then sent to a lightweight manifestation module to obtain conditional kernels to serve as the role of extracting semantics from the target domain for better adaptation. Subsequently, conditional kernels are integrated into global alignment to support the class-specific adaptation in a well-designed Conditional Kernel guided Alignment (CKA) module. Meanwhile, rich knowledge of the unbiased paradigm is transferred to the target domain with a novel Graph-based Semantic Transfer (GST) mechanism, yielding the adaptation in the category-based feature space. Comprehensive experiments conducted on three adaptation benchmarks demonstrate that SCAN outperforms existing works by a large margin.
Wuyang Li, Xinyu Liu 0001, Xiwen Yao, Yixuan Yuan
AAAI1
2022 SIGMA: Semantic-complete Graph Matching for Domain Adaptive Object Detection
abstract
Domain Adaptive Object Detection (DAOD) leverages a labeled domain to learn an object detector generalizing to a novel domain free of annotations. Recent advances align class-conditional distributions by narrowing down cross-domain prototypes (class centers). Though great success, they ignore the significant within-class variance and the domain-mismatched semantics within the training batch, leading to a sub-optimal adaptation. To overcome these challenges, we propose a novel SemantIc-complete Graph MAtching (SIGMA) framework for DAOD, which completes mismatched semantics and reformulates the adaptation with graph matching. Specifically, we design a Graph-embedded Semantic Completion module (GSC) that completes mis-matched semantics through generating hallucination graph nodes in missing categories. Then, we establish cross-image graphs to model class-conditional distributions and learn a graph-guided memory bank for better semantic completion in turn. After representing the source and target data as graphs, we reformulate the adaptation as a graph matching problem, i.e., finding well-matched node pairs across graphs to reduce the domain gap, which is solved with a novel Bipartite Graph Matching adaptor (BGM). In a nutshell, we utilize graph nodes to establish semantic-aware node affinity and leverage graph edges as quadratic constraints in a structure-aware matching loss, achieving fine-grained adaptation with a node-to-node graph matching. Extensive experiments verify that SIGMA outperforms existing works significantly. Our code is available at https://github.com/CityU-AIM-Group/SIGMA.
Wuyang Li, Xinyu Liu 0001, Yixuan Yuan
CVPR1
2022 Towards Robust Adaptive Object Detection under Noisy Annotations
abstract
Domain Adaptive Object Detection (DAOD) models a joint distribution of images and labels from an annotated source domain and learns a domain-invariant transformation to estimate the target labels with the given target domain images. Existing methods assume that the source domain labels are completely clean, yet large-scale datasets often contain error-prone annotations due to instance ambiguity, which may lead to a biased source distribution and severely degrade the performance of the domain adaptive detector de facto. In this paper, we represent the first effort to formulate noisy DAOD and propose a Noise Latent Transferability Exploration (NLTE) framework to address this issue. It is featured with 1) Potential Instance Mining (PIM), which leverages eligible proposals to recapture the miss-annotated instances from the background; 2) Morphable Graph Relation Module (MGRM), which models the adaptation feasibility and transition probability of noisy samples with relation matrices; 3) Entropy-Aware Gradient Reconcilement (EAGR), which incorporates the semantic information into the discrimination process and enforces the gradients provided by noisy and clean samples to be consistent towards learning domain-invariant representations. A thorough evaluation on benchmark DAOD datasets with noisy source annotations validates the effectiveness of NLTE. In particular, NLTE improves the mAP by 8.4% under 60% corrupted annotations and even approaches the ideal upper bound of training on a clean source dataset.11Code is available at https://github.com/CityU-AIM-Group/NLTE.
Xinyu Liu 0001, Wuyang Li, Qiushi Yang, Baopu Li, Yixuan Yuan
CVPR2
2022 Intervention & Interaction Federated Abnormality Detection with Noisy Clients
Xinyu Liu 0001, Wuyang Li, Yixuan Yuan
MICCAI (8)2
2021 HTD: Heterogeneous Task Decoupling for Two-Stage Object Detection
abstract
Decoupling the sibling head has recently shown great potential in relieving the inherent task-misalignment problem in two-stage object detectors. However, existing works design similar structures for the classification and regression, ignoring task-specific characteristics and feature demands. Besides, the shared knowledge that may benefit the two branches is neglected, leading to potential excessive decoupling and semantic inconsistency. To address these two issues, we propose Heterogeneous task decoupling (HTD) framework for object detection, which utilizes a Progressive Graph (PGraph) module and a Border-aware Adaptation (BA) module for task-decoupling. Specifically, we first devise a Semantic Feature Aggregation (SFA) module to aggregate global semantics with image-level supervision, serving as the shared knowledge for the task-decoupled framework. Then, the PGraph module performs progressive graph reasoning, including local spatial aggregation and global semantic interaction, to enhance semantic representations of region proposals for classification. The proposed BA module integrates multi-level features adaptively, focusing on the low-level border activation to obtain representations with spatial and border perception for regression. Finally, we utilize the aggregated knowledge from SFA to keep the instance-level semantic consistency (ISC) of decoupled frameworks. Extensive experiments demonstrate that HTD outperforms existing detection works by a large margin, and achieves single-model 50.4%AP and 33.2% APs on COCO test-dev set using ResNet-101-DCN backbone, which is the best entry among state-of-the-arts under the same configuration. Our code is available at https://github.com/CityU-AIM-Group/HTD.
Wuyang Li, Zhen Chen 0013, Baopu Li, Dingwen Zhang, Yixuan Yuan
IEEE Trans. Image Process.1