VLDB 2026 Research / reviewers in the wild / expert
Xinyu Liu 0001
dblp:98/738-1
· DBLP profile ↗
35ranked-venue papers
15as first author
35since 2021 · last 2026
0000-0002-5180-6958ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 8 first-author · 21 since 2021Artificial intelligence and machine learning · 17 · 7 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 7 first-author · 16 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Harnessing Lightweight Transformer With Contextual Synergic Enhancement for Efficient 3D Medical Image SegmentationabstractTransformers have shown remarkable performance in 3D medical image segmentation, but their high computational requirements and need for large amounts of labeled data limit their applicability. To address these challenges, we consider two crucial aspects: model efficiency and data efficiency. Specifically, we propose Light-UNETR, a lightweight transformer designed to achieve model efficiency. Light-UNETR features a Lightweight Dimension Reductive Attention (LIDR) module, which reduces spatial and channel dimensions while capturing both global and local features via multi-branch attention. Additionally, we introduce a Compact Gated Linear Unit (CGLU) to selectively control channel interaction with minimal parameters. Furthermore, we introduce a Contextual Synergic Enhancement (CSE) learning strategy, which aims to boost the data efficiency of Transformers. It first leverages the extrinsic contextual information to support the learning of unlabeled data with Attention-Guided Replacement, then applies Spatial Masking Consistency that utilizes intrinsic contextual information to enhance the spatial context reasoning for unlabeled data. Extensive experiments on various benchmarks demonstrate the superiority of our approach in both performance and efficiency. For example, with only 10% labeled data on the Left Atrial Segmentation dataset, our method surpasses BCP by 1.43% Jaccard while drastically reducing the FLOPs by 90.8% and parameters by 85.8%. Xinyu Liu 0001, Zhen Chen 0013, Wuyang Li, Chenxin Li, Yixuan Yuan |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | Fiber HGNN: Heterogeneous Graph Neural Network for Fiber Tract SegmentationabstractFiber tract segmentation is crucial for clinical applications such as brain function interpretation and surgical planning. Existing methods typically adopt either a cortical-parcellation-based or fiber clustering approach, but fail to simultaneously integrate heterogeneous information (e.g., streamline shape, point position, anatomical priors). In this work, we propose Fiber HGNN, a novel heterogeneous graph neural network that explicitly models and integrates heterogeneous information of fibers for accurate fiber tract segmentation. We construct a heterogeneous graph comprising three types of nodes: streamline, fiber keypoint and anatomical region. Specifically, fiber keypoints are representative points sampled along each streamline to characterize local geometric features, while anatomical regions provide contextual priors derived from brain atlas. This design enables the network to jointly capture the complementary information of streamline shape, local geometry, and anatomical priors, thus facilitating the learning of more discriminative feature representations. To further leverage implicit anatomical connectivity, we design a Metapath-guided Heterogeneous Information Aggregation (MHIA) network. By analyzing the spatial relationships between streamline keypoints and anatomical regions, the heterogeneous graph is decomposed into anatomical subgraphs for each streamline. In each subgraph, heterogeneous information from metapath-linked nodes is aggregated to obtain the final fiber representation. We evaluate the effectiveness of our framework on the HCP105 and TractoInferno datasets. The experimental results demonstrate that our method significantly outperforms previous state-of-the-art methods. The source code is available at https://github.com/CUHK-AIM-Group/Fiber-HGNN. Cheng Wang 0043, Wuyang Li, Xinyu Liu 0001, Yifan Liu 0010, Jian Cheng 0002, Yixuan Yuan |
IEEE Trans. Medical Imaging | 3 |
| 2025 | U-KAN Makes Strong Backbone for Medical Image Segmentation and GenerationabstractU-Net has become a cornerstone in various visual applications such as image segmentation and diffusion probability models. While numerous innovative designs and improvements have been introduced by incorporating transformers or MLPs, the networks are still limited to linearly modeling patterns as well as the deficient interpretability. To address these challenges, our intuition is inspired by the impressive results of the Kolmogorov-Arnold Networks (KANs) in terms of accuracy and interpretability, which reshape the neural network learning via the stack of non-linear learnable activation functions derived from the Kolmogorov-Anold representation theorem. Specifically, in this paper, we explore the untapped potential of KANs in improving backbones for vision tasks. We investigate, modify and re-design the established U-Net pipeline by integrating the dedicated KAN layers on the tokenized intermediate representation, termed U-KAN. Rigorous medical image segmentation benchmarks verify the superiority of UKAN by higher accuracy even with less computation cost. We further delved into the potential of U-KAN as an alternative U-Net noise predictor in diffusion models, demonstrating its applicability in generating task-oriented model architectures. Chenxin Li, Xinyu Liu 0001, Wuyang Li, Cheng Wang 0043, Hengyu Liu 0007, Yifan Liu 0010, Zhen Chen 0013, Yixuan Yuan |
AAAI | 2 |
| 2025 | Track Any Anomalous Object: A Granular Video Anomaly Detection PipelineabstractVideo anomaly detection (VAD) is crucial in scenarios such as surveillance and autonomous driving, where timely detection of unexpected activities is essential. Albeit existing methods have primarily focused on detecting anomalous objects in videos—either by identifying anomalous frames or objects—they often neglect finer-grained analysis, such as anomalous pixels, which limits their ability to capture a broader range of anomalies. To address this challenge, we propose an innovative VAD framework called Track Any Anomalous Object (TAO), which introduces a Granular Video Anomaly Detection Framework that, for the first time, integrates the detection of multiple fine-grained anomalous objects into a unified framework. Unlike methods that assign anomaly scores to every pixel at each moment, our approach transforms the problem into pixel-level tracking of anomalous objects. By linking anomaly scores to subsequent tasks such as image segmentation and video tracking, our method eliminates the need for threshold selection and achieves more precise anomaly localization, even in long and challenging video sequences. Experiments on extensive datasets demonstrate that TAO achieves state-of-the-art performance, setting a new progress for VAD by providing a practical, granular, and holistic solution. For more information, visit the project page at: https://tao-25.github.io/ Yuzhi Huang, Chenxin Li, Zixu Lin, Yunlong Lin, Hengyu Liu 0007, Wuyang Li, Xinyu Liu 0001, Jiechao Gao, Yue Huang 0001, Xinghao Ding, Yixuan Yuan |
CVPR | 8 |
| 2025 | InfoBridge: Balanced Multimodal Integration through Conditional Dependency Modeling
Chenxin Li, Yifan Liu 0010, Panwang Pan, Hengyu Liu 0007, Xinyu Liu 0001, Wuyang Li, Cheng Wang 0043, Weihao Yu 0004, Yiyang Lin, Yixuan Yuan |
ICCV | 5 |
| 2025 | MedVSR: Medical Video Super-Resolution with Cross State-Space PropagationabstractHigh-resolution (HR) medical videos are vital for accurate diagnosis, yet are hard to acquire due to hardware limitations and physiological constraints. Clinically, the collected low-resolution (LR) medical videos present unique challenges for video super-resolution (VSR) models, including camera shake, noise, and abrupt frame transitions, which result in significant optical flow errors and alignment difficulties. Additionally, tissues and organs exhibit continuous and nuanced structures, but current VSR models are prone to introducing artifacts and distorted features that can mislead doctors. To this end, we propose MedVSR, a tailored framework for medical VSR. It first employs Cross State-Space Propagation (CSSP) to address the imprecise alignment by projecting distant frames as control matrices within state-space models, enabling the selective propagation of consistent and informative features to neighboring frames for effective alignment. Moreover, we design an Inner State-Space Reconstruction (ISSR) module that enhances tissue structures and reduces artifacts with joint long-range spatial feature learning and large-kernel short-range information aggregation. Experiments across four datasets in diverse medical scenarios, including endoscopy and cataract surgeries, show that MedVSR significantly outperforms existing VSR models in reconstruction performance and efficiency. Code released at https://github.com/CUHK-AIM-Group/MedVSR. Xinyu Liu 0001, Guolei Sun, Cheng Wang 0043, Yixuan Yuan, Ender Konukoglu |
ICCV | 1 |
| 2025 | GaussianReg: Rapid 2D/3D Registration for Emergency Surgery Via Explicit 3D Modeling with Gaussian Primitives
Weihao Yu 0004, Xiaoqing Guo, Xinyu Liu 0001, Yifan Liu 0010, Hao Zheng 0008, Yawen Huang, Yixuan Yuan |
ICCV | 3 |
| 2025 | EndoGen: Conditional Autoregressive Endoscopic Video Generation
Xinyu Liu 0001, Hengyu Liu 0007, Cheng Wang 0043, Tianming Liu 0001, Yixuan Yuan |
MICCAI (10) | 1 |
| 2025 | STEAM: Self-supervised TEeth Analysis and Modeling for Point Cloud Segmentation
Yifan Liu 0010, Chen Yang 0026, Weihao Yu 0005, Xinyu Liu 0001, Hui Chen 0032, Max Q.-H. Meng, Yixuan Yuan |
MICCAI (9) | 4 |
| 2025 | UN-SAM: Domain-adaptive self-prompt segmentation for universal nuclei images
Zhen Chen 0013, Qing Xu 0014, Xinyu Liu 0001, Yixuan Yuan |
Medical Image Anal. | 3 |
| 2025 | FM-APP: Foundation Model for Any Phenotype Prediction via fMRI to sMRI Knowledge TransferabstractPredicting individual-level non-neuroimaging phenotypes (e.g., fluid intelligence) using brain imaging data is a fundamental goal of neuroscience. Recent research has focused on utilizing high-cost functional magnetic resonance imaging (fMRI) to predict phenotypes seen during training. However, these methods 1) only consider predicting seen phenotypes, failing to achieve zero-shot inference for unseen phenotypes; 2) overlook the knowledge transfer from fMRI to structural MRI (sMRI), missing out on utilizing cost-effective sMRI for accurate predictions. To address these challenges, we propose a Foundational Model for Any Phenotype Prediction via fMRI to sMRI knowledge transfer (FM-APP), consisting of a Phenotypes Text Memory Bank (PTMB) module, Any Phenotype Prediction (APP) module, and fMRI to sMRI Knowledge Transfer (F2SKT) module. Our proposed FM-APP adapts to downstream tasks by generating regressor parameters instead of fine-tuning the model itself. Specifically, to retain important clues from seen phenotype descriptions, PTMB utilizes the BiomedCLIP model to store semantic features of seen phenotypes. To achieve any phenotype prediction, the APP introduces a regressor synthesizer for zero-shot inference. Additionally, to improve sMRI prediction accuracy while preserving its cost advantage, the F2SKT uses the PTMB to construct phenotype active maps, guiding adaptive knowledge transfer from fMRI to sMRI. Experiments on the Human Connectome Project (HCP) and HCP Aging datasets demonstrate our approach outperforms state-of-the-art methods, showcasing strong zero-shot inference capabilities and providing a novel framework for analyzing brain structure and phenotypes. Our code: https://github.com/ZhibinHe/FM-APP. Wuyang Li, Yifan Liu 0010, Xinyu Liu 0001, Junwei Han 0001, Yixuan Yuan |
IEEE Trans. Medical Imaging | 4 |
| 2025 | ToothMaker: Realistic Panoramic Dental Radiograph Generation via Disentangled ControlabstractGenerating high-fidelity dental radiographs is essential for training diagnostic models. Despite the development of numerous methods for other medical data, generative approaches in dental radiology remain unexplored. Due to the intricate tooth structures and specialized terminology, these methods often yield ambiguous tooth regions and incorrect dental concepts when applied to dentistry. In this paper, we take the first attempt to investigate diffusion-based teeth X-ray image generation and propose ToothMaker, a novel framework specifically designed for the dental domain. Firstly, to synthesize X-ray images that possess accurate tooth structures and realistic radiological styles simultaneously, we design control-disentangled fine-tuning (CDFT) strategy. Specifically, we present two separate controllers to handle style and layout control respectively, and introduce a gradient-based decoupling method that optimizes each using their corresponding disentangled gradients. Secondly, to enhance model's understanding of dental terminology, we propose prior-disentangled guidance module (PDGM), enabling precise synthesis of dental concepts. It utilizes large language model to decompose dental terminology into a series of meta-knowledge elements and performs interactions and refinements through hypergraph neural network. These elements are then fed into the network to guide the generation of dental concepts. Extensive experiments demonstrate the high fidelity and diversity of the images synthesized by our approach. By incorporating the generated data, we achieve substantial performance improvements on downstream segmentation and visual question answering tasks, indicating that our method can greatly reduce the reliance on manually annotated data. Code will be public available at https://github.com/CUHK-AIM-Group/ToothMaker. Weihao Yu 0004, Xiaoqing Guo, Wuyang Li, Xinyu Liu 0001, Hui Chen 0032, Yixuan Yuan |
IEEE Trans. Medical Imaging | 4 |
| 2024 | PV-SSM: Exploring Pure Visual State Space Model for High-dimensional Medical Data AnalysisabstractDespite previous endeavors to utilize Convolutional Neural Networks and Transformers as base networks for medical image analysis, their architectures still harbor inherent limitations: either an inability to model long-range dependencies or colossal computational consumption due to global self-attention. Recently, State Space Models (SSMs) have exhibited impressive capabilities in modeling long-term dependencies with satisfactory linear computational complexity. Nevertheless, extant medical visual SSMs are constrained by their limited capacity to capture inter-patch relationships and inefficient modeling due to the introduction of additional depth convolutions to handle high-dimensional data. In this paper, we propose a novel, Pure Visual State Space Model (PV-SSM) for high-dimensional medical data analysis. Different from prior medical visual SSMs, our proposed framework does not involve any convolutional or global attention operations while leverages a series of Pure-SSM blocks that employ a novel parallel-SSM mechanism to simultaneously extract feature data across different dimensions. Furthermore, we propose a learnable Parameterized Positional Encoding, which incorporates absolute positional information into patch features, effectively endowing inter-patch relationships with stronger inferential capabilities. We conducted extensive validation on various modalities of medical imaging data. Experimental results demonstrate superior performance and efficacy of our model against existing models. Our codes are available at https://github.com/chengwang96/PV-SSM Cheng Wang 0043, Xinyu Liu 0001, Chenxin Li, Yifan Liu 0010, Yixuan Yuan |
BIBM | 2 |
| 2024 | CLIFF: Continual Latent Diffusion for Open-Vocabulary Object Detection
Wuyang Li, Xinyu Liu 0001, Jiayi Ma 0001, Yixuan Yuan |
ECCV (55) | 2 |
| 2024 | GTP-4o: Modality-Prompted Heterogeneous Graph Learning for Omni-Modal Biomedical Representation
Chenxin Li, Xinyu Liu 0001, Cheng Wang 0043, Yifan Liu 0010, Weihao Yu 0005, Yixuan Yuan |
ECCV (4) | 2 |
| 2024 | 👦 Endora: Video Generation Models as Endoscopy Simulators
Chenxin Li, Hengyu Liu 0007, Yifan Liu 0010, Brandon Yushan Feng, Wuyang Li, Xinyu Liu 0001, Zhen Chen 0013, Yixuan Yuan |
MICCAI (6) | 6 |
| 2024 | From Static to Dynamic Diagnostics: Boosting Medical Image Analysis via Motion-Informed Generative Videos
Wuyang Li, Xinyu Liu 0001, Qiushi Yang, Yixuan Yuan |
MICCAI (3) | 2 |
| 2024 | MOST: Multi-formation Soft Masking for Semi-supervised Medical Image Segmentation
Xinyu Liu 0001, Zhen Chen 0013, Yixuan Yuan |
MICCAI (11) | 1 |
| 2024 | DiffRect: Latent Diffusion Label Rectification for Semi-supervised Medical Image Segmentation
Xinyu Liu 0001, Wuyang Li, Yixuan Yuan |
MICCAI (12) | 1 |
| 2024 | Flaws can be Applause: Unleashing Potential of Segmenting Ambiguous Objects in SAMabstractAs the vision foundation models like the Segment Anything Model (SAM) demonstrate potent universality, they also present challenges in giving ambiguous and uncertain predictions. Significant variations in the model output and granularity can occur with simply subtle changes in the prompt, contradicting the consensus requirement for the robustness of a model. While some established works have been dedicated to stabilizing and fortifying the prediction of SAM, this paper takes a unique path to explore how this flaw can be inverted into an advantage when modeling inherently ambiguous data distributions. We introduce an optimization framework based on a conditional variational autoencoder, which jointly models the prompt and the granularity of the object with a latent probability distribution. This approach enables the model to adaptively perceive and represent the real ambiguous label distribution, taming SAM to produce a series of diverse, convincing, and reasonable segmentation outputs controllably. Extensive experiments on several practical deployment scenarios involving ambiguity demonstrates the exceptional performance of our framework. Project page: \url{https://a-sa-m.github.io/}. Chenxin Li, Yuzhi Huang, Wuyang Li, Hengyu Liu 0007, Xinyu Liu 0001, Qing Xu 0014, Zhen Chen 0013, Yue Huang 0001, Yixuan Yuan |
NeurIPS | 5 |
| 2024 | Decoupled Unbiased Teacher for Source-Free Domain Adaptive Medical Object DetectionabstractSource-free domain adaptation (SFDA) aims to adapt a lightweight pretrained source model to unlabeled new domains without the original labeled source data. Due to the privacy of patients and storage consumption concerns, SFDA is a more practical setting for building a generalized model in medical object detection. Existing methods usually apply the vanilla pseudo-labeling technique, while neglecting the bias issues in SFDA, leading to limited adaptation performance. To this end, we systematically analyze the biases in SFDA medical object detection by constructing a structural causal model (SCM) and propose an unbiased SFDA framework dubbed decoupled unbiased teacher (DUT). Based on the SCM, we derive that the confounding effect causes biases in the SFDA medical object detection task at the sample level, feature level, and prediction level. To prevent the model from emphasizing easy object patterns in the biased dataset, a dual invariance assessment (DIA) strategy is devised to generate counterfactual synthetics. The synthetics are based on unbiased invariant samples in both discrimination and semantic perspectives. To alleviate overfitting to domain-specific features in SFDA, we design a cross-domain feature intervention (CFI) module to explicitly deconfound the domain-specific prior with feature intervention and obtain unbiased features. Besides, we establish a correspondence supervision prioritization (CSP) strategy for addressing the prediction bias caused by coarse pseudo-labels by sample prioritizing and robust box supervision. Through extensive experiments on multiple SFDA medical object detection scenarios, DUT yields superior performance over previous state-of-the-art unsupervised domain adaptation (UDA) and SFDA counterparts, demonstrating the significance of addressing the bias issues in this challenging task. The code is available at https://github.com/CUHK-AIM-Group/Decoupled-Unbiased-Teacher. Xinyu Liu 0001, Wuyang Li, Yixuan Yuan |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | EfficientViT: Memory Efficient Vision Transformer with Cascaded Group AttentionabstractVision transformers have shown great success due to their high model capabilities. However, their remarkable performance is accompanied by heavy computation costs, which makes them unsuitable for real-time applications. In this paper, we propose a family of high-speed vision transformers named Efficient ViT. We find that the speed of existing transformer models is commonly bounded by memory inefficient operations, especially the tensor reshaping and element-wise functions in MHSA. Therefore, we design a new building block with a sandwich layout, i.e., using a single memory-bound MHSA between efficient FFN layers, which improves memory efficiency while enhancing channel communication. Moreover, we discover that the attention maps share high similarities across heads, leading to computational redundancy. To address this, we present a cascaded group attention module feeding attention heads with different splits of the full feature, which not only saves computation cost but also improves attention diversity. Comprehensive experiments demonstrate EfficientViT outperforms existing efficient models, striking a good trade-off between speed and accuracy. For instance, our EfficientViT-M5 surpasses MobileNetV3-Large by 1.9% in accuracy, while getting 40.4% and 45.2% higher throughput on Nvidia V100 GPU and Intel Xeon CPU, respectively. Compared to the recent efficient model MobileViT-XXS, EfficientViT-M2 achieves 1.8% superior accuracy, while running$5.8\times/3.7\times$faster on the GPU/CPU, and$7.4\times faster$when converted to ONNX format. Code and models are available at here. Xinyu Liu 0001, Houwen Peng, Ningxin Zheng, Yuqing Yang 0001, Han Hu 0001, Yixuan Yuan |
CVPR | 1 |
| 2023 | Generalized Gradient Flow Based Saliency for Pruning Deep Convolutional Neural Networks
Xinyu Liu 0001, Baopu Li, Zhen Chen 0013, Yixuan Yuan |
Int. J. Comput. Vis. | 1 |
| 2023 | SIGMA++: Improved Semantic-Complete Graph Matching for Domain Adaptive Object DetectionabstractDomain Adaptive Object Detection (DAOD) generalizes the object detector from an annotated domain to a label-free novel one. Recent works estimate prototypes (class centers) and minimize the corresponding distances to adapt the cross-domain class conditional distribution. However, this prototype-based paradigm 1) fails to capture the class variance with agnostic structural dependencies, and 2) ignores the domain-mismatched classes with a sub-optimal adaptation. To address these two challenges, we propose an improved SemantIc-complete Graph MAtching framework, dubbed SIGMA++, for DAOD, completing mismatched semantics and reformulating adaptation with hypergraph matching. Specifically, we propose a Hypergraphical Semantic Completion (HSC) module to generate hallucination graph nodes in mismatched classes. HSC builds a cross-image hypergraph to model class conditional distribution with high-order dependencies and learns a graph-guided memory bank to generate missing semantics. After representing the source and target batch with hypergraphs, we reformulate domain adaptation with a hypergraph matching problem, i.e., discovering well-matched nodes with homogeneous semantics to reduce the domain gap, which is solved with a Bipartite Hypergraph Matching (BHM) module. Graph nodes are used to estimate semantic-aware affinity, while edges serve as high-order structural constraints in a structure-aware matching loss, achieving fine-grained adaptation with hypergraph matching. The applicability of various object detectors verifies the generalization of SIGMA++, and extensive experiments on nine benchmarks show its state-of-the-art performance on both AP$_{50}$and adaptation gains. Wuyang Li, Xinyu Liu 0001, Yixuan Yuan |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | SCAN++: Enhanced Semantic Conditioned Adaptation for Domain Adaptive Object DetectionabstractDomain Adaptive Object Detection (DAOD) transfers an object detector from the labeled source domain to a novel unlabelled target domain. Recent advances bridge the domain gap by aligning category-agnostic feature distribution and minimizing the domain discrepancy for adapting semantic distribution. Though great success, these methods model domain discrepancy with prototypes within a batch, yielding a biased estimation of domain-level statistics. Moreover, the category-agnostic alignment leads to the disagreement of the cross-domain semantic distribution with inevitable classification errors. To address these two issues, we propose an enhanced Semantic Conditioned AdaptatioN (SCAN++) framework, which leverages unbiased semantics for DAOD. Specifically, in the source domain, we design the conditional kernel to sample Pixel of Interests (PoIs), and aggregate PoIs with a cross-image graph to estimate an unbiased semantic sequence. Conditioned on the semantic sequence, we further update the parameter of the conditional kernel in a semantic conditioned manifestation module, and establish a novel conditional graph in the target domain to model unlabeled semantics. After modeling the semantic distribution in both domains, we integrate the conditional kernel into adversarial alignment to achieve semantic-aware adaptation in a Conditional Kernel guided Alignment (CKA) module. Meanwhile, the Semantic Sequence guided Transport (SST) module is proposed to transfer reliable semantic knowledge to the target domain through solving the cross-domain Optimal Transport (OT) assignment, achieving unbiased adaptation at the semantic level. Comprehensive experiments on four adaptation scenarios demonstrate that SCAN++ achieves state-of-the-art results. The code is available athttps://github.com/CityU-AIM-Group/SCAN/tree/SCAN++. Wuyang Li, Xinyu Liu 0001, Yixuan Yuan |
IEEE Trans. Multim. | 2 |
| 2022 | SCAN: Cross Domain Object Detection with Semantic Conditioned AdaptationabstractThe domain gap severely limits the transferability and scalability of object detectors trained in a specific domain when applied to a novel one. Most existing works bridge the domain gap by minimizing the domain discrepancy in the category space and aligning category-agnostic global features. Though great success, these methods model domain discrepancy with prototypes within a batch, yielding a biased estimation of domain-level distribution. Besides, the category-agnostic alignment leads to the disagreement of class-specific distributions in the two domains, further causing inevitable classification errors. To overcome these two challenges, we propose a novel Semantic Conditioned AdaptatioN (SCAN) framework such that well-modeled unbiased semantics can support semantic conditioned adaptation for precise domain adaptive object detection. Specifically, class-specific semantics crossing different images in the source domain are graphically aggregated as the input to learn an unbiased semantic paradigm incrementally. The paradigm is then sent to a lightweight manifestation module to obtain conditional kernels to serve as the role of extracting semantics from the target domain for better adaptation. Subsequently, conditional kernels are integrated into global alignment to support the class-specific adaptation in a well-designed Conditional Kernel guided Alignment (CKA) module. Meanwhile, rich knowledge of the unbiased paradigm is transferred to the target domain with a novel Graph-based Semantic Transfer (GST) mechanism, yielding the adaptation in the category-based feature space. Comprehensive experiments conducted on three adaptation benchmarks demonstrate that SCAN outperforms existing works by a large margin. Wuyang Li, Xinyu Liu 0001, Xiwen Yao, Yixuan Yuan |
AAAI | 2 |
| 2022 | SIGMA: Semantic-complete Graph Matching for Domain Adaptive Object DetectionabstractDomain Adaptive Object Detection (DAOD) leverages a labeled domain to learn an object detector generalizing to a novel domain free of annotations. Recent advances align class-conditional distributions by narrowing down cross-domain prototypes (class centers). Though great success, they ignore the significant within-class variance and the domain-mismatched semantics within the training batch, leading to a sub-optimal adaptation. To overcome these challenges, we propose a novel SemantIc-complete Graph MAtching (SIGMA) framework for DAOD, which completes mismatched semantics and reformulates the adaptation with graph matching. Specifically, we design a Graph-embedded Semantic Completion module (GSC) that completes mis-matched semantics through generating hallucination graph nodes in missing categories. Then, we establish cross-image graphs to model class-conditional distributions and learn a graph-guided memory bank for better semantic completion in turn. After representing the source and target data as graphs, we reformulate the adaptation as a graph matching problem, i.e., finding well-matched node pairs across graphs to reduce the domain gap, which is solved with a novel Bipartite Graph Matching adaptor (BGM). In a nutshell, we utilize graph nodes to establish semantic-aware node affinity and leverage graph edges as quadratic constraints in a structure-aware matching loss, achieving fine-grained adaptation with a node-to-node graph matching. Extensive experiments verify that SIGMA outperforms existing works significantly. Our code is available at https://github.com/CityU-AIM-Group/SIGMA. Wuyang Li, Xinyu Liu 0001, Yixuan Yuan |
CVPR | 2 |
| 2022 | Towards Robust Adaptive Object Detection under Noisy AnnotationsabstractDomain Adaptive Object Detection (DAOD) models a joint distribution of images and labels from an annotated source domain and learns a domain-invariant transformation to estimate the target labels with the given target domain images. Existing methods assume that the source domain labels are completely clean, yet large-scale datasets often contain error-prone annotations due to instance ambiguity, which may lead to a biased source distribution and severely degrade the performance of the domain adaptive detector de facto. In this paper, we represent the first effort to formulate noisy DAOD and propose a Noise Latent Transferability Exploration (NLTE) framework to address this issue. It is featured with 1) Potential Instance Mining (PIM), which leverages eligible proposals to recapture the miss-annotated instances from the background; 2) Morphable Graph Relation Module (MGRM), which models the adaptation feasibility and transition probability of noisy samples with relation matrices; 3) Entropy-Aware Gradient Reconcilement (EAGR), which incorporates the semantic information into the discrimination process and enforces the gradients provided by noisy and clean samples to be consistent towards learning domain-invariant representations. A thorough evaluation on benchmark DAOD datasets with noisy source annotations validates the effectiveness of NLTE. In particular, NLTE improves the mAP by 8.4% under 60% corrupted annotations and even approaches the ideal upper bound of training on a clean source dataset.11Code is available at https://github.com/CityU-AIM-Group/NLTE. Xinyu Liu 0001, Wuyang Li, Qiushi Yang, Baopu Li, Yixuan Yuan |
CVPR | 1 |
| 2022 | Intervention & Interaction Federated Abnormality Detection with Noisy Clients
Xinyu Liu 0001, Wuyang Li, Yixuan Yuan |
MICCAI (8) | 1 |
| 2022 | Semi-supervised Medical Image Classification with Temporal Knowledge-Aware Regularization
Qiushi Yang, Xinyu Liu 0001, Zhen Chen 0013, Bulat Ibragimov, Yixuan Yuan |
MICCAI (8) | 2 |
| 2022 | Global Context Parallel Attention for Anchor-Free Instance Segmentation in Remote Sensing ImagesabstractSegmenting objects in optical remote sensing images has always been a hot topic for remote sensing image researchers. However, many previous works used segmentation algorithms designed for common objects without modification, leading to slow and poor results. In this work, we exploit self-attention mechanism into anchor-free segmentation architectures to improve the segmentation accuracy for objects in high-resolution remote sensing images. The proposed module integrates the self-attention mechanism, namely the global context parallel attention module (GC-PAM). It is composed of a parallel global context channel self-attention block and a spatial self-attention block. By implementing our GC-PAM in an anchor-free network, the channel-wise and spatial-wise weights are both reassigned, which can improve the segmentation accuracy significantly. Xinyu Liu 0001, Xiaoguang Di |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | A Source-Free Domain Adaptive Polyp Detection Framework With Style Diversification FlowabstractThe automatic detection of polyps across colonoscopy and Wireless Capsule Endoscopy (WCE) datasets is crucial for early diagnosis and curation of colorectal cancer. Existing deep learning approaches either require mass training data collected from multiple sites or use unsupervised domain adaptation (UDA) technique with labeled source data. However, these methods are not applicable when the data is not accessible due to privacy concerns or data storage limitations. Aiming to achieve source-free domain adaptive polyp detection, we propose a consistency based model that utilizes Source Model as Proxy Teacher (SMPT) with only a transferable pretrained model and unlabeled target data. SMPT first transfers the stored domain-invariant knowledge in the pretrained source model to the target model via Source Knowledge Distillation (SKD), then uses Proxy Teacher Rectification (PTR) to rectify the source model with temporal ensemble of the target model. Moreover, to alleviate the biased knowledge caused by domain gaps, we propose Uncertainty-Guided Online Bootstrapping (UGOB) to adaptively assign weights for each target image regarding their uncertainty. In addition, we design Source Style Diversification Flow (SSDF) that gradually generates diverse style images and relaxes style-sensitive channels based on source and target information to enhance the robustness of the model towards style variation. The capacities of SMPT and SSDF are further boosted with iterative optimization, constructing a stronger framework SMPT++ for cross-domain polyp detection. Extensive experiments are conducted on five distinct polyp datasets under two types of cross-domain settings. Our proposed method shows the state-of-the-art performance and even outperforms previous UDA approaches that require the source data by a large margin. The source code is available at github.com/CityU-AIM-Group/SFPolypDA. Xinyu Liu 0001, Yixuan Yuan |
IEEE Trans. Medical Imaging | 1 |
| 2021 | Exploring Gradient Flow Based Saliency for DNN Model CompressionabstractModel pruning aims to reduce the deep neural network (DNN) model size or computational overhead. Traditional model pruning methods such as l-1 pruning that evaluates the channel significance for DNN pay too much attention to the local analysis of each channel and make use of the magnitude of the entire feature while ignoring its relevance to the batch normalization (BN) and ReLU layer after each convolutional operation. To overcome these problems, we propose a new model pruning method from a new perspective of gradient flow in this paper. Specifically, we first theoretically analyze the channel's influence based on Taylor expansion by integrating the effects of BN layer and ReLU activation function. Then, the incorporation of the first-order Talyor polynomial of the scaling parameter and the shifting parameter in the BN layer is suggested to effectively indicate the significance of a channel in a DNN. Comprehensive experiments on both image classification and image denoising tasks demonstrate the superiority of the proposed novel theory and scheme. Code is available at https://github.com/CityU-AIM-Group/GFBS. Xinyu Liu 0001, Baopu Li, Zhen Chen 0013, Yixuan Yuan |
ACM Multimedia | 1 |
| 2021 | TanhExp: A smooth activation function with high convergence speed for lightweight neural networksabstractAbstract Lightweight or mobile neural networks used for real‐time computer vision tasks contain fewer parameters than normal networks, which lead to a constrained performance. Herein, a novel activation function named as Tanh Exponential Activation Function (TanhExp) is proposed which can improve the performance for these networks on image classification task significantly. The definition of TanhExp is f ( x ) = x tanh( e x ). The simplicity, efficiency, and robustness of TanhExp on various datasets and network models is demonstrated and TanhExp outperforms its counterparts in both convergence speed and accuracy. Its behaviour also remains stable even with noise added and dataset altered. It is shown that without increasing the size of the network, the capacity of lightweight neural networks can be enhanced by TanhExp with only a few training epochs and no extra parameters added. Xinyu Liu 0001, Xiaoguang Di |
IET Comput. Vis. | 1 |
| 2021 | Consolidated domain adaptive detection and localization framework for cross-device colonoscopic images
Xinyu Liu 0001, Xiaoqing Guo, Yixuan Yuan |
Medical Image Anal. | 1 |