EDBT 2026 Demo / reviewers in the wild / expert
Jie Liu 0044
dblp:03/2134-44
· DBLP profile ↗
38ranked-venue papers
9as first author
37since 2021 · last 2026
0000-0002-1327-1315ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 25 · 6 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 3 first-author · 12 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Benchmarking Egocentric Clinical Intent Understanding Capability for Medical Multimodal Large Language ModelsabstractShaonan Liu, Guo Yu, Xiaoling Luo, Shiyi Zheng, Jie Liu, Wenting Chen, Linlin Shen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Shaonan Liu, Xiaoling Luo 0001, Shiyi Zheng, Jie Liu 0044, Wenting Chen, LinLin Shen |
ACL (1) | 5 |
| 2026 | Task-Aware LLM Routing with Multi-Level Task-Profile-Guided Data Synthesis for Cold-Start ScenariosabstractLarge language models (LLMs) exhibit substantial variability in performance and computational cost across tasks and queries, motivating routing systems that select models to meet user-specific cost--performance trade-offs. However, existing routers generalize poorly in cold-start scenarios where in-domain training data is unavailable. We address this limitation with a multi-level task-profile--guided data synthesis framework that constructs a hierarchical task taxonomy and produces diverse question--answer pairs to approximate the test-time query distribution. Building on this, we introduce TRouter, a task-type--aware router approach that models query-conditioned cost and performance via latent task-type variables, with prior regularization derived from the synthesized task taxonomy. This design enhances TRouter's routing utility under both cold-start and in-domain settings. Across multiple benchmarks, we show that our synthesis framework alleviates cold-start issues and that TRouter delivers effective LLM routing. ©2026 Association for Computational Linguistics Hui Liu 0036, Kecheng Chen, Jie Liu 0044, Wenya Wang 0001, Haoliang Li |
ACL (1) | 4 |
| 2026 | Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language ModelsabstractWenxuan Wang, Zizhan Ma, Guo Yu, Yiu-Fai Cheung, Meidan Ding, Jie Liu, Wenting Chen, Linlin Shen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Wenxuan Wang 0001, Zizhan Ma, Yiu-Fai Cheung, Meidan Ding, Jie Liu 0044, Wenting Chen, LinLin Shen |
ACL (1) | 6 |
| 2026 | DP-DGAD: A Generalist Dynamic Graph Anomaly Detector with Dynamic PrototypesabstractDynamic graph anomaly detection (DGAD) is essential for iden- tifying anomalies in evolving graphs across domains such as fi- nance and social networks. Recently, generalist graph anomaly detection (GAD) models have shown promising results. They are pretrained on multiple source datasets and generalize across do- mains. While effective on static graphs, they struggle to capture evolving anomalies in dynamic graphs. Moreover, the continuous emergence of new domains and the lack of labeled data further challenge generalist DGAD. Effective cross-domain DGAD requires both domain-specific and domain-agnostic anomalous patterns. Importantly, these patterns evolve temporally within and across domains. Building on these insights, we propose a DGAD model with Dynamic Prototypes (DP) to capture evolving domain-specific and domain-agnostic patterns. Firstly, DP-DGAD extracts dynamic prototypes, i.e., evolving representations of normal and anomalous patterns, from temporal ego-graphs and stores them in a memory buffer. The buffer is selectively updated to retain general, domain- agnostic patterns while incorporating new domain-specific ones. Then, an anomaly scorer compares incoming data with dynamic prototypes to flag both general and domain-specific anomalies. Fi- nally, DP-DGAD employs confidence detection guided memory buffer updating for effective adaptation to target domain. Extensive experiments demonstrate state-of-the-art performance across ten real-world datasets from different domains. Jialun Zheng, Jie Liu 0044, Jiannong Cao 0001, Xiao Wang 0017, Hanchen Yang 0002, Yankai Chen 0001 |
WWW | 2 |
| 2026 | Enhancing 3D medical multi-modal large language models with integrated human body priors for computed tomography
Leilei Zeng, Jie Liu 0044, Wenting Chen, Chenyang Lyu, Wenxi Li, Shaonan Liu, Xiande Zhou, LinLin Shen |
Pattern Recognit. | 2 |
| 2025 | DAMPER: A Dual-Stage Medical Report Generation Framework with Coarse-Grained MeSH Alignment and Fine-Grained Hypergraph MatchingabstractMedical report generation is crucial for clinical diagnosis and patient management, summarizing diagnoses and recommendations based on medical imaging. However, existing work often overlook the clinical pipeline involved in report writing, where physicians typically conduct an initial quick review followed by a detailed examination. Moreover, current alignment methods may lead to misaligned relationships. To address these issues, we propose DAMPER, a dual-stage framework for medical report generation that mimics the clinical pipeline of report writing in two stages. In the first stage, a MeSH-Guided Coarse-Grained Alignment (MCG) stage that aligns chest X-ray (CXR) image features with medical subject headings (MeSH) features to generate a rough keyphrase representation of the overall impression. In the second stage, a Hypergraph-Enhanced Fine-Grained Alignment (HFG) stage that constructs hypergraphs for image patches and report annotations, modeling high-order relationships within each modality and performing hypergraph matching to capture semantic correlations between image regions and textual phrases. Finally,the coarse-grained visual features, generated MeSH representations, and visual hypergraph features are fed into a report decoder to produce the final medical report. Extensive experiments on public datasets demonstrate the effectiveness of DAMPER in generating comprehensive and accurate medical reports, outperforming state-of-the-art methods across various evaluation metrics. Wenting Chen, Jie Liu 0044, Qisheng Lu, Xiaoling Luo 0001, LinLin Shen |
AAAI | 3 |
| 2025 | Asclepius: A Spectrum Evaluation Benchmark for Medical Multi-Modal Large Language ModelsabstractThe significant breakthroughs of Medical Multi-Modal Large Language Models (Med-MLLMs) renovate modern healthcare with robust information synthesis and medical decision support. However, these models are often evaluated on benchmarks that are unsuitable for the Med-MLLMs due to the complexity of real-world diagnostics across diverse specialties. To address this gap, we introduce Asclepius, a novel Med-MLLM benchmark that comprehensively assesses Med-MLLMs in terms of: distinct medical specialties (cardiovascular, gas-troenterology, etc.) and different diagnostic capacities (perception, disease analysis, etc.). Grounded in 3 proposed core principles, Asclepius ensures a comprehensive evaluation by encompassing 15 medical specialties, stratifying into 3 main categories and 8 sub-categories of clinical tasks, and exempting overlap with existing VQA dataset. We further provide an in-depth analysis of 6 Med-MLLMs and compare them with 3 human specialists, providing insights into their competencies and limitations in various medical contexts. Our work not only advances the understanding of Med-MLLMs' capabilities but also sets a precedent for future evaluations and the safe deployment of these models in clinical environments. © 2025 Association for Computational Linguistics. Jie Liu 0044, Wenxuan Wang 0001, Yihang Su, Yudi Zhang 0005, Cheng-Yi Li, Wenting Chen, Xiaohan Xing, Kao-Jung Chang, LinLin Shen, Michael R. Lyu |
ACL (1) | 1 |
| 2025 | Q-PART: Quasi-Periodic Adaptive Regression with Test-time Training for Pediatric Left Ventricular Ejection Fraction RegressionabstractIn this work, we address the challenge of adaptive pediatric Left Ventricular Ejection Fraction (LVEF) assessment. While Test-time Training (TTT) approaches show promise for this task, they suffer from two significant limitations. Existing TTT works are primarily designed for classification tasks rather than continuous value regression, and they lack mechanisms to handle the quasi-periodic nature of cardiac signals. To tackle these issues, we propose a novel Quasi-Periodic Adaptive Regression with Test-time Training (Q-PART) framework. In the training stage, the proposed Quasi-Period Network decomposes the echocardiogram into periodic and aperiodic components within latent space by combining parameterized helix trajectories with Neural Controlled Differential Equations. During inference, our framework further employs a variance minimization strategy across image augmentations that simulate common quality issues in echocardiogram acquisition, along with differential adaptation rates for periodic and aperiodic components. Theoretical analysis is provided to demonstrate that our variance minimization objective effectively bounds the regression error under mild conditions. Furthermore, extensive experiments across three pediatric age groups demonstrate that Q-PART not only significantly outperforms existing approaches in pediatric LVEF prediction, but also exhibits strong clinical screening capability with high mAUROC scores (up to 0.9747) and maintains gender-fair performance across all metrics, validating its robustness and practical utility in pediatric echocardiography analysis. The project can be found in Q-PART. Jie Liu 0044, Tiexin Qin, Hui Liu 0036, Yilei Shi, Lichao Mou, Xiao Xiang Zhu 0001, Shiqi Wang 0001, Haoliang Li |
CVPR | 1 |
| 2025 | Test-time Adaptation for Foundation Medical Segmentation Model without Parametric UpdatesabstractFoundation medical segmentation models, with MedSAM being the most popular, have achieved promising performance across organs and lesions. However, MedSAM still suffers from compromised performance on specific lesions with intricate structures and appearance, as well as bounding box prompt-induced perturbations. Although current test-time adaptation (TTA) methods for medical image segmentation may tackle this issue, partial (e.g., batch normalization) or whole parametric updates restrict their effectiveness due to limited update signals or catastrophic forgetting in large models. Meanwhile, these approaches ignore the computational complexity during adaptation, which is particularly significant for modern foundation models. To this end, our theoretical analyses reveal that directly refining image embeddings is feasible to approach the same goal as parametric updates under the MedSAM architecture, which enables us to realize high computational efficiency and segmentation performance without the risk of catastrophic forgetting. Under this framework, we propose to encourage maximizing factorized conditional probabilities of the posterior prediction probability using a proposed distribution-approximated latent conditional random field loss combined with an entropy minimization loss. Experiments show that we achieve about 3\% Dice score improvements across three datasets while reducing computational complexity by over 7 times. Kecheng Chen, Xinyu Luo, Tiexin Qin, Jie Liu 0044, Hui Liu 0036, Victor Ho-fun Lee, Hong Yan 0001, Haoliang Li |
ICCV | 4 |
| 2025 | Large Language Models for Lossless Image Compression: Next-Pixel Prediction in Language Space is All You NeedabstractWe have recently witnessed that ''Intelligence" and `''Compression" are the two sides of the same coin, where the language large model (LLM) with unprecedented intelligence is a general-purpose lossless compressor for various data modalities. This attribute is particularly appealing to the lossless image compression community, given the increasing need to compress high-resolution images in the current streaming media era. Consequently, a spontaneous envision emerges: Can the compression performance of the LLM elevate lossless image compression to new heights? However, our findings indicate that the naive application of LLM-based lossless image compressors suffers from a considerable performance gap compared with existing state-of-the-art (SOTA) codecs on common benchmark datasets. In light of this, we are dedicated to fulfilling the unprecedented intelligence (compression) capacity of the LLM for lossless image compression tasks, thereby bridging the gap between theoretical and practical compression performance. Specifically, we propose P -LLM, a next-pixel prediction-based LLM, which integrates various elaborated insights and methodologies, \textit{e.g.,} pixel-level priors, the in-context ability of LLM, and a pixel-level semantic preservation strategy, to enhance the understanding capacity of pixel sequences for better next-pixel predictions. Extensive experiments on benchmark datasets demonstrate that P-LLM can beat SOTA classical and learned codecs. Kecheng Chen, Hui Liu 0036, Jie Liu 0044, Yibing Liu, Shiqi Wang 0001, Hong Yan 0001, Haoliang Li |
NeurIPS | 4 |
| 2025 | MedChain: Bridging the Gap Between LLM Agents and Clinical Practice with Interactive SequenceabstractClinical decision making (CDM) is a complex, dynamic process crucial to healthcare delivery, yet it remains a significant challenge for artificial intelligence systems. While Large Language Model (LLM)-based agents have been tested on general medical knowledge using licensing exams and knowledge question-answering tasks, their performance in the CDM in real-world scenarios is limited due to the lack of comprehensive benchmark that mirror actual medical practice. To address this gap, we present MedChain, a dataset of 12,163 clinical cases that covers five key stages of clinical workflow. MedChain distinguishes itself from existing benchmarks with three key features of real-world clinical practice: personalization, interactivity, and sequentiality. Further, to tackle real-world CDM challenges, we also propose MedChain-Agent, an AI system that integrates a feedback mechanism and a MedCase-RAG module to learn from previous cases and adapt its responses. MedChain-Agent demonstrates remarkable adaptability in gathering information dynamically and handling sequential clinical tasks, significantly outperforming existing approaches. The relevant dataset and code will be released upon acceptance of this paper. Jie Liu 0044, Wenxuan Wang 0001, Zizhan Ma, Guolin Huang, Yihang Su, Kao-Jung Chang, Haoliang Li, LinLin Shen, Michael R. Lyu, Wenting Chen |
NeurIPS | 1 |
| 2025 | Bi-VLGM: Bi-Level Class-Severity-Aware Vision-Language Graph Matching for Text Guided Medical Image SegmentationabstractAbstract Medical reports containing specific diagnostic results and additional information not present in medical images can be effectively employed to assist image understanding tasks, and the modality gap between vision and language can be bridged by vision-language matching (VLM). However, current vision-language models distort the intra-model relation and only include class information in reports that is insufficient for segmentation task. In this paper, we introduce a novel Bi-level class-severity-aware Vision-Language Graph Matching (Bi-VLGM) for text guided medical image segmentation, composed of a word-level VLGM module and a sentence-level VLGM module, to exploit the class-severity-aware relation among visual-textual features. In word-level VLGM, to mitigate the distorted intra-modal relation during VLM, we reformulate VLM as graph matching problem and introduce a vision-language graph matching (VLGM) to exploit the high-order relation among visual-textual features. Then, we perform VLGM between the local features for each class region and class-aware prompts to bridge their gap. In sentence-level VLGM, to provide disease severity information for segmentation task, we introduce a severity-aware prompting to quantify the severity level of disease lesion, and perform VLGM between the global features and the severity-aware prompts. By exploiting the relation between the local (global) and class (severity) features, the segmentation model can include the class-aware and severity-aware information to promote segmentation performance. Extensive experiments proved the effectiveness of our method and its superiority to existing methods. The source code will be released. Wenting Chen, Jie Liu 0044, Tianming Liu 0001, Yixuan Yuan |
Int. J. Comput. Vis. | 2 |
| 2025 | Unsupervised Domain Adaptation for Low-Dose CT Reconstruction via Bayesian Uncertainty AlignmentabstractLow-dose computed tomography (LDCT) image reconstruction techniques can reduce patient radiation exposure while maintaining acceptable imaging quality. Deep learning (DL) is widely used in this problem, but the performance of testing data (also known as target domain) is often degraded in clinical scenarios due to the variations that were not encountered in training data (also known as source domain). Unsupervised domain adaptation (UDA) of LDCT reconstruction has been proposed to solve this problem through distribution alignment. However, existing UDA methods fail to explore the usage of uncertainty quantification, which is crucial for reliable intelligent medical systems in clinical scenarios with unexpected variations. Moreover, existing direct alignment for different patients would lead to content mismatch issues. To address these issues, we propose to leverage a probabilistic reconstruction framework to conduct a joint discrepancy minimization between source and target domains in both the latent and image spaces. In the latent space, we devise a Bayesian uncertainty alignment to reduce the epistemic gap between the two domains. This approach reduces the uncertainty level of target domain data, making it more likely to render well-reconstructed results on target domains. In the image space, we propose a sharpness-aware distribution alignment (SDA) to achieve a match of second-order information, which can ensure that the reconstructed images from the target domain have similar sharpness to normal-dose CT (NDCT) images from the source domain. Experimental results on two simulated datasets and one clinical low-dose imaging dataset show that our proposed method outperforms other methods in quantitative and visualized performance. Kecheng Chen, Jie Liu 0044, Renjie Wan, Victor Ho-fun Lee, Varut Vardhanabhuti, Hong Yan 0001, Haoliang Li |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | 💎 GEM: Context-Aware Gaze EstiMation with Visual Search Behavior Matching for Chest Radiograph
Shaonan Liu, Wenting Chen, Jie Liu 0044, Xiaoling Luo 0001, LinLin Shen |
MICCAI (1) | 3 |
| 2024 | Infproto-Powered Adaptive Classifier and Agnostic Feature Learning for Single Domain Generalization in Medical ImagesabstractAbstract Designing a single domain generalization (DG) framework that generalizes from one source domain to arbitrary unseen domains is practical yet challenging in medical image segmentation, mainly due to the domain shift and limited source domain information. To tackle these issues, we reason that domain-adaptive classifier learning and domain-agnostic feature extraction are key components in single DG, and further propose an adaptive infinite prototypes (InfProto) scheme to facilitate the learning of the two components. InfProto harnesses high-order statistics and infinitely samples class-conditional instance-specific prototypes to form the classifier for discriminability enhancement. We then introduce probabilistic modeling and provide a theoretic upper bound to implicitly perform the infinite prototype sampling in the optimization of InfProto. Incorporating InfProto, we design a hierarchical domain-adaptive classifier to elasticize the model for varying domains. This classifier infinitely samples prototypes from the instance and mini-batch data distributions, forming the instance-level and mini-batch-level domain-adaptive classifiers, thereby generalizing to unseen domains. To extract domain-agnostic features, we assume each instance in the source domain is a micro source domain and then devise three complementary strategies, i.e., instance-level infinite prototype exchange, instance-batch infinite prototype interaction, and consistency regularization, to constrain outputs of the hierarchical domain-adaptive classifier. These three complementary strategies minimize distribution shifts among micro source domains, enabling the model to get rid of domain-specific characterizations and, in turn, concentrating on semantically discriminative features. Extensive comparison experiments demonstrate the superiority of our approach compared with state-of-the-art counterparts, and comprehensive ablation studies verify the effect of each proposed component. Notably, our method exhibits average improvements of 15.568% and 17.429% in dice on polyp and surgical instrument segmentation benchmarks. Xiaoqing Guo, Jie Liu 0044, Yixuan Yuan |
Int. J. Comput. Vis. | 2 |
| 2024 | Universal and extensible language-vision models for organ segmentation and tumor detection from abdominal computed tomography
Jie Liu 0044, Yixiao Zhang 0001, Kang Wang 0016, Mehmet Can Yavuz, Xiaoxi Chen, Yixuan Yuan, Haoliang Li, Yang Yang 0009, Alan L. Yuille, Yucheng Tang, Zongwei Zhou |
Medical Image Anal. | 1 |
| 2024 | MHD-Net: Memory-Aware Hetero-Modal Distillation Network for Thymic Epithelial Tumor Typing With Missing Pathology ModalityabstractFusing multi-modal radiology and pathology data with complementary information can improve the accuracy of tumor typing. However, collecting pathology data is difficult since it is high-cost and sometimes only obtainable after the surgery, which limits the application of multi-modal methods in diagnosis. To address this problem, we propose comprehensively learning multi-modal radiology-pathology data in training, and only using uni-modal radiology data in testing. Concretely, a Memory-aware Hetero-modal Distillation Network (MHD-Net) is proposed, which can distill well-learned multi-modal knowledge with the assistance of memory from the teacher to the student. In the teacher, to tackle the challenge in hetero-modal feature fusion, we propose a novel spatial-differentiated hetero-modal fusion module (SHFM) that models spatial-specific tumor information correlations across modalities. As only radiology data is accessible to the student, we store pathology features in the proposed contrast-boosted typing memory module (CTMM) that achieves type-wise memory updating and stage-wise contrastive memory boosting to ensure the effectiveness and generalization of memory items. In the student, to improve the cross-modal distillation, we propose a multi-stage memory-aware distillation (MMD) scheme that reads memory-aware pathology features from CTMM to remedy missing modal-specific information. Furthermore, we construct a Radiology-Pathology Thymic Epithelial Tumor (RPTET) dataset containing paired CT and WSI images with annotations. Experiments on the RPTET and CPTAC-LUAD datasets demonstrate that MHD-Net significantly improves tumor typing and outperforms existing multi-modal methods on missing modality situations. Huaqi Zhang, Jie Liu 0044, Weifan Liu, Zekuan Yu, Yixuan Yuan, Pengyu Wang 0005, Harry Qin |
IEEE J. Biomed. Health Informatics | 2 |
| 2024 | STAR-RL: Spatial-Temporal Hierarchical Reinforcement Learning for Interpretable Pathology Image Super-ResolutionabstractPathology image are essential for accurately interpreting lesion cells in cytopathology screening, but acquiring high-resolution digital slides requires specialized equipment and long scanning times. Though super-resolution (SR) techniques can alleviate this problem, existing deep learning models recover pathology image in a black-box manner, which can lead to untruthful biological details and misdiagnosis. Additionally, current methods allocate the same computational resources to recover each pixel of pathology image, leading to the sub-optimal recovery issue due to the large variation of pathology image. In this paper, we propose the first hierarchical reinforcement learning framework named Spatial-Temporal hierARchical Reinforcement Learning (STAR-RL), mainly for addressing the aforementioned issues in pathology image super-resolution problem. We reformulate the SR problem as a Markov decision process of interpretable operations and adopt the hierarchical recovery mechanism in patch level, to avoid sub-optimal recovery. Specifically, the higher-level spatial manager is proposed to pick out the most corrupted patch for the lower-level patch worker. Moreover, the higher-level temporal manager is advanced to evaluate the selected patch and determine whether the optimization should be stopped earlier, thereby avoiding the over-processed problem. Under the guidance of spatial-temporal managers, the lower-level patch worker processes the selected patch with pixel-wise interpretable actions at each time step. Experimental results on medical images degraded by different kernels show the effectiveness of STAR-RL. Furthermore, STAR-RL validates the promotion in tumor diagnosis with a large margin and shows generalizability under various degradations. The source code is available at https://github.com/CUHK-AIM-Group/STAR-RL. Wenting Chen, Jie Liu 0044, Tommy W. S. Chow, Yixuan Yuan |
IEEE Trans. Medical Imaging | 2 |
| 2023 | Adjustment and Alignment for Unbiased Open Set Domain AdaptationabstractOpen Set Domain Adaptation (OSDA) transfers the model from a label-rich domain to a label-free one containing novel-class samples. Existing OSDA works overlook abundant novel-class semantics hidden in the source domain, leading to a biased model learning and transfer. Although the causality has been studied to remove the semantic-level bias, the non-available novel-class samples result in the failure of existing causal solutions in OSDA. To break through this barrier, we propose a novel causality-driven solution with the unexplored front-door adjustment theory, and then implement it with a theoretically grounded framework, coined Adjustment and Alignment (ANNA), to achieve an unbiased OSDA. In a nutshell, ANNA consists of Front-Door Adjustment (FDA) to correct the biased learning in the source domain and Decoupled Causal Alignment (DCA) to transfer the model unbiasedly. On the one hand, FDA delves into fine-grained visual blocks to discover novel-class regions hidden in the base-class image. Then, it corrects the biased model optimization by implementing causal debiasing. On the other hand, DCA disentangles the base-class and novel-class regions with orthogonal masks, and then adapts the decoupled distribution for an unbiased model transfer. Extensive experiments show that ANNA achieves state-of-the-art results. The code is available at https://github.com/CityU-AIM-Group/Anna. Wuyang Li, Jie Liu 0044, Bo Han 0003, Yixuan Yuan |
CVPR | 2 |
| 2023 | CLIP-Driven Universal Model for Organ Segmentation and Tumor DetectionabstractAn increasing number of public datasets have shown a marked impact on automated organ segmentation and tumor detection. However, due to the small size and partially labeled problem of each dataset, as well as a limited investigation of diverse types of tumors, the resulting models are often limited to segmenting specific organs/tumors and ignore the semantics of anatomical structures, nor can they be extended to novel domains. To address these issues, we propose the CLIP-Driven Universal Model, which incorporates text embedding learned from Contrastive Language-Image Pre-training (CLIP) to segmentation models. This CLIP-based label encoding captures anatomical relationships, enabling the model to learn a structured feature embedding and segment 25 organs and 6 types of tumors. The proposed model is developed from an assembly of 14 datasets, using a total of 3,410 CT scans for training and then evaluated on 6,162 external CT scans from 3 additional datasets. We rank first on the Medical Segmentation Decathlon (MSD) public leaderboard and achieve state-of-the-art results on Beyond The Cranial Vault (BTCV). Additionally, the Universal Model is computationally more efficient (6× faster) compared with dataset-specific models, generalized better to CT scans from varying sites, and shows stronger transfer learning performance on novel tasks. Jie Liu 0044, Yixiao Zhang 0001, Jieneng Chen, Junfei Xiao, Yongyi Lu, Bennett A. Landman, Yixuan Yuan, Alan L. Yuille, Yucheng Tang, Zongwei Zhou |
ICCV | 1 |
| 2023 | Animal3D: A Comprehensive Dataset of 3D Animal Pose and ShapeabstractAccurately estimating the 3D pose and shape is an essential step towards understanding animal behavior, and can potentially benefit many downstream applications, such as wildlife conservation. However, research in this area is held back by the lack of a comprehensive and diverse dataset with high-quality 3D pose and shape annotations. In this paper, we propose Animal3D, the first comprehensive dataset for mammal animal 3D pose and shape estimation. Animal3D consists of 3379 images collected from 40 mammal species, high-quality annotations of 26 key-points, and importantly the pose and shape parameters of the SMAL [50] model. All annotations were labeled and checked manually in a multi-stage process to ensure highest quality results. Based on the Animal3D dataset, we benchmark representative shape and pose estimation models at: (1) supervised learning from only the Animal3D data, (2) synthetic to real transfer from synthetically generated images, and (3) fine-tuning human pose and shape estimation models. Our experimental results demonstrate that predicting the 3D shape and pose of animals across species remains a very challenging task, despite significant advances in human pose estimation. Our results further demonstrate that synthetic pre-training is a viable strategy to boost the model performance. Overall, Animal3D opens new directions for facilitating future research in animal 3D pose and shape estimation, and is publicly available. Jiacong Xu, Yi Zhang 0099, Wufei Ma, Artur Jesslen, Pengliang Ji, Qixin Hu, Qihao Liu, Jiahao Wang 0001, Wei Ji 0011, Chen Wang 0049, Xiaoding Yuan, Prakhar Kaushik, Guofeng Zhang 0020, Jie Liu 0044, Yushan Xie, Yawen Cui, Alan L. Yuille, Adam Kortylewski |
ICCV | 16 |
| 2023 | AbdomenAtlas-8K: Annotating 8, 000 CT Volumes for Multi-Organ Segmentation in Three WeeksabstractAnnotating medical images, particularly for organ segmentation, is laborious and time-consuming. For example, annotating an abdominal organ requires an estimated rate of 30-60 minutes per CT volume based on the expertise of an annotator and the size, visibility, and complexity of the organ. Therefore, publicly available datasets for multi-organ segmentation are often limited in data size and organ diversity. This paper proposes an active learning procedure to expedite the annotation process for organ segmentation and creates the largest multi-organ dataset (by far) with the spleen, liver, kidneys, stomach, gallbladder, pancreas, aorta, and IVC annotated in 8,448 CT volumes, equating to 3.2 million slices. The conventional annotation methods would take an experienced annotator up to 1,600 weeks (or roughly 30.8 years) to complete this task. In contrast, our annotation procedure has accomplished this task in three weeks (based on an 8-hour workday, five days a week) while maintaining a similar or even better annotation quality. This achievement is attributed to three unique properties of our method: (1) label bias reduction using multiple pre-trained segmentation models, (2) effective error detection in the model predictions, and (3) attention guidance for annotators to make corrections on the most salient errors. Furthermore, we summarize the taxonomy of common errors made by AI algorithms and annotators. This allows for continuous improvement of AI and annotations, significantly reducing the annotation costs required to create large-scale datasets for a wider variety of medical imaging tasks. Code and dataset are available at https://github.com/MrGiovanni/AbdomenAtlas Chongyu Qu, Tiezheng Zhang, Hualin Qiao, Jie Liu 0044, Yucheng Tang, Alan L. Yuille, Zongwei Zhou |
NeurIPS | 4 |
| 2023 | Handling Open-Set Noise and Novel Target Recognition in Domain Adaptive Semantic SegmentationabstractThis paper studies a practical domain adaptive (DA) semantic segmentation problem where only pseudo-labeled target data is accessible through a black-box model. Due to the domain gap and label shift between two domains, pseudo-labeled target data contains mixed closed-set and open-set label noises. In this paper, we propose a simplex noise transition matrix (SimT) to model the mixed noise distributions in DA semantic segmentation, and leverage SimT to handle open-set label noise and enable novel target recognition. When handling open-set noises, we formulate the problem as estimation of SimT. By exploiting computational geometry analysis and properties of segmentation, we design four complementary regularizers, i.e., volume regularization, anchor guidance, convex guarantee, and semantic constraint, to approximate the true SimT. Specifically, volume regularization minimizes the volume of simplex formed by rows of the non-square SimT, ensuring outputs of model to fit into the ground truth label distribution. To compensate for the lack of open-set knowledge, anchor guidance, convex guarantee, and semantic constraint are devised to enable the modeling of open-set noise distribution. The estimated SimT is utilized to correct noise issues in pseudo labels and promote the generalization ability of segmentation model on target domain data. In the task of novel target recognition, we first propose closed-to-open label correction (C2OLC) to explicitly derive the supervision signal for open-set classes by exploiting the estimated SimT, and then advance a semantic relation (SR) loss that harnesses the inter-class relation to facilitate the open-set class sample recognition in target domain. Extensive experimental results demonstrate that the proposed SimT can be flexibly plugged into existing DA methods to boost both closed-set and open-set class performance. Xiaoqing Guo, Jie Liu 0044, Tongliang Liu, Yixuan Yuan |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | GRAB-Net: Graph-Based Boundary-Aware Network for Medical Point Cloud SegmentationabstractPoint cloud segmentation is fundamental in many medical applications, such as aneurysm clipping and orthodontic planning. Recent methods mainly focus on designing powerful local feature extractors and generally overlook the segmentation around the boundaries between objects, which is extremely harmful to the clinical practice and degenerates the overall segmentation performance. To remedy this problem, we propose a GRAph-based Boundary-aware Network (GRAB-Net) with three paradigms, Graph-based Boundary-perception Module (GBM), Outer-boundary Context-assignment Module (OCM), and Inner-boundary Feature-rectification Module (IFM), for medical point cloud segmentation. Aiming to improve the segmentation performance around boundaries, GBM is designed to detect boundaries and interchange complementary information inside semantic and boundary features in the graph domain, where semantics-boundary correlations are modelled globally and informative clues are exchanged by graph reasoning. Furthermore, to reduce the context confusion that degenerates the segmentation performance outside the boundaries, OCM is proposed to construct the contextual graph, where dissimilar contexts are assigned to points of different categories guided by geometrical landmarks. In addition, we advance IFM to distinguish ambiguous features inside boundaries in a contrastive manner, where boundary-aware contrast strategies are proposed to facilitate the discriminative representation learning. Extensive experiments on two public datasets, IntrA and 3DTeethSeg, demonstrate the superiority of our method over state-of-the-art methods. Yifan Liu 0010, Wuyang Li, Jie Liu 0044, Hui Chen 0032, Yixuan Yuan |
IEEE Trans. Medical Imaging | 3 |
| 2022 | SimT: Handling Open-set Noise for Domain Adaptive Semantic SegmentationabstractThis paper studies a practical domain adaptive (DA) semantic segmentation problem where only pseudo-labeled target data is accessible through a black-box model. Due to the domain gap and label shift between two domains, pseudo-labeled target data contains mixed closed-set and open-set label noises. In this paper, we propose a simplex noise transition matrix (SimT) to model the mixed noise distributions in DA semantic segmentation and formulate the problem as estimation of SimT. By exploiting computational geometry analysis and properties of segmentation, we design three complementary regularizers, i.e. volume regularization, anchor guidance, convex guarantee, to approximate the true SimT. Specifically, volume regularization minimizes the volume of simplex formed by rows of the non-square SimT, which ensures outputs of segmentation model to fit into the ground truth label distribution. To compensate for the lack of open-set knowledge, anchor guidance and convex guarantee are devised to facilitate the modeling of open-set noise distribution and enhance the discriminative feature learning among closed-set and open-set classes. The estimated SimT is further utilized to correct noise issues in pseudo labels and promote the generalization ability of segmentation model on target domain data. Extensive experimental results demonstrate that the proposed SimT can be flexibly plugged into existing DA methods to boost the performance. The source code is available at https://github.com/CityU-AIM-Group/SimT. Xiaoqing Guo, Jie Liu 0044, Tongliang Liu, Yixuan Yuan |
CVPR | 2 |
| 2022 | Unknown-Oriented Learning for Open Set Domain Adaptation
Jie Liu 0044, Xiaoqing Guo, Yixuan Yuan |
ECCV (33) | 1 |
| 2022 | Promoting Saliency From Depth: Deep Unsupervised RGB-D Saliency Detection
Wei Ji 0011, Qi Bi, Chuan Guo 0002, Jie Liu 0044, Li Cheng 0001 |
ICLR | 5 |
| 2022 | Edge-Oriented Point-Cloud Transformer for 3D Intracranial Aneurysm Segmentation
Yifan Liu 0010, Jie Liu 0044, Yixuan Yuan |
MICCAI (5) | 2 |
| 2022 | Instance importance-Aware graph convolutional network for 3D medical diagnosis
Zhen Chen 0013, Jie Liu 0044, Meilu Zhu, Yat Ming Peter Woo, Yixuan Yuan |
Medical Image Anal. | 2 |
| 2022 | GNAS-U2Net: A New Optic Cup and Optic Disc Segmentation Architecture With Genetic Neural Architecture SearchabstractNeural architecture search (NAS) has made incredible progress in medical image segmentation tasks, due to its automatic design of the model. However, the search spaces studied in many existing studies are based on U-Net and its variants, which limits the potential of neural architecture search in modeling better architectures. In this study, we propose a new NAS architecture named GNAS-U2Net for the joint segmentation of optic cup and optic disc. This architecture is the first application of NAS in a two-level nested U-shaped structure. The best performance achieved by the joint segmentation model designed by NAS on the REFUGE dataset has an average DICE of 92.88%. Compared to U2-Net and other related work, the model has better performance and uses only 34.79M parameters. We then verify the generalization of the model on two datasets, namely the Drishti-GS dataset and the GAMMA dataset, for which we obtain an average DICE of 92.32% and 92.11% respectively. Junding Sun, Jie Liu 0044, Weifan Liu, Zekuan Yu |
IEEE Signal Process. Lett. | 3 |
| 2022 | Cross-Boosted Multi-Target Domain Adaptation for Multi-Modality Histopathology Image Translation and SegmentationabstractRecent digital pathology workflows mainly focus on mono-modality histopathology image analysis. However, they ignore the complementarity between Haematoxylin & Eosin (H&E) and Immunohistochemically (IHC) stained images, which can provide comprehensive gold standard for cancer diagnosis. To resolve this issue, we propose a cross-boosted multi-target domain adaptation pipeline for multi-modality histopathology images, which contains Cross-frequency Style-auxiliary Translation Network (CSTN) and Dual Cross-boosted Segmentation Network (DCSN). Firstly, CSTN achieves the one-to-many translation from fluorescence microscopy images to H&E and IHC images for providing source domain training data. To generate images with realistic color and texture, Cross-frequency Feature Transfer Module (CFTM) is developed to pertinently restructure and normalize high-frequency content and low-frequency style features from different domains. Then, DCSN fulfills multi-target domain adaptive segmentation, where a dual-branch encoder is introduced, and Bidirectional Cross-domain Boosting Module (BCBM) is designed to implement cross-modality information complementation through bidirectional inter-domain collaboration. Finally, we establish Multi-modality Thymus Histopathology (MThH) dataset, which is the largest publicly available H&E and IHC image benchmark. Experiments on MThH dataset and several public datasets show that the proposed pipeline outperforms state-of-the-art methods on both histopathology image translation and segmentation. Huaqi Zhang, Jie Liu 0044, Pengyu Wang 0005, Zekuan Yu, Weifan Liu |
IEEE J. Biomed. Health Informatics | 2 |
| 2022 | Semantic-Oriented Labeled-to-Unlabeled Distribution Translation for Image SegmentationabstractAutomatic medical image segmentation plays a crucial role in many medical applications, such as disease diagnosis and treatment planning. Existing deep learning based models usually regarded the segmentation task as pixel-wise classification and neglected the semantic correlations of pixels across different images, leading to vague feature distribution. Moreover, pixel-wise annotated data is rare in medical domain, and the scarce annotated data usually exhibits the biased distribution against the desired one, hindering the performance improvement under the supervised learning setting. In this paper, we propose a novel Labeled-to-unlabeled Distribution Translation (L2uDT) framework with Semantic-oriented Contrastive Learning (SoCL), mainly for addressing the aforementioned issues in medical image segmentation. In SoCL, a semantic grouping module is designed to cluster pixels into a set of semantically coherent groups, and a semantic-oriented contrastive loss is advanced to constrain group-wise prototypes, so as to explicitly learn a feature space with intra-class compactness and inter-class separability. We then establish a L2uDT strategy to approximate the desired data distribution for unbiased optimization, where we translate the labeled data distribution with the guidance of extensive unlabeled data. In particular, a bias estimator is devised to measure the distribution bias, then a gradual-paced shift is derived to progressively translate the labeled data distribution to unlabeled one. Both labeled and translated data are leveraged to optimize the segmentation model simultaneously. We illustrate the effectiveness of the proposed method on two benchmark datasets, EndoScene and PROSTATEx, and our method achieves state-of-the-art performance, which clearly demonstrates its effectiveness for medical image segmentation. The source code is available at https://github.com/CityU-AIM-Group/L2uDT. Xiaoqing Guo, Jie Liu 0044, Yixuan Yuan |
IEEE Trans. Medical Imaging | 2 |
| 2022 | Graph-Based Surgical Instrument Adaptive Segmentation via Domain-Common KnowledgeabstractUnsupervised domain adaptation (UDA), aiming to adapt the model to an unseen domain without annotations, has drawn sustained attention in surgical instrument segmentation. Existing UDA methods neglect the domain-common knowledge of two datasets, thus failing to grasp the inter-category relationship in the target domain and leading to poor performance. To address these issues, we propose a graph-based unsupervised domain adaptation framework, named Interactive Graph Network (IGNet), to effectively adapt a model to an unlabeled new domain in surgical instrument segmentation tasks. In detail, the Domain-common Prototype Constructor (DPC) is first advanced to adaptively aggregate the feature map into domain-common prototypes using the probability mixture model, and construct a prototypical graph to interact the information among prototypes from the global perspective. In this way, DPC can grasp the co-occurrent and long-range relationship for both domains. To further narrow down the domain gap, we design a Domain-common Knowledge Incorporator (DKI) to guide the evolution of feature maps towards domain-common direction via a common-knowledge guidance graph and category-attentive graph reasoning. At last, the Cross-category Mismatch Estimator (CME) is developed to evaluate the category-level alignment from a graph perspective and assign each pixel with different adversarial weights, so as to refine the feature distribution alignment. The extensive experiments on three types of tasks demonstrate the feasibility and superiority of IGNet compared with other state-of-the-art methods. Furthermore, ablation studies verify the effectiveness of each component of IGNet. The source code is available at https://github.com/CityU-AIM-Group/Prototypical-Graph-DA. Jie Liu 0044, Xiaoqing Guo, Yixuan Yuan |
IEEE Trans. Medical Imaging | 1 |
| 2021 | Dynamic Context-Sensitive Filtering Network for Video Salient Object DetectionabstractThe ability to capture inter-frame dynamics has been critical to the development of video salient object detection (VSOD). While many works have achieved great success in this field, a deeper insight into its dynamic nature should be developed. In this work, we aim to answer the following questions: How can a model adjust itself to dynamic variations as well as perceive fine differences in the real-world environment; How are the temporal dynamics well introduced into spatial information over time? To this end, we propose a dynamic context-sensitive filtering network (DCFNet) equipped with a dynamic context-sensitive filtering module (DCFM) and an effective bidirectional dynamic fusion strategy. The proposed DCFM sheds new light on dynamic filter generation by extracting location-related affinities between consecutive frames. Our bidirectional dynamic fusion strategy encourages the interaction of spatial and temporal information in a dynamic manner. Experimental results demonstrate that our proposed method can achieve state-of-the-art performance on most VSOD datasets while ensuring a real-time speed of 28 fps. The source code is publicly available at https://github.com/OIPLab-DUT/DCFNet. Miao Zhang 0004, Jie Liu 0044, Yongri Piao, Shunyu Yao 0004, Wei Ji 0011, Huchuan Lu, Zhongxuan Luo |
ICCV | 2 |
| 2021 | COINet: Adaptive Segmentation with Co-Interactive Network for Autonomous DrivingabstractSemantic segmentation serves as a cornerstone for safety autonomous driving and has been achieved remarkable progress at the price of dense annotations. Unsupervised domain adaptation was widely utilized to addresses this labor-intensive problem, which transfers the knowledge learned from labeled synthetic datset to real-world without any annotations. However, most existing adaptation works predict the segmentation results and domain identification results separately only with the last-layer feature, and ignore the intrinsic relationship among these two tasks. To address this issue, we present a CO-Interactive Network (COINet) for unsupervised adaptive segmentation. In particular, we propose a scale-aware distilled decoder to integrate multi-scale features dynamically through the designed inter-distilled module (IDM) and obtain fine-grained feature representations. A dual-task classifier is advanced with this decoder, to jointly predict the segmentation results and pixel-wise domain prediction results, which extracts shared complementary information for accurate segmentation. We further devise a co-interactive loss to explicitly model the intrinsic relationship among the segmentation and domain prediction, enabling the feature distribution alignment in pixel-level and an optimal segmentation decision boundary. We demonstrate the effectiveness of the proposed COINet on benchmark adaptation settings with extensive experimental and ablation results, and our model shows favorable performance against existing algorithms. Jie Liu 0044, Xiaoqing Guo, Baopu Li, Yixuan Yuan |
IROS | 1 |
| 2021 | Prototypical Interaction Graph for Unsupervised Domain Adaptation in Surgical Instrument Segmentation
Jie Liu 0044, Xiaoqing Guo, Yixuan Yuan |
MICCAI (3) | 1 |
| 2021 | MASG-GAN: A multi-view attention superpixel-guided generative adversarial network for efficient and simultaneous histopathology image segmentation and classification
Huaqi Zhang, Jie Liu 0044, Zekuan Yu, Pengyu Wang 0005 |
Neurocomputing | 2 |
| 2020 | Asymmetric Two-Stream Architecture for Accurate RGB-D Saliency Detection
Miao Zhang 0004, Sun Xiao Fei, Jie Liu 0044, Yongri Piao, Huchuan Lu |
ECCV (28) | 3 |