Zeju Li

dblp:205/4592 · DBLP profile ↗
← Back
31ranked-venue papers
11as first author
27since 2021 · last 2026
0000-0002-4608-2959ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 18 · 7 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 ODYSSEY: Open-World Quadrupeds Exploration and Manipulation for Long-Horizon Tasks
abstract
Language-guided long-horizon mobile manipulation has long been a grand challenge in embodied semantic reasoning, generalizable manipulation, and adaptive locomotion. Three fundamental limitations hinder progress: First, although large language models have shown promise in enhancing spatial reasoning and task planning through learned semantic priors, existing implementations remain confined to tabletop scenarios, failing to address the constrained perception and limited actuation ranges characteristic of mobile platforms. Second, current manipulation strategies exhibit insufficient generalization when confronted with the diverse object configurations encountered in open-world environments. Third, while crucial for practical deployment, the dual requirement of maintaining high platform maneuverability alongside precise end-effector control in unstructured settings remains understudied in the literature. In this work, we present ODYSSEY, a unified mobile manipulation framework for agile quadruped robots equipped with manipulators, which seamlessly integrates high-level task planning with low-level whole-body control. To address the challenge of egocentric perception in language-conditioned tasks, we introduce a hierarchical planner powered by a vision-language model, enabling long-horizon instruction decomposition and precise action execution. At the control level, our novel whole-body policy achieves robust coordination of locomotion and manipulation across challenging terrains. We further present the first comprehensive benchmark for long-horizon mobile manipulation, evaluating diverse indoor and outdoor scenarios. Through successful sim-to-real transfer, we demonstrate the system’s generalization and robustness in real-world deployments, underscoring the practicality of legged manipulators in unstructured environments. Our work advances the feasibility of generalized robotic assistants capable of complex, dynamic tasks.
Kaijun Wang, Liqin Lu, Jianuo Jiang, Zeju Li, Wancai Zheng, Hao Chen 0041, Chunhua Shen
AAAI5
2026 Solve-Detect-Verify: Inference-Time Scaling with Flexible Generative Verifier
abstract
Figure 1: The Solve-Detect-Verify (SDV) pipeline transforms linguistic signals into efficiency.Left: On AIME 2024, SDV achieves 83.3% accuracy (vs.63.3% for GenPRM) while using 6x fewer verification tokens by pruning redundant reasoning.Right: The pipeline is powered by FlexiVe , a unified verifier.Unlike process-based verifiers that incur per-step overhead, FlexiVe analyzes traces holistically.It employs a "pragmatic" consensus strategy: parallel "Fast Thinking" checks (∼0.1k tokens) provide an initial semantic intuition, escalating to deliberative "Slow Thinking" (∼4k tokens) only when the model exhibits verbalized uncertainty.
Jianyuan Zhong, Zeju Li, Xiangyu Wen 0001, Kezhi Li, Qiang Xu 0001
ACL (1)2
2026 Wavelet-inspired diffusion model with near-field constraint for real-time echocardiography dehazing
Xue Gao, Fangyan Tian, Fanggang Wu, Zeju Li, Yi Guo 0002, Yuanyuan Wang 0001
Medical Image Anal.7
2026 ColoDiff: Integrating Dynamic Consistency With Content Awareness for Colonoscopy Video Generation
abstract
Colonoscopy video generation delivers dynamic, information-rich data critical for diagnosing intestinal diseases, particularly in data-scarce scenarios. High-quality video generation demands temporal consistency and precise control over clinical attributes, but faces challenges from irregular intestinal structures, diverse disease representations, and various imaging modalities. To this end, we propose ColoDiff, a diffusion-based framework that generates dynamic-consistent and content-aware colonoscopy videos, aiming to alleviate data shortage and assist clinical analysis. At the inter-frame level, our TimeStream module decouples temporal dependency from video sequences through a cross-frame tokenization mechanism, enabling intricate dynamic modeling despite irregular intestinal structures. At the intra-frame level, our Content-Aware module incorporates noise-injected embeddings and learnable prototypes to realize precise control over clinical attributes, breaking through the coarse guidance of diffusion models. Additionally, ColoDiff employs a non-Markovian sampling strategy that cuts steps by over 90% for real-time generation. ColoDiff is evaluated across three public datasets and one hospital database, based on both generation metrics and downstream tasks including disease diagnosis, modality discrimination, bowel preparation scoring, and lesion segmentation. Extensive experiments show ColoDiff generates videos with smooth transitions and rich dynamics. ColoDiff also produces customized contents tailored for diverse tasks, e.g., colitis, polyps, and adenomas for diagnosis. Incorporating synthetic videos into training promotes discriminative representation learning and improves diagnosis accuracy by 7.1%. ColoDiff presents an effort in controllable colonoscopy video generation, revealing the potential of synthetic videos in complementing authentic representation and mitigating data scarcity in clinical settings.
Junhu Fu, Shuyu Liang, Wutong Li, Kehao Wang 0004, Shengli Lin, Pinghong Zhou, Zeju Li, Yuanyuan Wang 0001, Yi Guo 0002
IEEE Trans. Medical Imaging10
2025 Dyve: Thinking Fast and Slow for Dynamic Process Verification
abstract
Large Language Models (LLMs) have advanced significantly in complex reasoning, often leveraging external verifiers to improve multi-step process reliability.However, existing process verification methods face critical limitations: discriminative Process Reward Models (PRMs) often provide overly simplistic binary feedback and struggle with incomplete reasoning traces, while sophisticated Generative Reward Models (GenRMs) can be computationally expensive.Furthermore, curating quality supervision data for process verifier is of challenging.Therefore, we present Dyve, a dynamic process verifier that enhances reasoning error detection in LLMs by integrating fast (System 1) and slow (System 2) thinking, inspired by Kahneman's Systems Theory.Dyve adaptively applies immediate token-level confirmation for straightforward steps and comprehensive analysis for complex ones.To address data challenges and enable its adaptive fast and slow thinking, Dyve employs a novel step-wise consensus-filtered supervision strategy.This strategy leverages Monte Carlo estimation, LLM-as-a-Judge, and specialized reasoning models to extract the high-quality training signals from noisy rollouts.Experimental results on ProcessBench and the MATH dataset confirm that Dyve significantly outperforms existing process-based verifiers and boosts performance in Best-of-N settings, while maintaining computational efficiency through strategic resource allocation.Our code, data and model are released at: https://github.com/ staymylove/
Jianyuan Zhong, Zeju Li, Xiangyu Wen 0001, Qiang Xu 0001
EMNLP2
2025 DeepRTL: Bridging Verilog Understanding and Generation with a Unified Representation Model
abstract
Recent advancements in large language models (LLMs) have shown significant potential for automating hardware description language (HDL) code generation from high-level natural language instructions. While fine-tuning has improved LLMs' performance in hardware design tasks, prior efforts have largely focused on Verilog generation, overlooking the equally critical task of Verilog understanding. Furthermore, existing models suffer from weak alignment between natural language descriptions and Verilog code, hindering the generation of high-quality, synthesizable designs. To address these issues, we present DeepRTL, a unified representation model that excels in both Verilog understanding and generation. Based on CodeT5+, DeepRTL is fine-tuned on a comprehensive dataset that aligns Verilog code with rich, multi-level natural language descriptions. We also introduce the first benchmark for Verilog understanding and take the initiative to apply embedding similarity and GPT Score to evaluate the models' understanding capabilities. These metrics capture semantic similarity more accurately than traditional methods like BLEU and ROUGE, which are limited to surface-level n-gram overlaps. By adapting curriculum learning to train DeepRTL, we enable it to significantly outperform GPT-4 in Verilog understanding tasks, while achieving performance on par with OpenAI's o1-preview model in Verilog generation tasks.
Yi Liu 0081, Changran Xu, Yunhao Zhou, Zeju Li, Qiang Xu 0001
ICLR4
2025 Spatial 3D-LLM : Exploring Spatial Awareness in 3D Vision-Language Models
abstract
New era has unlocked exciting possibilities for extending Large Language Models (LLMs) to tackle 3D vision-language tasks. However, most existing 3D multimodal LLMs (MLLMs) rely on compressing holistic 3D scene information or segmenting independent objects to perform these tasks, which limits their spatial awareness due to insufficient representation of the richness inherent in 3D scenes. To overcome these limitations, we propose Spatial 3D-LLM, a 3D MLLM specifically designed to enhance spatial awareness for 3D vision-language tasks by enriching the spatial embeddings of 3D scenes. Spatial 3D-LLM integrates an LLM backbone with a progressive spatial awareness scheme that progressively captures spatial information as the perception field expands, generating location-enriched 3D scene embeddings to serve as visual prompts. Furthermore, we introduce two novel tasks: 3D object distance measurement and 3D layout editing, and construct a 3D instruction dataset, MODEL, to evaluate the model’s spatial awareness capabilities. Experimental results demonstrate that Spatial 3D-LLM achieves state-of-the-art performance across a wide range of 3D vision-language tasks, revealing the improvements stemmed from our progressive spatial awareness scheme of mining more profound spatial information. Our code is available at https://github.com/bjshuyuan/Spatial-3D-LLM.
Zeju Li, Yifan Xu 0028, Jiaxing Qi, Zhifei Yang 0004, Ruifei Ma, Xiangde Liu
ICME2
2025 VAP-Diffusion: Enriching Descriptions with MLLMs for Enhanced Medical Image Generation
Junhu Fu, Bowen Guo, Zeju Li, Yuanyuan Wang 0001, Yi Guo 0002
MICCAI (11)4
2025 Dependency Matters: Enhancing LLM Reasoning with Explicit Knowledge Grounding
abstract
Large language models (LLMs) often produce reasoning steps that are superficially coherent yet internally inconsistent, leading to unreliable outputs. Since such failures typically arise from implicit or poorly-grounded knowledge, we introduce \emph{Grounded Reasoning in Dependency (GRiD)}, a novel dependency-aware reasoning framework that explicitly grounds reasoning steps in structured knowledge. GRiD represents reasoning as a graph consisting of interconnected knowledge extraction nodes and reasoning nodes, enforcing logical consistency through explicit dependencies. Each reasoning step is validated via a lightweight, step-wise verifier that ensures logical correctness relative to its premises. Extensive experiments across diverse reasoning benchmarks—including StrategyQA, CommonsenseQA, GPQA, and TruthfulQA—demonstrate that GRiD substantially improves reasoning accuracy, consistency, and faithfulness compared to recent state-of-the-art structured reasoning methods. Notably, GRiD enhances performance even when applied purely as a lightweight verification module at inference time, underscoring its generalizability and practical utility. Code is available at: https://github.com/cure-lab/GRiD.
Xiangyu Wen 0001, Min Li 0019, Junhua Huang, Jianyuan Zhong, Zeju Li, Yongxiang Huang, Mingxuan Yuan, Qiang Xu 0001
NeurIPS6
2025 Argus: Leveraging Multiview Images for Improved 3-D Scene Understanding With Large Language Models
abstract
Advancements in foundation models have made it possible to conduct applications in various downstream tasks. Especially, the new era has witnessed a remarkable capability to extend large language models (LLMs) for tackling tasks of 3-D scene understanding. Current methods rely heavily on 3-D point clouds, but the 3-D point cloud reconstruction of an indoor scene often results in information loss. Some textureless planes or repetitive patterns are prone to omission and manifest as voids within the reconstructed 3-D point clouds. Besides, objects with complex structures tend to introduce distortion of details caused by misalignments between the captured images and the dense reconstructed point clouds. The 2-D multiview images present visual consistency with 3-D point clouds and provide more detailed representations of scene components, which can naturally compensate for these deficiencies. Based on these insights, we propose Argus, a novel 3-D multimodal framework that leverages multiview images for enhanced 3-D scene understanding with LLMs. In general, Argus can be treated as a 3-D large multimodal foundation model (3D-LMM) since it takes various modalities as input (text instructions, 2-D multiview images, and 3-D point clouds) and expands the capability of LLMs to tackle 3-D tasks. Argus involves fusing and integrating multiview images and camera poses into view-as-scene features, which interact with the 3-D features to create comprehensive and detailed 3-D-aware scene embeddings. Our approach compensates for the information loss while reconstructing 3-D point clouds and helps LLMs better understand the 3-D world. Extensive experiments demonstrate that our method outperforms existing 3D-LMMs in various downstream tasks.
Yifan Xu 0028, Hanqi Jiang, Ruifei Ma, Yiwei Li 0002, Zihao Wu 0001, Zeju Li, Xiangde Liu
IEEE Trans. Neural Networks Learn. Syst.8
2024 Exploring the Distributed Knowledge Congruence in Proxy-data-free Federated Distillation
abstract
Federated learning (FL) is a privacy-preserving machine learning paradigm in which the server periodically aggregates local model parameters from cli ents without assembling their private data. Constrained communication and personalization requirements pose severe challenges to FL. Federated distillation (FD) is proposed to simultaneously address the above two problems, which exchanges knowledge between the server and clients, supporting heterogeneous local models while significantly reducing communication overhead. However, most existing FD methods require a proxy dataset, which is often unavailable in reality. A few recent proxy-data-free FD approaches can eliminate the need for additional public data, but suffer from remarkable discrepancy among local knowledge due to client-side model heterogeneity, leading to ambiguous representation on the server and inevitable accuracy degradation. To tackle this issue, we propose a proxy-data-free FD algorithm based on distributed knowledge congruence (FedDKC). FedDKC leverages well-designed refinement strategies to narrow local knowledge differences into an acceptable upper bound, so as to mitigate the negative effects of knowledge incongruence. Specifically, from perspectives of peak probability and Shannon entropy of local knowledge, we design kernel-based knowledge refinement (KKR) and searching-based knowledge refinement (SKR) respectively, and theoretically guarantee that the refined-local knowledge can satisfy an approximately-similar distribution and be regarded as congruent. Extensive experiments conducted on three common datasets demonstrate that our proposed FedDKC significantly outperforms the state-of-the-art on various heterogeneous settings while evidently improving the convergence speed.
Yuwei Wang 0003, Min Liu 0001, Quyang Pan, Junbo Zhang 0004, Zeju Li, Qingxiang Liu 0004
ACM Trans. Intell. Syst. Technol.7
2024 Layer-Sensitive Neural Processing Architecture for Error-Tolerant Applications
abstract
Neural network (NN) operation has high requirements for storage resources and parallel computing, which bring huge challenges to the deployment of NNs in Internet-of-Things (IoT) devices. Consequently, this work proposed a low-power NN architecture, comprising an energy-efficient NN processor and a Cortex-M3 host processor to achieve state-of-the-art (SOTA) end-to-end inference at the edge. The innovations of this article are as follows: 1) to minimize the bit width of the weight while keeping the loss of accuracy within a small range, cross-layer error tolerance has been analyzed, and mixed precision quantization has been adopted for cross-layer mapping; 2) dynamic reconfigurable tensor processing unit (DR-TPU) with approximate computing has been proposed, which brings$1.45\times $computing energy reduction within 0.46% accurate loss in ResNet-50; and 3) a customized input feature map (IFM) reuse and over-writeback strategy has been adopted, eliminating the recurrent fetching from the on-chip and off-chip memories. The times of on-chip storage access can be reduced by 25%–60%, and the capacity of on-chip memory can be reduced to half of the original. The processor has been implemented at 28-nm CMOS technology. Combining the above work, the proposed architecture can achieve a 53.1% reduction of power and 17.2-TOPS/W energy efficiency.
Zeju Li, Qinfan Wang, Zihan Zou, Qiao Shen 0001, Na Xie, Hao Cai 0001, Hao Zhang 0111, Bo Liu 0019
IEEE Trans. Very Large Scale Integr. Syst.1
2023 Multi-Source Education Knowledge Graph Construction and Fusion for College Curricula
abstract
The field of education has undergone a significant transformation due to the rapid advancements in Artificial Intelligence (AI). Among the various AI technologies, Knowledge Graphs (KGs) using Natural Language Processing (NLP) have emerged as powerful visualization tools for integrating multi-faceted information. In the context of university education, the availability of numerous specialized courses and complicated learning resources often leads to inferior learning outcomes for students. In this paper, we propose an automated framework for knowledge extraction, visual KG construction, and graph fusion, tailored for the major of Electronic Information. Furthermore, we perform data analysis to investigate the correlation degree and relationship between courses, rank hot knowledge concepts, and explore the intersection of courses. Our objective is to enhance the learning efficiency of students and to explore new educational paradigms enabled by AI. The proposed framework is expected to enable students to better understand and appreciate the intricacies of their field of study by providing them with a comprehensive understanding of the relationships between the various concepts and courses.
Zeju Li, Linya Cheng, Chunhong Zhang, Xinning Zhu, Hui Zhao 0001
ICALT1
2023 Boosting Physical Layer Black-Box Attacks with Semantic Adversaries in Semantic Communications
abstract
End-to-end semantic communication (ESC) system is able to improve communication efficiency by only transmitting the semantics of the input rather than raw bits. Although promising, ESC has also been shown susceptible to the crafted physical layer adversarial perturbations due to the openness of wireless channels and the sensitivity of neural models. Previous works focus more on the physical layer white-box attacks, while the challenging black-box ones, as more practical adversaries in real-world cases, are still largely under-explored. To this end, we present SemBLK, a novel method that can learn to generate destructive physical layer semantic attacks for an ESC system under the black-box setting, where the adversaries are imperceptible to humans. Specifically, 1) we first introduce a surrogate semantic encoder and train its parameters by exploring a limited number of queries to an existing ESC system. 2) Equipped with such a surrogate encoder, we then propose a novel semantic perturbation generation method to learn to boost the physical layer attacks with semantic adversaries. Experiments on two public datasets show the effectiveness of our proposed SemBLK in attacking the ESC system under the black-box setting. Finally, we provide case studies to visually justify the superiority of our physical layer semantic perturbations.
Zeju Li, Xinghan Liu, Guoshun Nan, Jinfei Zhou, Xinchen Lyu, Qimei Cui, Xiaofeng Tao 0001
ICC1
2023 Robust Segmentation via Topology Violation Detection and Feature Synthesis
Liu Li 0001, Qiang Ma 0004, Cheng Ouyang, Zeju Li, Qingjie Meng, Mengyun Qiao, Vanessa Kyriakopoulou, Joseph V. Hajnal, Daniel Rueckert, Bernhard Kainz
MICCAI (4)4
2023 Learn2Reg: Comprehensive Multi-Task Medical Image Registration Challenge, Dataset and Evaluation in the Era of Deep Learning
abstract
Image registration is a fundamental medical image analysis task, and a wide variety of approaches have been proposed. However, only a few studies have comprehensively compared medical image registration approaches on a wide range of clinically relevant tasks. This limits the development of registration methods, the adoption of research advances into practice, and a fair benchmark across competing approaches. The Learn2Reg challenge addresses these limitations by providing a multi-task medical image registration data set for comprehensive characterisation of deformable registration algorithms. A continuous evaluation will be possible at https://learn2reg.grand-challenge.org. Learn2Reg covers a wide range of anatomies (brain, abdomen, and thorax), modalities (ultrasound, CT, MR), availability of annotations, as well as intra- and inter-patient registration evaluation. We established an easily accessible framework for training and validation of 3D registration methods, which enabled the compilation of results of over 65 individual method submissions from more than 20 unique teams. We used a complementary set of metrics, including robustness, accuracy, plausibility, and runtime, enabling unique insight into the current state-of-the-art of medical image registration. This paper describes datasets, tasks, evaluation methods and results of the challenge, as well as results of further analysis of transferability to new datasets, the importance of label supervision, and resulting bias. While no single approach worked best across all tasks, many methodological aspects could be identified that push the performance of medical image registration to new state-of-the-art performance. Furthermore, we demystified the common belief that conventional registration methods have to be much slower than deep-learning-based methods.
Alessa Hering, Lasse Hansen, Tony C. W. Mok, Albert C. S. Chung, Hanna Siebert, Stephanie Häger, Annkristin Lange, Sven Kuckertz, Stefan Heldmann, Wei Shao 0008, Sulaiman Vesal, Mirabela Rusu, Geoffrey A. Sonn, Théo Estienne, Maria Vakalopoulou, Luyi Han, Yunzhi Huang, Pew-Thian Yap, Mikael Brudfors, Yaël Balbastre, Samuel Joutard, Marc Modat, Gal Lifshitz, Dan Raviv, Jinxin Lv, Qiang Li 0018, Vincent Jaouen, Dimitris Visvikis, Constance Fourcade, Mathieu Rubeaux, Wentao Pan 0001, Zhe Xu 0012, Bailiang Jian, Francesca De Benetti, Marek Wodzinski, Niklas Gunnarsson, Jens Sjölund, Daniel Grzech, Huaqi Qiu, Zeju Li, Alexander Thorley, Jinming Duan 0001, Christoph Großbröhmer, Andrew Hoopes, Ingerid Reinertsen, Yiming Xiao 0001, Bennett A. Landman, Yuankai Huo, Keelin Murphy, Nikolas Leßmann, Bram van Ginneken, Adrian V. Dalca, Mattias P. Heinrich
IEEE Trans. Medical Imaging40
2023 Joint Optimization of Class-Specific Training- and Test-Time Data Augmentation in Segmentation
abstract
This paper presents an effective and general data augmentation framework for medical image segmentation. We adopt a computationally efficient and data-efficient gradient-based meta-learning scheme to explicitly align the distribution of training and validation data which is used as a proxy for unseen test data. We improve the current data augmentation strategies with two core designs. First, we learn class-specific training-time data augmentation (TRA) effectively increasing the heterogeneity within the training subsets and tackling the class imbalance common in segmentation. Second, we jointly optimize TRA and test-time data augmentation (TEA), which are closely connected as both aim to align the training and test data distribution but were so far considered separately in previous works. We demonstrate the effectiveness of our method on four medical image segmentation tasks across different scenarios with two state-of-the-art segmentation models, DeepMedic and nnU-Net. Extensive experimentation shows that the proposed data augmentation framework can significantly and consistently improve the segmentation performance when compared to existing solutions. Code is publicly available at https://github.com/ZerojumpLine/JCSAugment.
Zeju Li, Konstantinos Kamnitsas, Qi Dou 0001, Chen Qin, Ben Glocker
IEEE Trans. Medical Imaging1
2023 Context Label Learning: Improving Background Class Representations in Semantic Segmentation
abstract
Background samples provide key contextual information for segmenting regions of interest (ROIs). However, they always cover a diverse set of structures, causing difficulties for the segmentation model to learn good decision boundaries with high sensitivity and precision. The issue concerns the highly heterogeneous nature of the background class, resulting in multi-modal distributions. Empirically, we find that neural networks trained with heterogeneous background struggle to map the corresponding contextual samples to compact clusters in feature space. As a result, the distribution over background logit activations may shift across the decision boundary, leading to systematic over-segmentation across different datasets and tasks. In this study, we propose context label learning (CoLab) to improve the context representations by decomposing the background class into several subclasses. Specifically, we train an auxiliary network as a task generator, along with the primary segmentation model, to automatically generate context labels that positively affect the ROI segmentation accuracy. Extensive experiments are conducted on several challenging segmentation tasks and datasets. The results demonstrate that CoLab can guide the segmentation model to map the logits of background samples away from the decision boundary, resulting in significantly improved segmentation accuracy. Code is available at https://github.com/ZerojumpLine/CoLab.
Zeju Li, Konstantinos Kamnitsas, Cheng Ouyang, Chen Chen 0042, Ben Glocker
IEEE Trans. Medical Imaging1
2023 Causality-Inspired Single-Source Domain Generalization for Medical Image Segmentation
abstract
Deep learning models usually suffer from the domain shift issue, where models trained on one source domain do not generalize well to other unseen domains. In this work, we investigate the single-source domain generalization problem: training a deep network that is robust to unseen domains, under the condition that training data are only available from one source domain, which is common in medical imaging applications. We tackle this problem in the context of cross-domain medical image segmentation. In this scenario, domain shifts are mainly caused by different acquisition processes. We propose a simple causality-inspired data augmentation approach to expose a segmentation model to synthesized domain-shifted training examples. Specifically, 1) to make the deep model robust to discrepancies in image intensities and textures, we employ a family of randomly-weighted shallow networks. They augment training images using diverse appearance transformations. 2) Further we show that spurious correlations among objects in an image are detrimental to domain robustness. These correlations might be taken by the network as domain-specific clues for making predictions, and they may break on unseen domains. We remove these spurious correlations via causal intervention. This is achieved by resampling the appearances of potentially correlated objects independently. The proposed approach is validated on three cross-domain segmentation scenarios: cross-modality (CT-MRI) abdominal image segmentation, cross-sequence (bSSFP-LGE) cardiac MRI segmentation, and cross-site prostate MRI segmentation. The proposed approach yields consistent performance gains compared with competitive methods when tested on unseen domains.
Cheng Ouyang, Chen Chen 0042, Surui Li, Zeju Li, Chen Qin, Wenjia Bai, Daniel Rueckert
IEEE Trans. Medical Imaging4
2022 Tackling Long-Tailed Category Distribution Under Domain Shifts
Xiao Gu 0003, Yao Guo 0002, Zeju Li, Jianing Qiu, Qi Dou 0001, Yuxuan Liu 0013, Benny P. L. Lo, Guang-Zhong Yang
ECCV (23)3
2022 MaxStyle: Adversarial Style Composition for Robust Medical Image Segmentation
Chen Chen 0042, Zeju Li, Cheng Ouyang, Matthew Sinclair, Wenjia Bai, Daniel Rueckert
MICCAI (5)2
2022 Estimating Model Performance Under Domain Shifts with Class-Specific Confidence Scores
Zeju Li, Konstantinos Kamnitsas, Mobarakol Islam, Chen Chen 0042, Ben Glocker
MICCAI (8)1
2022 SemBAT: Physical Layer Black-box Adversarial Attacks for Deep Learning-based Semantic Communication Systems
abstract
Deep learning-based semantic communications (DLSC) replace the physical blocks in traditional communication systems as end-to-end neural networks. DLSC significantly boost communication efficiency by only transmitting the meaning of data, showing great potentials for applications like automatic driving, digital twin and smart health. However, DLSC are fragile to black-box adversarial attacks due to the openness of wireless channel and sensitivities of neural models. To this end, this paper proposes SemBAT, a novel approach for crafting physical layer black-box adversarial attacks for semantic communication systems. The key ingredients of our method include the training of surrogate encoder and generation of adversarial perturbations. Specifically, we train our surrogate encoder by directly estimating the gradients based on Jacobian-matrixs, and then generate the adversarial perturbations by the particle swarm optimizations. Extensive experiments on a public benchmark show the effectiveness of our proposed SemBAT. We observe that our SemBAT with black-box adversaries can sharply decrease the classification accuracy of the semantic communication system from 78.4% to 11.6%. Meanwhile, such attacks are also imperceptible in terms of image quality metrics measured by the Structural similarity index measure (SSIM) and Peak Signal to Noise Ratio(PSNR).
Zeju Li, Jinfei Zhou, Guoshun Nan, Zhichun Li, Qimei Cui, Xiaofeng Tao 0001
VTC Fall1
2022 Enhancing MR image segmentation with realistic adversarial data augmentation
abstract
The success of neural networks on medical image segmentation tasks typically relies on large labeled datasets for model training. However, acquiring and manually labeling a large medical image set is resource-intensive, expensive, and sometimes impractical due to data sharing and privacy issues. To address this challenge, we propose AdvChain, a generic adversarial data augmentation framework, aiming at improving both the diversity and effectiveness of training data for medical image segmentation tasks. AdvChain augments data with dynamic data augmentation, generating randomly chained photo-metric and geometric transformations to resemble realistic yet challenging imaging variations to expand training data. By jointly optimizing the data augmentation model and a segmentation network during training, challenging examples are generated to enhance network generalizability for the downstream task. The proposed adversarial data augmentation does not rely on generative networks and can be used as a plug-in module in general segmentation networks. It is computationally efficient and applicable for both low-shot supervised and semi-supervised learning. We analyze and evaluate the method on two MR image segmentation tasks: cardiac segmentation and prostate segmentation with limited labeled data. Results show that the proposed approach can alleviate the need for labeled data while improving model generalization ability, indicating its practical value in medical imaging applications.
Chen Chen 0042, Chen Qin, Cheng Ouyang, Zeju Li, Shuo Wang 0011, Huaqi Qiu, Liang Chen 0018, Giacomo Tarroni, Wenjia Bai, Daniel Rueckert
Medical Image Anal.4
2022 Breast Tumor Classification Based on MRI-US Images by Disentangling Modality Features
abstract
Dynamic contrast-enhanced magnetic resonance imaging (DCE-MRI) and ultrasound (US), which are two common modalities for clinical breast tumor diagnosis besides Mammograms, can provide different and complementary information for the same tumor regions. Although many machine learning methods have been proposed for breast tumor classification based on either single modality, it remains unclear how to further boost the classification performance by utilizing paired multi-modality information with different dimensions. In this paper, we propose MRI-US multi-modality network (MUM-Net) to classify breast tumor into different subtypes based on 3D MR and 2D US images. The key insight of MUM-Net is that we explicitly distill modality-agnostic features for tumor classification. Specifically, we first adopt a discrimination-adaption module to decompose features into modality-agnostic and modality-specific ones with min-max training strategies. Then, we propose a feature fusion module to increase the compactness of the modality-agnostic features by utilizing an affinity matrix with nearest neighbour selection. We build a paired MRI-US breast tumor classification dataset containing 502 cases with three clinical indicators to validate the proposed method. In three tasks including lymph node metastasis, histological grade and Ki-67 level, MUM-Net achieves AUC scores of 0.8581, 0.8965 and 0.8577, outperforming other counterparts which are based on single task or single modality by a wide margin. In addition, we find that the extracted modality-agnostic features can help the network focus on the tumor regions in both modalities.
Mengyun Qiao, Chencheng Liu, Zeju Li, Shichong Zhou, Cai Chang, Yajia Gu, Yi Guo 0002, Yuanyuan Wang 0001
IEEE J. Biomed. Health Informatics3
2021 DeepVolume: Brain Structure and Spatial Connection-Aware Network for Brain MRI Super-Resolution
abstract
Thin-section magnetic resonance imaging (MRI) can provide higher resolution anatomical structures and more precise clinical information than thick-section images. However, thin-section MRI is not always available due to the imaging cost issue. In multicenter retrospective studies, a large number of data are often in thick-section manner with different section thickness. The lack of thin-section data and the difference in section thickness bring considerable difficulties in the study based on the image big data. In this article, we introduce DeepVolume, a two-step deep learning architecture to address the challenge of accurate thin-section MR image reconstruction. The first stage is the brain structure-aware network, in which the thick-section MR images in axial and sagittal planes are fused by a multitask 3-D U-net with prior knowledge of brain volume segmentation, which encourages the reconstruction result to have correct brain structure. The second stage is the spatial connection-aware network, in which the preliminary reconstruction results are adjusted slice-by-slice by a recurrent convolutional network embedding convolutional long short-term memory (LSTM) block, which enhances the precision of the reconstruction by utilizing the previously unassessed sagittal information. We used 305 paired brain MRI samples with thickness of 1.0 mm and 6.5 mm in this article. Extensive experiments illustrate that DeepVolume can produce the state-of-the-art reconstruction results by embedding more anatomical knowledge. Furthermore, considering DeepVolume as an intermediate step, the practical and clinical value of our method is validated by applying the brain volume estimation and voxel-based morphometry. The results show that DeepVolume can provide much more reliable brain volume estimation in the normalized space based on the thick-section MR images compared with the traditional solutions.
Zeju Li, Jinhua Yu 0003, Yuanyuan Wang 0001, Hanzhang Zhou, Zhongwei Qiao
IEEE Trans. Cybern.1
2021 Analyzing Overfitting Under Class Imbalance in Neural Networks for Image Segmentation
abstract
Class imbalance poses a challenge for developing unbiased, accurate predictive models. In particular, in image segmentation neural networks may overfit to the foreground samples from small structures, which are often heavily under-represented in the training set, leading to poor generalization. In this study, we provide new insights on the problem of overfitting under class imbalance by inspecting the network behavior. We find empirically that when training with limited data and strong class imbalance, at test time the distribution of logit activations may shift across the decision boundary, while samples of the well-represented class seem unaffected. This bias leads to a systematic under-segmentation of small structures. This phenomenon is consistently observed for different databases, tasks and network architectures. To tackle this problem, we introduce new asymmetric variants of popular loss functions and regularization techniques including a large margin loss, focal loss, adversarial training, mixup and data augmentation, which are explicitly designed to counter logit shift of the under-represented classes. Extensive experiments are conducted on several challenging segmentation tasks. Our results demonstrate that the proposed modifications to the objective function can lead to significantly improved segmentation accuracy compared to baselines and alternative approaches.
Zeju Li, Konstantinos Kamnitsas, Ben Glocker
IEEE Trans. Medical Imaging1
2020 High-Resolution Chest X-Ray Bone Suppression Using Unpaired CT Structural Priors
abstract
There is clinical evidence that suppressing the bone structures in Chest X-rays (CXRs) improves diagnostic value, either for radiologists or computer-aided diagnosis. However, bone-free CXRs are not always accessible. We hereby propose a coarse-to-fine CXR bone suppression approach by using structural priors derived from unpaired computed tomography (CT) images. In the low-resolution stage, we use the digitally reconstructed radiograph (DRR) image that is computed from CT as a bridge to connect CT and CXR. We then perform CXR bone decomposition by leveraging the DRR bone decomposition model learned from unpaired CTs and domain adaptation between CXR and DRR. To further mitigate the domain differences between CXRs and DRRs and speed up the learning convergence, we perform all the aboved operations in Laplacian of Gaussian (LoG) domain. After obtaining the bone decomposition result in DRR, we upsample it to a high resolution, based on which the bone region in the original high-resolution CXR is cropped and processed to produce a high-resolution bone decomposition result. Finally, such a produced bone image is subtracted from the original high-resolution CXR to obtain the bone suppression result. We conduct experiments and clinical evaluations based on two benchmarking CXR databases to show that (i) the proposed method outperforms the state-of-the-art unsupervised CXR bone suppression approaches; (ii) the CXRs with bone suppression are instrumental to radiologists for reducing their false-negative rate of lung diseases from 15% to 8%; and (iii) state-of-the-art disease classification performances are achieved by learning a deep network that takes the original CXR and its bone-suppressed image as inputs.
Hu Han 0001, Zeju Li, Jingjing Lu, Shaohua Kevin Zhou
IEEE Trans. Medical Imaging3
2019 Overfitting of Neural Nets Under Class Imbalance: Analysis and Improvements for Segmentation
Zeju Li, Konstantinos Kamnitsas, Ben Glocker
MICCAI (3)1
2019 Encoding CT Anatomy Knowledge for Unpaired Chest X-ray Image Decomposition
Zeju Li, Hu Han 0001, Gonglei Shi, Jiannan Wang 0005, Shaohua Kevin Zhou
MICCAI (6)1
2018 Left Ventricle Segmentation via Optical-Flow-Net from Short-Axis Cine MRI: Preserving the Temporal Coherence of Cardiac Motion
Yuanyuan Wang 0001, Zeju Li, Rob J. van der Geest
MICCAI (4)3