VLDB 2026 Research / reviewers in the wild / expert
Mingxuan Gu
dblp:282/0002
· DBLP profile ↗
6ranked-venue papers
1as first author
6since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Vision transformer Hook for dense predictionsabstractPre-trained vision transformers (ViTs) have demonstrated remarkable capability in learning semantically rich image representations. However, their underlying plain architectures yield low-resolution feature maps, lacking essential fine-grained spatial details required for dense prediction tasks. To better transfer the learned visual features, we present ViT-Hook, a novel hybrid backbone compatible with plain ViTs that effectively bridges the gap between global semantic understanding and local spatial encodings. Specifically, our method aims to broaden the scope and impact of ViT from the following perspectives: (1) We propose a simple transformer-decoder-inspired hook module that receives hierarchical CNN features as spatial queries and interacts with expressive ViT features from large-scale pre-training, therefore instantiating general-purpose representations into task-suited ones. (2) ViT-Hook is a plug-and-play solution for powerful vision foundation models, such as DINOv2 and RADIO. In this case, we find that only partially fine-tuning several intermediate ViT layers can outperform previous full fine-tuning methods, while substantially reducing compute and memory burdens with most parameters frozen. (3) We evaluate ViT-Hook with various pre-trained sources on multiple dense prediction tasks, including semantic segmentation, instance segmentation, and object detection. Notably, tested on the unified UperNet and Mask R-CNN frameworks, our ViT-Hook surpasses state-of-the-art by a large margin, achieving 59.7 (+4.7) mIoU on ADE20K val, 55.0 (+3.6) box AP and 48.5 (+3.3) mask AP on COCO val2017. • We propose ViT-Hook, a hybrid backbone that effectively enhances Vision Transformer performance on various dense prediction tasks. • The proposed spatial query and hook modules are lightweight yet powerful, achieving competitive results compared to SoTA on widely used benchmarks. • We introduce a novel partial fine-tuning strategy, which outperforms full fine-tuning while using only a fraction of compute and memory. • We validate the generalizability of ViT-Hook on multiple types of upstream pre-training methods, including the most recent vision foundation models. Siyuan Mei, Mareike Thies, Yan Xia 0002, Yipeng Sun, Fei Wu 0025, Fuxin Fan, Mingxuan Gu, Chengze Ye, Yixing Huang, Vincent Christlein, Andreas K. Maier |
Pattern Recognit. | 7 |
| 2025 | A Gradient-Based Approach to Fast and Accurate Head Motion Compensation in Cone-Beam CTabstractCone-beam computed tomography (CBCT) systems, with their flexibility, present a promising avenue for direct point-of-care medical imaging, particularly in critical scenarios such as acute stroke assessment. However, the integration of CBCT into clinical workflows faces challenges, primarily linked to long scan duration resulting in patient motion during scanning and leading to image quality degradation in the reconstructed volumes. This paper introduces a novel approach to CBCT motion estimation using a gradient-based optimization algorithm, which leverages generalized derivatives of the backprojection operator for cone-beam CT geometries. Building on that, a fully differentiable target function is formulated which grades the quality of the current motion estimate in reconstruction space. We drastically accelerate motion estimation yielding a 19-fold speed-up compared to existing methods. Additionally, we investigate the architecture of networks used for quality metric regression and propose predicting voxel-wise quality maps, favoring autoencoder-like architectures over contracting ones. This modification improves gradient flow, leading to more accurate motion estimation. The presented method is evaluated through realistic experiments on head anatomy. It achieves a reduction in reprojection error from an initial average of 3mm to 0.61mm after motion compensation and consistently demonstrates superior performance compared to existing approaches. The analytic Jacobian for the backprojection operation, which is at the core of the proposed method, is made publicly available. In summary, this paper contributes to the advancement of CBCT integration into clinical workflows by proposing a robust motion estimation approach that enhances efficiency and accuracy, addressing critical challenges in time-sensitive scenarios. Mareike Thies, Fabian Wagner, Noah Maul, Manuela Goldmann, Linda-Sophie Schneider, Mingxuan Gu, Siyuan Mei, Lukas Folle, Alexander Preuhs, Michael Manhart 0001, Andreas K. Maier |
IEEE Trans. Medical Imaging | 7 |
| 2024 | Unsupervised Domain Adaptation Using Soft-Labeled Contrastive Learning with Reversed Monte Carlo Method for Cardiac Image Segmentation
Mingxuan Gu, Mareike Thies, Siyuan Mei, Fabian Wagner, Mingcheng Fan, Yipeng Sun, Zhaoya Pan, Sulaiman Vesal, Ronak Kosti, Dennis Possart, Jonas Utz, Andreas K. Maier |
MICCAI (9) | 1 |
| 2024 | Differentiable Score-Based Likelihoods: Learning CT Motion Compensation from Clean Images
Mareike Thies, Noah Maul, Siyuan Mei, Laura Pfaff, Nastassia Vysotskaya, Mingxuan Gu, Jonas Utz, Dennis Possart, Lukas Folle, Fabian Wagner, Andreas K. Maier |
MICCAI (7) | 6 |
| 2021 | Spatio-Temporal Multi-Task Learning for Cardiac MRI Left Ventricle QuantificationabstractQuantitative assessment of cardiac left ventricle (LV) morphology is essential to assess cardiac function and improve the diagnosis of different cardiovascular diseases. In current clinical practice, LV quantification depends on the measurement of myocardial shape indices, which is usually achieved by manual contouring of the endo- and epicardial. However, this process subjected to inter and intra-observer variability, and it is a time-consuming and tedious task. In this article, we propose a spatio-temporal multi-task learning approach to obtain a complete set of measurements quantifying cardiac LV morphology, regional-wall thickness (RWT), and additionally detecting the cardiac phase cycle (systole and diastole) for a given 3D Cine-magnetic resonance (MR) image sequence. We first segment cardiac LVs using an encoder-decoder network and then introduce a multitask framework to regress 11 LV indices and classify the cardiac phase, as parallel tasks during model optimization. The proposed deep learning model is based on the 3D spatio-temporal convolutions, which extract spatial and temporal features from MR images. We demonstrate the efficacy of the proposed method using cine-MR sequences of 145 subjects and comparing the performance with other state-of-the-art quantification methods. The proposed method obtained high prediction accuracy, with an average mean absolute error (MAE) of 129 mm2, 1.23 mm, 1.76 mm, Pearson correlation coefficient (PCC) of 96.4%, 87.2%, and 97.5% for LV and myocardium (Myo) cavity regions, 6 RWTs, 3 LV dimensions, and an error rate of 9.0% for phase classification. The experimental results highlight the robustness of the proposed method, despite varying degrees of cardiac morphology, image appearance, and low contrast in the cardiac MR sequences. Sulaiman Vesal, Mingxuan Gu, Andreas K. Maier, Nishant Ravikumar |
IEEE J. Biomed. Health Informatics | 2 |
| 2021 | Adapt Everywhere: Unsupervised Adaptation of Point-Clouds and Entropy Minimization for Multi-Modal Cardiac Image SegmentationabstractDeep learning models are sensitive to domain shift phenomena. A model trained on images from one domain cannot generalise well when tested on images from a different domain, despite capturing similar anatomical structures. It is mainly because the data distribution between the two domains is different. Moreover, creating annotation for every new modality is a tedious and time-consuming task, which also suffers from high inter- and intra- observer variability. Unsupervised domain adaptation (UDA) methods intend to reduce the gap between source and target domains by leveraging source domain labelled data to generate labels for the target domain. However, current state-of-the-art (SOTA) UDA methods demonstrate degraded performance when there is insufficient data in source and target domains. In this paper, we present a novel UDA method for multi-modal cardiac image segmentation. The proposed method is based on adversarial learning and adapts network features between source and target domain in different spaces. The paper introduces an end-to-end framework that integrates: a) entropy minimization, b) output feature space alignment and c) a novel point-cloud shape adaptation based on the latent features learned by the segmentation model. We validated our method on two cardiac datasets by adapting from the annotated source domain, bSSFP-MRI (balanced Steady-State Free Procession-MRI), to the unannotated target domain, LGE-MRI (Late-gadolinium enhance-MRI), for the multi-sequence dataset; and from MRI (source) to CT (target) for the cross-modality dataset. The results highlighted that by enforcing adversarial learning in different parts of the network, the proposed method delivered promising performance, compared to other SOTA methods. Sulaiman Vesal, Mingxuan Gu, Ronak Kosti, Andreas K. Maier, Nishant Ravikumar |
IEEE Trans. Medical Imaging | 2 |