EDBT 2026 Demo / reviewers in the wild / expert
Mingze Yuan
dblp:257/4827
· DBLP profile ↗
15ranked-venue papers
4as first author
12since 2021 · last 2025
0000-0001-5403-5979ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RODS: Robust Optimization Inspired Diffusion Sampling for Detecting and Reducing Hallucination in Generative ModelsabstractDiffusion models have achieved state-of-the-art performance in generative modeling, yet their sampling procedures remain vulnerable to hallucinations—often stemming from inaccuracies in score approximation. In this work, we reinterpret diffusion sampling through the lens of optimization and introduce RODS (Robust Optimization–inspired Diffusion Sampler), a novel method that detects and corrects high-risk sampling steps using geometric cues from the loss landscape. RODS enforces smoother sampling trajectories and \textit{adaptively} adjusts perturbations, reducing hallucinations without retraining and at minimal additional inference cost. Experiments on AFHQv2, FFHQ, and 11k-hands demonstrate that RODS maintains comparable image quality and preserves generation diversity. More importantly, it improves both sampling fidelity and robustness, detecting over 70\% of hallucinated samples and correcting more than 25\%, all while avoiding the introduction of new artifacts. We release our code at https://github.com/Yiqi-Verna-Tian/RODS. Yiqi Tian, Pengfei Jin, Mingze Yuan, Na Li 0002, Quanzheng Li |
NeurIPS | 3 |
| 2024 | CycleINR: Cycle Implicit Neural Representation for Arbitrary-Scale Volumetric Super-Resolution of Medical DataabstractIn the realm of medical 3D data, such as CT and MRI images, prevalent anisotropic resolution is characterized by high intra-slice but diminished inter-slice resolution. The lowered resolution between adjacent slices poses challenges, hindering optimal viewing experiences and impeding the development of robust downstream analysis algorithms. Various volumetric super-resolution algorithms aim to surmount these challenges, enhancing inter-slice resolution and overall 3D medical imaging quality. However, existing approaches confront inherent challenges: 1) often tailored to specific upsampling factors, lacking flexibility for diverse clinical scenarios; 2) newly generated slices frequently suffer from over-smoothing, degrading fine details, and leading to inter-slice inconsistency. In response, this study presents CycleINR, a novel enhanced Implicit Neural Representation model for 3D medical data volumetric super-resolution. Leveraging the continuity of the learned implicit function, the CycleINR model can achieve results with arbitrary up-sampling rates, eliminating the need for separate training. Additionally, we enhance the grid sampling in CycleINR with a local attention mechanism and mitigate over-smoothing by integrating cycleconsistent loss. We introduce a new metric, Slice-wise Noise Level Inconsistency (SNLI), to quantitatively assess inter-slice noise level inconsistency. The effectiveness of our approach is demonstrated through image quality evaluations on an in-house dataset and a downstream task analysis on the Medical Segmentation Decathlon liver tumor dataset. Wei Fang 0005, Yuxing Tang, Heng Guo 0008, Mingze Yuan, Tony C. W. Mok, Ke Yan 0006, Jiawen Yao, Xin Chen 0058, Zaiyi Liu, Le Lu 0001, Ling Zhang 0002, Minfeng Xu |
CVPR | 4 |
| 2023 | Devil is in the Queries: Advancing Mask Transformers for Real-world Medical Image Segmentation and Out-of-Distribution LocalizationabstractReal-world medical image segmentation has tremendous long-tailed complexity of objects, among which tail conditions correlate with relatively rare diseases and are clinically significant. A trustworthy medical AI algorithm should demonstrate its effectiveness on tail conditions to avoid clinically dangerous damage in these out-of-distribution (OOD) cases. In this paper, we adopt the concept of object queries in Mask Transformers to formulate semantic segmentation as a soft cluster assignment. The queries fit the feature-level cluster centers of inliers during training. Therefore, when performing inference on a medical image in real-world scenarios, the similarity between pixels and the queries detects and localizes OOD regions. We term this OOD localization as MaxQuery. Furthermore, the foregrounds of real-world medical images, whether OOD objects or inliers, are lesions. The difference between them is less than that between the foreground and background, possibly misleading the object queries to focus redundantly on the background. Thus, we propose a query-distribution (QD) loss to enforce clear boundaries between segmentation targets and other regions at the query level, improving the inlier segmentation and OOD indication. Our proposed framework is tested on two real-world segmentation tasks, i.e., segmentation of pancreatic and liver tumors, outperforming previous state-of-the-art algorithms by an average of 7.39% on AUROC, 14.69% on AUPR, and 13.79% on FPR95 for OOD localization. On the other hand, our framework improves the performance of inlier segmentation by an average of 5.27% DSC when compared with the leading baseline nnUNet. Mingze Yuan, Yingda Xia, Hexin Dong, Zifan Chen, Jiawen Yao, Mingyan Qiu, Ke Yan 0006, Xiaoli Yin, Xin Chen 0058, Zaiyi Liu, Bin Dong 0001, Jingren Zhou 0001, Le Lu 0001, Ling Zhang 0002, Li Zhang 0047 |
CVPR | 1 |
| 2023 | CancerUniT: Towards a Single Unified Model for Effective Detection, Segmentation, and Diagnosis of Eight Major Cancers Using a Large Collection of CT ScansabstractHuman readers or radiologists routinely perform full-body multi-organ multi-disease detection and diagnosis in clinical practice, while most medical AI systems are built to focus on single organs with a narrow list of a few diseases. This might severely limit AI’s clinical adoption. A certain number of AI models need to be assembled nontrivially to match the diagnostic process of a human reading a CT scan. In this paper, we construct a Unified Tumor Transformer (CancerUniT) model to jointly detect tumor existence & location and diagnose tumor characteristics for eight major cancers in CT scans. CancerUniT is a query-based Mask Transformer model with the output of multi-tumor prediction. We decouple the object queries into organ queries, tumor detection queries and tumor diagnosis queries, and further establish hierarchical relationships among the three groups. This clinically-inspired architecture effectively assists inter- and intra-organ representation learning of tumors and facilitates the resolution of these complex, anatomically related multi-organ cancer image reading tasks. CancerUniT is trained end-to-end using a curated large-scale CT images of 10,042 patients including eight major types of cancers and occurring non-cancer tumors (all are pathology-confirmed with 3D tumor masks annotated by radiologists). On the test set of 631 patients, CancerUniT has demonstrated strong performance under a set of clinically relevant evaluation metrics, substantially outperforming both multi-disease methods and an assembly of eight single-organ expert models in tumor detection, segmentation, and diagnosis. This moves one step closer towards a universal high performance cancer screening tool. Jieneng Chen, Yingda Xia, Jiawen Yao, Ke Yan 0006, Le Lu 0001, Fakai Wang, Bo Zhou 0009, Mingyan Qiu, Qihang Yu, Mingze Yuan, Wei Fang 0005, Yuxing Tang, Minfeng Xu, Xianghua Ye, Xiaoli Yin, Xin Chen 0058, Jingren Zhou 0001, Alan L. Yuille, Zaiyi Liu, Ling Zhang 0002 |
ICCV | 11 |
| 2023 | Improved Prognostic Prediction of Pancreatic Cancer Using Multi-phase CT by Integrating Neural Distance and Texture-Aware Transformer
Hexin Dong, Jiawen Yao, Yuxing Tang, Mingze Yuan, Yingda Xia, Jingren Zhou 0001, Bin Dong 0001, Le Lu 0001, Zaiyi Liu, Li Zhang 0047, Ling Zhang 0002 |
MICCAI (5) | 4 |
| 2023 | Cluster-Induced Mask Transformers for Effective Opportunistic Gastric Cancer Screening on Non-contrast CT Scans
Mingze Yuan, Yingda Xia, Xin Chen 0058, Jiawen Yao, Mingyan Qiu, Hexin Dong, Jingren Zhou 0001, Bin Dong 0001, Le Lu 0001, Li Zhang 0047, Zaiyi Liu, Ling Zhang 0002 |
MICCAI (5) | 1 |
| 2023 | Unsupervised Image Denoising with Score FunctionabstractThough achieving excellent performance in some cases, current unsupervised learning methods for single image denoising usually have constraints in applications. In this paper, we propose a new approach which is more general and applicable to complicated noise models. Utilizing the property of score function, the gradient of logarithmic probability, we define a solving system for denoising. Once the score function of noisy images has been estimated, the denoised result can be obtained through the solving system. Our approach can be applied to multiple noise models, such as the mixture of multiplicative and additive noise combined with structured correlation. Experimental results show that our method is comparable when the noise model is simple, and has good performance in complicated cases where other methods are not applicable or perform poorly. Yutong Xie 0004, Mingze Yuan, Bin Dong 0001, Quanzheng Li |
NeurIPS | 2 |
| 2023 | An Efficient and Fully Refined Deformation Extraction Method for Deriving Mining-Induced Subsidence by the Joint of Probability Integral Method and SBAS-InSARabstractThis study proposes a novel approach to derive the mining-induced goaf deformation, named the fully refined deformation extraction method (FRDEM). In this method, the Probability Integral Method (PIM) and the Small Baseline Subset Interferometric Synthetic Aperture Radar (SBAS-InSAR) technique are geographically integrated in a low-cost way to derive the fully refined goaf deformation and therefore to improve the detection ability for mining subsidence with fast deformation rate and large-gradient. To achieve this, we first established a functional relationship model to connect the 3-D parameters of PIM with SBAS-InSAR. Then, an improved genetic algorithm (IGA) was presented for parameter inversion, thus the optimal parameter set for PIM and the predicted displacements of the mining goaf was achieved economically. Afterward, we developed a geographically weighted data fusion model by presenting a data fusion strategy to eliminate the spatial heterogeneity in goaf boundary areas, and the fully refined goaf deformation field for the working face 2302 in the Guotun coal mining area was finally derived. Results demonstrate that FRDEM-derived displacements are highly consistent with field measurements, with RMSE decreasing to 0.053, 0.057, and 0.044 m respectively for the entire goaf, goaf center, and goaf boundary field, compared with those of ~0.092, ~0.095, and ~0.084 m for PIM and ~0.578, ~0.696, and ~0.046 m for SBAS-InSAR, respectively. This implies that the proposed FRDEM has significantly improved detection ability in deriving the mining-induced deformation, and thus can be a very promising tool to forecast and evaluate potential geohazards in the coal mining area. Hui Liu 0041, Mingze Yuan, Jinzheng Wang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | 360MonoDepth: High-Resolution 360° Monocular Depth Estimationabstract360° cameras can capture complete environments in a single shot, which makes 360° imagery alluring in many computer vision tasks. However, monocular depth estimation remains a challenge for 360° data, particularly for high resolutions like 2K (2048 × 1 024) and beyond that are important for novel-view synthesis and virtual reality applications. Current CNN-based methods do not support such high resolutions due to limited GPU memory. In this work, we propose aflexible framework for monocular depth estimation from high-resolution 360° images using tangent images. We project the 360° input image onto a set of tangent planes that produce perspective views, which are suitable for the latest, most accurate state-of-the-art perspective monocular depth estimators. To achieve globally consistent disparity estimates, we recombine the individual depth estimates using deformable multi-scale alignment followed by gradient-domain blending. The result is a dense, high-resolution 360° depth map with a high level of detail, also for outdoor scenes which are not supported by existing methods. Our source code and data are available at https://manurare.github.io/360monodepth/. Manuel Rey-Area, Mingze Yuan, Christian Richardt |
CVPR | 2 |
| 2022 | A Modified Nonlinear Chirp Scaling Algorithm for Highly Squinted SAR on Maneuvering PlatformabstractHighly squinted synthetic aperture radar (HS-SAR) data focusing is a challenging task due to the heavy coupling between the range and azimuth, which would lead to the failure of the traditional imaging algorithms. In order to overcome these issues, a modified nonlinear chirp scaling (CS) algorithm for HS-SAR is proposed in this manuscript. First, traditional linear range walk correction (LRWC), range compressing (RC), secondary range compressing (SRC), range curvature correction (RCC) are employed as the range processing. Subsequently, a modulation phase factor (MPF) is introduced to weaken the influence of spatial variant (SV). After that, a high order SV correction phase (SVCP) approach is derived to eliminate the azimuth dependence of Doppler parameters. In addition, the sub-aperture data is focused on the range time and azimuth frequency domain by SPECAN processing to avoid numerous zeros-padding operations. Real data processing is adopted to prove the efficiency and validity of the proposed method. Jun Wang 0150, Tinghao Zhang, Mingze Yuan, Yachao Li 0001 |
IGARSS | 3 |
| 2022 | Region-Aware Metric Learning for Open World Semantic Segmentation via Meta-Channel AggregationabstractAs one of the most challenging and practical segmentation tasks, open-world semantic segmentation requires the model to segment the anomaly regions in the images and incrementally learn to segment out-of-distribution (OOD) objects, especially under a few-shot condition. The current state-of-the-art (SOTA) method, Deep Metric Learning Network (DMLNet), relies on pixel-level metric learning, with which the identification of similar regions having different semantics is difficult. Therefore, we propose a method called region-aware metric learning (RAML), which first separates the regions of the images and generates region-aware features for further metric learning. RAML improves the integrity of the segmented anomaly regions. Moreover, we propose a novel meta-channel aggregation (MCA) module to further separate anomaly regions, forming high-quality sub-region candidates and thereby improving the model performance for OOD objects. To evaluate the proposed RAML, we have conducted extensive experiments and ablation studies on Lost And Found and Road Anomaly datasets for anomaly segmentation and the CityScapes dataset for incremental few-shot learning. The results show that the proposed RAML achieves SOTA performance in both stages of open world segmentation. Our code and appendix are available at https://github.com/czifan/RAML. Hexin Dong, Zifan Chen, Mingze Yuan, Yutong Xie 0004, Jie Zhao 0009, Fei Yu 0018, Bin Dong 0001, Li Zhang 0047 |
IJCAI | 3 |
| 2021 | 360° Optical Flow using Tangent Images
Mingze Yuan, Christian Richardt |
BMVC | 1 |
| 2020 | OmniPhotos: casual 360° VR photographyabstractVirtual reality headsets are becoming increasingly popular, yet it remains difficult for casual users to capture immersive 360° VR panoramas. State-of-the-art approaches require capture times of usually far more than a minute and are often limited in their supported range of head motion. We introduce OmniPhotos, a novel approach for quickly and casually capturing high-quality 360° panoramas with motion parallax. Our approach requires a single sweep with a consumer 360° video camera as input, which takes less than 3 seconds to capture with a rotating selfie stick or 10 seconds handheld. This is the fastest capture time for any VR photography approach supporting motion parallax by an order of magnitude. We improve the visual rendering quality of our OmniPhotos by alleviating vertical distortion using a novel deformable proxy geometry, which we fit to a sparse 3D reconstruction of captured scenes. In addition, the 360° input views significantly expand the available viewing area, and thus the range of motion, compared to previous approaches. We have captured more than 50 OmniPhotos and show video results for a large variety of scenes. We will make our code available. Tobias Bertel, Mingze Yuan, Reuben Lindroos, Christian Richardt |
ACM Trans. Graph. | 2 |
| 2018 | How to Support Fraction Learning with Math Game "Run Fraction": Theory, Design and Application
Junjie Shang 0001, Ruonan Hu, Sijie Ma, Jialing Zeng, Mingze Yuan, Jingang Sun |
ICCE | 6 |
| 2018 | Values and Design Strategies of Emotional Design in Educational Games
Mingze Yuan, Junjie Shang 0001 |
ICCE | 1 |