EDBT 2026 Demo / reviewers in the wild / expert
Haoxuan Che
dblp:227/7965
· DBLP profile ↗
14ranked-venue papers
6as first author
14since 2021 · last 2026
0000-0002-5844-1285ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MMMamba: A Versatile Cross-Modal in Context Fusion Framework for Pan-Sharpening and Zero-Shot Image EnhancementabstractPan-sharpening aims to generate high-resolution multispectral (HRMS) images by integrating a high-resolution panchromatic (PAN) image with its corresponding low-resolution multispectral (MS) image. To achieve effective fusion, it is crucial to fully exploit the complementary information between the two modalities. Traditional CNN-based methods typically rely on channel-wise concatenation with fixed convolutional operators, which limits their adaptability to diverse spatial and spectral variations. While cross-attention mechanisms enable global interactions, they are computationally inefficient and may dilute fine-grained correspondences, making it difficult to capture complex semantic relationships. Recent advances in the Multimodal Diffusion Transformer (MMDiT) architecture have demonstrated impressive success in image generation and editing tasks. Unlike cross-attention, MMDiT employs in-context conditioning to facilitate more direct and efficient cross-modal information exchange. In this paper, we propose MMMamba, a cross-modal in-context fusion framework for pan-sharpening, with the flexibility to support image super-resolution in a zero-shot manner. Built upon the Mamba architecture, our design ensures linear computational complexity while maintaining strong cross-modal interaction capacity. Furthermore, we introduce a novel multimodal interleaved (MI) scanning mechanism that facilitates effective information exchange between the PAN and MS modalities. Extensive experiments demonstrate the superior performance of our method compared to existing state-of-the-art (SOTA) techniques across multiple tasks and benchmarks. Yingying Wang 0005, Xuanhua He, Jialing Huang, Suiyun Zhang, Xinghao Ding, Haoxuan Che |
AAAI | 8 |
| 2026 | LLM-Driven Medical Report Generation via Communication-Efficient Heterogeneous Federated LearningabstractLarge Language Models (LLMs) have demonstrated significant potential in Medical Report Generation (MRG), yet their development requires large amounts of medical image-report pairs, which are commonly scattered across multiple centers. Centralizing these data is exceptionally challenging due to privacy regulations, thereby impeding model development and broader adoption of LLM-driven MRG models. To address this challenge, we present FedMRG, the first framework that leverages Federated Learning (FL) to enable privacy-preserving, multi-center development of LLM-driven MRG models, specifically designed to overcome the critical challenge of communication-efficient LLM training under multi-modal data heterogeneity. To start with, our framework tackles the fundamental challenge of communication overhead in federated LLM tuning by employing low-rank factorization to efficiently decompose parameter updates, significantly reducing gradient transmission costs and making LLM-driven MRG feasible in bandwidth-constrained FL settings. Furthermore, we observed the dual heterogeneity in MRG under the FL scenario: varying image characteristics across medical centers, as well as diverse reporting styles and terminology preferences. To address the data heterogeneity, we further enhance FedMRG with (1) client-aware contrastive learning in the MRG encoder, coupled with diagnosis-driven prompts, which capture both globally generalizable and locally distinctive features while maintaining diagnostic accuracy; and (2) a dual-adapter mutual boosting mechanism in the MRG decoder that harmonizes generic and specialized adapters to address variations in reporting styles and terminology. Through extensive evaluation of our established FL-MRG benchmark, we demonstrate the generalizability and adaptability of FedMRG, underscoring its potential in harnessing multi-center data and generating clinically accurate reports while maintaining communication efficiency. Haoxuan Che, Haibo Jin, Zhengrui Guo, Yi Lin 0009, Cheng Jin 0003, Hao Chen 0011 |
IEEE Trans. Medical Imaging | 1 |
| 2025 | GameGen-X: Interactive Open-world Game Video GenerationabstractWe introduce GameGen-$\mathbb{X}$, the first diffusion transformer model specifically designed for both generating and interactively controlling open-world game videos.
This model facilitates high-quality, open-domain generation by approximating various game elements, such as innovative characters, dynamic environments, complex actions, and diverse events.
Additionally, it provides interactive controllability, predicting and altering future content based on the current clip, thus allowing for gameplay simulation.
To realize this vision, we first collected and built an Open-World Video Game Dataset (OGameData) from scratch.
It is the first and largest dataset for open-world game video generation and control, which comprises over one million diverse gameplay video clips with informative captions.
GameGen-$\mathbb{X}$ undergoes a two-stage training process, consisting of pre-training and instruction tuning.
Firstly, the model was pre-trained via text-to-video generation and video continuation, enabling long-sequence open-domain game video generation with improved fidelity and coherence.
Further, to achieve interactive controllability, we designed InstructNet to incorporate game-related multi-modal control signal experts.
This allows the model to adjust latent representations based on user inputs, advancing the integration of character interaction and scene content control in video generation.
During instruction tuning, only the InstructNet is updated while the pre-trained foundation model is frozen, enabling the integration of interactive controllability without loss of diversity and quality of generated content.
GameGen-$\mathbb{X}$ contributes to advancements in open-world game design using generative models.
It demonstrates the potential of generative models to serve as auxiliary tools to traditional rendering techniques, demonstrating the potential for merging creative generation with interactive capabilities.
The project will be available at https://github.com/GameGen-X/GameGen-X. Haoxuan Che, Xuanhua He, Quande Liu, Cheng Jin 0003, Hao Chen 0011 |
ICLR | 1 |
| 2025 | Unpaired Optical Coherence Tomography Angiography Image Super-Resolution via Frequency-Aware Inverse-Consistency GANabstractFor optical coherence tomography angiography (OCTA) images, the limited scanning rate leads to a trade-off between field-of-view (FOV) and imaging resolution. Although larger FOV images may reveal more parafoveal vascular lesions, their application is hampered due to lower resolution. To increase the resolution, previous works only achieved satisfactory performance by using paired data for training, but real-world applications are limited by the challenge of collecting large-scale paired images. Thus, an unpaired approach is highly demanded. Generative Adversarial Network (GAN) has been commonly used in the unpaired setting, but it may struggle to accurately preserve fine-grained capillary details, which are critical biomarkers for OCTA. In this paper, our approach aspires to preserve these details by leveraging the frequency information, which represents details as high-frequencies (${\bm {hf}}$) and coarse-grained features as low-frequencies (${\bm {lf}}$). We propose a GAN-based unpaired super-resolution method for OCTA images and exceptionally emphasize ${\bm {hf}}$ fine capillaries through a dual-path generator. To facilitate a precise spectrum of the reconstructed image, we also propose a frequency-aware adversarial loss for the discriminator and introduce a frequency-aware focal consistency loss for end-to-end optimization. We collected a paired dataset for evaluation and showed that our method outperforms other state-of-the-art unpaired methods both quantitatively and visually. Haoxuan Che, An-ran Ran, Carol Y. Cheung, Hao Chen 0011 |
IEEE J. Biomed. Health Informatics | 3 |
| 2025 | FedDAG: Federated Domain Adversarial Generation Toward Generalizable Medical Image AnalysisabstractFederated domain generalization aims to train a global model from multiple source domains and ensure its generalization ability to unseen target domains. Due to the target domain being with unknown domain shifts, attempting to approximate these gaps by source domains may be the key to improving model generalization capability. Existing works mainly focus on sharing and recombining local domain-specific attributes to increase data diversity and simulate potential domain shifts. However, these methods may be insufficient since only the local attribute recombination can be hard to touch the out-of-distribution of global data. In this paper, we propose a simple-yet-efficient framework named Federated Domain Adversarial Generation (FedDAG). It aims to simulate the domain shift and improve the model generalization by adversarially generating novel domains different from local and global source domains. Specifically, it generates novel-style images by maximizing the instance-level feature discrepancy between original and generated images and trains a generalizable task model by minimizing their feature discrepancy. Further, we observed that FedDAG could cause different performance improvements for local models. It may be due to inherent data isolation and heterogeneity among clients, exacerbating the imbalance in their generalization contributions to the global model. Ignoring this imbalance can lead the global model's generalization ability to be sub-optimal, further limiting the novel domain generation procedure. Thus, to mitigate this imbalance, FedDAG hierarchically aggregates local models at the within-client and across-client levels by using the sharpness concept to evaluate client model generalization contributions. Extensive experiments across four medical benchmarks demonstrate FedDAG's ability to enhance generalization in federated medical scenarios. Haoxuan Che, Haibo Jin, Yong Xia 0001, Hao Chen 0011 |
IEEE Trans. Medical Imaging | 1 |
| 2025 | A Chain of Diagnosis Framework for Accurate and Explainable Radiology Report GenerationabstractDespite the progress of radiology report generation (RRG), existing works face two challenges: 1) The performances in clinical efficacy are unsatisfactory, especially for lesion attributes description; 2) the generated text lacks explainability, making it difficult for radiologists to trust the results. To address the challenges, we focus on a trustworthy RRG model, which not only generates accurate descriptions of abnormalities, but also provides basis of its predictions. To this end, we propose a framework named chain of diagnosis (CoD), which maintains a chain of diagnostic process for clinically accurate and explainable RRG. It first generates question-answer (QA) pairs via diagnostic conversation to extract key findings, then prompts a large language model with QA diagnoses for accurate generation. To enhance explainability, a diagnosis grounding module is designed to match QA diagnoses and generated sentences, where the diagnoses act as a reference. Moreover, a lesion grounding module is designed to locate abnormalities in the image, further improving the working efficiency of radiologists. To facilitate label-efficient training, we propose an omni-supervised learning strategy with clinical consistency to leverage various types of annotations from different datasets. Our efforts lead to 1) an omni-labeled RRG dataset with QA pairs and lesion boxes; 2) a evaluation tool for assessing the accuracy of reports in describing lesion location and severity; 3) extensive experiments to demonstrate the effectiveness of CoD, where it outperforms both specialist and generalist models consistently on two RRG benchmarks and shows promising explainability by accurately grounding generated sentences to QA diagnoses and images. Haibo Jin, Haoxuan Che, Sunan He, Hao Chen 0011 |
IEEE Trans. Medical Imaging | 2 |
| 2024 | PromptMRG: Diagnosis-Driven Prompts for Medical Report GenerationabstractAutomatic medical report generation (MRG) is of great research value as it has the potential to relieve radiologists from the heavy burden of report writing. Despite recent advancements, accurate MRG remains challenging due to the need for precise clinical understanding and disease identification. Moreover, the imbalanced distribution of diseases makes the challenge even more pronounced, as rare diseases are underrepresented in training data, making their diagnosis unreliable. To address these challenges, we propose diagnosis-driven prompts for medical report generation (PromptMRG), a novel framework that aims to improve the diagnostic accuracy of MRG with the guidance of diagnosis-aware prompts. Specifically, PromptMRG is based on encoder-decoder architecture with an extra disease classification branch. When generating reports, the diagnostic results from the classification branch are converted into token prompts to explicitly guide the generation process. To further improve the diagnostic accuracy, we design cross-modal feature enhancement, which retrieves similar reports from the database to assist the diagnosis of a query image by leveraging the knowledge from a pre-trained CLIP. Moreover, the disease imbalanced issue is addressed by applying an adaptive logit-adjusted loss to the classification branch based on the individual learning status of each disease, which overcomes the barrier of text decoder's inability to manipulate disease distributions. Experiments on two MRG benchmarks show the effectiveness of the proposed method, where it obtains state-of-the-art clinical efficacy performance on both datasets. Haibo Jin, Haoxuan Che, Yi Lin 0009, Hao Chen 0011 |
AAAI | 2 |
| 2024 | Deep Model Reference: Simple Yet Effective Confidence Estimation for Image Classification
Yuanhang Zheng, Yiqiao Qiu, Haoxuan Che, Hao Chen 0011, Wei-Shi Zheng 0001 |
MICCAI (10) | 3 |
| 2024 | Rethinking Self-Training for Semi-Supervised Landmark Detection: A Selection-Free ApproachabstractSelf-training is a simple yet effective method for semi-supervised learning, during which pseudo-label selection plays an important role for handling confirmation bias. Despite its popularity, applying self-training to landmark detection faces three problems: 1) The selected confident pseudo-labels often contain data bias, which may hurt model performance; 2) It is not easy to decide a proper threshold for sample selection as the localization task can be sensitive to noisy pseudo-labels; 3) coordinate regression does not output confidence, making selection-based self-training infeasible. To address the above issues, we propose Self-Training for Landmark Detection (STLD), a method that does not require explicit pseudo-label selection. Instead, STLD constructs a task curriculum to deal with confirmation bias, which progressively transitions from more confident to less confident tasks over the rounds of self-training. Pseudo pretraining and shrink regression are two essential components for such a curriculum, where the former is the first task of the curriculum for providing a better model initialization and the latter is further added in the later rounds to directly leverage the pseudo-labels in a coarse-to-fine manner. Experiments on three facial and one medical landmark detection benchmark show that STLD outperforms the existing methods consistently in both semi- and omni-supervised settings. The code is available at https://github.com/jhb86253817/STLD. Haibo Jin, Haoxuan Che, Hao Chen 0011 |
IEEE Trans. Image Process. | 2 |
| 2023 | Image Quality-aware Diagnosis via Meta-knowledge Co-embeddingabstractMedical images usually suffer from image degradation in clinical practice, leading to decreased performance of deep learning-based models. To resolve this problem, most previous works have focused on filtering out degradation-causing low-quality images while ignoring their potential value for models. Through effectively learning and leveraging the knowledge of degradations, models can better resist their adverse effects and avoid misdiagnosis. In this paper, we raise the problem of image quality-aware diagnosis, which aims to take advantage of low-quality images and image quality labels to achieve a more accurate and robust diagnosis. However, the diversity of degradations and superficially unrelated targets between image quality assessment and disease diagnosis makes it still quite challenging to effectively leverage quality labels to assist diagnosis. Thus, to tackle these issues, we propose a novel meta-knowledge co-embedding network, consisting of two subnets: Task Net and Meta Learner. Task Net constructs an explicit quality information utilization mechanism to enhance diagnosis via knowledge co-embedding features, while Meta Learner ensures the effectiveness and constrains the semantics of these features via meta-learning and joint-encoding masking. Superior performance on five datasets with four widely-used medical imaging modalities demonstrates the effectiveness and generalizability of our method. Haoxuan Che, Siyu Chen 0045, Hao Chen 0011 |
CVPR | 1 |
| 2023 | Towards Generalizable Diabetic Retinopathy Grading in Unseen Domains
Haoxuan Che, Yuhan Cheng, Haibo Jin, Hao Chen 0011 |
MICCAI (5) | 1 |
| 2023 | Unsupervised Domain Adaptation for Anatomical Landmark Detection
Haibo Jin, Haoxuan Che, Hao Chen 0011 |
MICCAI (1) | 2 |
| 2022 | Learning Robust Representation for Joint Grading of Ophthalmic Diseases via Adaptive Curriculum and Feature Disentanglement
Haoxuan Che, Haibo Jin, Hao Chen 0011 |
MICCAI (3) | 1 |
| 2021 | Pain-FL: Personalized Privacy-Preserving Incentive for Federated LearningabstractFederated learning (FL) is a privacy-preserving distributed machine learning framework, which involves training statistical models over a number of mobile users (i.e., workers) while keeping data localized. However, recent works have demonstrated that workers engaged in FL are still susceptible to advanced inference attacks when sharing model updates or gradients, which would discourage them from participating. Most of the existing incentive mechanisms for FL mainly account for workers’ resource cost, while the cost incurred by potential privacy leakage resulting from inference attacks has rarely been incorporated. To address these issues, in this paper, we propose a contract-based personalized privacy-preserving incentive for FL, named Pain-FL, to provide customized payments for workers with different privacy preferences as compensation for privacy leakage cost while ensuring satisfactory convergence performance of FL models. The core idea of Pain-FL is that each worker agrees on a customized contract, which specifies a kind of privacy-preserving level (PPL) and the corresponding payment, with the server in each round of FL. Then, the worker perturbs her calculated stochastic gradients to be uploaded with that PPL in exchange for that payment. In particular, we respectively derive a set of optimal contracts analytically under both complete and incomplete information models, which could optimize the convergence performance of the finally learned global model, while bearing some desired economic properties, i.e., budget feasibility, individual rationality, and incentive compatibility. An exhaustive experimental evaluation of Pain-FL is conducted, and the results corroborate its practicability and effectiveness. Peng Sun 0003, Haoxuan Che, Zhibo Wang 0001, Yuwei Wang 0001, Liantao Wu, Huajie Shao |
IEEE J. Sel. Areas Commun. | 2 |