EDBT 2026 Demo / reviewers in the wild / expert
Hongyan Xu 0002
dblp:41/5438-2
· DBLP profile ↗
12ranked-venue papers
7as first author
12since 2021 · last 2026
0000-0003-3846-5236ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 6 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ROVER: Robust Generative Continual Identity Unlearning Against Relearning AttacksabstractRecent generative unlearning models synthesize high quality samples while protecting private information by unlearning the identity. However, existing generative identity unlearning methods face two challenges in multi-identity unlearning: 1) identity conflicts, which cause conflicts of model parameters in the continuous erasure of multiple identities; 2) fragile unlearning, where the model's unlearning ability deteriorates or fails under malicious attacks. In this paper, we introduce a critical yet under-explored task called robust multi-identity unlearning, with the goals of resolving identity conflicts to achieve interference-free unlearning and protecting against malicious attacks to achieve robust unlearning. To satisfy these goals, we propose a novel framework, RObust generatiVE continual identity unlearning against Relearning attacks (ROVER). By filtering unlearning requests with latent similarity, our method effectively isolates benign unlearning from malicious attacks to preserve identity removal integrity. Meanwhile, residual orthogonal resonator resolves identity conflicts in the continuous erasure of multiple identities, preserving stability in benign continual unlearning. Moreover, we introduce the phantom guard network to block malicious attacks by absorbing adversarial gradients, ensuring irreversible identity unlearning. The extensive experiments demonstrate that our proposed method achieves state-of-the-art performance on the task of robust multi-identity unlearning against relearning attacks. Tairan Huang 0001, Qiang Chen 0016, Beibei Hu, Yunlong Zhao 0003, Hongyan Xu 0002, Xiu Su |
AAAI | 5 |
| 2026 | STAG: Biologically guided spatial transcriptomics prediction via hypergraph learningabstractSpatial transcriptomics (ST) enables spatially resolved gene expression profiling within intact tissue sections. However, its widespread adoption is constrained by the high cost and low throughput of current sequencing-based protocols. This has motivated growing interest in computationally predicting gene expression directly from routinely acquired histology images. Existing methods are largely restricted to isolated 2D tissue slices and fail to capture richer spatial relationships or structured dependencies among spot-level gene expression profiles. In this paper, we propose STAG, a dual-branch framework for gene-aware expression prediction and spatial context modeling. A Query branch predicts ST expression for an individual target spot, while a Neighbor branch acts as an auxiliary branch to model structured relationships among multiple spots. By leveraging hypergraph learning, the Neighbor branch captures higher-order spatial and molecular dependencies, enabling unified modeling of both intra-slice and inter-slice relationships. This design supports standard 2D settings (a single slice) and naturally extends to 3D scenarios when adjacent tissue sections are available. Moreover, STAG leverages gene semantic information as biological guidance by encoding gene names with a foundation model, enabling coordinated gene-aware interactions beyond independent gene prediction. STAG achieves an average gain of 5.16% in PCC@250 across six datasets. Under highly variable gene selection, STAG maintains the lowest RMSE and highest PCC@50 across three datasets. The effectiveness of the learned representations is further demonstrated in pseudo-3D prediction and downstream cancer classification tasks. Code is available at https://github.com/MCPathology/STAG. Mingcheng Qu, Yuchuan Zhao, Donglin Di, Xiu Su, Hongyan Xu 0002, Yang Song 0001, Lei Fan 0007 |
Medical Image Anal. | 6 |
| 2025 | TinyMIG: Transferring Generalization from Vision Foundation Models to Single-Domain Medical ImagingabstractMedical imaging faces significant challenges in single-domain generalization (SDG) due to the diversity of imaging devices and the variability among data collection centers. To address these challenges, we propose \textbf{TinyMIG}, a framework designed to transfer generalization capabilities from vision foundation models to medical imaging SDG. TinyMIG aims to enable lightweight specialized models to mimic the strong generalization capabilities of foundation models in terms of both global feature distribution and local fine-grained details during training. Specifically, for global feature distribution, we propose a Global Distribution Consistency Learning strategy that mimics the prior distributions of the foundation model layer by layer. For local fine-grained details, we further design a Localized Representation Alignment method, which promotes semantic alignment and generalization distillation between the specialized model and the foundation model. These mechanisms collectively enable the specialized model to achieve robust performance in diverse medical imaging scenarios. Extensive experiments on large-scale benchmarks demonstrate that TinyMIG, with extremely low computational cost, significantly outperforms state-of-the-art models, showcasing its superior SDG capabilities. All the code and model weights will be publicly available. Hongyan Xu 0002, Yichao Cao, Xiu Su, Tianfa Li, Shan An, Haogang Zhu |
ICML | 2 |
| 2025 | UniMRG: Refining Medical Semantic Understanding Across Modalities via LLM-Orchestrated Synergistic Evolution
Hongyan Xu 0002, Arcot Sowmya, Ian Katz, Dadong Wang |
MICCAI (5) | 1 |
| 2025 | Addressing Granularity-induced Semantic Drift in OvOD via Graph-guided semantically consistent representationabstractOpen-vocabulary object detection (OvOD) uses Vision-Language Models (VLMs) to detect arbitrary categories specified by natural language. However, existing methods often struggle with performance instability caused by granularity-induced semantic drift, which arises from misaligned label embeddings across varying levels of specificity. In this paper, we propose GraSecon, a Graph-guided Semantically Consistent representation framework that enhances zero-shot detection robustness without requiring additional training. We construct a hierarchical Fine-grained Semantic Graph enriched with visually grounded attributes from large language models (LLMs). This graph captures hierarchical, sibling and cross-level relations, enabling controlled Laplacian refinement to harmonize the embedding space and improve visual-semantic alignment. To strengthen fine-grained discriminability, we introduce a Key Semantic Node Mining module that identifies and anchors semantically sensitive nodes, ensuring robust feature representation. Furthermore, our Semantic Relevance-Driven Laplacian Propagation adaptively propagates information, promoting coherent and context-aware embedding alignment across granularities. Extensive experiments on the iNatLoc and FSOD datasets demonstrate that GraSecon outperforms prior SOTA methods, achieving average mAP50 improvements of 6.5% and 5.4%. Code is publicly available at: https://github.com/minoslab-csu/GraSecon. Hongyan Xu 0002, Zhongze Wu, Ang He, Xi Lin 0003, Xiu Su |
ACM Multimedia | 1 |
| 2024 | Beyond the Limit of Weight-Sharing: Pioneering Space-Evolving NAS with Large Language ModelsabstractLarge language models (LLMs) offer impressive performance across diverse fields, but their increasing complexity raises both design costs and the need for specialized expertise. These challenges are intensified for Neural Architecture Search (NAS) methods reliant on weight-sharing techniques. This paper introduces GNAS, a new NAS method that boosts the search process with the aid of LLMs for efficient model discovery. With insights from existing architectures, GNAS swiftly identifies superior models that can adapt to changing resource constraints. We provide a mathematical framework to facilitate the transfer of knowledge across different model sizes, thereby improving search efficiency. Our experiments conducted on ImageNet, NAS-Bench-Macro, and ChannelBench-Macro confirm the effectiveness of GNAS across both CNN and Transformer architectures. Xiu Su, Shan You, Hongyan Xu 0002, Xiuxing Li, Chang Xu 0002 |
ICASSP | 3 |
| 2024 | TCNAS: Transformer Architecture Evolving in Code Clone DetectionabstractCode clone detection aims at finding code fragments with syntactic or semantic similarity. Most of current approaches mainly focus on detecting syntactic similarity while ignoring semantic long-term context alignment, and these detection methods encode the source code using human-designed models, a process which requires both expert input and a significant cost of time for experimentation and refinement. To address these challenges, we introduce the Transformer Code Neural Architecture Search (TCNAS), an approach designed to optimize transformer-based architectures for detection. In TCNAS, all channels are trained and evaluated equitably to enhance search efficiency. Besides, we introduce the dataflow of the code by extracting the semantic information from the code fragments. TCNAS facilitates the discovery of an optimal model structure geared towards the detection, eliminating the need for manual design. The searched optimal architecture is utilized to detect the code pairs. We conduct various empirical experiments on the benchmark, which covering all four types of code clone detection. The results demonstrate our approach consistently yields competitive detection scores across a range of evaluations. Hongyan Xu 0002, Xiaohuan Pei, Xiu Su, Shan You, Chang Xu 0002 |
ICASSP | 1 |
| 2024 | SCD-NAS: Towards Zero-Cost Training in Melanoma DiagnosisabstractDiagnosing melanoma remains challenging despite advances in Convolutional Neural Networks (CNNs) for skin cancer detection. Their application in clinical settings is often limited by differences between natural and clinical images. To address this, we introduce the Skin Cancer Detection Neural Architecture Search (SCD-NAS) framework. In our method, Large Language Model (LLM) is leveraged as a proxy, which helps SCD-NAS achieve cost-free training. Additionally, to maximize the benefits of various architectural design spaces, we introduce a Search Space Expansion (SSE) methodology. This effectively combines the merits of diverse architectural configurations, thereby enhancing model performance. We conducted experiments on the ISIC 2020, MedMNISTv2, CIFAR-10 and CIFAR-100 datasets. Our SCD-NAS-derived ResNet50 model achieved an Area Under the Curve (AUC) of 91.23% on the ISIC 2020 dataset, improving the baseline by 5.93%. It also exceeded the CIFAR-10 benchmark by 2.45% in accuracy. Hongyan Xu 0002, Xiu Su, Arcot Sowmya, Ian Katz, Dadong Wang |
ICME | 1 |
| 2024 | Detecting Any instruction-to-answer interaction relationship: Universal Instruction-to-Answer Navigator for Med-VQAabstractMedical Visual Question Answering (Med-VQA) interprets complex medical imagery using user instructions for precise diagnostics, yet faces challenges due to diverse, inadequately annotated images. In this paper, we introduce the Universal Instruction-Vision Navigator (Uni-Med) framework for extracting instruction-to-answer relationships, facilitating the understanding of visual evidence behind responses. Specifically, we design the Instruct-to-Answer Clues Interpreter (IAI) to generate visual explanations based on the answers and mark the core part of instructions with "real intent" labels. The IAI-Med VQA dataset, produced using IAI, is now publicly available to advance Med-VQA research. Additionally, our Token-Level Cut-Mix module dynamically aligns visual explanations with image patches, ensuring answers are traceable and learnable. We also implement intention-guided attention to minimize non-core instruction interference, sharpening focus on ’real intent’. Extensive experiments on SLAKE datasets show Uni-Med’s superior accuracies (87.52% closed, 86.12% overall), outperforming MedVInT-PMC-VQA by 1.22% and 0.92%. Code and dataset are available at: https://github.com/zhongzee/Uni-Med-master. Zhongze Wu, Hongyan Xu 0002, Yitian Long, Shan You, Xiu Su, Yueyi Luo, Chang Xu 0002 |
ICML | 2 |
| 2023 | Detection of Basal Cell Carcinoma in Whole Slide Images
Hongyan Xu 0002, Dadong Wang, Arcot Sowmya, Ian Katz |
MICCAI (6) | 1 |
| 2022 | Data Agnostic Filter Gating For Efficient Deep NetworksabstractFilter pruning is essential for deploying a well-trained CNN model on edge computation devices with a target computation budget (e.g., FLOPs). Current filter pruning methods mainly focus on leveraging feature maps to analyze the importance of filters, and prune those with less impact on the value of the CNN’s loss function, thereby ignoring the variance of input batches to differences in sparse structure over the filters. In this paper, we propose a data-agnostic filter pruning method that uses an auxiliary network named Dagger module to induce pruning with the pre-trained weights as input. Besides, to help prune filters with a preset FLOPs constraint, we utilize an explicit FLOPs-aware regularisation mechanism to directly promote pruning filters toward the target FLOPs. Experimental results on CIFAR-10 and ImageNet datasets show that the proposed filter pruning method surpasses the state-of-the-art. Hongyan Xu 0002, Xiu Su, Shan You, Tao Huang 0020, Fei Wang 0032, Chen Qian 0006, Changshui Zhang, Chang Xu 0002, Dadong Wang, Arcot Sowmya |
ICASSP | 1 |
| 2022 | Multi-scale alignment and Spatial ROI Module for COVID-19 DiagnosisabstractCoronavirus Disease 2019 (COVID-19) has spread globally and become a health crisis faced by humanity since first reported. Radiology imaging technologies such as computer tomography (CT) and chest X-ray imaging (CXR) are effective tools for diagnosing COVID-19. However, in CT and CXR images, the infected area occupies only a small part of the image. Some common deep learning methods that integrate large-scale receptive fields may cause the loss of image detail, resulting in the omission of the region of interest (ROI) in COVID-19 images and are therefore not suitable for further processing. To this end, we propose a deep spatial pyramid pooling (D-SPP) module to integrate contextual information over different resolutions, aiming to extract information under different scales of COVID-19 images effectively. Besides, we propose a COVID-19 infection detection (CID) module to draw attention to the lesion area and remove interference from irrelevant information. Extensive experiments on four CT and CXR datasets have shown that our method produces higher accuracy of detecting COVID-19 lesions in CT and CXR images. It can be used as a computer-aided diagnosis tool to help doctors effectively diagnose and screen for COVID-19. Hongyan Xu 0002, Dadong Wang, Arcot Sowmya |
IJCNN | 1 |