VLDB 2026 Research / reviewers in the wild / expert
Jianan Fan
dblp:248/7360
· DBLP profile ↗
14ranked-venue papers
5as first author
14since 2021 · last 2026
0000-0003-4424-9572ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 5 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ART-ASyn: Anatomy-aware Realistic Texture-based Anomaly Synthesis Framework for Chest X-RaysabstractUnsupervised anomaly detection aims to identify anomalies without pixel-level annotations. Synthetic anomaly-based methods exhibit a unique capacity to introduce controllable irregularities with known masks, enabling explicit supervision during training. However, existing methods often produce synthetic anomalies that are visually distinct from real pathological patterns and ignore anatomical structure. This paper presents a novel Anatomy-aware Realistic Texture-based Anomaly Synthesis framework (ART-ASyn) for chest X-rays that generates realistic and anatomically consistent lung opacity related anomalies using texture-based augmentation guided by our proposed Progressive Binary Thresholding Segmentation method (PBTSeg) for lung segmentation. The generated paired samples of synthetic anomalies and their corresponding precise pixel-level anomaly mask for each normal sample enable explicit segmentation supervision. In contrast to prior work limited to one-class classification, ART-ASyn is further evaluated for zero-shot anomaly segmentation, demonstrating generalizability on an unseen dataset without target-domain annotations. Code availability is available at https://github.com/angelacao-hub/ART-ASyn. Qinyi Cao, Jianan Fan, Tom Weidong Cai |
WACV | 2 |
| 2026 | Gene-DML: Dual-Pathway Multi-Level Discrimination for Gene Expression Prediction from Histopathology ImagesabstractAccurately predicting gene expression from histopathology images offers a scalable and non-invasive approach to molecular profiling, with significant implications for precision medicine and computational pathology. However, existing methods often underutilize the cross-modal representation alignment between histopathology images and gene expression profiles across multiple representational levels, thereby limiting their prediction performance. To address this, we propose Gene-DML, a unified framework that structures latent space through Dual-pathway Multi-Level discrimination to enhance correspondence between morphological and transcriptional modalities. The multi-scale instance-level discrimination pathway aligns hierarchical histopathology representations extracted at local, neighbor, and global levels with gene expression profiles, capturing scale-aware morphological-transcriptional relationships. In parallel, the cross-level instance-group discrimination pathway enforces structural consistency between individual (image/gene) instances and modality-crossed (gene/image, respectively) groups, strengthening the alignment across modalities. By jointly modeling fine-grained and structural-level discrimination, Gene-DML is able to learn robust cross-modal representations, enhancing both predictive accuracy and generalization across diverse biological contexts. Extensive experiments on public spatial transcriptomics datasets demonstrate that Gene-DML achieves state-of-the-art performance in gene expression prediction. The code and processed datasets are available at https://github.com/YXSong000/Gene-DML. Yaxuan Song, Jianan Fan, Hang Chang, Tom Weidong Cai |
WACV | 2 |
| 2026 | Beyond benchmarks of IUGC: Rethinking requirements of deep learning method for intrapartum ultrasound biometry from fetal ultrasound videos
Jieyun Bai, Yitong Tang, Zhuonan Liang, Jianan Fan, Lisa Mcguire, Jillian Clarke, Tom Weidong Cai, Jacqueline Spurway, Yubo Tang, Shiye Wang, Wenda Shen, Wangwang Yu, Philippe Zhang, Weili Jiang, Salem Muhsin Ali Binqahal Al Nasim, Arsen Abzhanov, Numan Saeed, Mohammad Yaqub, Zunhui Xia, Hongxing Li 0001, Libin Lan, Jayroop Ramesh, Valentin Bacher, Mark Eid, Hoda Kalabizadeh, Christian Rupprecht 0001, Ana I. L. Namburete, Pak-Hei Yeung, Madeleine K. Wyburd, Nicola K. Dinsdale, Assanali Serikbey, Jiankai Li, Sung-Liang Chen, Zicheng Hu, Nana Liu, Yian Deng, Wenfeng Zhang, Mai Tuyet Nhi, Gregor Koehler, Rapheal Stock, Klaus H. Maier-Hein, Marawan Elbatel, Xiaomeng Li 0001, Saad Slimani, Victor M. Campello, Benard Ohene Botwe, Isaac Khobo, Zhenyan Han, Hongying Hou, Di Qiu, Gongning Luo, Dong Ni 0001, Yaosheng Lu, Karim Lekadir, Shuo Li 0001 |
Medical Image Anal. | 6 |
| 2026 | VGM-UNet: A hybrid visual graph deformable mamba with fourier neural operator U-Net for medical image segmentationabstractDeep learning methods have demonstrated remarkable advancements in medical image segmentation. However, achieving high accuracy remains a prominent challenge. In this paper, we present a new architecture, named VGM-UNet, that improves U-shaped segmentation models in performance and expressiveness by introducing the Structured State Space Duality algorithm to combine Graph Neural Networks, sparse attention, and Mamba-2, into U-Net and yield the best of these designs. This is accomplished through three primary modifications: we first adopt a novel approach, constructing a 2D State Space Model and an eight-way multi-scanning module, thereby creating the Vision Mamba-2, which serves as the foundation for building a hierarchical visual backbone that can be directly applied to the graph structure of image patches. Then, based on the Fast Fourier transform, we construct a Feed-Forward Network module, as a complement to Mamba, to model channel contents and improve the accuracy of capturing small objects. Moreover, on the basis of the modular architecture, we build a simple yet powerful U-shaped hybrid network, which simplifies the model design and enhances the model's expressiveness. These changes alleviate the limitations of conventional U-shaped architectures in accuracy improvements and achieve impressive results. We validate VGM-UNet through extensive experiments, demonstrating that our model outperforms existing state-of-the-art models in terms of segmentation accuracy on the Synapse and ACDC benchmark datasets. The experimental results also indicate that the Visual Graph State Space module can be conveniently applied to various medical image segmentation tasks. Jianan Fan, Tom Weidong Cai |
Neural Networks | 2 |
| 2026 | Cell as Point: One-stage framework for efficient cell trackingabstractConventional multi-stage cell tracking approaches rely heavily on detection or segmentation in each frame as a prerequisite, requiring substantial resources for high-quality segmentation masks and increasing the overall prediction time. To address these limitations, we propose CAP , a novel end-to-end one-stage framework that reimagines cell tracking by treating C ell a s P oint. Unlike traditional methods, CAP eliminates the need for explicit detection or segmentation, instead jointly tracking cells for sequences in one stage by leveraging the inherent correlations among their trajectories. This simplification reduces both labeling requirements and pipeline complexity. However, directly processing the entire sequence in one stage poses challenges related to data imbalance in capturing cell division events and long sequence inference. To solve these challenges, CAP introduces two key innovations: (1) adaptive event-guided (AEG) sampling, which prioritizes cell division events to mitigate the occurrence imbalance of cell events, and (2) the rolling-as-window (RAW) inference strategy, which ensures continuous and stable tracking of newly emerging cells over extended sequences. By removing the dependency on segmentation-based preprocessing while addressing the challenges of imbalanced occurrence of cell events and long-sequence tracking, CAP demonstrates promising cell tracking performance and is 8 to 32 times more efficient than existing methods. The code and model checkpoints are available at https://github.com/YXSong000/CAP . Yaxuan Song, Jianan Fan, Heng Huang 0001, Tom Weidong Cai |
Pattern Recognit. | 2 |
| 2026 | MIRROR: Multi-Modal Pathological Self-Supervised Representation Learning via Modality Alignment and RetentionabstractHistopathology and transcriptomics are fundamental modalities in cancer diagnostics, encapsulating the morphological and molecular characteristics of the disease. Multi-modal self-supervised learning has demonstrated remarkable potential in learning pathological representations by integrating diverse data sources. Conventional multi-modal integration methods primarily emphasize modality alignment, while paying insufficient attention to retaining the modality-specific intrinsic structures. However, unlike conventional scenarios where multi-modal inputs often share highly overlapping features, histopathology and transcriptomics exhibit pronounced heterogeneity, offering orthogonal yet complementary insights. Histopathology data provides morphological and spatial context, elucidating tissue architecture and cellular topology, whereas transcriptomics data delineates molecular signatures through quantifying gene expression patterns. This inherent disparity introduces a major challenge in aligning these modalities while maintaining modality-specific fidelity. To address these challenges, we present MIRROR, a novel multi-modal representation learning framework designed to foster both modality alignment and retention. MIRROR employs dedicated encoders to extract comprehensive feature representations for each modality, which is further complemented by a modality alignment module to achieve seamless integration between phenotype patterns and molecular profiles. Furthermore, a modality retention module safeguards unique attributes from each modality, while a style clustering module mitigates redundancy and enhances disease-relevant information by modeling and aligning consistent pathological signatures within a clustering space. Extensive evaluations on The Cancer Genome Atlas (TCGA) cohorts for cancer subtyping and survival analysis highlight MIRROR's superior performance, demonstrating its effectiveness in constructing comprehensive oncological feature representations and benefiting the cancer diagnosis. Code is available at https://github.com/TianyiFranklinWang/MIRROR. Jianan Fan, Dingxin Zhang 0001, Dongnan Liu, Yong Xia 0001, Heng Huang 0001, Tom Weidong Cai |
IEEE Trans. Medical Imaging | 2 |
| 2025 | ScSAM: Debiasing Morphology and Distributional Variability in Subcellular Semantic SegmentationabstractThe significant morphological and distributional variability among subcellular components poses a long-standing challenge for learning-based organelle segmentation models, significantly increasing the risk of biased feature learning. Existing methods often rely on single mapping relationships, overlooking feature diversity and thereby inducing biased training. Although the Segment Anything Model (SAM) provides rich feature representations, its application to subcellular scenarios is hindered by two key challenges: (1) The variability in subcellular morphology and distribution creates gaps in the label space, leading the model to learn spurious or biased features. (2) SAM focuses on global contextual understanding and often ignores fine-grained spatial details, making it challenging to capture subtle structural alterations and cope with skewed data distributions. To address these challenges, we introduce ScSAM, a method that enhances feature robustness by fusing pre-trained SAM with Masked Autoencoder (MAE)-guided cellular prior knowledge to alleviate training bias from data imbalance. Specifically, we design a feature alignment and fusion module to align pre-trained embeddings to the same feature space and efficiently combine different representations. Moreover, we present a cosine similarity matrix-based class prompt encoder to activate class-specific features to recognize subcellular categories. Extensive experiments on diverse subcellular image datasets demonstrate that ScSAM outperforms state-of-the-art methods. Jianan Fan, Dongnan Liu, Hang Chang, Gerald J. Shami, Filip Braet, Tom Weidong Cai |
ECAI | 2 |
| 2025 | AMNCutter: Affinity-Attention-Guided Multi-View Normalized Cutter for Unsupervised Surgical Instrument SegmentationabstractSurgical instrument segmentation (SIS) is pivotal for robotic-assisted minimally invasive surgery, assisting surgeons by identifying surgical instruments in endoscopic video frames. Recent unsupervised surgical instrument segmentation (USIS) methods primarily rely on pseudo-labels derived from low-level features such as color and optical flow, but these methods show limited effective-ness and generalizability in complex and unseen endo-scopic scenarios. In this work, we propose a label-free unsupervised model featuring a novel module named Multi-View Normalized Cutter (m-NCutter). Different from previous USIS works, our model is trained using a graph-cutting loss function that leverages patch affini-ties for supervision, eliminating the need for pseudo-labels. The framework adaptively determines which affini-ties from which levels should be prioritized. Therefore, the low- and high-level features and their affinities are effectively integrated to train a label-free unsupervised model, showing superior effectiveness and generalization abil-ity. We conduct comprehensive experiments across mul-tiple SIS datasets to validate our approach's state-of-the-art (SOTA) performance, robustness, and exceptional potential as a pre-trained model. Our code is released at https://github.com/MingyuShengSMYIAMNCutter. Mingyu Sheng, Jianan Fan, Dongnan Liu, Ron Kikinis, Tom Weidong Cai |
WACV | 2 |
| 2025 | A Small Target Recognition Method Based on an Improved YOLOv8abstractThe recognition of small-size targets presents a significant challenge in computer vision, as their reduced dimensions often lead to diminished detection accuracy. To address this issue, this paper proposes an enhanced small-size target recognition method based on an improved version of You Only Look Once Version 8 (YOLOv8). The proposed improvements include integrating the dynamic attention mechanism of BiFormer (Vision Transformer with Bi-level Routing Attention), which leverages the sparsity of dynamic and query perception to enable more flexible and adaptive content perception. Additionally, the Weighted Intersection over Union (WIOU) loss function is introduced to address the imbalance in Bounding Box Regression (BBR) between samples, enhancing the overall accuracy of the model. Furthermore, a specialized detection head for small targets and a confidence-adaptive module are added at the detection head’s end, improving feature extraction and continuous tracking capabilities for small targets, especially under conditions of low visibility and target occlusion. Experimental results demonstrate that the improved model significantly enhances the detection of incomplete and small-sized targets, providing robust performance in scenarios with occlusion and reduced visibility. This study emphasizes the potential of the enhanced YOLOv8 model in real-world applications, providing new improvement ideas for the development of occluded and small target recognition. Sai Biao Jiang, Jianan Fan, Zhijin Sun, Kaitao Deng, Yanbing Huang |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2025 | On Structuring Hyperspherical Manifold for Probing Novel Biomedical EntitiesabstractThe insufficient high-throughput modeling capability for high-dimensional, multiscale, and nonlinear real-world observations and measurements stands as one of the major impediments for modern science advancements. In this regard, machine learning holds tremendous promise for transforming the fundamental practice of scientific discovery by virtue of its data-driven disposition. With the ever-increasing stream of research data collection, it would be appealing to automate the exploration of patterns and insights from observational data for discovering novel classes of phenotypes and entities. However, in the discipline of biomedical investigation, the cumulative data is intrinsically subjected to non-i.i.d. distribution and severe biases amongst different clusters, inducing disorganization and ambiguity in the learned representation space. To contend with the inherent challenges, in this paper, we present a geometry- constrained probabilistic modeling treatment on hyperspherical manifolds. It firstly parameterizes the approximated posterior of instance-wise embedding as a marginal von MisesFisher distribution to account for the interference of distributional latent shift, and thereafter incorporates a suite of critical inductive biases to organically shape the layout of tailored embedding space. Together, these advancements offer a systematic solution to regularize the uncontrollable risk for unseen class learning and prospecting. Furthermore, we propose a spectral graph-theoretic method to efficiently estimate the number of potential novel classes and endow the prediction with adorable taxonomy adaptability. Through extensive experiments under various settings, we demonstrate the effectiveness and general applicability of the proposed methods in recognizing and structurally phenotyping novel visual concepts. Jianan Fan, Dongnan Liu, Hang Chang, Heng Huang 0001, Tom Weidong Cai |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Seeing Unseen: Discover Novel Biomedical Concepts via Geometry-Constrained Probabilistic ModelingabstractMachine learning holds tremendous promise for trans-forming the fundamental practice of scientific discovery by virtue of its data-driven nature. With the ever-increasing stream of research data collection, it would be appealing to autonomously explore patterns and insights from obser-vational data for discovering novel classes of phenotypes and concepts. However, in the biomedical domain, there are several challenges inherently presented in the cumu-lated data which hamper the progress of novel class dis-covery. The non-i.i.d. data distribution accompanied by the severe imbalance among different groups of classes es-sentially leads to ambiguous and biased semantic represen-tations. In this work, we present a geometry-constrained probabilistic modeling treatment to resolve the identified is-sues. First, we propose to parameterize the approximated posterior of instance embedding as a marginal von Mises-Fisher distribution to account for the interference of distri-butional latent bias. Then, we incorporate a suite of critical geometric properties to impose proper constraints on the layout of constructed embedding space, which in turn min-imizes the uncontrollable risk for unknown class learning and structuring. Furthermore, a spectral graph-theoretic method is devised to estimate the number of potential novel classes. It inherits two intriguing merits compared to exis-tent approaches, namely high computational efficiency and flexibility for taxonomy-adaptive estimation. Extensive ex-periments across various biomedical scenarios substantiate the effectiveness and general applicability of our method. Jianan Fan, Dongnan Liu, Hang Chang, Heng Huang 0001, Tom Weidong Cai |
CVPR | 1 |
| 2024 | Revisiting Adaptive Cellular Recognition Under Domain Shifts: A Contextual Correspondence View
Jianan Fan, Dongnan Liu, Canran Li, Hang Chang, Heng Huang 0001, Filip Braet, Tom Weidong Cai |
ECCV (73) | 1 |
| 2024 | Learning to Generalize over Subpartitions for Heterogeneity-Aware Domain Adaptive Nuclei SegmentationabstractAbstract Annotation scarcity and cross-modality/stain data distribution shifts are two major obstacles hindering the application of deep learning models for nuclei analysis, which holds a broad spectrum of potential applications in digital pathology. Recently, unsupervised domain adaptation (UDA) methods have been proposed to mitigate the distributional gap between different imaging modalities for unsupervised nuclei segmentation in histopathology images. However, existing UDA methods are built upon the assumption that data distributions within each domain should be uniform. Based on the over-simplified supposition, they propose to align the histopathology target domain with the source domain integrally, neglecting severe intra-domain discrepancy over subpartitions incurred by mixed cancer types and sampling organs. In this paper, for the first time, we propose to explicitly consider the heterogeneity within the histopathology domain and introduce open compound domain adaptation (OCDA) to resolve the crux. In specific, a two-stage disentanglement framework is proposed to acquire domain-invariant feature representations at both image and instance levels. The holistic design addresses the limitations of existing OCDA approaches which struggle to capture instance-wise variations. Two regularization strategies are specifically devised herein to leverage the rich subpartition-specific characteristics in histopathology images and facilitate subdomain decomposition. Moreover, we propose a dual-branch nucleus shape and structure preserving module to prevent nucleus over-generation and deformation in the synthesized images. Experimental results on both cross-modality and cross-stain scenarios over a broad range of diverse datasets demonstrate the superiority of our method compared with state-of-the-art UDA and OCDA methods. Graphical abstract Jianan Fan, Dongnan Liu, Hang Chang, Tom Weidong Cai |
Int. J. Comput. Vis. | 1 |
| 2023 | Taxonomy Adaptive Cross-Domain Adaptation in Medical Imaging via Optimization Trajectory DistillationabstractThe success of automated medical image analysis depends on large-scale and expert-annotated training sets. Unsupervised domain adaptation (UDA) has been raised as a promising approach to alleviate the burden of labeled data collection. However, they generally operate under the closed-set adaptation setting assuming an identical label set between the source and target domains, which is over-restrictive in clinical practice where new classes commonly exist across datasets due to taxonomic inconsistency. While several methods have been presented to tackle both domain shifts and incoherent label sets, none of them take into account the common characteristics of the two issues and consider the learning dynamics along network training. In this work, we propose optimization trajectory distillation, a unified approach to address the two technical challenges from a new perspective. It exploits the low-rank nature of gradient space and devises a dual-stream distillation algorithm to regularize the learning dynamics of insufficiently annotated domain and classes with the external guidance obtained from reliable sources. Our approach resolves the issue of inadequate navigation along network optimization, which is the major obstacle in the taxonomy adaptive cross-domain adaptation scenario. We evaluate the proposed method extensively on several tasks towards various endpoints with clinical and open-world significance. The results demonstrate its effectiveness and improvements over previous methods. Code is available at https://github.com/camwew/TADA-MI. Jianan Fan, Dongnan Liu, Hang Chang, Heng Huang 0001, Tom Weidong Cai |
ICCV | 1 |