EDBT 2026 Demo / reviewers in the wild / expert
Jun Wang 0121
dblp:125/8189-121
· DBLP profile ↗
15ranked-venue papers
7as first author
14since 2021 · last 2026
0000-0003-1581-8369ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Advancing federated domain generalization in ophthalmology: Vision enhancement and consistency assurance for multicenter fundus image segmentation
Yang Zhao 0019, Xianxun Zhu, Jun Wang 0121, Yan Liu 0052 |
Pattern Recognit. | 5 |
| 2026 | Improving Medical Visual Representation Learning With Pathological-Level Cross-Modal Alignment and Correlation ExplorationabstractLearning medical visual representations from image-report pairs through joint learning has garnered increasing research attention due to its potential for transferring acquired knowledge to various downstream medical tasks. Previous works have predominantly focused on instance-wise or token-wise cross-modal alignment, often neglecting the importance of pathological-level consistency. This paper presents a novel framework PLACE that promotes the Pathological-Level Alignment and enriches the fine-grained details via Correlation Exploration without additional human annotations. Specifically, we propose a novel pathological-level cross-modal alignment (PCMA) approach to maximize the consistency of pathology observations from both images and reports. To facilitate this, a Visual Pathology Observation Extractor is introduced to extract visual pathological observation representations from localized tokens. The PCMA module operates independently of any external disease annotations, enhancing the generalizability and robustness of our methods. Furthermore, we design a proxy task that enforces the model to identify correlations among image patches, thereby enriching the fine-grained details crucial for various downstream tasks. Experimental results demonstrate that our proposed framework achieves new state-of-the-art performance on multiple downstream tasks, including classification, image-to-text retrieval, semantic segmentation, object detection and report generation. Jun Wang 0121, Lixing Zhu, Xiaohan Yu 0001, Abhir Bhalerao, Yulan He 0001 |
IEEE J. Biomed. Health Informatics | 1 |
| 2025 | LlmLink: Dual LLMs for Dynamic Entity Linking on Long Narratives with Collaborative Memorisation and Prompt OptimisationabstractWe address the task of CoREFerence resolution (CoREF) in chunked long narratives. Existing approaches remain either focused on supervised fine-tuning or limited to one-off prediction, which poses a challenge where the context is long. We develop a dynamic approach to cope with this: by deploying dual Large Language Models (LLMs), we assign specialised LLMs to local named entity recognition and distant CoREF tasks, respectively, while ensuring their exchange of information. Utilising our novel memorisation schemes, the coreference resolution LLM would memorise characters and their associated descriptions, thereby reducing token consumption compared with storing previous messages. To alleviate hallucinations of LLMs, we employ an automatic prompt optimisation method, with the LLM ranker modified to leverage annotations. Our approach achieves performance gains over other LLM-based models and fine-tuning approaches on long narrative datasets, significantly reducing the resources required for inference and training. Lixing Zhu, Jun Wang 0121, Yulan He 0001 |
COLING | 2 |
| 2025 | COCO-OLAC: A Benchmark for Occluded Panoptic Segmentation and Image UnderstandingabstractTo help address the occlusion problem in panoptic segmentation and image understanding, this paper proposes a new large-scale dataset named COCO-OLAC (COCO Occlusion Labels for All Computer Vision Tasks), which is derived from the COCO dataset by manually labelling images into three perceived occlusion levels. Using COCO-OLAC, we systematically assess and quantify the impact of occlusion on panoptic segmentation on samples having different levels of occlusion. Comparative experiments with SOTA panoptic models demonstrate that the presence of occlusion significantly affects performance, with higher occlusion levels resulting in notably poorer performance. Additionally, we propose a straightforward yet effective method as an initial attempt to leverage the occlusion annotation using contrastive learning to render a model that learns a more robust representation capturing different severities of occlusion. Experimental results demonstrate that the proposed approach boosts the performance of the baseline model and achieves SOTA performance on the proposed COCO-OLAC dataset.1 Jun Wang 0121, Abhir Bhalerao |
ICASSP | 2 |
| 2025 | Deep Learning for Multivariate Time Series Imputation: A SurveyabstractMissing values are ubiquitous in multivariate time series (MTS) data, posing significant challenges for accurate analysis and downstream applications. In recent years, deep learning-based methods have successfully handled missing data by leveraging complex temporal dependencies and learned data distributions. In this survey, we provide a comprehensive summary of deep learning approaches for multivariate time series imputation (MTSI) tasks. We propose a novel taxonomy that categorizes existing methods based on two key perspectives: imputation uncertainty and neural network architecture. Furthermore, we summarize existing MTSI toolkits with a particular emphasis on the PyPOTS Ecosystem, which provides an integrated and standardized foundation for MTSI research. Finally, we discuss key challenges and future research directions, which give insight for further MTSI research. This survey aims to serve as a valuable resource for researchers and practitioners in the field of time series analysis and missing data imputation tasks. A well-maintained MTSI paper and tool list is available at https://github.com/WenjieDu/Awesome_Imputation. Jun Wang 0121, Yiyuan Yang, Linglong Qian, Keli Zhang, Yuxuan Liang 0002, Qingsong Wen |
IJCAI | 1 |
| 2025 | How Deep is Your Guess? A Fresh Perspective on Deep Learning for Medical Time-Series ImputationabstractWe present a comprehensive analysis of deep learning approaches for Electronic Health Record (EHR) time-series imputation, examining how the interplay between architectural and framework design decisions gives rise to higher-level properties of a given deep imputer model and distinct biases towards complex data characteristics. Our investigation reveals the varying capabilities of deep imputers in capturing complex spatio-temporal dependencies within EHRs, and that the effectiveness of the model depends on how its combined biases align with the characteristics of the medical time series. Our experimental evaluation challenges common assumptions about model complexity, demonstrating that larger models do not necessarily improve performance. Rather, carefully designed architectures can better capture the complex patterns inherent in clinical data. The study highlights the need for imputation approaches that prioritise clinically meaningful data reconstruction over statistical accuracy. Our experiments further reveal up to 20% in variations of imputation performance based on preprocessing and implementation choices, emphasising the need for standardised benchmarking methodologies. Finally, we identify critical gaps between current deep imputation methods and medical requirements, highlighting the importance of integrating clinical insights to achieve more reliable imputation approaches for healthcare applications. Linglong Qian, Hugh Logan Ellis, Tao Wang 0036, Jun Wang 0121, Robin Mitra, Richard J. B. Dobson, Zina M. Ibrahim |
IEEE J. Biomed. Health Informatics | 4 |
| 2024 | CAMANet: Class Activation Map Guided Attention Network for Radiology Report GenerationabstractRadiology report generation (RRG) has gained increasing research attention because of its huge potential to mitigate medical resource shortages and aid the process of disease decision making by radiologists. Recent advancements in Radiology Report Generation (RRG) are largely driven by improving a model's capabilities in encoding single-modal feature representations, while few studies explicitly explore the cross-modal alignment between image regions and words. Radiologists typically focus first on abnormal image regions before composing the corresponding text descriptions, thus cross-modal alignment is of great importance to learn a RRG model which is aware of abnormalities in the image. Motivated by this, we propose a Class Activation Map guided Attention Network (CAMANet) which explicitly promotes cross-modal alignment by employing aggregated class activation maps to supervise cross-modal attention learning, and simultaneously enrich the discriminative information. Experimental results demonstrate that CAMANet outperforms previous SOTA methods on two commonly used RRG benchmarks. Jun Wang 0121, Abhir Bhalerao, Terry Yin, Simon See, Yulan He 0001 |
IEEE J. Biomed. Health Informatics | 1 |
| 2023 | Optimizing Vision Transformers for Medical Image SegmentationabstractFor medical image semantic segmentation (MISS), Vision Transformers have emerged as strong alternatives to convolutional neural networks thanks to their inherent ability to capture long-range correlations. However, existing research uses off-the-shelf vision Transformer blocks based on linear projections and feature processing which lack spatial and local context to refine organ boundaries. Furthermore, Transformers do not generalize well on small medical imaging datasets and rely on large-scale pre-training due to limited inductive biases. To address these problems, we demonstrate the design of a compact and accurate Transformer network for MISS, CS-Unet, which introduces convolutions in a multi-stage design for hierarchically enhancing spatial and local modeling ability of Transformers. This is mainly achieved by our well-designed Convolutional Swin Transformer (CST) block which merges convolutions with Multi-Head Self-Attention and Feed-Forward Networks for providing inherent localized spatial context and inductive biases. Experiments demonstrate CS-Unet without pre-training out- performs other counterparts by large margins on multi-organ and cardiac datasets with fewer parameters and achieves state-of-the-art performance. Our code is available at Github1. Qianying Liu, Chaitanya Kaul, Jun Wang 0121, Christos Anagnostopoulos 0001, Roderick Murray-Smith, Fani Deligianni |
ICASSP | 3 |
| 2023 | CLE-ViT: Contrastive Learning Encoded Transformer for Ultra-Fine-Grained Visual CategorizationabstractUltra-fine-grained visual classification (ultra-FGVC) targets at classifying sub-grained categories of fine-grained objects. This inevitably requires discriminative representation learning within a limited training set. Exploring intrinsic features from the object itself, e.g., predicting the rotation of a given image, has demonstrated great progress towards learning discriminative representation. Yet none of these works consider explicit supervision for learning mutual information at instance level. To this end, this paper introduces CLE-ViT, a novel contrastive learning encoded transformer, to address the fundamental problem in ultra-FGVC. The core design is a self-supervised module that performs self-shuffling and masking and then distinguishes these altered images from other images. This drives the model to learn an optimized feature space that has a large inter-class distance while remaining tolerant to intra-class variations. By incorporating this self-supervised module, the network acquires more knowledge from the intrinsic structure of the input data, which improves the generalization ability without requiring extra manual annotations. CLE-ViT demonstrates strong performance on 7 publicly available datasets, demonstrating its effectiveness in the ultra-FGVC task. The code is available at https://github.com/Markin-Wang/CLEViT. Xiaohan Yu 0001, Jun Wang 0121, Yongsheng Gao 0001 |
IJCAI | 2 |
| 2023 | Mix-ViT: Mixing attentive vision transformer for ultra-fine-grained visual categorization
Xiaohan Yu 0001, Jun Wang 0121, Yang Zhao 0019, Yongsheng Gao 0001 |
Pattern Recognit. | 2 |
| 2022 | Cross-Modal Prototype Driven Network for Radiology Report Generation
Jun Wang 0121, Abhir Bhalerao, Yulan He 0001 |
ECCV (35) | 1 |
| 2022 | PGTRNET: Two-Phase Weakly Supervised Object Detection with Pseudo Ground Truth RefinementabstractCurrent state-of-the-art weakly supervised object detection (WSOD) studies mainly follow a two-stage training strategy which integrates a fully supervised detector (FSD) with a pure WSOD model. There are two main problems hindering the performance of the two-phase WSOD approaches, i.e., insufficient learning problem and strict reliance between the FSD and the pseudo ground truth (PGT) generated by the WSOD model. This paper proposes pseudo ground truth refinement network (PGTRNet), a simple yet effective method with-out introducing any extra learnable parameters, to cope with these problems. PGTRNet utilizes multiple bounding boxes to establish the PGT, mitigating the insufficient learning problem. Besides, we propose a novel online PGT refinement approach to steadily improve the quality of PGT by fully taking advantage of the power of FSD during the second-phase training, decoupling the first and second-phase models. Elaborate experiments are conducted on the PASCAL VOC 2007 benchmark to verify the effectiveness of our methods. Experimental results demonstrate that PGTRNet boosts the backbone model by 2.1% mAP and achieves the state-of-the-art performance. Jun Wang 0121, Hefeng Zhou, Xiaohan Yu 0001 |
ICASSP | 1 |
| 2021 | Feature Fusion Vision Transformer for Fine-Grained Visual Categorization
Jun Wang 0121, Xiaohan Yu 0001, Yongsheng Gao 0001 |
BMVC | 1 |
| 2021 | Mask Guided Attention For Fine-Grained Patchy Image ClassificationabstractIn this work, we present a novel mask guided attention (MGA) method for fine-grained patchy image classification. The key challenge of fine-grained patchy image classification lies in two folds, ultra-fine-grained inter-category variances among objects and very few data available for training. This motivates us to consider employing more useful supervision signal to train a discriminative model within limited training samples. Specifically, the proposed MGA integrates a pre-trained semantic segmentation model that produces auxiliary supervision signal, i.e., patchy attention mask, enabling a discriminative representation learning. The patchy attention mask drives the classifier to filter out the insignificant parts of images (e.g., common features between different categories), which enhances the robustness of MGA for the fine-grained patchy image classification. We verify the effectiveness of our method on three publicly available patchy image datasets. Experimental results demonstrate that our MGA method achieves superior performance on three datasets compared with the state-of-the-art methods. In addition, our ablation study shows that MGA improves the accuracy by 2.25% and 2% on the SoyCultivarVein and BtfPIS datasets, indicating its practicality towards solving the fine-grained patchy image classification. Jun Wang 0121, Xiaohan Yu 0001, Yongsheng Gao 0001 |
ICIP | 1 |
| 2019 | Nurse scheduling problem based on hydrologic cycle optimizationabstractBuilding the work timetables for staff in healthcare institutions is known to be a highly constrained and NP-hard problem. In this research, a mathematical programming model, maximizing nurses' preference for work shifts and rest days while minimizing hospital operating costs, is proposed to solve the nurse scheduling problem (NSP) optimally. Then, we apply a new optimization algorithm-HCOMA, HCO based memetic algorithm, combining entropy-based decision-making mechanism and local search, to heuristically solve the NSP. In the global search, the entropy is calculated to assess population diversity following by every specified iteration. By analyzing the change of diversity, the population can identify the stagnation of search and perform local search at the best time. In summary, the local search includes three core parts: Meta-Lamarckian learning strategy, cooling schedule and Metropolis Criterion. Three neighborhood structures are utilized to exchange or reset the nurse's shifts, expanding the feasible solution area of the search and generating high-quality solutions. The Meta-Lamarckian learning strategy is used to automatically choose the best search structure based on their performance. The performance of HCOMA was tested with sufficient experimentations. The test problems were generated based on the actual situation of a hospital, including an instance and 30 random problems. The results indicate that the proposed algorithm was superior to the standard HCO and three well-known evolutionary algorithms in solution quality and convergence rate. Qianying Liu, Ben Niu 0002, Jun Wang 0121, Hong Wang 0016, Li Li 0004 |
CEC | 3 |