Zihao Jin

dblp:189/5822 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
9since 2021 · last 2026
0009-0000-4903-7312ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Computer networks · 2 · 1 since 2021Security and privacy · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 GEMA-Score: Granular Explainable Multi-Agent Scoring Framework for Radiology Report Evaluation
abstract
Automatic medical report generation has the potential to support clinical diagnosis, reduce the workload of radiologists, and demonstrate potential for enhancing diagnostic consistency. However, current evaluation metrics often fail to reflect the clinical reliability of generated reports. Overlap-based methods overlook fine-grained details (e.g., location, severity), diagnostic metrics are constrained by fixed vocabularies. Some diagnostic metrics are limited by fixed vocabularies or templates, reducing their ability to capture diverse clinical expressions. LLM-based metrics lack interpretable reasoning, limiting trust in clinical settings. Therefore, we propose a Granular Explainable Multi-Agent Score (GEMA-Score) in this paper, which conducts both objective quantification and subjective evaluation through a large language model-based multi-agent workflow. Our GEMA-Score parses structured reports and employs stable calculations through interactive exchanges of information among agents to assess disease diagnosis, location, severity, and uncertainty. Additionally, an LLM-based scoring agent evaluates completeness, readability, and clinical terminology while providing explanatory feedback. Extensive experiments show that GEMA-Score achieves the highest correlation with human experts on public datasets (Kendall = 0.69 on ReXVal; 0.45 on RadEvalX), demonstrating improved clinical scoring reliability.
Zhenxuan Zhang, Kinhei Lee, Peiyuan Jing, Weihang Deng, Huichi Zhou, Zihao Jin, Zhifan Gao, Dominic C. Marshall, Yingying Fang, Guang Yang 0006
AAAI6
2026 Predicting diabetic macular edema treatment responses using OCT: Dataset and methods of APTOS competition
abstract
• First challenge to focus on pre-treatment stratification for diabetic macular edema (DME): This study pioneers the use of pre-treatment OCT biomarkers to predict individual responses to anti-VEGF therapy, advancing the concept of personalized medicine in DME management. • Large-scale, publicly accessible OCT dataset: The competition provides one of the most comprehensive open-access DME datasets to date, comprising tens of thousands of OCT images from 2,000 patients, significantly addressing the field’s data scarcity. The dataset includes both per-eye and per-scan annotations for several critical retinal biomarkers, enabling a wide range of supervised learning applications. • Benchmark for future research: With 170 registered teams and 41 finalists, the challenge fostered broad engagement. The best team achieved an AUC of 80.06%, highlighting the feasibility of accurate outcome prediction. The challenge provides standardized evaluation metrics and a curated leaderboard, laying the groundwork for reproducible and comparable AI model development in ophthalmic treatment prediction. Diabetic macular edema (DME) significantly contributes to visual impairment in diabetic patients. Treatment responses to intravitreal therapies vary, highlighting the need for patient stratification to predict therapeutic benefits and enable personalized strategies. To our knowledge, this study is the first to explore pre-treatment stratification for predicting DME treatment responses. To advance this research, we organized the 2nd Asia-Pacific Tele-Ophthalmology Society (APTOS) Big Data Competition in 2021. The competition focused on improving predictive accuracy for anti-VEGF therapy responses using ophthalmic OCT images. We provided a dataset containing tens of thousands of OCT images from 2,000 patients with labels across four sub-tasks. This paper details the competition’s structure, dataset, leading methods, and evaluation metrics. The competition attracted strong scientific community participation, with 170 teams initially registering and 41 reaching the final round. The top-performing team achieved an AUC of 80.06%, highlighting the potential of AI in personalized DME treatment and clinical decision-making.
Weiyi Zhang 0004, Peranut Chotcomwongse, Yinwen Li, Pusheng Xu, Ruijie Yao, Lianhao Zhou, Qiping Zhou, Shoujin Huang, Zihao Jin, Florence H. T. Chung, Yalin Zheng, Mingguang He, Danli Shi, Paisan Ruamviboonsuk
Medical Image Anal.12
2025 Cyclic Vision-Language Manipulator: Towards Reliable and Fine-Grained Image Interpretation for Automated Report Generation
abstract
Despite significant advancements in automated report generation, the opaqueness of text interpretability continues to cast doubt on the reliability of the content produced. This paper introduces a novel approach to identify specific image features in X-ray images that influence the outputs of report generation models. Specifically, we propose Cyclic Vision-Language Manipulator (CVLM), a module to generate a manipulated X-ray from an original X-ray and its report from a designated report generator. The essence of CVLM is that cycling manipulated X-rays to the report generator produces altered reports aligned with the alterations pre-injected into the reports for X-ray generation, achieving the term ``cyclic manipulation''. This process allows direct comparison between original and manipulated X-rays, clarifying the critical image features driving changes in reports and enabling model users to assess the reliability of the generated texts. Empirical evaluations demonstrate that CVLM can identify more precise and reliable features compared to existing explanation methods, significantly enhancing the transparency and applicability of AI-generated reports.
Yingying Fang, Zihao Jin, Shaojie Guo, Jinda Liu, Zhiling Yue, Yijian Gao, Junzhi Ning, Simon Walsh, Guang Yang 0006
IJCAI2
2025 Informer learning framework based on secondary decomposition for multi-step forecast of ultra-short term wind speed
Zihao Jin, Xiaomengting Fu, Ling Xiang, Guopeng Zhu, Aijun Hu
Eng. Appl. Artif. Intell.1
2025 Power Allocation Based on Federated Multiagent Deep Reinforcement Learning for NOMA Maritime Networks
abstract
The absence of established communication infrastructures at sea presents a significant hurdle for the development of the Internet of Things (IoT) in maritime environments. In this article, we propose a novel solution: a nonorthogonal multiple access (NOMA) maritime network employing power allocation facilitated by federated multiagent deep reinforcement learning (DRL). Conventional single-agent DRL algorithms encounter challenges, such as dimensionality explosion in real-world scenarios. We address this issue by extending these algorithms to incorporate multiple agents. Furthermore, to mitigate risks associated with centralized training, such as data leakage and network attacks, we adopt a federated learning (FL) framework for distributed training across multiple agents. By uploading only select parameters during training and keeping data locally, FL not only enhances algorithm convergence speed but also bolsters data privacy. Specifically, we introduce the federated multiagent deep Q-network (FLMADQN) algorithm tailored for power allocation in NOMA maritime networks. Our algorithm aims to maximize system throughput, optimize data transfer rates, and expedite convergence. Through extensive computer simulations, we validate the efficacy of our proposed approach. Results demonstrate that the FLMADQN algorithm significantly outperforms the traditional DRL algorithm, DQN, in NOMA maritime environments, improving average system throughput by 20.12%, peak system throughput by 45.10%, and system spectral efficiency by 48.33%. Moreover, FLMADQN exhibits a twofold increase in convergence speed compared to MADQN without FL and is 2.5 times faster than DQN. Our findings underscore the potential of federated multiagent DRL in advancing communication systems for maritime IoT applications.
Yakai Zhang, Zihao Jin, Qinyu Zhang 0001
IEEE Internet Things J.4
2024 DiffExplainer: Unveiling Black Box Models Via Counterfactual Generation
Yingying Fang, Shuang Wu 0002, Zihao Jin, Caiwen Xu, Simon Walsh, Guang Yang 0006
MICCAI (10)3
2024 Diff3Dformer: Leveraging Slice Sequence Diffusion for Enhanced 3D CT Classification with Transformer Networks
Zihao Jin, Yingying Fang, Caiwen Xu, Simon Walsh, Guang Yang 0006
MICCAI (1)1
2023 A Security Study about Electron Applications and a Programming Methodology to Tame DOM Functionalities
Zihao Jin, Shuo Chen 0001, Hai-Xin Duan, Jianjun Chen 0005
NDSS1
2022 Timing-Based Browsing Privacy Vulnerabilities Via Site Isolation
abstract
Chromium’s site isolation ensures that different sites are rendered by different processes, which is a vision that academic researchers set forth over a decade ago. The journey from academic prototypes to the commercial availability represents a holistic rethinking about the security architecture for modern browsers. In this paper, we emphasize that the timing issues under site isolation need a thorough study. Specifically, we show that site isolation enables a realistic timing attack, which allows the attacker to identify which websites in a given target-sites set are loaded into the browser, as well as the website the user is currently interacting with. Through these vulnerabilities, the user’s site-visit behavior is leaked to the attacker. Our evaluation using Alexa Top 3000 websites gives very high vulnerability percentages – 99%, 99% and 95% for our three key metrics of vulnerabilities. Moreover, the attack is very robust without any special assumption, so will be effective if deployed in the field. The main challenge revealed by our work is the tension between the scarcity of processes and the obligation to isolate cross-site frames in different processes. We are working with the Google Chrome team and Microsoft Edge team to propose and evaluate mitigation options.
Zihao Jin, Ziqiao Kong, Shuo Chen 0001, Hai-Xin Duan
SP1
2019 Trace-based Behaviour Analysis of Network Servers
abstract
Analysing software and networks can be done using established tools, such as debuggers and packet analysers, but using established tools to analyse network software is difficult and impractical because of the sheer detail the tools present and the performance overheads they typically impose. This makes it difficult to precisely diagnose performance anomalies in network software to identify their causes (is it a DoS attack or a bug?) and determine what needs to be fixed.We present Flowdar: a practical tool for analysing software traces to produce intuitive summaries of network software behaviour by abstracting unimportant details and demultiplexing traces into different sessions' subtraces. Flowdar can use existing state-of-the-art tracing tools for lower overhead during trace gathering for offline analysis. Using Flowdar we can drill down when diagnosing performance anomalies without getting overwhelmed in detail or burdening the system being observed.We show that Flowdar can be applied to existing real-world software and can digest complex behaviour into an intuitive visualisation.
Nik Sultana, Achala Rao, Zihao Jin, Pardis Pashakhanloo, Henry Zhu, Vinod Yegneswaran, Boon Thau Loo
CNSM3
2019 Hashtray: Turning the tables on Scalable Client Classification
Nik Sultana, Pardis Pashakhanloo, Zihao Jin, Achala Rao, Boon Thau Loo
IM3