Luis Filipe Nakayama

dblp:358/7248 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2025
0000-0002-6847-6748ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Trustworthy machine learning · 77% Language models and text generation · 23%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Medical and health informatics · 100%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
fairness
0.912025
Multi-OphthaLingua: A Multilingual Benchmark for Assessing and Debiasing LLM Ophthalmological QA in LMICs · AAAI 2025
Medical and health informatics › biomedical natural language processing
medical question answering
0.912025
Multi-OphthaLingua: A Multilingual Benchmark for Assessing and Debiasing LLM Ophthalmological QA in LMICs · AAAI 2025
Natural language and speech › Language models and text generation
retrieval-augmented generation
0.312025
Multi-OphthaLingua: A Multilingual Benchmark for Assessing and Debiasing LLM Ophthalmological QA in LMICs · AAAI 2025

Methods — techniques the papers use, named apart from their topics

self-verification · 1.7retrieval-augmented generation · 1.7chain-of-thought · 1.7
YearPublicationVenuePosition
2025 Multi-OphthaLingua: A Multilingual Benchmark for Assessing and Debiasing LLM Ophthalmological QA in LMICs
abstract
Current ophthalmology clinical workflows are plagued by over-referrals, long waits, and complex and heterogeneous medical records. Large language models (LLMs) present a promising solution to automate various procedures such as triaging, preliminary tests like visual acuity assessment, and report summaries. However, LLMs have demonstrated significantly varied performance across different languages in natural language question-answering tasks, potentially exacerbating healthcare disparities in Low and Middle-Income Countries (LMICs). This study introduces the first multilingual ophthalmological question-answering benchmark with manually curated questions parallel across languages, allowing for direct cross-lingual comparisons. Our evaluation of 6 popular LLMs across 7 different languages reveals substantial bias across different languages, highlighting risks for clinical deployment of LLMs in LMICs. Existing debiasing methods such as Translation Chain-of-Thought or Retrieval-augmented generation (RAG) by themselves fall short of closing this performance gap, often failing to improve performance across all languages and lacking specificity for the medical domain. To address this issue, We propose CLARA (Cross-Lingual Reflective Agentic system), a novel inference time de-biasing method leveraging retrieval augmented generation and self-verification. Our approach not only improves performance across all languages but also significantly reduces the multilingual bias gap, facilitating equitable LLM application across the globe.
David S. Restrepo, Chenwei Wu 0006, Zhengxu Tang, Zitao Shuai, Thao Nguyen Minh Phan, Jun-En Ding, Cong-Tinh Dao, Jack Gallifant, Robyn Gayle Dychiao, Jose Carlo Artiaga, André Hiroshi Bando, Carolina Pelegrini Barbosa Gracitelli, Vincenz Ferrer, Leo A. Celi, Danielle S. Bitterman, Michael G. Morley, Luis Filipe Nakayama
AAAI17
2025 Enhancing AI-based diabetic retinopathy screening in low- and middle-income countries with synthetic data
abstract
AI-based DR screening is promising in low- and middle-income countries (LMICs), where limited human resources constrain access to specialist-led programs. However, current systems often degrade under real-world image-quality variations, especially with portable devices that are vital for low- and middle-income countries. This study aims to develop Retsyn, a synthetic-data augmentation framework that improves screening robustness across devices and imaging conditions. RetSyn leverages advanced diffusion models to generate synthetic retinal images with diverse device and imaging quality characteristics. To address the challenges of (1) portable device data scarcity, (2) disease and quality distribution imbalance, and (3) varying image quality, RetSyn uses class and quality-conditioned diffusion for controllable synthesis, a group-balanced loss to increase coverage of minority (quality, disease) pairs, and a Direct Preference Optimization alignment step with a small paired smartphone–tabletop set. The synthesized images are then used to augment classifier training. The effectiveness of RetSyn-generated images was evaluated by training retinal diagnosis models on a combination of real and synthetic data. RetSyn yields consistent gains in-domain and out-of-domain. On low-quality tabletop images, F1 improves from 0.781 to 0.874 (binary) and 0.607 to 0.703 (three-class), while AUROC reaches 0.982 and 0.951, respectively. On out-of-domain portable images, RetSyn attains AUROC 0.813/F1 0.703 (binary) and AUROC 0.804/F1 0.609 (three-class), exceeding group-robustness baselines such as GroupDRO (binary: AUROC 0.786/F1 0.626; three-class: AUROC 0.789/F1 0.544). RetSyn presents an effective and scalable synthetic data framework that significantly enhances the robustness and generalizability of AI-based DR screening models in LMICs. By addressing the critical challenges posed by varying image quality and device characteristics, RetSyn facilitates more reliable deployment of AI diagnostics in underserved regions. Additionally, the release of the first publicly available paired smartphone-tabletop retinal image dataset will support further research into cross-device DR screening solutions. • RetSyn generates synthetic retinal images to enhance AI robustness across varying image qualities and devices. • Group-balanced training and preference optimization enable diverse medical image synthesis with minimal cross-device paired data. • Releases the first publicly available paired smartphone-tabletop retinal dataset. • Performance improved significantly on low-quality and portable device photographs. • Achieves first clinically acceptable results for smartphone-based DR screening in LMICs.
Zitao Shuai, Chenwei Wu 0006, Zhengxu Tang, David S. Restrepo, Michael G. Morley, Luis Filipe Nakayama
J. Biomed. Informatics6
2024 Classification of Keratitis from Eye Corneal Photographs using Deep Learning
abstract
Keratitis is an inflammatory corneal condition responsible for 10% of visual impairment in low- and middle-income countries (LMICs), with bacteria, fungi, or amoeba as the most common infection etiologies. While an accurate and timely diagnosis is crucial for the selected treatment and the patients’ sight outcomes, due to the high cost and limited availability of laboratory diagnostics in LMICs, diagnosis is often made by clinical observation alone, despite its lower accuracy. In this study, we investigate and compare different deep learning approaches to diagnose the source of infection: 1) three separate binary models for infection type predictions; 2) a multitask model with a shared backbone and three parallel classification layers (Multitask V1); and, 3) a multitask model with a shared backbone and a multi-head classification layer (Multitask V2). We used a private Brazilian cornea dataset to conduct the empirical evaluation. We achieved the best results with Multitask V2, with an area under the receiver operating characteristic curve (AUROC) confidence intervals of 0.7413-0.7740 (bacteria), 0.83950.8725 (fungi), and 0.9448-0.9616 (amoeba). A statistical analysis of the impact of patient features on models’ performance revealed that sex significantly affects amoeba infection prediction, and age seems to affect fungi and bacteria predictions.
Maria Miguel Beirão, Tiago Gonçalves 0001, Camila Kase, Luis Filipe Nakayama, Denise de Freitas, Jaime S. Cardoso 0001
BIBM5