EDBT 2026 Demo / reviewers in the wild / expert
Yafei Ou
dblp:244/7904
· DBLP profile ↗
16ranked-venue papers
5as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Prompt-driven multi-instance CLIP: Aligning heterogeneous modalities with missing data tolerance for multi-modal medical analysis
Yafei Ou, Cenyang Zheng, Xun Gong 0002 |
Expert Syst. Appl. | 2 |
| 2026 | Dual-branch decoupling framework for medical one-class classification
Xun Gong 0002, Lingsong Huang, Yafei Ou |
Pattern Recognit. | 4 |
| 2025 | BLS-GAN: A Deep Layer Separation Framework for Eliminating Bone Overlap in Conventional RadiographsabstractConventional radiography is the widely used imaging technology in diagnosing, monitoring, and prognosticating musculoskeletal (MSK) diseases because of its easy availability, versatility, and cost-effectiveness. Bone overlaps are prevalent in conventional radiographs, and can impede the accurate assessment of bone characteristics by radiologists or algorithms, posing significant challenges to conventional clinical diagnosis and computer-aided diagnosis. This work initiated the study of a challenging scenario - bone layer separation in conventional radiographs, in which separate overlapped bone regions enable the independent assessment of the bone characteristics of each bone layer and lay the groundwork for MSK disease diagnosis and its automation. This work proposed a Bone Layer Separation GAN (BLS-GAN) framework that can produce high-quality bone layer images with reasonable bone characteristics and texture. This framework introduced a reconstructor based on conventional radiography imaging principles, which achieved efficient reconstruction and mitigates the recurrent calculations and training instability issues caused by soft tissue in the overlapped regions. Additionally, pre-training with synthetic images was implemented to enhance the stability of both the training process and the results. The generated images passed the visual Turing test, and improved performance in downstream tasks. This work affirms the feasibility of extracting bone layer images from conventional radiographs, which holds promise for leveraging layer separation technology to facilitate more comprehensive analytical research in MSK diagnosis, monitoring, and prognosis. Haolin Wang 0007, Yafei Ou, Prasoon Ambalathankandy, Gen Ota, Pengyu Dai, Masayuki Ikebe, Kenji Suzuki 0001, Tamotsu Kamishima |
AAAI | 2 |
| 2025 | Bi-Directional Mutual Supervision for Adversarial Domain Adaptation in Medical ImagingabstractAdversarial domain adaptation is widely adopted for knowledge transfer in unsupervised medical imaging scenarios. However, existing methods often suffer from rigid alignment strategies, limiting their efficacy in complex clinical dataset shifts caused by heterogeneous scanners. To address this, we propose a bi-directional mutually-supervised (BDMS) framework to find an optimal intermediate distribution that combines the prior distributions of both domains. First, feature extractors are established for both source and target domains to fit their respective posterior distributions. Then, through mutual supervision, a multiple bidirectional state is reached. In order to ensure improved classification performance and the acquisition of distinct information in each domain, we propose the utilization of mutually learning classifiers. These classifiers are designed to stabilize dynamic changes during training and expedite convergence. Additionally, we introduce a novel conditional consistency loss to regulate the conditional distributions and further enhance the overall performance. Extensive experiments on multi-center medical datasets (breast ultrasound and endoscopic ultrasound) and natural image benchmarks (Office-31) demonstrate that BDMS achieves state-of-the-art performance, validates the universality and robustness of our approach. Xun Gong 0001, Yafei Ou |
BIBM | 4 |
| 2025 | AP-DPM: A Dual-Path Merging Network Via Adversarial Anatomical Prior Guidance for Wrist Bone SegmentationabstractAccurate segmentation of wrist bones from conventional radiographs remains a significant challenge due to severe anatomical overlap and blurred bone boundaries, particularly in patients with rheumatoid arthritis. To address these issues, we propose AP-DPM, a novel Dual-Path Prior-constrained Merging model with adversarial anatomical priors. AP-DPM employs a dual-path architecture to separately predict complete bone masks and overlapping regions, which are subsequently integrated through a residual merging network. To enhance anatomical plausibility, we incorporate an adversarial prior guided by a pre-trained discriminator. Extensive experiments on the publicly available RAM-W600 dataset demonstrate that AP-DPM outperforms state-of-the-art methods across seven quantitative metrics, achieving superior performance particularly in diagnostically critical overlapping regions. Ablation studies further validate that both the dualpath structure for enhanced focusing on overlap regions and the adversarial anatomical prior contribute significantly to performance gains, enhancing local boundary sensitivity and global structural consistency. These results highlight the potential of AP-DPM to improve automated radiographic assessment of rheumatoid arthritis progression. Code is available at https://github.com/YSongxiao/AP-DPM Songxiao Yang, Haolin Wang 0007, Masayuki Ikebe, Tamotsu Kamishima, Yafei Ou, Masatoshi Okutomi |
BIBM | 5 |
| 2025 | CRESSim-MPM: A Material Point Method Library for Surgical Soft Body Simulation with Cutting and SuturingabstractA number of recent studies have focused on developing surgical simulation platforms to train machine learning (ML) agents or models with synthetic data for surgical assistance. While existing platforms excel at tasks such as rigid body manipulation and soft body deformation, they struggle to simulate more complex soft body behaviors like cutting and suturing. A key challenge lies in modeling soft body fracture and splitting using the finite-element method (FEM), which is the predominant approach in current platforms. Additionally, the two-way suture needle/thread contact inside a soft body is further complicated when using FEM. In this work, we use the material point method (MPM) for such challenging simulations and propose new rigid geometries and soft-rigid contact methods specifically designed for them. We introduce CRESSim-MPM, a GPU-accelerated MPM library that integrates multiple MPM solvers and incorporates surgical geometries for cutting and suturing, serving as a specialized physics engine for surgical applications. It is further integrated into Unity, requiring minimal modifications to existing projects for soft body simulation. We demonstrate the simulator’s capabilities in real-time simulation of cutting and suturing on soft tissue and provide an initial performance evaluation of different MPM solvers when simulating varying numbers of particles. The source code is available at https://github.com/yafei-ou/CRESSim-MPM. Yafei Ou, Mahdi Tavakoli |
IROS | 1 |
| 2025 | GoCa: Trustworthy Multi-modal RAG with Explicit Thinking Distillation for Reliable Decision-Making in Med-LVLMs
Pengyu Dai, Yafei Ou, Yuqiao Yang, Ze Jin, Kenji Suzuki 0001 |
MICCAI (14) | 2 |
| 2025 | Layer Separation: Towards Adjustable Joint Space Width Images SynthesisabstractRheumatoid arthritis (RA) is a chronic autoimmune disease characterized by joint inflammation and progressive structural damage. Joint space width (JSW) is a critical indicator in conventional radiography (CR) for evaluating disease progression, which has become a prominent research topic in computer-aided diagnostic (CAD) systems. However, deep learning-based radiological CAD systems for JSW analysis face significant challenges in data quality, including data imbalance, limited variety, and annotation difficulties. This work introduced a challenging image synthesis scenario and proposed Layer Separation Networks (LSN) to accurately separate the soft tissue layer, the upper bone layer, and the lower bone layer in conventional radiographs of finger joints. Using these layers, the adjustable JSW images can be synthesized to address data quality challenges and achieve ground truth (GT) generation. Experimental results demonstrated that LSN-based synthetic images closely resemble real radiographs, and significantly enhanced the performance in downstream tasks. The code and dataset are available at: https://github.com/pokeblow/LSN. Haolin Wang 0007, Yafei Ou, Prasoon Ambalathankandy, Gen Ota, Pengyu Dai, Masayuki Ikebe, Kenji Suzuki 0001, Tamotsu Kamishima |
ACM Multimedia | 2 |
| 2025 | RAM-W600: A Multi-Task Wrist Dataset and Benchmark for Rheumatoid ArthritisabstractRheumatoid arthritis (RA) is a common autoimmune disease that has been the focus of research in computer-aided diagnosis (CAD) and disease monitoring. In clinical settings, conventional radiography (CR) is widely used for the screening and evaluation of RA due to its low cost and accessibility. The wrist is a critical region for the diagnosis of RA. However, CAD research in this area remains limited, primarily due to the challenges in acquiring high-quality instance-level annotations. (i) The wrist comprises numerous small bones with narrow joint spaces, complex structures, and frequent overlaps, requiring detailed anatomical knowledge for accurate annotation. (ii) Disease progression in RA often leads to osteophyte, bone erosion (BE), and even bony ankylosis, which alter bone morphology and increase annotation difficulty, necessitating expertise in rheumatology.This work presents a multi-task dataset for wrist bone in CR, including two tasks: (i) wrist bone instance segmentation and (ii) Sharp/van der Heijde (SvdH) BE scoring, which is the first public resource for wrist bone instance segmentation. This dataset comprises 1048 wrist conventional radiographs of 388 patients from six medical centers, with pixel-level instance segmentation annotations for 618 images and SvdH BE scores for 800 images. This dataset can potentially support a wide range of research tasks related to RA, including joint space narrowing (JSN) progression quantification, BE detection, bone deformity evaluation, and osteophyte detection. It may also be applied to other wrist-related tasks, such as carpal bone fracture localization.We hope this dataset will significantly lower the barrier to research on wrist RA and accelerate progress in CAD research within the RA-related domain.Benchmark & Code: https://github.com/YSongxiao/RAM-W600Data & Dataset Card: https://huggingface.co/datasets/TokyoTechMagicYang/RAM-W600 Songxiao Yang, Haolin Wang 0007, Tamotsu Kamishima, Masayuki Ikebe, Yafei Ou, Masatoshi Okutomi |
NeurIPS | 7 |
| 2025 | Cycle-VQA: A Cycle-Consistent Framework for Robust Medical Visual Question Answering
Xun Gong 0002, Cenyang Zheng, Xuli Tan, Yafei Ou |
Pattern Recognit. | 6 |
| 2025 | Learning Autonomous Surgical Irrigation and Suction With the da Vinci Research Kit Using Reinforcement LearningabstractThe irrigation-suction process is a common procedure to rinse and clean up the surgical field in minimally invasive surgery (MIS). In this process, surgeons first irrigate liquid, typically saline, into the surgical scene for rinsing and diluting the contaminant, and then suction the liquid out of the surgical field. While recent advances have shown promising results in the application of reinforcement learning (RL) for automating surgical subtasks, fewer studies have explored the automation of fluid-related tasks. In this work, we explore the automation of both steps in the irrigation-suction procedure and train two vision-based RL agents to complete irrigation and suction autonomously. To achieve this, a platform is developed for creating simulated surgical robot learning environments and for training agents, and two simulated learning environments are built for irrigation and suction with visually plausible fluid rendering capabilities. With techniques such as domain randomization (DR) and imitation learning, two agents are trained in the simulator and transferred to the real world. Individual evaluations of both agents show satisfactory real-world results. With an initial amount of around 5 grams of contaminants, the irrigation agent ultimately achieved an average of 2.21 grams remaining after a manual suction. As a comparison, fully manual operation by a human results in 1.90 grams remaining. The suction agent achieved 2.64 and 2.24 grams of liquid remaining across two trial groups with more than 20 and 30 grams of initial liquid in the container. Fully autonomous irrigation-suction trials reduce the contaminant in the container from around 5 grams to an average of 2.42 grams, although yielding a higher total weight remaining (4.40) due to residual liquid not suctioned. Further information about the project is available at https://tbs-ualberta.github.io/CRESSim/. Yafei Ou, Mahdi Tavakoli |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2024 | Tri-VQA: Triangular Reasoning Medical Visual Question Answering for Multi-Attribute AnalysisabstractMed-VQA, intersecting medicine and VQA, is a challenging field offering benefits like patient engagement and clinical collaboration. However, current joint embedding methods lack explanation of reasoning, undermining VQA credibility. In this paper, motivated by causal effect, we propose a novel Triangular Reasoning VQA (Tri-VQA) framework, which constructs reverse causal questions from the perspective of "Why this answer?" to elucidate the source of the answer and stimulate more reasonable forward reasoning processes. We evaluate our method on EUS multi-attribute annotated datasets from five centers and medical VQA datasets, demonstrating superior performance. Our codes and pre-trained models are available at https://github.com/hahaha111111/Tri-VQA. Xun Gong 0002, Cenyang Zheng, Yafei Ou |
BIBM | 4 |
| 2024 | SaSaMIM: Synthetic Anatomical Semantics-Aware Masked Image Modeling for Colon Tumor Segmentation in Non-contrast Abdominal Computed Tomography
Pengyu Dai, Yafei Ou, Yuqiao Yang, Dichao Liu, Masahiro Hashimoto, Masahiro Jinzaki, Mototaka Miyake, Kenji Suzuki 0001 |
MICCAI (11) | 2 |
| 2024 | Corrections to "A Sub-Pixel Accurate Quantification of Joint Space Narrowing Progression in Rheumatoid Arthritis"abstractPresents corrections to the article "A Sub-Pixel Accurate Quantification of Joint Space Narrowing Progression in Rheumatoid Arthritis". Yafei Ou, Prasoon Ambalathankandy, Ryunosuke Furuya, Seiya Kawada, Tianyu Zeng, Yujie An, Tamotsu Kamishima, Kenichi Tamura, Masayuki Ikebe |
IEEE J. Biomed. Health Informatics | 1 |
| 2023 | A Sub-Pixel Accurate Quantification of Joint Space Narrowing Progression in Rheumatoid ArthritisabstractRheumatoid arthritis (RA) is a chronic autoimmune disease that primarily affects peripheral synovial joints, like fingers, wrists and feet. Radiology plays a critical role in the diagnosis and monitoring of RA. Limited by the current spatial resolution of radiographic imaging, joint space narrowing (JSN) progression of RA for the same reason above can be less than one pixel per year with universal spatial resolution. Insensitive monitoring of JSN can hinder the radiologist/rheumatologist from making a proper and timely clinical judgment. In this paper, we propose a novel and sensitive method that we call partial image phase-only correlation which aims to automatically quantify JSN progression in the early RA. The majority of the current literature utilizes the mean error, root-mean-square deviation and standard deviation to report the accuracy at pixel level. Our work measures JSN progression between a baseline and its follow-up finger joint images by using the phase spectrum in the frequency domain. Using this study, the mean error can be reduced to 0.0130 mm when applied to phantom radiographs with ground truth, and 0.0519 mm standard deviation for clinical radiography. With the sub-pixel accuracy far beyond usual manual measurements, we are optimistic that the proposed work is a promising scheme for automatically quantifying JSN progression. Yafei Ou, Prasoon Ambalathankandy, Ryunosuke Furuya, Seiya Kawada, Tianyu Zeng, Yujie An, Tamotsu Kamishima, Kenichi Tamura, Masayuki Ikebe |
IEEE J. Biomed. Health Informatics | 1 |
| 2022 | Real-Time Tone Mapping: A Survey and Cross-Implementation Hardware BenchmarkabstractThe rising demand for high quality display has ensued active research in high dynamic range (HDR) imaging, which has the potential to replace the standard dynamic range imaging. This is due to HDR’s features like accurate reproducibility of a scene with its entire spectrum of visible lighting and color depth. But this capability comes with expensive capture, display, storage and distribution resource requirements. Also, display of HDR images/video content on an ordinary display device with limited dynamic range requires some form of adaptation. Many adaptation algorithms, widely known as tone mapping (TM) operators, have been studied and proposed in the last few decades. In this article, we present a comprehensive survey of 60 TM algorithms that have been implemented on hardware for acceleration and real-time performance. In this state-of-the-art survey, we will discuss those TM algorithms which have been implemented on GPU, FPGA, and ASIC in terms of their hardware specifications and performance. Output image quality is an important metric for TM algorithms. From our literature survey we found that, various objective quality metrics have been used to demonstrate the quality of those algorithms hardware implementation. We have compiled those metrics used in this survey, and analyzed the relationship between hardware cost, image quality and computational efficiency. Currently, machine learning-based (ML) algorithms have become an important tool to solve many image processing tasks, and this article concludes with a discussion on the future research directions to realize ML-based TM operators on hardware. Yafei Ou, Prasoon Ambalathankandy, Shinya Takamaeda-Yamazaki, Masato Motomura, Tetsuya Asai, Masayuki Ikebe |
IEEE Trans. Circuits Syst. Video Technol. | 1 |