EDBT 2026 Demo / reviewers in the wild / expert
Nan Xi
dblp:245/7558
· DBLP profile ↗
12ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0002-7334-7772ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Chain-of-Look Spatial Reasoning for Dense Surgical Instrument CountingabstractAccurate counting of surgical instruments in Operating Rooms (OR) is a critical prerequisite for ensuring patient safety during surgery. Despite recent progress of large visual-language models and agentic AI, accurately counting such instruments remains highly challenging, particularly in dense scenarios where instruments are tightly clustered. To address this problem, we introduce Chain-of-Look, a novel visual reasoning framework that mimics the sequential human counting process by enforcing a structured visual chain, rather than relying on classic object detection which is unordered. This visual chain guides the model to count along a coherent spatial trajectory, improving accuracy in complex scenes. To further enforce the physical plausibility of the visual chain, we introduce the neighboring loss function, which explicitly models the spatial constraints inherent to densely packed surgical instruments. We also present SurgCount-HD, a new dataset comprising 1,464 high-density surgical instrument images. Extensive experiments demonstrate that our method outperforms state-of-the-art approaches for counting (e.g., CountGD, REC) as well as Multimodality Large Language Models (e.g., Qwen, ChatGPT) in the challenging task of dense surgical instrument counting. The code and dataset is available at https://github.com/rishi1134/CoLSR.git Rishikesh Bhyri, Brian R. Quaranto, Junsong Yuan 0001, Peter C. W. Kim, Nan Xi |
WACV | 5 |
| 2026 | HolistAno: Retinal anomaly detection with holistic feature modelingabstractEarly detection of retinal lesions is critical for preventing vision loss. While supervised learning has shown promise, existing methods often rely on extensive labeled data, which is costly and difficult to obtain in medical applications. Unsupervised anomaly detection provides an attractive alternative by requiring only healthy retinal images and no abnormal annotations. However, current methods face significant challenges in modeling the complex structures of normal retinal anatomy, learning discriminative features for detecting subtle lesions, and capturing multi-scale features to handle anomalies of varying sizes – highlighting the need for holistic feature modeling that comprehensively represents both retinal anatomy and pathology. To address these challenges, we propose HolistAno, a novel unsupervised anomaly detection framework with holistic retinal modeling. HolistAno adopts a two-stage network architecture, incorporating a novel anomaly generator and a Balanced Mamba Scale Fusion (BMSF) module to effectively learn comprehensive retinal feature representations. This enables accurate detection of subtle lesions, diverse lesion types, and anomalies across multiple scales. Extensive experiments on five benchmark datasets demonstrate that HolistAno achieves state-of-the-art performance in both anomaly classification and localization tasks, with superior generalization and robustness across multiple datasets and cross-dataset scenarios compared to existing methods. Jingqi Niu, Kang Dang, Nan Xi, Junsong Yuan 0001, Yanjing Liu, Mian Zhou, Jionglong Su |
Expert Syst. Appl. | 3 |
| 2025 | dFLMoE: Decentralized Federated Learning via Mixture of Experts for Medical Data AnalysisabstractFederated learning has wide applications in the medical field. It enables knowledge sharing among different healthcare institutes while protecting patients’ privacy. However, existing federated learning systems are typically centralized, requiring clients to upload client-specific knowledge to a central server for aggregation. This centralized approach would integrate the knowledge from each client into a centralized server, and the knowledge would be already undermined during the centralized integration before it reaches back to each client. Besides, the centralized approach also creates a dependency on the central server, which may affect training stability if the server malfunctions or connections are unstable. To address these issues, we propose a decentralized federated learning framework named dFLMoE. In our framework, clients directly exchange lightweight head models with each other. After exchanging, each client treats both local and received head models as individual experts, and utilizes a client-specific Mixture of Experts (MoE) approach to make collective decisions. This design not only reduces the knowledge damage with client-specific aggregations but also removes the dependency on the central server to enhance the robustness of the framework. We validate our framework on multiple medical tasks, demonstrating that our method evidently outperforms state-of-the-art approaches under both model homogeneity and heterogeneity settings. Luyuan Xie, Tianyu Luan, Wenyuan Cai, Guochen Yan, Nan Xi, Yuejian Fang, Qingni Shen, Zhonghai Wu, Junsong Yuan 0001 |
CVPR | 6 |
| 2025 | PathDiff: Histopathology Image Synthesis with Unpaired Text and Mask Conditions
Mahesh Bhosale, Abdul Wasi, Yuanhao Zhai 0001, Yunjie Tian, Samuel P. Border, Nan Xi, Pinaki Sarder, Junsong Yuan 0001, David S. Doermann |
ICCV | 6 |
| 2025 | PP-Motion: Physical-Perceptual Fidelity Evaluation for Human Motion Generation
Sihan Zhao, Zixuan Wang 0026, Tianyu Luan, Jia Jia 0001, Wentao Zhu 0004, Jiebo Luo 0001, Junsong Yuan 0001, Nan Xi |
ACM Multimedia | 8 |
| 2024 | Interaction-Centric Spatio-Temporal Context Reasoning for Multi-person Video HOI Recognition
Yisong Wang 0005, Nan Xi, Jingjing Meng, Junsong Yuan 0001 |
ECCV (32) | 2 |
| 2023 | Open Set Video HOI detection from Action-centric Chain-of-Look PromptingabstractHuman-Object Interaction (HOI) detection is essential for understanding and modeling real-world events. Existing works on HOI detection mainly focus on static images and a closed setting, where all HOI classes are provided in the training set. In comparison, detecting HOIs in videos in open set scenarios is more challenging. First, under open set circumstances, HOI detectors are expected to hold strong generalizability to recognize unseen HOIs not included in the training data. Second, accurately capturing temporal contextual information from videos is difficult, but it is crucial for detecting temporal-related actions such as open, close, pull, push. To this end, we propose ACoLP, a model of Action-centric Chain-of-Look Prompting for open set video HOI detection. ACoLP regards actions as the carrier of semantics in videos, which captures the essential semantic information across frames. To make the model generalizable on unseen classes, inspired by the chain-of-thought prompting in natural language processing, we introduce the chain-of-look prompting scheme that decomposes prompt generation from large-scale vision-language model into a series of intermediate visual reasoning steps. Consequently, our model captures complex visual reasoning processes underlying the HOI events in videos, providing essential guidance for detecting unseen classes. Extensive experiments on two video HOI datasets, VidHOI and CAD120, demonstrate that ACoLP achieves competitive performance compared with the state-of-the-art methods in the conventional closed setting, and outperforms existing methods by a large margin in the open set setting. Our code is avaliable at https://github.com/southnx/ACoLP. Nan Xi, Jingjing Meng, Junsong Yuan 0001 |
ICCV | 1 |
| 2023 | Source-Free Domain Adaptation for Medical Image Segmentation via Prototype-Anchored Feature Alignment and Contrastive Learning
Qinji Yu, Nan Xi, Junsong Yuan 0001, Kang Dang |
MICCAI (7) | 2 |
| 2023 | Chain-of-Look Prompting for Verb-centric Surgical Triplet Recognition in Endoscopic VideosabstractSurgical triplet recognition aims to recognize surgical activities as triplets (i.e., ), which provides fine-grained information essential for surgical scene understanding. Existing methods for surgical triplet recognition rely on compositional methods that recognize the instrument, verb, and target simultaneously. In contrast, our method, called chain-of-look prompting, casts the problem of surgical triplet recognition as visual prompt generation from large-scale vision-language (VL) models, and explicitly decomposes the task into a series of video reasoning processes. Chain-of-Look prompting is inspired by: (1) the chain-of-thought prompting in natural language processing, which divides a problem into a sequence of intermediate reasoning steps; (2) the inter-dependency between motion and visual appearance in the human vision system. Since surgical activities are conveyed by the actions of physicians, we regard the verbs as the carrier of semantics in surgical endoscopic videos. Additionally, we utilize the BioMed large language model to calibrate the generated visual prompt features for surgical scenarios. Our approach captures the visual reasoning processes underlying surgical activities and achieves better performance compared to the state-of-the-art methods on the largest surgical triplet recognition dataset, CholecT50. The code is available at https://github.com/southnx/CoLSurgical. Nan Xi, Jingjing Meng, Junsong Yuan 0001 |
ACM Multimedia | 1 |
| 2022 | Forest Graph Convolutional Network for Surgical Action Triplet Recognition in Endoscopic VideosabstractRecognizing surgical activities in endoscopic videos is of vital importance for developing context-aware decision support in the operating room. In this work, we model each surgical activity as an action triplet, consisting of the surgical instrument, the action, and the target organ that the instrument is interacting with. The goal is to recognize these action triplets from endoscopic videos. However, correctly recognizing fine-grained activity triplets is challenging because of the long-tail distribution of the triplet classes and the complex associations between triplets as well as within each triplet. In addition, multiple triplets may appear in a given video frame. To address these challenges, we propose a new model for surgical action triplet recognition based on a classification forest and Graph Convolutional Network (GCN), which we call Forest GCN. The classification forest is employed to calibrate fine-grained triplet classifiers by the upstream parent classifiers to suppress noisy logits of the triplet classes in the long tail. And stacked GCNs are designed to model the dependencies between triplet classes while leveraging the language embedding. Experiments on the endoscopic video dataset, CholecT50, demonstrate that our proposed method outperforms current state-of-the-art methods on surgical action triplet recognition. Nan Xi, Jingjing Meng, Junsong Yuan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Learning to Detect Monoclonal Protein in Electrophoresis ImagesabstractMonoclonal protein (M-protein) detection with elec-trophoresis is of vital importance for the diagnosis of lympho-proliferative processes and monoclonal gammopathies (MGs). Although identifying M-proteins are key for the diagnosis and monitoring of these disorders, it requires specialized knowledge and is time consuming and labor intensive. Despite existing powerful machine learning methods, it often requires to obtain large number of labeled data for training, which is difficult to obtain. Besides, electrophoresis image quality could vary dramatically, affecting the proper identification of M-protein. To address these challenges, we propose to represent electrophoresis images using Gaussian Mixture Model (GMM) and leverage peak detection method to identify visual features for M-protein detection. Utilizing random forest classifier, our method can work with a small amount of labeled data to train the model and is not sensitive to samples of varying quality. Furthermore, with extracted image features, it is possible for specially trained technologists and pathologists to understand and check the decision process of the learned model. Extensive experiments indicate our proposed method achieves satisfactory results on test data, demonstrating the effectiveness and robustness of the proposed model for M-protein detection. Sabrina Racine-Brzostek, Nan Xi, Jiwen Luo, Junsong Yuan 0001 |
VCIP | 3 |
| 2020 | Understanding the Political Ideology of Legislators from Social Media Images
Nan Xi, Marcus Liou, Zachary C. Steinert-Threlkeld, Jason Anastasopoulos, Jungseock Joo |
ICWSM | 1 |