Jinlin Wu

dblp:123/7200 · DBLP profile ↗
← Back
36ranked-venue papers
6as first author
32since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 22 · 3 first-author · 19 since 2021Artificial intelligence and machine learning · 20 · 5 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 10 since 2021Systems, architecture and hardware · 3 · 3 since 2021Computer networks · 2 · 2 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 6DAttack: Backdoor Attacks in the 6DoF Pose Estimation
abstract
Recent advances in deep learning have enabled highly accurate six-degree-of-freedom (6DoF) object pose estimation, leading to its widespread use in real-world applications such as robotics, augmented reality, virtual reality, and autonomous systems. However, backdoor attacks pose a major security risk to deep learning models. By injecting malicious triggers into training data, an attacker can cause a model to perform normally on benign inputs but behave incorrectly under specific conditions. While most research on backdoor attacks has focused on 2D vision tasks, their impact on 6DoF pose estimation remains largely unexplored. Furthermore, unlike traditional backdoors that only change the object class, backdoors against 6DoF pose estimation must additionally control continuous pose parameters, such as translation and rotation, making existing 2D backdoor attack methods not directly applicable to this setting. To address this gap, we propose a novel backdoor attack framework (6DAttack) that exposes vulnerabilities in 6DoF pose estimation. 6DAttack uses synthetic and real 3D objects of varying shapes as triggers and assigns target poses to induce controlled erroneous pose outputs while maintaining normal behavior on clean inputs. We evaluated this attack on multiple models (including PVNet, DenseFusion, and PoseDiffusion) and datasets (including LINEMOD, YCB-Video, and CO3D). Experimental results demonstrate that 6DAttack achieves extremely high attack success rates (ASRs) without compromising performance on legitimate tasks. Across various models and objects, the backdoored models achieve up to 100% ADD accuracy on clean data, while also achieving 100% ASR under trigger conditions. The accuracy of controlled erroneous pose output is also extremely high, with triggered samples achieving 97.70% ADD-P. These results demonstrate that the backdoor can be reliably implanted and activated, achieving a high ASR under trigger conditions while maintaining a negligible impact on benign data. Furthermore, we evaluate a representative defense and show that it remains ineffective under 6DAttack. Overall, our findings reveal a potentially serious and previously underexplored threat to modern 6DoF pose estimation models.
Jihui Guo, Zongmin Zhang, Zhen Sun 0001, Jinlin Wu, Xinlei He 0001
AAAI5
2026 MedLA: A Logic-Driven Multi-Agent Framework for Complex Medical Reasoning with Large Language Models
abstract
Answering complex medical questions requires not only domain expertise and patient-specific information, but also structured and multi-perspective reasoning. Existing multi-agent approaches often rely on fixed roles or shallow interaction prompts, limiting their ability to detect and resolve fine-grained logical inconsistencies. To address this, we propose MedLA, a logic-driven multi-agent framework built on large language models. Each agent organizes its reasoning process into an explicit logical tree based on syllogistic triads (major premise, minor premise, and conclusion), enabling transparent inference and premise-level alignment. Agents engage in a multi-round, graph-guided discussion to compare and iteratively refine their logic trees, achieving consensus through error correction and contradiction resolution. We demonstrate that MedLA consistently outperforms both static role-based systems and single-agent baselines on challenging benchmarks such as MedDDx and standard medical QA tasks. Furthermore, MedLA scales effectively across both open-source and commercial LLM backbones, achieving state-of-the-art performance and offering a generalizable paradigm for trustworthy medical reasoning.
Fan Zhang 0010, Jinlin Wu, Guohui Fan, Zelin Zang
AAAI4
2026 EndoChat: Grounded multimodal large language model for endoscopic surgery
abstract
Recently, Multimodal Large Language Models (MLLMs) have demonstrated their immense potential in computer-aided diagnosis and decision-making. In the context of robotic-assisted surgery, MLLMs can serve as effective tools for surgical training and guidance. However, there is still a deficiency of MLLMs specialized for surgical scene understanding in endoscopic procedures. To this end, we present EndoChat, an MLLM tailored to address various dialogue paradigms and subtasks in understanding endoscopic procedures. To train our EndoChat, we construct the Surg-396K dataset through a novel pipeline that systematically extracts surgical information and generates structured annotations based on large-scale endoscopic surgery datasets. Furthermore, we introduce a multi-scale visual token interaction mechanism and a visual contrast-based reasoning mechanism to enhance the model's representation learning and reasoning capabilities. Our model achieves state-of-the-art performance across five dialogue paradigms and seven surgical scene understanding tasks. Additionally, we conduct evaluations with professional surgeons, who provide positive feedback on the majority of conversation cases generated by EndoChat. Overall, these results demonstrate that EndoChat has the potential to advance training and automation in robotic-assisted surgery. Our dataset and model are publicly available at https://github.com/gkw0010/EndoChat.
Guankun Wang, Long Bai 0008, Kun Yuan 0004, Zhen Li 0026, Tianxu Jiang, Xiting He, Jinlin Wu, Zhen Chen 0018, Zhen Lei 0001, Hongbin Liu 0001, Fan Zhang 0016, Nicolas Padoy, Nassir Navab, Hongliang Ren 0001
Medical Image Anal.8
2026 SA-Person: Text-Based Person Retrieval With Scene-Aware Re-Ranking
Yingjia Xu, Jinlin Wu, Daming Gao, Zhen Chen 0018, Yang Yang 0062, Min Cao 0005, Mang Ye, Zhen Lei 0001
IEEE Trans. Inf. Forensics Secur.2
2026 Procedure-Aware Hierarchical Alignment for Open Surgery Video-Language Pretraining
abstract
Recent advances in surgical robotics and computer vision have greatly improved intelligent systems' autonomy and perception in the operating room (OR), especially in endoscopic and minimally invasive surgeries. However, for open surgery, which is still the predominant form of surgical intervention worldwide, there has been relatively limited exploration due to its inherent complexity and the lack of large-scale, diverse datasets. To close this gap, we present OpenSurgery, by far the largest video-text pretraining and evaluation dataset for open surgery understanding. OpenSurgery consists of two subsets: OpenSurgery-Pretrain and OpenSurgery-EVAL. OpenSurgery-Pretrain consists of 843 publicly available open surgery videos for pretraining, spanning 102 hours and encompassing over 20 distinct surgical types. OpenSurgery-EVAL is a benchmark dataset for evaluating model performance in open surgery understanding, comprising 280 training and 120 test videos, totaling 49 hours. Each video in OpenSurgery is meticulously annotated by expert surgeons at three hierarchical levels of video, operation, and frame to ensure both high quality and strong clinical applicability. Next, we propose the Hierarchical Surgical Knowledge Pretraining (HierSKP) framework to facilitate large-scale multimodal representation learning for open surgery understanding. HierSKP leverages a granularity-aware contrastive learning strategy and enhances procedural comprehension by constructing hard negative samples and incorporating a Dynamic Time Warping (DTW)-based loss to capture fine-grained temporal alignment of visual semantics. Extensive experiments show that HierSKP achieves state-of-the-art performance on OpenSurgegy-EVAL across multiple tasks, including operation recognition, temporal action localization, and zero-shot cross-modal retrieval. This demonstrates its strong generalizability for further advances in open surgery understanding.
Boqiang Xu, Jinlin Wu, Jian Liang 0001, Zhenan Sun, Hongbin Liu 0001, Jiebo Luo 0001, Zhen Lei 0001
IEEE Trans. Image Process.2
2026 Multi-View Images Suffice 3D Reasoning Through Chain-of-Thought Selection and Question-Guided Fusion
abstract
3D reasoning is crucial in areas like robotics and autonomous driving. Due to the high cost of 3D data acquisition, some recent methods attempt to enable LLMs to perform 3D reasoning through multi-view images, thereby transferring the powerful 2D reasoning capabilities of LLMs to 3D environments. However, these methods face challenges: either they use redundant views that contain many perspectives irrelevant to the question, or they rely on globally aggregated multi-view representations, losing the fine-grained vision-language correlations. To tackle these challenges, we propose 3DMulti-LLM, which mainly consists of three components: a COT selector, a question-guided fusion block, and pre-trained LLMs. Specifically, first, the COT selector leverages the powerful chain-of-thought reasoning capabilities of LLMs to identify question-related multi-view images. In this way, 3DMulti-LLM can eliminate a substantial amount of interference from unnecessary viewpoints. Then, we propose a question-guided fusion block for integrating multi-view features via question-guided interaction among various viewpoints. Finally, the pre-trained LLMs are utilized to reason in 3D scenes directly through multi-view features. Notably, our approach understands the 3D scene solely through multi-view images, without requiring the input of point cloud information or additional 3D feature extraction. Through our experiments, 3DMulti-LLM achieves impressive performance and surpasses existing 3D-input-free methods by + 12.2% and + 7.1% on ScanQA and 3DMV-VQA datasets, respectively.
Boqiang Xu, Jinlin Wu, Wei Zhang 0255, Chenyang Su, Jian Liang 0001, Zhenan Sun, Zhen Lei 0001
IEEE Trans. Image Process.2
2025 Universal Domain Adaptive Object Detection via Dual Probabilistic Alignment
abstract
Domain Adaptive Object Detection (DAOD) transfers knowledge from a labeled source domain to an unannotated target domain under closed-set assumption. Universal DAOD (UniDAOD) extends DAOD to handle open-set, partial-set, and closed-set domain adaptation. In this paper, we first unveil two issues: domain-private category alignment is crucial for global-level features, and the domain probability heterogeneity of features across different levels. To address these issues, we propose a novel Dual Probabilistic Alignment (DPA) framework to model domain probability as Gaussian distribution, enabling the heterogeneity domain distribution sampling and measurement. The DPA consists of three tailored modules: the Global-level Domain Private Alignment (GDPA), the Instance-level Domain Shared Alignment (IDSA), and the Private Class Constraint (PCC). GDPA utilizes the global-level sampling to mine domain-private category samples and calculate alignment weight through a cumulative distribution function to address the global-level private category alignment. IDSA utilizes instance-level sampling to mine domain-shared category samples and calculates alignment weight through Gaussian distribution to conduct the domain-shared category domain alignment to address the feature heterogeneity. The PCC aggregates domain-private category centroids between feature and probability spaces to mitigate negative transfer. Extensive experiments demonstrate that our DPA outperforms state-of-the-art UniDAOD and DAOD methods across various datasets and scenarios, including open, partial, and closed sets.
Yuanfan Zheng, Jinlin Wu, Wuyang Li, Zhen Chen 0018
AAAI2
2025 StreamWMR: A Streaming Framework for Real-time 3D Whole-body Mesh Recovery
abstract
3D whole-body mesh recovery aims to extract parameters for the human body, hands, and head from a single human image. Most applications related to human mesh recovery, such as physical fitness motion capture and operating room motion capture, necessitate real-time video stream processing. However, existing methods ignore the video processing and often require significant computational resources, making real-time performance unattainable and greatly limiting their practicality. Moreover, noticeable misalignments are often observed when concatenating them back to the body and reprojecting them onto the image. In this paper, we propose a streaming framework for whole-body mesh recovery in the video. First, we simplify pose regression by leveraging the root nodes of the hands and head to locate each component. Second, for temporal optimization, we incorporate attention mechanisms related to keypoint velocity to incorporate information from previous frames and achieve more stable and smooth motions. Finally, we propose a multi-view projection loss to eliminate the ambiguity caused by inaccurate 3D regression and pose estimation in computing reprojection errors. The combination of these enables our method to achieve real-time inference speed while maintaining accuracy and stability.
Xiangyu Zhu 0001, Jinlin Wu, Zidu Wang, Shukai Chen, Dong Yi, Zhen Lei 0001
IJCB3
2025 PRO-VPT: Distribution-Adaptive Visual Prompt Tuning via Prompt Relocation
abstract
Visual prompt tuning (VPT), i.e., fine-tuning some lightweight prompt tokens, provides an efficient and effective approach for adapting pre-trained models to various downstream tasks. However, most prior art indiscriminately uses a fixed prompt distribution across different tasks, neglecting the importance of each block varying depending on the task. In this paper, we introduce adaptive distribution optimization (ADO) by tackling two key questions: (1) How to appropriately and formally define ADO, and (2) How to design an adaptive distribution strategy guided by this definition? Through empirical analysis, we first confirm that properly adjusting the distribution significantly improves VPT performance, and further uncover a key insight that a nested relationship exists between ADO and VPT. Based on these findings, we propose a new VPT framework, termed PRO-VPT (iterative Prompt RelOcation-based VPT), which adaptively adjusts the distribution built upon a nested optimization formulation. Specifically, we develop a prompt relocation strategy derived from this formulation, comprising two steps: pruning idle prompts from prompt-saturated blocks, followed by allocating these prompts to the most prompt-needed blocks. By iteratively performing prompt relocation and VPT, our proposal can adaptively learn the optimal prompt distribution in a nested optimization-based manner, thereby unlocking the full potential of VPT. Extensive experiments demonstrate that our proposal significantly outperforms advanced VPT methods, e.g., PRO-VPT surpasses VPT by 1.6 pp and 2.0 pp average accuracy, leading prompt-based methods to state-of-the-art performance on VTAB-1k and FGVC benchmarks. The code is available at https://github.com/ckshang/PRO-VPT.
Chikai Shang, Mengke Li 0001, Yiqun Zhang 0006, Zhen Chen 0018, Jinlin Wu, Fangqing Gu, Yang Lu 0009, Yiu-Ming Cheung
ICCV5
2025 SurgPLAN++: Universal Surgical Phase Localization Network for Online and Offline Inference
abstract
Surgical phase recognition is critical for assisting surgeons in understanding surgical videos. Existing studies focused more on online surgical phase recognition, by leveraging preceding frames to predict the current frame. Despite great progress, they formulated the task as a series of frame-wise classification, which resulted in a lack of global context of the entire procedure and incoherent predictions. Moreover, besides online analysis, accurate offline surgical phase recognition is also in significant clinical need for retrospective analysis, and existing online algorithms do not fully analyze the entire video, thereby limiting accuracy in offline analysis. To over-come these challenges and enhance both online and offline inference capabilities, we propose a universal Surgical Phase LocalizAtion Network, named SurgPLAN++, with the principle of temporal detection. To ensure a global understanding of the surgical procedure, we devise a phase localization strategy for SurgPLAN ++ to predict phase segments across the entire video through phase proposals. For online analysis, to generate high-quality phase proposals, SurgPLAN++ incorporates a data augmentation strategy to extend the streaming video into a pseudo-complete video through mirroring, center-duplication, and down-sampling. For offline analysis, SurgPLAN++ capi-talizes on its global phase prediction framework to continu-ously refine preceding predictions during each online inference step, thereby significantly improving the accuracy of phase recognition. We perform extensive experiments to validate the effectiveness, and our SurgPLAN++ achieves remarkable performance in both online and offline modes, which outper-forms state-of-the-art methods. The source code is available at https://github.com/franciszchenlSurgPLAN-Plus.
Zhen Chen 0018, Xingjian Luo, Jinlin Wu, Long Bai 0008, Zhen Lei 0001, Hongliang Ren 0001, Sébastien Ourselin, Hongbin Liu 0001
ICRA3
2025 Advancing Dense Endoscopic Reconstruction with Gaussian Splatting-Driven Surface Normal-Aware Tracking and Mapping
abstract
Simultaneous Localization and Mapping (SLAM) is essential for precise surgical interventions and robotic tasks in minimally invasive procedures. While recent advancements in 3D Gaussian Splatting (3DGS) have improved SLAM with high-quality novel view synthesis and fast rendering, these systems struggle with accurate depth and surface reconstruction due to multi-view inconsistencies. Simply incorporating SLAM and 3DGS leads to mismatches between the reconstructed frames. In this work, we present Endo-2DTAM, a real-time endoscopic SLAM system with 2D Gaussian Splatting (2DGS) to address these challenges. Endo-2DTAM incorporates a surface normal-aware pipeline, which consists of tracking, mapping, and bundle adjustment modules for geometrically accurate reconstruction. Our robust tracking module combines point-topoint and point-to-plane distance metrics, while the mapping module utilizes normal consistency and depth distortion to enhance surface reconstruction quality. We also introduce a pose-consistent strategy for efficient and geometrically coherent keyframe sampling. Extensive experiments on public endoscopic datasets demonstrate that Endo-2DTAM achieves an RMSE of$1.87 \pm 0.63 \mathbf{m m}$for depth reconstruction of surgical scenes while maintaining computationally efficient tracking, high-quality visual appearance, and real-time rendering. Our code will be released at github.com/lastbasket/Endo-2DTAM.
Yiming Huang 0007, Beilei Cui, Long Bai 0008, Zhen Chen 0018, Jinlin Wu, Zhen Li 0026, Hongbin Liu 0001, Hongliang Ren 0001
ICRA5
2025 Endo-4DGX: Robust Endoscopic Scene Reconstruction and Illumination Correction with Gaussian Splatting
Yiming Huang 0007, Long Bai 0008, Beilei Cui, Yanheng Li 0002, Tong Chen 0011, Jie Wang 0097, Jinlin Wu, Zhen Lei 0001, Hongbin Liu 0001, Hongliang Ren 0001
MICCAI (9)7
2025 Reconstructing 3D Hand-Instrument Interaction from a Single 2D Image in Medical Scenes
Xiangyu Zhu 0001, Jinlin Wu, Ming Feng, Zelin Zang, Hongbin Liu 0001, Zhen Lei 0001
MICCAI (10)3
2025 PitVis-2023 challenge: Workflow recognition in videos of endoscopic pituitary surgery
abstract
The field of computer vision applied to videos of minimally invasive surgery is ever-growing. Workflow recognition pertains to the automated recognition of various aspects of a surgery, including: which surgical steps are performed; and which surgical instruments are used. This information can later be used to assist clinicians when learning the surgery or during live surgery. The Pituitary Vision (PitVis) 2023 Challenge tasks the community to step and instrument recognition in videos of endoscopic pituitary surgery. This is a particularly challenging task when compared to other minimally invasive surgeries due to: the smaller working space, which limits and distorts vision; and higher frequency of instrument and step switching, which requires more precise model predictions. Participants were provided with 25-videos, with results presented at the MICCAI-2023 conference as part of the Endoscopic Vision 2023 Challenge in Vancouver, Canada, on 08-Oct-2023. There were 18-submissions from 9-teams across 6-countries, using a variety of deep learning models. The top performing model for step recognition utilised a transformer based architecture, uniquely using an autoregressive decoder with a positional encoding input. The top performing model for instrument recognition utilised a spatial encoder followed by a temporal encoder, which uniquely used a 2-layer temporal architecture. In both cases, these models outperformed purely spatial based models, illustrating the importance of sequential and temporal information. This PitVis-2023 therefore demonstrates state-of-the-art computer vision models in minimally invasive surgery are transferable to a new dataset. Benchmark results are provided in the paper, and the dataset is publicly available at: https://doi.org/10.5522/04/26531686.
Adrito Das, Danyal Z. Khan, Dimitris Psychogyios, John G. Hanrahan, Francisco Vasconcelos 0001, You Pang, Zhen Chen 0018, Jinlin Wu, Xiaoyang Zou, Guoyan Zheng, Abdul Qayyum 0002, Moona Mazher, Muhammad Imran Razzak, Tianbin Li, Jin Ye 0002, Junjun He, Szymon Plotka, Joanna Kaleta, Amine Yamlahi, Antoine Jund, Patrick Godau, Satoshi Kondo, Satoshi Kasai, Kousuke Hirasawa, Dominik Rivoir, Stefanie Speidel, Alejandra Pérez, Santiago Rodríguez, Pablo Andrés Arbeláez, Danail Stoyanov, Hani J. Marcus, Sophia Bano
Medical Image Anal.9
2025 BronchoTrack: Airway Lumen Tracking for Branch-Level Bronchoscopic Localization
abstract
Localizing the bronchoscope in real time is essential for ensuring intervention quality. However, most existing vision-based methods struggle to balance between speed and generalization. To address these challenges, we present BronchoTrack, an innovative real-time framework for accurate branch-level localization, encompassing lumen detection, tracking, and airway association. To achieve real-time performance, we employ benchmark light weight detector for efficient lumen detection. We firstly introduce multi-object tracking to bronchoscopic localization, mitigating temporal confusion in lumen identification caused by rapid bronchoscope movement and complex airway structures. To ensure generalization across patient cases, we propose a training-free detection-airway association method based on a semantic airway graph that encodes the hierarchy of bronchial tree structures. Experiments on 11 patient datasets demonstrate BronchoTrack's localization accuracy of 81.72%, while accessing up to the 6th generation of airways. Furthermore, we tested BronchoTrack in an in-vivo animal study using a porcine model, where it localized the bronchoscope into the 8th generation airway successfully. Experimental evaluation underscores BronchoTrack's real-time performance in both satisfying accuracy and generalization, demonstrating its potential for clinical applications.
Qingyao Tian, Huai Liao, Bingyu Yang, Jinlin Wu, Jian Chen 0036, Lujie Li, Hongbin Liu 0001
IEEE Trans. Medical Imaging5
2024 Compositional Inversion for Stable Diffusion Models
abstract
Inversion methods, such as Textual Inversion, generate personalized images by incorporating concepts of interest provided by user images. However, existing methods often suffer from overfitting issues, where the dominant presence of inverted concepts leads to the absence of other desired concepts. It stems from the fact that during inversion, the irrelevant semantics in the user images are also encoded, forcing the inverted concepts to occupy locations far from the core distribution in the embedding space. To address this issue, we propose a method that guides the inversion process towards the core distribution for compositional embeddings. Additionally, we introduce a spatial regularization approach to balance the attention on the concepts being composed. Our method is designed as a post-training approach and can be seamlessly integrated with other inversion methods. Experimental results demonstrate the effectiveness of our proposed approach in mitigating the overfitting problem and generating more diverse and balanced compositions of concepts in the synthesized images. The source code is available at https://github.com/zhangxulu1996/Compositional-Inversion.
Xulu Zhang, Xiaoyong Wei, Jinlin Wu, Zhaoxiang Zhang 0001, Zhen Lei 0001, Qing Li 0001
AAAI3
2024 SurgFC: Multimodal Surgical Function Calling Framework on the Demand of Surgeons
abstract
The surgical intervention is crucial to patient healthcare, and many studies have developed advanced algorithms to provide understanding and decision-making assistance for surgeons. Despite great progress, these algorithms are developed for a single specific task and scenario, and in practice require the manual combination of different functions, thus limiting the applicability. Thus, an intelligent surgical assistant is expected to accurately understand the surgeon’s intentions and accordingly conduct the specific tasks to support the surgical process. In this work, by improving advanced multimodal large language models (MLLMs), we propose a multimodal Surgical Function Calling (SurgFC) framework that can accurately understand the surgeon’s intention and complete a series of surgical understanding tasks, e.g., surgical scene analysis, surgical instrument detection, and segmentation on demand. Specifically, to achieve superior surgical multimodal understanding, we devise a mixture-of-projectors (MOP) module to align the surgical MLLM in SurgFC to balance the natural and surgical knowledge. Moreover, we devise a surgical Function-Calling Tuning strategy to enable the SurgFC to understand surgical intentions, and thus make a series of surgical function calls on demand to meet the needs of the surgeons. Extensive experiments on neurosurgery data confirm that our SurgFC can understand the surgeon’s intention more accurately than the existing MLLM, resulting in overwhelming performance in textual analysis and visual tasks. The source code is available at https://github.com/franciszchen/SurgFC.
Zhen Chen 0018, Xingjian Luo, Jinlin Wu, Danny T. M. Chan, Zhen Lei 0001, Sébastien Ourselin, Hongbin Liu 0001
BIBM3
2024 SurgBox: Agent-Driven Operating Room Sandbox with Surgery Copilot
abstract
Surgical interventions, particularly in neurology, represent complex and high-stakes scenarios that impose substantial cognitive burdens on surgical teams. Although deliberate education and practice can enhance cognitive capabilities, surgical training opportunities remain limited due to patient safety concerns. To address these cognitive challenges in surgical training and operation, we propose SurgBox, an agent-driven sandbox framework to systematically enhance the cognitive capabilities of surgeons in immersive surgical simulations. Specifically, our SurgBox leverages large language models (LLMs) with tailored Retrieval-Augmented Generation (RAG) to authentically replicate various surgical roles, enabling realistic training environments for deliberate practice. In particular, we devise Surgery Copilot, an AI-driven assistant to actively coordinate the surgical information stream and support clinical decision-making, thereby diminishing the cognitive workload of surgical teams during surgery. By incorporating a novel Long-Short Memory mechanism, our Surgery Copilot can effectively balance immediate procedural assistance with comprehensive surgical knowledge. Extensive experiments using real neurosurgical procedure records validate our SurgBox framework in both enhancing surgical cognitive capabilities and supporting clinical decision-making. By providing an integrated solution for training and operational support to address cognitive challenges, our SurgBox framework advances surgical education and practice, potentially transforming surgical outcomes and healthcare quality. The code is available at https://github.com/franciszchen/SurgBox.
Jinlin Wu, Xusheng Liang, Xuexue Bai, Zhen Chen 0018
IEEE Big Data1
2024 Expanding Scene Graph Boundaries: Fully Open-Vocabulary Scene Graph Generation via Visual-Concept Alignment and Retention
Zuyao Chen, Jinlin Wu, Zhen Lei 0001, Zhaoxiang Zhang 0001, Chang Wen Chen
ECCV (66)2
2024 UNO Arena for Evaluating Sequential Decision-Making Capability of Large Language Models
abstract
Zhanyue Qin, Haochuan Wang, Deyuan Liu, Ziyang Song, Cunhang Fan, Zhao Lv, Jinlin Wu, Zhen Lei, Zhiying Tu, Dianhui Chu, Xiaoyan Yu, Dianbo Sui. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Zhanyue Qin, Deyuan Liu, Cunhang Fan, Zhao Lv, Jinlin Wu, Zhen Lei 0001, Zhiying Tu, Dianbo Sui
EMNLP7
2024 Optimal Location of Electric Vehicle Charging Stations Using Geospatial Big Data
abstract
This study focuses on optimizing the location of electric vehicle (EV) charging stations in the state of Georgia using geospatial big data. With the growing concern about climate change and the need for a transition to sustainable energy, the adoption of EVs is rapidly increasing, necessitating the development of a robust charging infrastructure. We use historical data and predictive modeling to forecast EV adoption by 2030, assessing future charging demand through travel survey data to understand spatial and temporal travel patterns. A Mixed-Integer Programming (MIP) approach, combined with a genetic algorithm, is employed to optimize the layout of charging stations, balancing multiple objectives such as minimizing the number of stations, maximizing coverage, ensuring social equity, and balancing urban-rural distribution. This framework aims to support efficient and equitable development of charging infrastructure in Georgia, ultimately facilitating the wider adoption of electric vehicles.
Mei-Po Kwan, Jinlin Wu
SIGSPATIAL/GIS3
2024 PWISeg: Weakly-Supervised Surgical Instrument Instance Segmentation
abstract
AI-assisted operating room scene understanding is essential for the next generation of surgical interventions. Surgical instrument localization plays an important role in this context. However, existing instrument localization methods primarily focus on surgical instrument localization in endoscopy images and struggle with occlusions in broader operating room scenarios. In this work, we propose a weakly supervised instance segmentation framework, Pixel-driven Weakly-supervised Instance Segmentation (PWISeg), to solve the occluded instrument localization with low-cost annotations. Specifically, We utilize the projection relationship between the bounding box and the surgical instrument mask as a supervision signal to train PWISeg to predict coarse masks of surgical instruments. Then, we use the annotation of a few pixel points to train PWISeg to predict accurate masks of surgical instruments. To extensively validate the effectiveness, we collect and release a high-quality dataset, Surg-Inst that covers real-world hard cases of overlapping, dense placement, and various levels of instrument occlusion. Experiments demonstrate that our PWISeg achieves a remarkable performance advantage over state-of-the-art methods on both Surg-Inst and public HOSPI-Tools datasets.
Zhen Sun 0001, Huan Xu 0003, Jinlin Wu, Zhen Chen 0018, Hongbin Liu 0001, Zhen Lei 0001
ICIP3
2024 ASI-Seg: Audio-Driven Surgical Instrument Segmentation with Surgeon Intention Understanding
abstract
Surgical instrument segmentation is crucial in surgical scene understanding, thereby facilitating surgical safety. Existing algorithms directly detected all instruments of predefined categories in the input image, lacking the capability to segment specific instruments according to the surgeon’s intention. During different stages of surgery, surgeons exhibit varying preferences and focus toward different surgical instruments. Therefore, an instrument segmentation algorithm that adheres to the surgeon’s intention can minimize distractions from irrelevant instruments and assist surgeons to a great extent. The recent Segment Anything Model (SAM) reveals the capability to segment objects following prompts, but the manual annotations for prompts are impractical during the surgery. To address these limitations in operating rooms, we propose an audio-driven surgical instrument segmentation framework, named ASI-Seg, to accurately segment the required surgical instruments by parsing the audio commands of surgeons. Specifically, we propose an intention-oriented multimodal fusion to interpret the segmentation intention from audio commands and retrieve relevant instrument details to facilitate segmentation. Moreover, to guide our ASI-Seg segment of the required surgical instruments, we devise a contrastive learning prompt encoder to effectively distinguish the required instruments from the irrelevant ones. Therefore, our ASI-Seg promotes the workflow in the operating rooms, thereby providing targeted support and reducing the cognitive load on surgeons. Extensive experiments are performed to validate the ASI-Seg framework, which reveals remarkable advantages over classical state-of-the-art and medical SAMs in both semantic segmentation and intention-oriented segmentation. The source code is available at https://github.com/Zonmgin-Zhang/ASI-Seg.
Zhen Chen 0018, Zongming Zhang, Wenwu Guo, Xingjian Luo, Long Bai 0008, Jinlin Wu, Hongliang Ren 0001, Hongbin Liu 0001
IROS6
2024 EndoUIC: Promptable Diffusion Transformer for Unified Illumination Correction in Capsule Endoscopy
Long Bai 0008, Tong Chen 0011, Qiaozhi Tan, Wan Jun Nah, Yanheng Li 0002, Zhicheng He 0010, Sishen Yuan, Zhen Chen 0018, Jinlin Wu, Mobarakol Islam, Zhen Li 0026, Hongbin Liu 0001, Hongliang Ren 0001
MICCAI (7)9
2024 Transforming Surgical Interventions with Embodied Intelligence for Ultrasound Robotics
Huan Xu 0003, Jinlin Wu, Guanglin Cao, Zhen Chen 0018, Zhen Lei 0001, Hongbin Liu 0001
MICCAI (6)2
2024 Generative Active Learning for Image Synthesis Personalization
abstract
This paper presents a pilot study that explores the application of active learning, traditionally studied in the context of discriminative models, to generative models. We specifically focus on image synthesis personalization tasks. The primary challenge in conducting active learning on generative models lies in the open-ended nature of querying, which differs from the closed form of querying in discriminative models that typically target a single concept. We introduce the concept of anchor directions to transform the querying process into a semi-open problem. We propose a direction-based uncertainty sampling strategy to enable generative active learning and tackle the exploitation-exploration dilemma. Extensive experiments are conducted to validate the effectiveness of our approach, demonstrating that an open-source model can achieve superior performance compared to closed-source models developed by large companies, such as Google's StyleDrop. The source code is available at https://github.com/zhangxulu1996/GAL4Personalization.
Xulu Zhang, Wengyu Zhang, Xiaoyong Wei, Jinlin Wu, Zhaoxiang Zhang 0001, Zhen Lei 0001, Qing Li 0001
ACM Multimedia4
2024 Deep Learning Based Occluded Person Re-Identification: A Survey
abstract
Occluded person re-identification (Re-ID) focuses on addressing the occlusion problem when retrieving the person of interest across non-overlapping cameras. With the increasing demand for intelligent video surveillance and the application of person Re-ID technology, the real-world occlusion problem draws considerable interest from researchers. Although a large number of occluded person Re-ID methods have been proposed, there are few surveys that focus on occlusion. To fill this gap and help boost future research, this article provides a systematic survey of occluded person Re-ID. In this work, we review recent deep learning based occluded person Re-ID research. First, we summarize the main issues caused by occlusion as four groups: position misalignment, scale misalignment, noisy information, and missing information. Second, we categorize existing methods into six solution groups: matching, image transformation, multi-scale features, attention mechanism, auxiliary information, and contextual recovery. We also discuss the characteristics of each approach, as well as the issues they address. Furthermore, we present the performance comparison of recent occluded person Re-ID methods on four public datasets: Partial-ReID, Partial-iLIDS, Occluded-ReID, and Occluded-DukeMTMC. We conclude the study with thoughts on promising future research directions.
Yunjie Peng, Jinlin Wu, Boqiang Xu, Chunshui Cao, Xu Liu 0008, Zhenan Sun, Zhiqiang He 0002
ACM Trans. Multim. Comput. Commun. Appl.2
2023 Self-similarity Driven Scale-invariant Learning for Weakly Supervised Person Search
abstract
Weakly supervised person search aims to jointly detect and match persons with only bounding box annotations. Existing approaches typically focus on improving the features by exploring the relations of persons. However, scale variation problem is a more severe obstacle and under-studied that a person often owns images with different scales (resolutions). For one thing, small-scale images contain less information of a person, thus affecting the accuracy of the generated pseudo labels. For another, different similarities between cross-scale images of a person increase the difficulty of matching. In this paper, we address it by proposing a novel one-step framework, named Self-similarity driven Scale-invariant Learning (SSL). Scale invariance can be explored based on the self-similarity prior that it shows the same statistical properties of an image at different scales. To this end, we introduce a Multi-scale Exemplar Branch to guide the network in concentrating on the foreground and learning scale-invariant features by hard exemplars mining. To enhance the discriminative power of the learned features, we further introduce a dynamic pseudo label prediction that progressively seeks true labels for training. Experimental results on two standard benchmarks, i.e., PRW and CUHK-SYSU datasets, demonstrate that the proposed method can solve scale variation problem effectively and perform favorably against state-of-the-art methods. Code is available at https://github.com/Wangbenzhi/SSL.git.
Benzhi Wang, Yang Yang 0062, Jinlin Wu, Guo-Jun Qi, Zhen Lei 0001
ICCV3
2023 Camera-aware representation learning for person re-identification
Jinlin Wu, Zhen Lei 0001, Yang Yang 0062, Shukai Chen, Stan Z. Li
Neurocomputing1
2023 Color-Unrelated Head-Shoulder Networks for Fine-Grained Person Re-identification
abstract
Person re-identification (re-id) attempts to match pedestrian images with the same identity across non-overlapping cameras. Existing methods usually study person re-id by learning discriminative features based on the clothing attributes (e.g., color, texture). However, the clothing appearance is not sufficient to distinguish different persons especially when they are in similar clothes, which is known as the fine-grained (FG) person re-id problem. By contrast, this paper proposes to exploit the color-unrelated feature along with the head-shoulder feature for FG person re-id. Specifically, a color-unrelated head-shoulder network (CUHS) is developed, which is featured in three aspects: (1) It consists of a lightweight head-shoulder segmentation layer for localizing the head-shoulder region and learning the corresponding feature. (2) It exploits instance normalization (IN) for learning color-unrelated features. (3) As IN inevitably reduces inter-class differences, we propose to explore richer visual cues for IN by an attention exploration mechanism to ensure high discrimination. We evaluate our model on the FG-reID, Market1501, and DukeMTMC-reID datasets, and the results show that CUHS surpasses previous methods on both the FG and conventional person re-id problems.
Boqiang Xu, Jian Liang 0001, Lingxiao He, Jinlin Wu, Zhenan Sun
ACM Trans. Multim. Comput. Commun. Appl.4
2022 CAViT: Contextual Alignment Vision Transformer for Video Object Re-identification
Jinlin Wu, Lingxiao He, Wu Liu 0005, Yang Yang 0062, Zhen Lei 0001, Tao Mei 0001, Stan Z. Li
ECCV (14)1
2021 Relationship quality and supply chain quality performance: The effect of supply chain integration in hotel industry
abstract
Abstract This study investigates the relationship between relationship quality and supply chain quality performance in hotel supply chain, through the mediating effect of supply chain integration. A questionnaire survey is used to collect data relating to the research hypotheses. Structural equation model technique is suited for our research goals, and the SmartPLS software is implemented to test the conceptual model. The results show that relationship quality has a direct and positive impact on supply chain quality performance; but after introducing mediating variable—supply chain integration, relationship quality indirectly affects supply chain quality performance through supply chain integration. This means that a good relationship quality can promote supply chain integration and ultimately improve supply chain quality performance.
Shanghong Le, Jinlin Wu, Jianlan Zhong
Comput. Intell.2
2020 An end-to-end exemplar association for unsupervised person Re-identification
Jinlin Wu, Yang Yang 0062, Zhen Lei 0001, Jinqiao Wang, Stan Z. Li, Prayag Tiwari, Hari Mohan Pandey
Neural Networks1
2019 Unsupervised Graph Association for Person Re-Identification
abstract
In this paper, we propose an unsupervised graph association (UGA) framework to learn the underlying view-invariant representations from the video pedestrian tracklets. The core points of UGA are mining the underlying cross-view associations and reducing the damage of noise associations. To this end, UGA is adopts a two-stage training strategy: (1) intra-camera learning stage and (2) intercamera learning stage. The former learns the intra-camera representation for each camera. While the latter builds a cross-view graph (CVG) to associate different cameras. By doing this, we can learn view-invariant representation for all person. Extensive experiments and ablation studies on seven re-id datasets demonstrate the superiority of the proposed UGA over most state-of-the-art unsupervised and domain adaptation re-id methods.
Jinlin Wu, Yang Yang 0062, Zhen Lei 0001, Shengcai Liao, Stan Z. Li
ICCV1
2019 Clustering and Dynamic Sampling Based Unsupervised Domain Adaptation for Person Re-Identification
abstract
Person Re-Identification (Re-ID) has witnessed great improvements due to the advances of the deep convolutional neural networks (CNN). Despite this, existing methods mainly suffer from the poor generalization ability to unseen scenes because of the different characteristics between different domains. To address this issue, a Clustering and Dynamic Sampling (CDS) method is proposed in this paper, which tries to transfer the useful knowledge of existing labeled source domain to the unlabeled target one. Specifically, to improve the discriminability of CNN model on source domain, we use the commonly shared pedestrian attributes (e.g., gender, hat and clothing color etc.) to enrich the information and resort to the margin-based softmax (e.g., A-Softmax) loss to train the model. For the unlabeled target domain, we iteratively cluster the samples into several centers and dynamically select informative ones from each center to fine-tune the source-domain model. Extensive experiments on DukeMTMC-reID and Market-1501 datasets show that the proposed method greatly improves the state of the arts in unsupervised domain adaptation.
Jinlin Wu, Shengcai Liao, Zhen Lei 0001, Xiaobo Wang 0001, Yang Yang 0062, Stan Z. Li
ICME1
2017 Multi-modality Network with Visual and Geometrical Information for Micro Emotion Recognition
abstract
Micro emotion recognition is a very challenging problem because of the subtle appearance variants among different facial expression classes. To deal with the mentioned problem, we proposed a multi-modality convolutional neural networks (CNNs) based on visual and geometrical information in this paper. The visual face image and structured geometry are embedded into a unified network and the recognition accuracy can be benefic from the fused information. The proposed network includes two branches. The first branch is used to extract visual feature from color face images, and another branch is used to extract the geometry feature from 68 facial landmarks. Then, both visual and geometry features are concatenated into a long vector. Finally, the concatenated vector is fed to the hinge loss layer. Compared with the CNN architecture only used face images, our method is more effective and has got better performance. In the final testing phase of Micro Emotion Challenge1, our method has got the first place with the misclassification of 80.212137.
Jianzhu Guo, Jinlin Wu, Jun Wan 0001, Xiangyu Zhu 0001, Zhen Lei 0001, Stan Z. Li
FG3