EDBT 2026 Demo / reviewers in the wild / expert
Jilan Xu
dblp:232/2004
· DBLP profile ↗
32ranked-venue papers
5as first author
31since 2021 · last 2026
0000-0002-8208-2067ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 4 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 3 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding
Guo Chen 0006, Yifei Huang 0002, Jilan Xu, Baoqi Pei, Jiahao Wang 0005, Zhe Chen 0017, Tong Lu 0002, Limin Wang 0002 |
Int. J. Comput. Vis. | 3 |
| 2025 | Minimizing Disparities between Real and Pseudo Queries for Unsupervised Visual GroundingabstractVisual grounding involves the identification and localization of image regions given textual descriptions. To reduce the manual labeling effort on region-text pairs, unsupervised visual grounding aims to generate pseudo bounding box and query pairs for training grounding models. However, there exists significant disparities between real and pseudo queries in terms of object, attribute distributions, and textual formats, limiting the generalization performance of unsupervised grounding methods. To address this challenge, we propose a novel unsupervised visual grounding framework. During training, we prompt Multimodal Large Language Models to generate pseudo queries, in which the entities are beyond the object detector’s pre-defined limited categories, and are associated with richer attributes. We further devise a Modifier Tree structure to bridge the gap of textual format between real and pseudo queries. Extensive experiments demonstrate that our method significantly outperforms state-of-the-art unsupervised approaches on public benchmark datasets, particularly when dealing with complex queries. Changkai Ji, Jilan Xu, Yanhao Zhu, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003 |
ICASSP | 3 |
| 2025 | Learning Streaming Video Representation via Multitask TrainingabstractUnderstanding continuous video streams plays a fundamental role in real-time applications including embodied AI and autonomous driving. Unlike offline video understanding, streaming video understanding requires the ability to process video streams frame by frame, preserve historical information, and make low-latency decisions. To address these challenges, our main contributions are three-fold. (i) We develop a novel streaming video backbone, termed as StreamFormer, by incorporating causal temporal attention into a pre-trained vision transformer. This enables efficient streaming video processing while maintaining image representation capability. (ii) To train StreamFormer, we propose to unify diverse spatial-temporal video understanding tasks within a multitask visual-language alignment framework. Hence, StreamFormer learns global semantics, temporal dynamics, and fine-grained spatial relationships simultaneously. (iii) We conduct extensive experiments on online action detection, online video instance segmentation, and video question answering. StreamFormer achieves competitive results while maintaining efficiency, demonstrating its potential for real-time applications. Yibin Yan, Jilan Xu, Shangzhe Di, Yudi Shi, Qirui Chen, Yifei Huang 0002, Weidi Xie |
ICCV | 2 |
| 2025 | CG-Bench: Clue-grounded Question Answering Benchmark for Long Video UnderstandingabstractThe existing video understanding benchmarks for multimodal large language models (MLLMs) mainly focus on short videos. The few benchmarks for long video understanding often rely on multiple-choice questions (MCQs). Due to the limitations of MCQ evaluations and the advanced reasoning abilities of MLLMs, models can often answer correctly by combining short video insights with elimination, without truly understanding the content. To bridge this gap, we introduce CG-Bench, a benchmark for clue-grounded question answering in long videos. CG-Bench emphasizes the model's ability to retrieve relevant clues, enhancing evaluation credibility. It includes 1,219 manually curated videos organized into 14 primary, 171 secondary, and 638 tertiary categories, making it the largest benchmark for long video analysis. The dataset features 12,129 QA pairs in three question types: perception, reasoning, and hallucination. To address the limitations of MCQ-based evaluation, we develop two novel clue-based methods: clue-grounded white box and black box evaluations, assessing whether models generate answers based on accurate video understanding. We evaluated multiple closed-source and open-source MLLMs on CG-Bench. The results show that current models struggle significantly with long videos compared to short ones, and there is a notable gap between open-source and commercial models. We hope CG-Bench will drive the development of more reliable and capable MLLMs for long video comprehension. Guo Chen 0006, Yifei Huang 0002, Baoqi Pei, Jilan Xu, Yuping He, Tong Lu 0002, Yali Wang 0001, Limin Wang 0002 |
ICLR | 5 |
| 2025 | Modeling Fine-Grained Hand-Object Dynamics for Egocentric Video Representation LearningabstractIn egocentric video understanding, the motion of hands and objects as well as their interactions play a significant role by nature.
However, existing egocentric video representation learning methods mainly focus on aligning video representation with high-level narrations, overlooking the intricate dynamics between hands and objects.
In this work, we aim to integrate the modeling of fine-grained hand-object dynamics into the video representation learning process.
Since no suitable data is available, we introduce HOD, a novel pipeline employing a hand-object detector and a large language model to generate high-quality narrations with detailed descriptions of hand-object dynamics.
To learn these fine-grained dynamics, we propose EgoVideo, a model with a new lightweight motion adapter to capture fine-grained hand-object motion information.
Through our co-training strategy, EgoVideo effectively and efficiently leverages the fine-grained hand-object dynamics in the HOD data.
Extensive experiments demonstrate that our method achieves state-of-the-art performance across multiple egocentric downstream tasks, including improvements of 6.3% in EK-100 multi-instance retrieval, 5.7% in EK-100 classification, and 16.3% in EGTEA classification in zero-shot settings. Furthermore, our model exhibits robust generalization capabilities in hand-object interaction and robot manipulation tasks. Baoqi Pei, Yifei Huang 0002, Jilan Xu, Guo Chen 0006, Yuping He, Lijin Yang, Yali Wang 0001, Weidi Xie, Yu Qiao 0001, Fei Wu 0001, Limin Wang 0002 |
ICLR | 3 |
| 2025 | EgoExo-Gen: Ego-centric Video Prediction by Watching Exo-centric VideosabstractGenerating videos in the first-person perspective has broad application prospects in the field of augmented reality and embodied intelligence.
In this work, we explore the cross-view video prediction task, where given an exo-centric video, the first frame of the corresponding ego-centric video, and textual instructions, the goal is to generate future frames of the ego-centric video.
Inspired by the notion that hand-object interactions (HOI) in ego-centric videos represent the primary intentions and actions of the current actor, we present EgoExo-Gen that explicitly models the hand-object dynamics for cross-view video prediction.
EgoExo-Gen consists of two stages. First, we design a cross-view HOI mask prediction model that anticipates the HOI masks in future ego-frames by modeling the spatio-temporal ego-exo correspondence.
Next, we employ a video diffusion model to predict future ego-frames using the first ego-frame and textual instructions, while incorporating the HOI masks as structural guidance to enhance prediction quality.
To facilitate training, we develop a fully automated pipeline to generate pseudo HOI masks for both ego- and exo-videos by exploiting vision foundation models.
Extensive experiments demonstrate that our proposed EgoExo-Gen achieves better prediction performance compared to previous video prediction models on the public Ego-Exo4D and H2O benchmark datasets, with the HOI masks significantly improving the generation of hands and interactive objects in the ego-centric videos. Jilan Xu, Yifei Huang 0002, Baoqi Pei, Junlin Hou, Qingqiu Li, Guo Chen 0006, Yuejie Zhang, Rui Feng 0001, Weidi Xie |
ICLR | 1 |
| 2025 | TGSAM-2: Text-Guided Medical Image Segmentation Using Segment Anything Model 2
Runtian Yuan, Ling Zhou 0002, Jilan Xu, Qingqiu Li, Mohan Chen 0001, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003 |
MICCAI (10) | 3 |
| 2025 | Text-Promptable Propagation for Referring Medical Image Sequence SegmentationabstractReferring Medical Image Sequence Segmentation (Ref-MISS) is a novel and challenging task that aims to segment anatomical structures in medical image sequences (e.g., endoscopy, ultrasound, CT, and MRI) based on natural language descriptions. Existing 2D and 3D segmentation models struggle to explicitly track objects of interest across medical image sequences, and lack support for interactive, text-driven guidance. To address these limitations, we propose Text-Promptable Propagation (TPP), which enables the recognition of referred objects through cross-modal referring interaction, and maintains continuous tracking across the sequence via Transformer-based triple propagation, using text embeddings as queries. To support this task, we curate a large-scale benchmark, Ref-MISS-Bench, which covers 4 imaging modalities and 20 different organs and lesions. Experimental results on this benchmark demonstrate that TPP consistently outperforms state-of-the-art methods in both medical segmentation and referring video object segmentation. Code and data are available at https://github.com/yuanruntian/TPP. Runtian Yuan, Mohan Chen 0001, Jilan Xu, Ling Zhou 0002, Qingqiu Li, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003 |
ACM Multimedia | 3 |
| 2025 | EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMsabstractTransferring and integrating knowledge across first-person (egocentric) and third-person (exocentric) viewpoints is intrinsic to human intelligence, enabling humans to learn from others and convey insights from their own experiences. Despite rapid progress in multimodal large language models (MLLMs), their ability to perform such cross-view reasoning remains unexplored. To address this, we introduce EgoExoBench, the first benchmark for egocentric exocentric video understanding and reasoning. Built from publicly available datasets, EgoExoBench comprises over 7300 question–answer pairs spanning eleven sub-tasks organized into three core challenges: semantic alignment, viewpoint association, and temporal reasoning. We evaluate 13 state-of-the-art MLLMs and find that while these models excel on single-view tasks, they struggle to align semantics across perspectives, accurately associate views, and infer temporal dynamics in the ego-exo context. We hope EgoExoBench can serve as a valuable resource for research on embodied agents and intelligent assistants seeking human-like cross-view intelligence. Yuping He, Yifei Huang 0002, Guo Chen 0006, Baoqi Pei, Jilan Xu, Tong Lu 0002, Jiangmiao Pang |
NeurIPS | 5 |
| 2025 | AOR: Anatomical Ontology-Guided Reasoning for Medical Large Multimodal Model in Chest X-Ray InterpretationabstractChest X-rays (CXRs) are the most frequently performed imaging examinations in clinical settings. Recent advancements in Medical Large Multimodal Models (MLMMs) have enabled automated CXR interpretation, improving diagnostic accuracy and efficiency. However, despite their strong visual understanding, current MLMMs still face two major challenges: (1) insufficient region-level understanding and interaction, and (2) limited accuracy and interpretability due to single-step prediction. In this paper, we address these challenges by empowering MLMMs with anatomy-centric reasoning capabilities to enhance their interactivity and explainability. Specifically, we propose an Anatomical Ontology-Guided Reasoning (AOR) framework that accommodates both textual and optional visual prompts, centered on region-level information to enable multimodal multi-step reasoning. We also develop AOR-Instruction, a large instruction dataset for MLMMs training, under the guidance of expert physicians. Our experiments demonstrate AOR's superior performance in both Visual Question Answering (VQA) and report generation tasks. Code and data are available at: https://github.com/Liqq1/AOR. Qingqiu Li, Zihang Cui, Seongsu Bae, Jilan Xu, Runtian Yuan, Yuejie Zhang, Rui Feng 0001, Quanli Shen, Shang Gao 0003, Junjun He |
NeurIPS | 4 |
| 2025 | EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoTabstractEgocentric video reasoning centers on an unobservable agent behind the camera who dynamically shapes the environment, requiring inference of hidden intentions and recognition of fine-grained interactions. This core challenge limits current multimodal large language models (MLLMs), which excel at visible event reasoning but lack embodied, first-person understanding. To bridge this gap, we introduce EgoThinker, a novel framework that endows MLLMs with robust egocentric reasoning capabilities through spatio-temporal chain-of-thought supervision and a two-stage learning curriculum. First, we introduce EgoRe-5M, a large-scale egocentric QA dataset constructed from 13M diverse egocentric video clips. This dataset features multi-minute segments annotated with detailed CoT rationales and dense hand–object grounding. Second, we employ SFT on EgoRe-5M to instill reasoning skills, followed by reinforcement fine-tuning (RFT) to further enhance spatio-temporal localization. Experimental results show that EgoThinker outperforms existing methods across multiple egocentric benchmarks, while achieving substantial improvements in fine-grained spatio-temporal localization tasks. Baoqi Pei, Yifei Huang 0002, Jilan Xu, Yuping He, Guo Chen 0006, Fei Wu 0001, Jiangmiao Pang, Yu Qiao 0001 |
NeurIPS | 3 |
| 2025 | Guiding Audio-Visual Question Answering with Collective Question ReasoningabstractAbstract Audio-Visual Question Answering (AVQA) requires the model to answer questions with complex dynamic audio-visual information. Prior works on this task mainly consider only using single question-answer pairs during training, overlooking the rich semantic associations between questions. In this work, we propose a novel Collective Question-Guided Network (CoQo), which accepts multiple question-answer pairs as input and leverages the reasoning over these questions to assist the model training process. The core module is the proposed Question Guided Transformer (QGT), which uses collective question reasoning to perform question-guided feature extraction. Since multiple question-answer pairs are not always available, especially during inference, our QGT uses a set of learnable tokens to learn the collective information from multiple questions during training. At inference time, these learnable tokens bring additional reasoning information even when only one question is used as input. We employ QGT in both spatial and temporal dimensions to extract question-related features effectively and efficiently. To better capture detailed audio-visual associations, we train the model in a finer level by distinguishing feature pairs of different questions within the same video. Extensive experiments demonstrate that our method can achieve state-of-the-art performance on three AVQA datasets while reducing training time significantly. We also observe strong performances of our method on three VQA benchmarks. Detailed ablation studies further confirm the effectiveness of our proposed collective question reasoning scheme, both quantitatively and qualitatively. Baoqi Pei, Yifei Huang 0002, Guo Chen 0006, Jilan Xu, Yali Wang 0001, Limin Wang 0002, Tong Lu 0002, Yu Qiao 0001, Fei Wu 0001 |
Int. J. Comput. Vis. | 4 |
| 2025 | QMix: Quality-Aware Learning With Mixed Noise for Robust Retinal Disease DiagnosisabstractDue to the complex nature of medical image acquisition and annotation, medical datasets inevitably contain noise. This adversely affects the robustness and generalization of deep neural networks. Previous noise learning methods mainly considered noise arising from images being mislabeled, i.e., label noise, assuming all mislabeled images were of high quality. However, medical images can also suffer from severe data quality issues, i.e., data noise, where discriminative visual features for disease diagnosis are missing. In this paper, we propose QMix, a noise learning framework that learns a robust disease diagnosis model under mixed noise scenarios. QMix alternates between sample separation and quality-aware semi-supervised training in each epoch. The sample separation phase uses a joint uncertainty-loss criterion to effectively separate (1) correctly labeled images, (2) mislabeled high-quality images, and (3) mislabeled low-quality images. The semi-supervised training phase then learns a robust disease diagnosis model from the separated samples. Specifically, we propose a sample-reweighing loss to mitigate the effect of mislabeled low-quality images during training, and a contrastive enhancement loss to further distinguish them from correctly labeled images. QMix achieved state-of-the-art performance on six public retinal image datasets and exhibited significant improvements in robustness against mixed noise. Code will be available upon acceptance. Junlin Hou, Jilan Xu, Rui Feng 0001, Hao Chen 0011 |
IEEE Trans. Medical Imaging | 2 |
| 2024 | MVBench: A Comprehensive Multi-modal Video Understanding BenchmarkabstractWith the rapid development of Multi-modal Large language Models (MLLMs), a number of diagnostic bench-marks have recently emerged to evaluate the comprehension capabilities of these models. However, most bench-marks predominantly assess spatial understanding in the static image tasks, while overlooking temporal understanding in the dynamic video tasks. To alleviate this issue, we introduce a comprehensive Multi-modal Video understanding Benchmark, namely MVBench, which covers 20 chal-lenging video tasks that cannot be effectively solved with a single frame. Specifically, we first introduce a novel static-to-dynamic method to define these temporal-related tasks. By transforming various static tasks into dynamic ones, we enable the systematic generation of video tasks that require a broad spectrum of temporal skills, ranging from perception to cognition. Then, guided by the task definition, we au-tomatically convert public video annotations into multiple-choice QA to evaluate each task. On one hand, such a distinct paradigm allows us to build MVBench efficiently, without much manual intervention. On the other hand, it guarantees evaluation fairness with ground-truth video an-notations, avoiding the biased scoring of LLMs. More-over, we further develop a robust video MLLM baseline, i.e., VideoChat2, by progressive multi-modal training with di-verse instruction-tuning data. The extensive results on our MVBench reveal that, the existing MLLMs are far from sat-isfactory in temporal understanding, while our VideoChat2 largely surpasses these leading models by over 15% on MVBench. All models and data are available at https://github.com/OpenGVLab/Ask-Anything. Kunchang Li 0002, Yali Wang 0001, Yinan He, Yizhuo Li 0001, Yi Wang 0074, Yi Liu 0081, Zun Wang 0001, Jilan Xu, Guo Chen 0006, Ping Lou, Limin Wang 0002, Yu Qiao 0001 |
CVPR | 8 |
| 2024 | EgoExoLearn: A Dataset for Bridging Asynchronous Ego- and Exo-centric View of Procedural Activities in Real WorldabstractBeing able to map the activities of others into one's own point of view is a fundamental human skill even from a very early age. Taking a step toward understanding this human ability, we introduce EgoExoLearn, a large-scale dataset that emulates the human demonstration following process, in which individuals record egocentric videos as they execute tasks guided by exocentric-view demonstration videos. Focusing on the potential applications in daily assistance and professional support, EgoExoLearn contains egocentric and demonstration video data spanning 120 hours captured in daily life scenarios and specialized laboratories. Along with the videos we record high-quality gaze data and provide detailed multimodal annotations, formulating a playground for modeling the human ability to bridge asynchronous procedural actions from different viewpoints. To this end, we present benchmarks such as crossview association, cross-view action planning, and crossview referenced skill assessment, along with detailed analysis. We expect EgoExoLearn can serve as an important resource for bridging the actions across views, thus paving the way for creating AI agents capable of seamlessly learning by observing humans in the real world. The dataset and benchmark codes are available at https://github.com/OpenGVLab/EgoExoLearn. Yifei Huang 0002, Guo Chen 0006, Jilan Xu, Mingfang Zhang 0002, Lijin Yang, Baoqi Pei, Hongjie Zhang 0002, Lu Dong 0005, Yali Wang 0001, Limin Wang 0002, Yu Qiao 0001 |
CVPR | 3 |
| 2024 | Retrieval-Augmented Egocentric Video CaptioningabstractUnderstanding human actions from videos offirst-person view poses significant challenges. Most prior approaches explore representation learning on egocentric videos only, while overlooking the potential benefit of exploiting existing large-scale third-person videos. In this paper, (1) we develop EgoInstructor, a retrieval-augmented multimodal captioning model that automatically retrieves semantically relevant third-person instructional videos to enhance the video captioning of egocentric videos, (2) for training the cross-view retrieval module, we devise an au-tomatic pipeline to discover ego-exo video pairs from distinct large-scale egocentric and exocentric datasets, (3) we train the cross-view retrieval module with a novel EgoEx-oNCE loss that pulls egocentric and exocentric video features closer, by aligning them to shared text features that describe similar actions, (4) through extensive experiments, our cross-view retrieval module demonstrates superior performance across seven benchmarks. Regarding egocen-tric video captioning, EgoInstructor exhibits significant improvements by leveraging third-person videos as references. Jilan Xu, Yifei Huang 0002, Junlin Hou, Guo Chen 0006, Yuejie Zhang, Rui Feng 0001, Weidi Xie |
CVPR | 1 |
| 2024 | InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Yi Wang 0074, Kunchang Li 0002, Xinhao Li 0004, Jiashuo Yu, Yinan He, Guo Chen 0006, Baoqi Pei, Rongkun Zheng, Zun Wang 0001, Yansong Shi, Tianxiang Jiang, Jilan Xu, Hongjie Zhang 0002, Yifei Huang 0002, Yu Qiao 0001, Yali Wang 0001, Limin Wang 0002 |
ECCV (85) | 13 |
| 2024 | Concept-Attention Whitening for Interpretable Skin Lesion Diagnosis
Junlin Hou, Jilan Xu |
MICCAI (10) | 2 |
| 2024 | Anatomical Structure-Guided Medical Vision-Language Pre-training
Qingqiu Li, Xiaohan Yan, Jilan Xu, Runtian Yuan, Yuejie Zhang, Rui Feng 0001, Quanli Shen |
MICCAI (11) | 3 |
| 2024 | Does Video-Text Pretraining Help Open-Vocabulary Online Action Detection?abstractVideo understanding relies on accurate action detection for temporal analysis. However, existing mainstream methods have limitations in real-world applications due to their offline and closed-set evaluation approaches, as well as their dependence on manual annotations. To address these challenges and enable real-time action understanding in open-world scenarios, we propose OV-OAD, a zero-shot online action detector that leverages vision-language models and learns solely from text supervision. By introducing an object-centered decoder unit into a Transformer-based model, we aggregate frames with similar semantics using video-text correspondence. Extensive experiments on four action detection benchmarks demonstrate that OV-OAD outperforms other advanced zero-shot methods. Specifically, it achieves 37.5\% mean average precision on THUMOS’14 and 73.8\% calibrated average precision on TVSeries. This research establishes a robust baseline for zero-shot transfer in online action detection, enabling scalable solutions for open-world temporal understanding. The code will be available for download at \url{https://github.com/OpenGVLab/OV-OAD}. Yi Wang 0074, Jilan Xu, Yinan He, Zifan Song, Limin Wang 0002, Yu Qiao 0001, Cairong Zhao |
NeurIPS | 3 |
| 2023 | Enhanced Knowledge Injection for Radiology Report GenerationabstractAutomatic generation of radiology reports holds crucial clinical value, as it can alleviate substantial workload on radiologists and remind less experienced ones of potential anomalies. Despite the remarkable performance of various image captioning methods in the natural image field, generating accurate reports for medical images still faces challenges, i.e., disparities in visual and textual data, and lack of accurate domain knowledge. To address these issues, we propose an enhanced knowledge injection framework, which utilizes two branches to extract different types of knowledge. The Weighted Concept Knowledge (WCK) branch is responsible for introducing clinical medical concepts weighted by TF-IDF scores. The Multimodal Retrieval Knowledge (MRK) branch extracts triplets from similar reports, emphasizing crucial clinical information related to entity positions and existence. By integrating this finer-grained and well-structured knowledge with the current image, we are able to leverage the multi-source knowledge gain to ultimately facilitate more accurate report generation. Extensive experiments have been conducted on two public benchmarks, demonstrating that our method achieves superior performance over other state-of-the-art methods. Ablation studies further validate the effectiveness of two extracted knowledge sources. Qingqiu Li, Jilan Xu, Runtian Yuan, Mohan Chen 0001, Yuejie Zhang, Rui Feng 0001, Shang Gao 0003 |
BIBM | 2 |
| 2023 | Semi-MedSeq: Semi-supervised Semantic Segmentation for Medical Image SequencesabstractIn clinical practice, medical imaging techniques include 2D video-based examinations that capture sequential scans, and 3D volumetric imaging that forms a comprehensive 3D representation from a stack of 2D slices. The medical image sequences produced by the above techniques provide valuable spatio-temporal characteristics for analysis and segmentation, but the annotation of image sequences is extremely time-consuming and labor-intensive. To exploit the coherence and address the scarcity of labeled data, we propose a novel semi-supervised semantic segmentation framework for medical image sequences, which consists of a conditional network and a denoising network. Specifically, we embed a Sequential Feature Reconstruction module into both networks. This module reconstructs the target frame from contiguous frames and captures their shared visual features. Guided by the context-enhancing information from the conditioning network, the denoising network suppresses background noise via a Diffusion-based Noise Elimination module. Extensive experiments are conducted on 2D and 3D tasks, including cardiac segmentation, polyp segmentation, placenta vessel segmentation and abdomen multi-organ segmentation. The results show our method is superior to existing semi-supervised methods and exhibits advantages over fully-supervised medical image segmentation methods with only 1/2 labeled data, validating its effectiveness and generalization ability. Runtian Yuan, Jilan Xu, Qingqiu Li, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003 |
BIBM | 2 |
| 2023 | Learning Open-Vocabulary Semantic Segmentation Models From Natural Language SupervisionabstractThis paper considers the problem of open-vocabulary semantic segmentation (OVS), that aims to segment objects of arbitrary classes beyond a pre-defined, closed-set categories. The main contributions are as follows: First, we propose a transformer-based model for OVS, termed as OVSegmentor, which only exploits web-crawled imagetext pairs for pre-training without using any mask annotations. OVSegmentor assembles the image pixels into a set of learnable group tokens via a slotattention based binding module, then aligns the group tokens to corresponding caption embeddings. Second, we propose two proxy tasks for training, namely masked entity completion and cross-image mask consistency. The former aims to infer all masked entities in the caption given group tokens, that enables the model to learn fine-grained alignment between visual groups and text entities. The latter enforces consistent mask predictions between images that contain shared entities, encouraging the model to learn visual invariance. Third, we construct CC4M dataset for pre-training by filtering CC12M with frequently appeared entities, which significantly improves training efficiency. Fourth, we perform zero-shot transfer on four benchmark datasets, PASCAL VOC, PASCAL Context, COCO Object, and ADE20K. OVSegmentor achieves superior results over state-of-the-art approaches on PASCAL VOC using only 3% data (4M vs 134M) for pre-training. Jilan Xu, Junlin Hou, Yuejie Zhang, Rui Feng 0001, Yi Wang 0074, Yu Qiao 0001, Weidi Xie |
CVPR | 1 |
| 2023 | Diabetic Retinopathy Grading with Weakly-Supervised Lesion PriorsabstractExplicit information of lesions can provide visual instructions for diabetic retinopathy (DR) grading on fundus images. However, pixel-level lesion annotations are extremely difficult and time-consuming to acquire. In this work, we propose a novel weakly-supervised lesion-aware network for DR grading, which enhances the discriminative features with lesion priors by only image-level supervision. Specifically, we design a lesion attention module that generates lesion activation maps by introducing an auxiliary task of binary DR identification. Lesion activation maps are utilized to assist the network to focus on the most relevant regions for boosting DR grading performance. Besides, we particularly devise an adaptive joint loss to balance the DR identification and DR grading tasks dynamically. Extensive results on the public DR dataset demonstrate the superiority and generality of our proposed lesion-aware network. The interpretability of generated lesion activation maps is also verified by the comparison with ground truth segmentation masks. Junlin Hou, Jilan Xu, Rui Feng 0001, Yue Zhang 0004, Haidong Zou, Lina Lu, Wenwen Xue |
ICASSP | 3 |
| 2023 | SCSGNet: Spatial-Correlated and Shape-Guided Network for Breast Mass SegmentationabstractAutomatic and accurate breast mass segmentation plays a crucial role in the early diagnosis of breast cancer. However, it has been a challenging task for two main reasons: (1) Breast masses are diverse; and (2) The boundaries of masses are ambiguous. To address these problems, we propose a Spatial-Correlated and Shape-Guided Network (SCSGNet), which combines global context extraction with local boundary refinement. Specifically, the high-level features are aggregated to produce a global map as the initial guidance area, and a Series-Parallel Feature Fusion (SPFF) module is added to capture masses of different shapes and sizes. Besides, we design a Dynamic Long-range Correlation Capture (DLCC) module to capture the spatial correlation of masses at different positions. Finally, we devise a Triplet Attention Guide (TAG) module to iteratively update the feature map and refine the boundary. Experiments on two public datasets demonstrate that our method achieves superior performance over other state-of-the-art methods. Qingqiu Li, Jilan Xu, Runtian Yuan, Yuejie Zhang, Rui Feng 0001 |
ICASSP | 2 |
| 2022 | Single-Modality Endoscopic Polyp Segmentation via Random Color Reversal Synthesis and Two-Branched LearningabstractEndoscopic polyp segmentation plays a fundamental role in the diagnosis and treatment of colorectal cancer. However, polyp segmentation often suffers from limited accuracy due to its large variations in appearance, blurry boundary and severe imbalanced illumination. In this paper, we propose a novel Translation Assisted Segmentation Network (TASNet) for polyp segmentation of single-modality endoscopic images. It consists of two branches, i.e. an image-to-image translation branch and an image segmentation branch. These two branches communicate via a shared encoder. For the image-to-image translation branch, a Color Reversal Strategy is established to treat the original image as source image and synthesize target images. Moreover, we introduce a Random Color Reversal Synthesis module for progressive segmentation. Extensive experiments show that our framework achieves superior performance than state-of-the-art methods on five widely-used endoscopic image datasets. Mingzhu Chen, Jilan Xu, Runtian Yuan, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003 |
BIBM | 3 |
| 2022 | Cross-Field Transformer for Diabetic Retinopathy Grading on Two-field Fundus ImagesabstractAutomatic diabetic retinopathy (DR) grading based on fundus photography has been widely explored to benefit the routine screening and early treatment. Existing researches generally focus on single-field fundus images, which have limited field of view for precise eye examinations. In clinical applications, ophthalmologists adopt two-field fundus photography as the dominating tool, where the information from each field (i.e., macula-centric and optic disc-centric) is highly correlated and complementary, and benefits comprehensive decisions. However, automatic DR grading based on two-field fundus photography remains a challenging task due to the lack of publicly available datasets and effective fusion strategies. In this work, we first construct a new benchmark dataset (DRTiD) for DR grading, consisting of 3,100 two-field fundus images. To the best of our knowledge, it is the largest public DR dataset with diverse and high-quality two-field images. Then, we propose a novel DR grading approach, namely Cross-Field Transformer (CrossFiT), to capture the correspondence between two fields as well as the long-range spatial correlations within each field. Considering the inherent two-field geometric constraints, we particularly define aligned position embeddings to preserve relative consistent position in fundus. Besides, we perform masked cross-field attention during interaction to filter the noisy relations between fields. Extensive experiments on our DRTiD dataset and a public DeepDRiD dataset demonstrate the effectiveness of our CrossFiT network. The new dataset and the source code of CrossFiT will be publicly available at https://github.com/DU-VTS/DRTiD. Junlin Hou, Jilan Xu, Yuejie Zhang, Haidong Zou, Lina Lu, Wenwen Xue, Rui Feng 0001 |
BIBM | 2 |
| 2022 | MedSeq: Semantic Segmentation for Medical Image SequencesabstractMedical image segmentation plays a critical role in computer-aided diagnosis, while the diversity and complexity of medical images make it difficult to segment precisely. In practice, medical images of specific modalities (e.g. Magnetic Resonance Imaging, Colonoscopy and Ultrasonography) are collected as sequences independently for every patient. However, 1) there exists few works exploiting sequence information among successive frames, neglecting inter-frame relationships that are useful to locate target objects; 2) the performance of medical image segmentation is limited to the low contrast or blurry boundary of medical images, and intra-frame dependencies are not fully explored. Thus in this paper, we propose MedSeq for segmenting objects of interest in medical image sequences. Following the “locate-then-refine” paradigm, we locate target regions by modeling cross-frame relationships and then perform refinement on coarse masks. More specifically, we design a Cross-frame Attention module to learn correlations among frames, taking advantages of their similar appearances. For refinement, we propose a novel Boundary-aware Transformer to improve the segmentation of boundary patches. Extensive experiments are conducted on benchmark datasets of Cardiac Segmentation and Video Polyp Segmentation. Our method achieves superior performance over the state-of-the-art methods. Runtian Yuan, Jilan Xu, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003 |
BIBM | 2 |
| 2022 | CREAM: Weakly Supervised Object Localization via Class RE-Activation MappingabstractWeakly Supervised Object Localization (WSOL) aims to localize objects with image-level supervision. Existing works mainly rely on Class Activation Mapping (CAM) de-rived from a classification model. However, CAM-based methods usually focus on the most discriminative parts of an object (i.e., incomplete localization problem). In this paper, we empirically prove that this problem is associated with the mixup of the activation values between less discrimi-native foreground regions and the background. To address it, we propose Class RE-Activation Mapping (CREAM), a novel clustering-based approach to boost the activation values of the integral object regions. To this end, we in-troduce class-specific foreground and background context embeddings as cluster centroids. A CAM-guided momen-tum preservation strategy is developed to learn the context embeddings during training. At the inference stage, the re-activation mapping is formulated as a parameter es-timation problem under Gaussian Mixture Model, which can be solved by deriving an unsupervised Expectation- Maximization based soft-clustering algorithm. By simply integrating CREAM into various WSOL approaches, our method significantly improves their performance. CREAM achieves the state-of-the-art performance on CUB, ILSVRC and OpenImages benchmark datasets. Code will be avail-able at https://github.com/lazzcharles/CREAM. Jilan Xu, Junlin Hou, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Xuequan Lu, Shang Gao 0003 |
CVPR | 1 |
| 2022 | TCCNet: Temporally Consistent Context-Free Network for Semi-supervised Video Polyp SegmentationabstractAutomatic video polyp segmentation (VPS) is highly valued for the early diagnosis of colorectal cancer. However, existing methods are limited in three respects: 1) most of them work on static images, while ignoring the temporal information in consecutive video frames; 2) all of them are fully supervised and easily overfit in presence of limited annotations; 3) the context of polyp (i.e., lumen, specularity and mucosa tissue) varies in an endoscopic clip, which may affect the predictions of adjacent frames. To resolve these challenges, we propose a novel Temporally Consistent Context-Free Network (TCCNet) for semi-supervised VPS. It contains a segmentation branch and a propagation branch with a co-training scheme to supervise the predictions of unlabeled image. To maintain the temporal consistency of predictions, we design a Sequence-Corrected Reverse Attention module and a Propagation-Corrected Reverse Attention module. A Context-Free Loss is also proposed to mitigate the impact of varying contexts. Extensive experiments show that even trained under 1/15 label ratio, TCCNet is comparable to the state-of-the-art fully supervised methods for VPS. Also, TCCNet surpasses existing semi-supervised methods for natural image and other medical image segmentation tasks. Jilan Xu, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Xuequan Lu, Shang Gao 0003 |
IJCAI | 2 |
| 2021 | Periphery-aware COVID-19 diagnosis with contrastive representation enhancement
Junlin Hou, Jilan Xu, Longquan Jiang 0003, Shanshan Du, Rui Feng 0001, Yuejie Zhang, Xiangyang Xue 0001 |
Pattern Recognit. | 2 |
| 2020 | Data-Efficient Histopathology Image Analysis with Deformation Representation LearningabstractHistopathological examination of tissue biopsies plays a fundamental role in disease assessment. Automatic histopathology image analysis requires substantial task-specific annotations, which are often expensive and laborious in realworld scenarios. This insufficient annotation of data limits the generalization ability of supervised learning models. To address this challenge, we propose a self-supervised Deformation Representation Learning (DRL) framework to learn semantic features from unlabeled data. As a novel paradigm, our approach utilizes deformation as supervisory signals based on two critical features, i.e., local structure heterogeneity and global context homogeneity. Given an original histopathology image and its deformed counterpart, there exists a moderate difference in local structures. In contrast, due to the transformation-invariance, both images share a similar global context compared with other images. Specifically, an encoder network is trained to distinguish the local inconsistency by measuring the mutual information and maintain the global consistency with noise contrastive estimation. Extensive experiments on public histopathology image datasets show that the learned representations are generalizable for various downstream tasks, such as transfer learning on segmentation and semi-supervised classification. Our approach achieves superior results over other self-supervised methods and the ImageNet pre-trained model, and it reveals the ability as a novel pre-training scheme in histopathology image analysis. Jilan Xu, Junlin Hou, Yuejie Zhang, Rui Feng 0001, Chunyang Ruan, Tao Zhang 0022, Weiguo Fan |
BIBM | 1 |