VLDB 2026 Research / reviewers in the wild / expert
Zhen Li 0026
dblp:74/2397-26
· DBLP profile ↗
117ranked-venue papers
6as first author
103since 2021 · last 2027
0000-0002-7669-2686ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 78 · 5 first-author · 70 since 2021Graphics, computer vision, multimedia, augmented reality and games · 71 · 3 first-author · 59 since 2021Applied, interdisciplinary, general and emerging computing · 30 · 1 first-author · 25 since 2021Systems, architecture and hardware · 6 · 6 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | MedDATP: Adapting CLIP for few-shot medical image classification via domain adapter and task prompts
Zixun Zhang, Yuncheng Jiang 0002, Jun Wei 0006, Huazhu Fu, Shuguang Cui, Tao Luo 0014, Zhen Li 0026 |
Expert Syst. Appl. | 7 |
| 2026 | Composition-Incremental Learning for Compositional GeneralizationabstractCompositional generalization has achieved substantial progress in computer vision on pre-collected training data. Nonetheless, real-world data continually emerges, with possible compositions being nearly infinite, long-tailed, and not entirely visible. Thus, an ideal model is supposed to gradually improve the capability of compositional generalization in an incremental manner. In this paper, we explore Composition-Incremental Learning for Compositional Generalization (CompIL) in the context of the compositional zero-shot learning (CZSL) task, where models need to continually learn new compositions, intending to improve their compositional generalization capability progressively. To quantitatively evaluate CompIL, we develop a benchmark construction pipeline leveraging existing datasets, yielding MIT-States-CompIL and C-GQA-CompIL. Furthermore, we propose a pseudo-replay framework utilizing a visual synthesizer to synthesize visual representations of learned compositions and a linguistic primitive distillation mechanism to maintain aligned primitive representations across the learning process. Extensive experiments demonstrate the effectiveness of the proposed framework. Zhen Li 0026, Yuwei Wu 0001, Chenchen Jing, Che Sun, Chuanhao Li 0001, Yunde Jia |
AAAI | 1 |
| 2026 | DriveFlow: Rectified Flow Adaptation for Robust 3D Object Detection in Autonomous DrivingabstractIn autonomous driving, vision-centric 3D object detection recognizes and localizes 3D objects from RGB images. However, due to high annotation costs and diverse outdoor scenes, training data often fails to cover all possible test scenarios, known as the out-of-distribution (OOD) issue. Training-free image editing offers a promising solution for improving model robustness by training data enhancement without any modifications to pre-trained diffusion models. Nevertheless, inversion-based methods often suffer from limited effectiveness and inherent inaccuracies, while recent rectified-flow-based approaches struggle to preserve objects with accurate 3D geometry. In this paper, we propose DriveFlow, a Rectified Flow Adaptation method for training data enhancement in autonomous driving based on pre-trained Text-to-Image flow models. Based on frequency decomposition, DriveFlow introduces two strategies to adapt noise-free editing paths derived from text-conditioned velocities. 1) High-Frequency Foreground Preservation: DriveFlow incorporates a high-frequency alignment loss for foreground to maintain precise 3D object geometry. 2) Dual-Frequency Background Optimization: DriveFlow also conducts dual-frequency optimization for background, balancing editing flexibility and semantic consistency. Comprehensive experiments validate the effectiveness and efficiency of DriveFlow, demonstrating comprehensive performance improvements on all categories across OOD scenarios. Yiming Yang 0001, Chaoda Zheng, Yifan Zhang 0004, Shuaicheng Niu, Zilu Guo, Gui Gui, Shuguang Cui, Zhen Li 0026 |
AAAI | 10 |
| 2026 | MDK12-Bench: A Multi-Discipline Benchmark for Evaluating Reasoning in Multimodal Large Language ModelsabstractMultimodal large language models (MLLMs), which integrate language and visual cues for problem-solving, are crucial for advancing artificial general intelligence (AGI). However, current benchmarks for measuring the intelligence of MLLMs suffer from limited scale, narrow coverage, and unstructured knowledge, offering only static and undifferentiated evaluations. To bridge this gap, we introduce MDK12-Bench, a large-scale multidisciplinary benchmark built from real-world K–12 exams spanning six disciplines with 141K instances and 6,225 knowledge points organized in a six-layer taxonomy. Covering five question formats with difficulty and year annotations, it enables comprehensive evaluation to capture the extent to which MLLMs perform over four dimensions: 1) difficulty levels, 2) temporal (cross-year) shifts, 3) contextual shifts, and 4) knowledge-driven reasoning. We propose a novel dynamic evaluation framework that introduces unfamiliar visual, textual, and question form shifts to challenge model generalization while improving benchmark objectivity and longevity by mitigating data contamination. We further evaluate knowledge-point reference-augmented generation (KP-RAG) to examine the role of knowledge in reasoning. Key findings reveal limitations in current MLLMs in multiple aspects and provide guidance for enhancing model reasoning, robustness, and AI-assisted education. Xiaopeng Peng 0001, Fanrui Zhang, Zhaopan Xu, Jiaxin Ai, Yansheng Qiu, Wangbo Zhao, Jiajun Song, Chuanhao Li 0001, Weidong Tang, Zhen Li 0026, Haoquan Zhang, Zizhen Li, Xiaofeng Mao, Yukang Feng, Kai Wang 0036, Xiaojun Chang, Wenqi Shao, Yang You 0001, Kaipeng Zhang |
AAAI | 11 |
| 2026 | PiSA: A Self-Augmented Data Engine and Training Strategy for 3D Understanding with Large Modelsabstract3D Multimodal Large Language Models (MLLMs) have recently made substantial advancements. However, their potential remains untapped, primarily due to the limited quantity and suboptimal quality of 3D datasets. Current approaches attempt to transfer knowledge from 2D MLLMs to expand 3D instruction data, but still face modality and domain gaps. To this end, we introduce PiSA-Engine (Point-Self-Augmented-Engine), a new framework for generating instruction point-language datasets enriched with 3D spatial semantics. We observe that existing 3D MLLMs offer a comprehensive understanding of point clouds for annotation, while 2D MLLMs excel at cross-validation by providing complementary information. By integrating holistic 2D and 3D insights from off-the-shelf MLLMs, PiSA-Engine enables a continuous cycle of high-quality data generation. We select PointLLM as the baseline and adopt this co-evolution training framework to develop an enhanced 3D MLLM, termed PointLLM-PiSA. Additionally, we identify limitations in previous 3D benchmarks, which often feature coarse language captions and insufficient category diversity, resulting in inaccurate evaluations. To address this gap, we further introduce PiSA-Bench, a comprehensive 3D benchmark covering six key aspects with detailed and diverse labels. Experimental results demonstrate PointLLM-PiSA’s state-of-the-art performance in zero-shot 3D object captioning and generative classification on our PiSA-Bench, achieving significant improvements of 46.45% (+8.33%) and 63.75% (+16.25%), respectively. Project page: CG-ops/PiSA. Zilu Guo, Zhihao Yuan, Chaoda Zheng, Pengshuo Qiu, Dongzhi Jiang, Renrui Zhang, Chun-Mei Feng 0001, Zhen Li 0026 |
WACV | 9 |
| 2026 | Adaptive Pruning for Large Language Models With Structural Importance AwarenessabstractThe recent advancements in large language models (LLMs) have significantly enhanced language understanding and content generation capabilities. However, the deployment of LLMs on resource-constrained Internet of Things (IoT) devices remains challenging due to their substantial computational and storage requirements. To address this issue, we propose a novel LLM pruning method, termed structurally-aware adaptive pruning (SAAP), to reduce computational and storage costs for LLMs while maintaining model performance. Specifically, SAAP first leverages maximum likelihood estimation to calibrate traditional structural importance metrics for LLM pruning. Next, it employs a Bayesian fusion approach to address the predictive uncertainty in multi-granularity metrics, enabling accurate assessments of structural importance for LLMs. Then, SAAP introduces a cross-layer importance alignment mechanism based on quantile mapping, which normalizes layer-wise importance scores to ensure consistent pruning from a global perspective. Furthermore, SAAP develops an efficient block-wise fine-tuning strategy for enhancing the performance of the LLM after pruning. To validate the effectiveness of SAAP, we conduct extensive experiments on nine open-source LLMs across two representative tasks—language modeling and zero-shot classification. Experimental results show that SAAP consistently outperforms several baseline methods, achieving accuracy improvements of 2.5%, 2.63%, and 2.44% on LLaMA-7B, Vicuna-7B, and LLaMA-13B when the pruning ratio is 50%. Finally, SAAP is implemented on a testbed—NVIDIA Jetson AGX Orin 32GB Developer Kit. Test results demonstrate that compared to the foundation LLM, SAAP enhances the inference speed by 86.86% at a pruning ratio of 50%, highlighting its potential for practical deployment on resource-constrained IoT devices. Jinke Ren, Yatong Han, Yushan Sun, Ruichen Zhang 0001, Zhen Li 0026, Dusit Niyato, Shuguang Cui |
IEEE Internet Things J. | 7 |
| 2026 | EndoChat: Grounded multimodal large language model for endoscopic surgeryabstractRecently, Multimodal Large Language Models (MLLMs) have demonstrated their immense potential in computer-aided diagnosis and decision-making. In the context of robotic-assisted surgery, MLLMs can serve as effective tools for surgical training and guidance. However, there is still a deficiency of MLLMs specialized for surgical scene understanding in endoscopic procedures. To this end, we present EndoChat, an MLLM tailored to address various dialogue paradigms and subtasks in understanding endoscopic procedures. To train our EndoChat, we construct the Surg-396K dataset through a novel pipeline that systematically extracts surgical information and generates structured annotations based on large-scale endoscopic surgery datasets. Furthermore, we introduce a multi-scale visual token interaction mechanism and a visual contrast-based reasoning mechanism to enhance the model's representation learning and reasoning capabilities. Our model achieves state-of-the-art performance across five dialogue paradigms and seven surgical scene understanding tasks. Additionally, we conduct evaluations with professional surgeons, who provide positive feedback on the majority of conversation cases generated by EndoChat. Overall, these results demonstrate that EndoChat has the potential to advance training and automation in robotic-assisted surgery. Our dataset and model are publicly available at https://github.com/gkw0010/EndoChat. Guankun Wang, Long Bai 0008, Kun Yuan 0004, Zhen Li 0026, Tianxu Jiang, Xiting He, Jinlin Wu, Zhen Chen 0018, Zhen Lei 0001, Hongbin Liu 0001, Fan Zhang 0016, Nicolas Padoy, Nassir Navab, Hongliang Ren 0001 |
Medical Image Anal. | 5 |
| 2026 | Interpretable General Image Fusion via Scalable Autoregressive ModelingabstractExisting image fusion methods have developed increasingly sophisticated network architectures for exploiting modality-shared and modality-specific features. However, despite these advancements in feature extraction, most methods ultimately rely on relatively simple implicit or explicit fusion strategies, which can compromise interpretability and limit fusion accuracy. In this paper, we incorporate visual autoregressive modeling to bridge the gap between implicit feature extraction and explicit modality fusion. First, the proposed approach conducts a low-to-high resolution autoregressive objective with modality-specific features, introducing a scalable feature autoregressive mechanism. It aggregates local and global contextual dependencies while enhancing implicit cross-scale interaction. Furthermore, to promote the consistency and complementarity across modalities, we embed an explicit high-order fusion strategy within the progressive modality-specific feature extraction process. This integration facilitates a next-scale synergistic relationship between implicit learning and explicit fusion. Our High-order Feature AutoRegressive Fusion framework (HFARFusion) provides a robust and interpretable solution for general image fusion tasks, effectively balancing fusion performance and transparency through the strengths of autoregressive learning. Extensive experiments demonstrate the outstanding performance of the proposed method in several classical fusion tasks, including infrared-visible, medical, multi-focus, and multi-exposure image fusion. Our code is available at https://github.com/happysbn/HFARFusion. Jingwei Xin, Boneng Shi, Zhen Li 0026, Xuehao Song, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 3 |
| 2025 | VQA4CIR: Boosting Composed Image Retrieval with Visual Question AnsweringabstractAlbeit progress has been made in Composed Image Retrieval (CIR), we empirically find that a certain percentage of failure retrieval results are not consistent with their relative captions. To address this issue, this work provides a Visual Question Answering (VQA) perspective to boost the performance of CIR. The resulting VQA4CIR is a post-processing approach and can be directly plugged into existing CIR methods. Given the top-C retrieved images by a CIR method, VQA4CIR aims to decrease the adverse effect of the failure retrieval results being inconsistent with the relative caption. To find the retrieved images inconsistent with the relative caption, we resort to the "QA generation → VQA" self-verification pipeline. For QA generation, we suggest fine-tuning LLM (e.g., LLaMA) to generate several pairs of questions and answers from each relative caption. We then fine-tune LVLM (e.g., LLaVA) to obtain the VQA model. By feeding the retrieved image and question to the VQA model, one can find the images inconsistent with relative caption when the answer by VQA is inconsistent with the answer in the QA pair. Consequently, the CIR performance can be boosted by modifying the ranks of inconsistently retrieved images. Experimental results show that our proposed method outperforms state-of-the-art CIR methods on the CIRR and Fashion-IQ datasets. Chun-Mei Feng 0001, Yang Bai 0011, Tao Luo 0014, Zhen Li 0026, Salman Khan 0001, Wangmeng Zuo, Rick Siow Mong Goh, Yong Liu 0026 |
AAAI | 4 |
| 2025 | Consistency of Compositional Generalization Across Multiple LevelsabstractCompositional generalization is the capability of a model to understand novel compositions composed of seen concepts. There are multiple levels of novel compositions including phrase-phrase level, phrase-word level, and word-word level. Existing methods achieve promising compositional generalization, but the consistency of compositional generalization across multiple levels of novel compositions remains unexplored. The consistency refers to that a model should generalize to a phrase-phrase level novel composition, and phrase-word/word-word level novel compositions that can be derived from it simultaneously. In this paper, we propose a meta-learning based framework, for achieving consistent compositional generalization across multiple levels. The basic idea is to progressively learn compositions from simple to complex for consistency. Specifically, we divide the original training set into multiple validation sets based on compositional complexity, and introduce multiple meta-weight-nets to generate sample weights for samples in different validation sets. To fit the validation sets in order of increasing compositional complexity, we optimize the parameters of each meta-weight-net independently and sequentially in a multilevel optimization manner. We build a GQA-CCG dataset to quantitatively evaluate the consistency. Experimental results on visual question answering and temporal video grounding, demonstrate the effectiveness of the proposed framework. Chuanhao Li 0001, Zhen Li 0026, Chenchen Jing, Xiaomeng Fan, Wenbo Ye, Yuwei Wu 0001, Yunde Jia |
AAAI | 2 |
| 2025 | Topo2Seq: Enhanced Topology Reasoning via Topology Sequence LearningabstractExtracting lane topology from perspective views (PV) is crucial for planning and control in autonomous driving. This approach extracts potential drivable trajectories for self-driving vehicles without relying on high-definition (HD) maps. However, the unordered nature and weak long-range perception of the DETR-like framework can result in misaligned segment endpoints and limited topological prediction capabilities. Inspired by the learning of contextual relationships in language models, the connectivity relations in roads can be characterized as explicit topology sequences. In this paper, we introduce Topo2Seq, a novel approach for enhancing topology reasoning via topology sequences learning. The core concept of Topo2Seq is a randomized order prompt-to-sequence learning between lane segment decoder and topology sequence decoder. The dual-decoder branches simultaneously learn the lane topology sequences extracted from the Directed Acyclic Graph (DAG) and the lane graph containing geometric information. Randomized order prompt-to-sequence learning extracts unordered key points from the lane graph predicted by the lane segment decoder, which are then fed into the prompt design of the topology sequence decoder to reconstruct an ordered and complete lane graph. In this way, the lane segment decoder learns powerful long-range perception and accurate topological reasoning from the topology sequence decoder. Notably, topology sequence decoder is only introduced during training and does not affect the inference efficiency. Experimental evaluations on the OpenLane-V2 dataset demonstrate the state-of-the-art performance of Topo2Seq in topology reasoning. Yiming Yang 0001, Yueru Luo, Bingkun He, Erlong Li, Zhipeng Cao 0002, Chao Zheng 0004, Shuqi Mei, Zhen Li 0026 |
AAAI | 8 |
| 2025 | VesSAM: Efficient Multi-Prompting for Segmenting Complex VesselabstractPrecise vessel segmentation is vital for clinical applications such as diagnosis and surgical planning but remains challenging due to thin, branching geometries and low texture contrast. Although foundation models such as the Segment Anything Model (SAM) show strong performance in general segmentation tasks, they remain suboptimal for vascular structures. In this work, we present VesSAM, a powerful and efficient framework tailored for 2D vessel segmentation. VesSAM integrates three core modules: a convolutional adapter that enhances local texture features, a multi-prompt encoder that fuses anatomical cues via hierarchical cross-attention, and a lightweight mask decoder that reduces jagged artifacts. We also introduce an automated pipeline to generate structured multi-prompt annotations, and curate a diverse benchmark dataset spanning 8 datasets across 5 imaging modalities. Extensive experiments show that VesSAM surpasses state-of-the-art PEFT-based SAM variants by over$\text{1 0 \%}$Dice and 13% IoU, while maintaining competitive accuracy to fully fine-tuned methods with far fewer parameters. VesSAM also generalizes well to out-of-distribution (OoD) settings, outperforming all baselines in average OoD Dice and IoU. Suzhong Fu, Jingqi Dong, Yiming Yang 0001, Yao Zhu 0003, Min Chang Jordan Ren, Delin Deng, Angelica I. Avilés-Rivero, Shuguang Cui, Zhen Li 0026 |
BIBM | 11 |
| 2025 | VisionPAD: A Vision-Centric Pre-training Paradigm for Autonomous DrivingabstractThis paper introduces VisionPAD, a novel self-supervised pre-training paradigm designed for vision-centric algorithms in autonomous driving. In contrast to previous approaches that employ neural rendering with explicit depth supervision, VisionPAD utilizes more efficient 3D Gaussian Splatting to reconstruct multi-view representations using only images as supervision. Specifically, we introduce a self-supervised method for voxel velocity estimation. By warping voxels to adjacent frames and supervising the rendered outputs, the model effectively learns motion cues in the sequential data. Furthermore, we adopt a multi-frame photometric consistency approach to enhance geometric perception. It projects adjacent frames to the current frame based on rendered depths and relative poses, boosting the 3D geometric representation through pure image supervision. Extensive experiments on autonomous driving datasets demonstrate that VisionPAD significantly improves performance in 3D object detection, occupancy prediction and map segmentation, surpassing state-of-the-art pre-training strategies by a considerable margin. Haiming Zhang 0001, Wending Zhou, Yiyao Zhu, Xu Yan 0005, Jiantao Gao, Dongfeng Bai, Yingjie Cai, Shuguang Cui, Zhen Li 0026 |
CVPR | 10 |
| 2025 | DriveGEN: Generalized and Robust 3D Detection in Driving via Controllable Text-to-Image Diffusion GenerationabstractIn autonomous driving, vision-centric 3D detection aims to identify 3D objects from images. However, high data collection costs and diverse real-world scenarios limit the scale of training data. Once distribution shifts occur between training and test data, existing methods often suffer from performance degradation, known as Out-of-Distribution (OOD) problems. To address this, controllable Text-to-Image (T2I) diffusion offers a potential solution for training data enhancement, which is required to generate diverse OOD scenarios with precise 3D object geometry. Nevertheless, existing controllable T2I approaches are restricted by the limited scale of training data or struggle to preserve all annotated 3D objects. In this paper, we present DriveGEN, a method designed to improve the robustness of 3D detectors in Driving via Training-Free Controllable Text-to-Image Diffusion Generation. Without extra diffusion model training, DriveGEN consistently preserves objects with precise 3D geometry across diverse OOD generations, consisting of 2 stages: 1) Self-Prototype Extraction: We empirically find that self-attention features are semantic-aware but require accurate region selection for 3D objects. Thus, we extract precise object features via layouts to capture 3D object geometry, termed self-prototypes. 2) Prototype-Guided Diffusion: To preserve objects across various OOD scenarios, we perform semantic-aware feature alignment and shallow feature alignment during denoising. Extensive experiments demonstrate our effectiveness in improving 3D detection. The code is available at github.com/Hongbin98/DriveGEN. Zilu Guo, Yifan Zhang 0004, Shuaicheng Niu, Ruimao Zhang, Shuguang Cui, Zhen Li 0026 |
CVPR | 8 |
| 2025 | DSPNet: Dual-vision Scene Perception for Robust 3D Question Answeringabstract3D Question Answering (3D QA) requires the model to comprehensively understand its situated 3D scene described by the text, then reason about its surrounding environment and answer a question under that situation. However, existing methods usually rely on global scene perception from pure 3D point clouds and overlook the importance of rich local texture details from multi-view images. Moreover, due to the inherent noise in camera poses and complex occlusions, there exists significant feature degradation and reduced feature robustness problems when aligning 3D point cloud with multi-view images. In this paper, we propose a Dual-vision Scene Perception Network (DSPNet), to comprehensively integrate multi-view and point cloud features to improve robustness in 3D QA. Our Text-guided Multi-view Fusion (TGMF) module prioritizes image views that closely match the semantic content of the text. To adaptively fuse back-projected multi-view images with point cloud features, we design the Adaptive Dual-vision Perception (ADVP) module, enhancing 3D scene comprehension. Additionally, our Multimodal Context-guided Reasoning (MCGR) module facilitates robust reasoning by integrating contextual information across visual and linguistic modalities. Experimental results on SQA3D and ScanQA datasets demonstrate the superiority of our DSPNet. Codes will be available at https://github.com/LZ-CH/DSPNet. Jingzhou Luo, Yang Liu 0267, Zhen Li 0026, Yaowei Wang 0001, Guanbin Li, Liang Lin 0004 |
CVPR | 4 |
| 2025 | Empowering Large Language Models with 3D Situation AwarenessabstractDriven by the great success of Large Language Models (LLMs) in the 2D image domain, their application in 3D scene understanding has emerged as a new trend. A key difference between 3D and 2D is that the situation of an egocentric observer in 3D scenes can change, resulting in different descriptions (e.g., "left" or "right"). However, current LLM-based methods overlook the egocentric perspective and use datasets from a global viewpoint. To address this issue, we propose a novel approach to automatically generate a situation-aware dataset by leveraging the scanning trajectory during data collection and utilizing Vision-Language Models (VLMs) to produce high-quality captions and question-answer pairs. Furthermore, we introduce a situation grounding module to explicitly predict the position and orientation of the observer’s viewpoint, thereby enabling LLMs to ground situation descriptions in 3D scenes. We evaluate our approach on several benchmarks, demonstrating that our method effectively enhances the 3D situational awareness of LLMs while significantly expanding existing datasets and reducing manual effort. Zhihao Yuan, Yibo Peng, Jinke Ren, Yinghong Liao, Yatong Han, Chun-Mei Feng 0001, Hengshuang Zhao, Guanbin Li, Shuguang Cui, Zhen Li 0026 |
CVPR | 10 |
| 2025 | Lumina-Image 2.0: a Unified and Efficient Image Generative FrameworkabstractWe introduce Lumina-Image 2.0, an advanced text-to-image generation framework that achieves significant progress compared to previous work, Lumina-Next. Lumina-Image 2.0 is built upon two key principles: (1) Unification - it adopts a unified architecture (Unified Next-DiT) that treats text and image tokens as a joint sequence, enabling natural cross-modal interactions and allowing seamless task expansion. Besides, since high-quality captioners can provide semantically well-aligned text-image training pairs, we introduce a unified captioning system, Unified Captioner (UniCap), specifically designed for T2I generation tasks. UniCap excels at generating comprehensive and accurate captions, accelerating convergence and enhancing prompt adherence. (2) Efficiency - to improve the efficiency of our proposed model, we develop multi-stage progressive training strategies and introduce inference acceleration techniques without compromising image quality. Extensive evaluations on academic benchmarks and public text-to-image arenas show that Lumina-Image 2.0 delivers strong performances even with only 2.6B parameters, highlighting its scalability and design efficiency. We have released our training details, code, and models at https://github.com/Alpha-VLLM/Lumina-Image-2.0. Le Zhuo, Yi Xin 0003, Ruoyi Du, Zhen Li 0026, Yiting Lu, Xinyue Li 0001, Will Beddow, Erwann Millon, Victor Perez 0005, Wenhai Wang, Yu Qiao 0001, Bo Zhang 0069, Xiaohong Liu 0001, Hongsheng Li 0001, Chang Xu 0002, Peng Gao 0007 |
ICCV | 5 |
| 2025 | Advancing Dense Endoscopic Reconstruction with Gaussian Splatting-Driven Surface Normal-Aware Tracking and MappingabstractSimultaneous Localization and Mapping (SLAM) is essential for precise surgical interventions and robotic tasks in minimally invasive procedures. While recent advancements in 3D Gaussian Splatting (3DGS) have improved SLAM with high-quality novel view synthesis and fast rendering, these systems struggle with accurate depth and surface reconstruction due to multi-view inconsistencies. Simply incorporating SLAM and 3DGS leads to mismatches between the reconstructed frames. In this work, we present Endo-2DTAM, a real-time endoscopic SLAM system with 2D Gaussian Splatting (2DGS) to address these challenges. Endo-2DTAM incorporates a surface normal-aware pipeline, which consists of tracking, mapping, and bundle adjustment modules for geometrically accurate reconstruction. Our robust tracking module combines point-topoint and point-to-plane distance metrics, while the mapping module utilizes normal consistency and depth distortion to enhance surface reconstruction quality. We also introduce a pose-consistent strategy for efficient and geometrically coherent keyframe sampling. Extensive experiments on public endoscopic datasets demonstrate that Endo-2DTAM achieves an RMSE of$1.87 \pm 0.63 \mathbf{m m}$for depth reconstruction of surgical scenes while maintaining computationally efficient tracking, high-quality visual appearance, and real-time rendering. Our code will be released at github.com/lastbasket/Endo-2DTAM. Yiming Huang 0007, Beilei Cui, Long Bai 0008, Zhen Chen 0018, Jinlin Wu, Zhen Li 0026, Hongbin Liu 0001, Hongliang Ren 0001 |
ICRA | 6 |
| 2025 | ETSM: Automating Dissection Trajectory Suggestion and Confidence Map-Based Safety Margin Prediction for Robot-Assisted Endoscopic Submucosal DissectionabstractRobot-assisted Endoscopic Submucosal Dissection (ESD) improves the surgical procedure by providing a more comprehensive view through advanced robotic instruments and bimanual operation, thereby enhancing dissection efficiency and accuracy. Accurate prediction of dissection trajectories is crucial for better decision-making, reducing intraoperative errors, and improving surgical training. Nevertheless, predicting these trajectories is challenging due to variable tumor margins and dynamic visual conditions. To address this issue, we create the ESD Trajectory and Confidence Map-based Safety Margin (ETSM) dataset with 1849 short clips, focusing on submucosal dissection with a dual-arm robotic system. We also introduce a framework that combines optimal dissection trajectory prediction with a confidence map-based safety margin, providing a more secure and intelligent decision-making tool to minimize surgical risks for ESD procedures. Additionally, we propose the Regression-based Confidence Map Prediction Network (RCMNet), which utilizes a regression approach to predict confidence maps for dissection areas, thereby delineating various levels of safety margins. We evaluate our RCMNet using three distinct experimental setups: in-domain evaluation, robustness assessment, and out-of-domain evaluation. Experimental results show that our approach excels in the confidence map-based safety margin prediction task, achieving a mean absolute error (MAE) of only 3.18. To the best of our knowledge, this is the first study to apply a regression approach for visual guidance concerning delineating varying safety levels of dissection areas. Our approach bridges gaps in current research by improving prediction accuracy and enhancing the safety of the dissection process, showing great clinical significance in practice. The dataset and code are available at https://github.com/FrankMOWJ/RCMNet. Mengya Xu, Wenjin Mo, Guankun Wang, Huxin Gao, An Wang 0007, Long Bai 0008, Chaoyang Lyu, Xiaoxiao Yang, Zhen Li 0026, Hongliang Ren 0001 |
ICRA | 9 |
| 2025 | Multi-Sourced Compositional Generalization in Visual Question AnsweringabstractCompositional generalization is the ability of generalizing novel compositions from seen primitives, and has received much attention in vision-and-language (V&L) recently. Due to the multi-modal nature of V&L tasks, the primitives composing compositions source from different modalities, resulting in multi-sourced novel compositions. However, the generalization ability over multi-sourced novel compositions, i.e., multi-sourced compositional generalization (MSCG) remains unexplored. In this paper, we explore MSCG in the context of visual question answering (VQA), and propose a retrieval-augmented training framework to enhance the MSCG ability of VQA models by learning unified representations for primitives from different modalities. Specifically, semantically equivalent primitives are retrieved for each primitive in the training samples, and the retrieved features are aggregated with the original primitive to refine the model. This process helps the model learn consistent representations for the same semantic primitives across different modalities. To evaluate the MSCG ability of VQA models, we construct a new GQA-MSCG dataset based on the GQA dataset, in which samples include three types of novel compositions composed of primitives from different modalities. The GQA-MSCG dataset is available at https://github.com/NeverMoreLCH/MSCG. Chuanhao Li 0001, Wenbo Ye, Zhen Li 0026, Yuwei Wu 0001, Yunde Jia |
IJCAI | 3 |
| 2025 | CLEA: Closed-Loop Embodied Agent for Enhancing Task Execution in Dynamic EnvironmentsabstractLarge Language Models (LLMs) exhibit remarkable capabilities in the hierarchical decomposition of complex tasks through semantic reasoning. However, their application in embodied systems faces challenges in ensuring reliable execution of subtask sequences and achieving one-shot success in long-term task completion. To address these limitations in dynamic environments, we propose Closed-Loop Embodied Agent (CLEA)—a novel architecture incorporating four specialized open-source LLMs with functional decoupling for closed-loop task management. The framework features two core innovations: (1) Interactive task planner that dynamically generates executable subtasks based on the environmental memory, and (2) Multimodal execution critic employing an evaluation framework to conduct a probabilistic assessment of action feasibility, triggering hierarchical re-planning mechanisms when environmental perturbations exceed preset thresholds. To validate CLEA’s effectiveness, we conduct experiments in a real environment with manipulable objects, using two heterogeneous robots for object search, manipulation, and search-manipulation integration tasks. Across 12 task trials, CLEA outperforms the baseline model, achieving a 67.3% improvement in success rate and a 52.8% increase in task completion rate. These results demonstrate that CLEA significantly enhances the robustness of task planning and execution in dynamic environments. Our code is available at https://sp4595.github.io/CLEA/. Mingcong Lei, Ge Wang 0007, Zhixin Mai, Yao Guo 0002, Zhen Li 0026, Shuguang Cui, Yatong Han, Jinke Ren |
IROS | 7 |
| 2025 | CoPESD: A Multi-Level Surgical Motion Dataset for Training Large Vision-Language Models to Co-Pilot Endoscopic Submucosal Dissection
Guankun Wang, Han Xiao 0010, Renrui Zhang, Huxin Gao, Long Bai 0008, Xiaoxiao Yang, Zhen Li 0026, Hongsheng Li 0001, Hongliang Ren 0001 |
ACM Multimedia | 7 |
| 2025 | Sekai: A Video Dataset towards World ExplorationabstractVideo generation techniques have made remarkable progress, promising to be the foundation of interactive world exploration.However, existing video generation datasets are not well-suited for world exploration training as they suffer from some limitations: limited locations, short duration, static scenes, and a lack of annotations about exploration and the world.In this paper, we introduce Sekai (meaning "world" in Japanese), a high-quality first-person view worldwide video dataset with rich annotations for world exploration. It consists of over 5,000 hours of walking or drone view (FPV and UVA) videos from over 100 countries and regions across 750 cities. We develop an efficient and effective toolbox to collect, pre-process and annotate videos with location, scene, weather, crowd density, captions, and camera trajectories.Comprehensive analyses and experiments demonstrate the dataset’s scale, diversity, annotation quality, and effectiveness for training video generation models.We believe Sekai will benefit the area of video generation and world exploration, and motivate valuable applications. Zhen Li 0026, Chuanhao Li 0001, Xiaofeng Mao, Shaoheng Lin, Ming Li 0010, Shitian Zhao, Zhaopan Xu, Xinyue Li 0001, Yukang Feng, Zizhen Li, Fanrui Zhang, Jiaxin Ai, Yuwei Wu 0001, Tong He 0001, Yunde Jia, Kaipeng Zhang |
NeurIPS | 1 |
| 2025 | AffordBot: 3D Fine-grained Embodied Reasoning via Multimodal Large Language ModelsabstractEffective human-agent collaboration in physical environments requires understanding not only what to act upon, but also where the actionable elements are and how to interact with them. Existing approaches often operate at the object level or disjointedly handle fine-grained affordance reasoning, lacking coherent, instruction-driven grounding and reasoning. In this work, we introduce a new task: Fine-grained 3D Embodied Reasoning, which requires an agent to predict, for each referenced affordance element in a 3D scene, a structured triplet comprising its spatial location, motion type, and motion axis, based on a task instruction. To solve this task, we propose AffordBot, a novel framework that integrates Multimodal Large Language Models (MLLMs) with a tailored chain-of-thought (CoT) reasoning paradigm. To bridge the gap between 3D input and 2D-compatible MLLMs, we render surround-view images of the scene and project 3D element candidates into these views, forming a rich visual representation aligned with the scene geometry. Our CoT pipeline begins with an active perception stage, prompting the MLLM to select the most informative viewpoint based on the instruction, before proceeding with step-by-step reasoning to localize affordance elements and infer plausible interaction motions. Evaluated on the SceneFun3D dataset, AffordBot achieves state-of-the-art performance, demonstrating strong generalization and physically grounded reasoning with only 3D point cloud input and MLLMs. Our code is available at [https://github.com/hannahwxy/AffordBot](https://github.com/hannahwxy/AffordBot). Xun Yang 0001, Yanlong Xu, Zhen Li 0026, Na Zhao 0004 |
NeurIPS | 5 |
| 2025 | SQS: Enhancing Sparse Perception Models via Query-based Splatting in Autonomous DrivingabstractSparse Perception Models (SPMs) adopt a query-driven paradigm that forgoes explicit dense BEV or volumetric construction, enabling highly efficient computation and accelerated inference. In this paper, we introduce SQS, a novel query-based splatting pre-training specifically designed to advance SPMs in autonomous driving. SQS introduces a plug-in module that predicts 3D Gaussian representations from sparse queries during pre-training, leveraging self-supervised splatting to learn fine-grained contextual features through the reconstruction of multi-view images and depth maps. During fine-tuning, the pre-trained Gaussian queries are seamlessly integrated into downstream networks via query interaction mechanisms that explicitly connect pre-trained queries with task-specific queries, effectively accommodating the diverse requirements of occupancy prediction and 3D object detection. Extensive experiments on autonomous driving benchmarks demonstrate that SQS delivers considerable performance gains across multiple query-based 3D perception tasks, notably in occupancy prediction and 3D object detection, outperforming prior state-of-the-art pre-training approaches by a significant margin (i.e., +1.3 mIoU on occupancy prediction and +1.0 NDS on 3D detection). Haiming Zhang 0001, Yiyao Zhu, Wending Zhou, Xu Yan 0005, Yingjie Cai, Shuguang Cui, Zhen Li 0026 |
NeurIPS | 8 |
| 2025 | Diffusion-Enhanced Test-Time Adaptation with Text and Image Augmentation
Chun-Mei Feng 0001, Yuanyang He, Jian Zou 0005, Salman Khan 0001, Huan Xiong, Zhen Li 0026, Wangmeng Zuo, Rick Siow Mong Goh, Yong Liu 0026 |
Int. J. Comput. Vis. | 6 |
| 2025 | Self-randomized focuses effectively boost metric-based few-shot classifiers
Zhen Li 0026, Zhongyuan Liu, Dongliang Chang, Aneeshan Sain, Zhanyu Ma, Jing-Hao Xue, Yi-Zhe Song |
Pattern Recognit. | 1 |
| 2025 | V²-SfMLearner: Learning Monocular Depth and Ego-Motion for Multimodal Wireless Capsule EndoscopyabstractDeep learning can predict depth maps and capsule ego-motion from capsule endoscopy videos, aiding in 3D scene reconstruction and lesion localization. However, the collisions of the capsule endoscopies within the gastrointestinal tract cause vibration perturbations in the training data. Existing solutions focus solely on vision-based processing, neglecting other auxiliary signals like vibrations that could reduce noise and improve performance. Therefore, we propose V2-SfMLearner, a multimodal approach integrating vibration signals into vision-based depth and capsule motion estimation for monocular capsule endoscopy. We construct a multimodal capsule endoscopy dataset containing vibration and visual signals, and our artificial intelligence solution develops an unsupervised method using vision-vibration signals, effectively eliminating vibration perturbations through multimodal learning. Specifically, we carefully design a vibration network branch and a Fourier fusion module, to detect and mitigate vibration noises. The fusion framework is compatible with popular vision-only algorithms. Extensive validation on the multimodal dataset demonstrates superior performance and robustness against vision-only algorithms. Without the need for large external equipment, our V2-SfMLearner has the potential for integration into clinical capsule robots, providing real-time and dependable digestive examination tools. The findings show promise for practical implementation in clinical settings, enhancing the diagnostic capabilities of doctors. Note to Practitioners—This paper is motivated by the problem of estimating the depth and ego-motion information for the wireless capsule endoscopy in the human gastrointestinal tract to realize accurate, efficient, robust, and real-time inspection. Our estimation method does not engage any external localization equipment. Instead, inspired by the existing research on integrating capsule endoscopy and inertial measurement units, we introduce vibration signals into vision-based depth and ego-motion estimation approaches, improving the accuracy and robustness of the estimation results based on multimodal learning methods. Research on capsule robots or computer vision can readily be combined with our framework for various clinical and industrial applications. Long Bai 0008, Beilei Cui, Yanheng Li 0002, Shilong Yao, Sishen Yuan, Yanan Wu 0003, Yang Zhang 0053, Max Q.-H. Meng, Zhen Li 0026, Weiping Ding 0001, Hongliang Ren 0001 |
IEEE Trans Autom. Sci. Eng. | 10 |
| 2025 | Highlighted Diffusion Model as Plug-In Priors for Polyp SegmentationabstractAutomated polyp segmentation from colonoscopy images is crucial for colorectal cancer diagnosis. The accuracy of such segmentation, however, is challenged by two main factors. First, the variability in polyps' size, shape, and color, coupled with the scarcity of well-annotated data due to the need for specialized manual annotation, hampers the efficacy of existing deep learning methods. Second, concealed polyps often blend with adjacent intestinal tissues, leading to poor contrast that challenges segmentation models. Recently, diffusion models have been explored and adapted for polyp segmentation tasks. However, the significant domain gap between RGB-colonoscopy images and grayscale segmentation masks, along with the low efficiency of the diffusion generation process, hinders the practical implementation of these models. To mitigate these challenges, we introduce the Highlighted Diffusion Model Plus (HDM+), a two-stage polyp segmentation framework. This framework incorporates the Highlighted Diffusion Model (HDM) to provide explicit semantic guidance, thereby enhancing segmentation accuracy. In the initial stage, the HDM is trained using highlighted ground-truth data, which emphasizes polyp regions while suppressing the background in the images. This approach reduces the domain gap by focusing on the image itself rather than on the segmentation mask. In the subsequent second stage, we employ the highlighted features from the trained HDM's U-Net model as plug-in priors for polyp segmentation, rather than generating highlighted images, thereby increasing efficiency. Extensive experiments conducted on six polyp segmentation benchmarks demonstrate the effectiveness of our approach. Yuncheng Jiang 0002, Shuangyi Tan, Si-Qi Liu 0003, Zhen Li 0026, Guanbin Li |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | Boosting 3D Object Detection via Self-Distilling Introspective Dataabstract3D object detection is a fundamental yet critical task for autonomous driving. In this paper, we investigate a novel self-distilling paradigm by proposing Self-distilling Introspective Data (SID) to boost the accuracy of 3D object detection in both LiDAR-based and LiDAR-Camera-based scenarios. The proposed SID significantly improves the applicability of the distillation approach since it does not require extra training data or complex teacher network design. Specifically, we first employ an introspective data augmentation method to enrich object-aware information in sparse point clouds through geometric or semantic injection. We then utilize this enhanced data to train a robust teacher model. In contrast to traditional distillation that relies on larger models to enhance the representations of smaller ones, the teacher model in SID shares the same architecture as the student model but exhibits exceptionally high discriminative ability. This enables the effective transfer of rich feature representations to the student model. Rooted on such a scheme, when conducting LiDAR-based detectors, SID significantly enhances the semantic representation capabilities of sparse point clouds. Additionally, in the LiDAR-Camera-based setting, SID also effectively supervises the fusion of the two modalities at the feature level, ensuring more reasonable cross-modal learning. Extensive experiments show the proposed SID improves a variety of detectors. For the LiDAR-based detector, the SID gains 2.31% mAP improvements for the hard objects in KITTI, while 1.76% NDS improvements on nuScenes. For the LiDAR-Camera-based detectors, the SID boosts the detection accuracy significantly, with 1.5% mAP promotion on KITTI and 2.15% NDS improvements on the nuScenes benchmark. Chaoqun Wang 0012, Yiran Qin, Zijian Kang, Ningning Ma, Yukai Shi, Zhen Li 0026, Ruimao Zhang |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2025 | Toward Fine-Grained 3-D Visual Grounding Through Referring Textual PhrasesabstractRecent progress in 3-D scene understanding has explored visual grounding [3D visual grounding (3DVG)] to localize a target object through a language description. However, existing methods only consider the dependency between the entire sentence and the target object, ignoring fine-grained relationships between contexts and nontarget ones. In this article, we extend 3DVG to a more fine-grained task, called 3D phrase-aware grounding (3DPAG). The 3DPAG task aims to localize the target objects in a 3-D scene by explicitly identifying all phrase-related objects and then conducting the reasoning according to contextual phrases. To tackle this problem, we manually labeled about 227 K phrase-level annotations using a self-developed platform, from 88 K sentences of widely used 3DVG datasets, i.e., Natural Reference in 3-D (Nr3D), Spatial Reference in 3-D (Sr3D), and ScanRefer. By tapping on our datasets, we can extend previous 3DVG methods to the fine-grained phrase-aware scenario. It is achieved through the proposed novel phrase-object alignment (POA) optimization and phrase-specific pretraining (PSP), boosting conventional 3DVG performance as well. Extensive results confirm significant improvements, i.e., previous state-of-the-art method achieves 3.9%, 3.5%, and 4.6% overall accuracy gains on Nr3D, Sr3D, and ScanRefer, respectively. Our datasets and platform are released in https://github.com/CurryYuan/PhraseRefer. Zhihao Yuan, Xu Yan 0005, Xuhao Li, Yao Guo 0002, Shuguang Cui, Zhen Li 0026 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2025 | GSmoothFace: Generalized Smooth Talking Face Generation via Fine Grained 3D Face GuidanceabstractAlthough existing speech-driven talking face generation methods achieve significant progress, they are far from real-world application due to the avatar-specific training demand and unstable lip movements. To address the above issues, we propose the GSmoothFace, a novel two-stage generalized talking face generation model guided by a fine-grained 3D face model, which can synthesize smooth lip dynamics while preserving the speaker's identity. Our proposed GSmoothFace model mainly consists of the Audio to Expression Prediction (A2EP) module and the Target Adaptive Face Translation (TAFT) module. Specifically, we first develop the A2EP module to predict expression parameters synchronized with the driven speech. It uses a transformer to capture the long-term audio context and learns the parameters from the fine-grained 3D facial vertices, resulting in accurate and smooth lip-synchronization performance. Afterward, the well-designed TAFT module, empowered by Morphology Augmented Face Blending (MAFB), takes the predicted expression parameters and target video as inputs to modify the facial region of the target video without distorting the background content. The TAFT effectively exploits the identity appearance and background context in the target video, which makes it possible to generalize to different speakers without retraining. Both quantitative and qualitative experiments confirm the superiority of our method in terms of realism, lip-synchronization, and visual quality. Haiming Zhang 0001, Zhihao Yuan, Chaoda Zheng, Xu Yan 0005, Baoyuan Wang, Guanbin Li, Shuguang Cui, Zhen Li 0026 |
IEEE Trans. Vis. Comput. Graph. | 9 |
| 2024 | CrossBind: Collaborative Cross-Modal Identification of Protein Nucleic-Acid-Binding ResiduesabstractAccurate identification of protein nucleic acid binding residues poses a significant challenge with important implications for various biological processes and drug design. Many typical computational methods for protein analysis rely on a single model that could ignore either the semantic context of the protein or the global 3D geometric information. Consequently, these approaches may result in incomplete or inaccurate protein analysis. To address the above issue, in this paper, we present CrossBind, a novel collaborative cross modal approach for identifying binding residues by exploiting both protein geometric structure and its sequence prior knowledge extracted from a large scale protein language model. Specifically, our multi modal approach leverages a contrastive learning technique and atom wise attention to capture the positional relationships between atoms and residues, thereby incorporating fine grained local geometric knowledge, for better binding residue prediction. Extensive experimental results demonstrate that our approach outperforms the next best state of the art methods, GraphSite and GraphBind, on DNA and RNA datasets by 10.8/17.3% in terms of the harmonic mean of precision and recall (F1 Score) and 11.9/24.8% in Matthews correlation coefficient (MCC), respectively. We release the code at https://github.com/BEAM-Labs/CrossBind. Linglin Jing, Yifan Wang 0008, Zhigang Ji, Hui Fang 0003, Zhen Li 0026 |
AAAI | 8 |
| 2024 | X4D-SceneFormer: Enhanced Scene Understanding on 4D Point Cloud Videos through Cross-Modal Knowledge TransferabstractThe field of 4D point cloud understanding is rapidly developing with the goal of analyzing dynamic 3D point cloud sequences. However, it remains a challenging task due to the sparsity and lack of texture in point clouds. Moreover, the irregularity of point cloud poses a difficulty in aligning temporal information within video sequences. To address these issues, we propose a novel cross-modal knowledge transfer framework, called X4D-SceneFormer. This framework enhances 4D-Scene understanding by transferring texture priors from RGB sequences using a Transformer architecture with temporal relationship mining. Specifically, the framework is designed with a dual-branch architecture, consisting of an 4D point cloud transformer and a Gradient-aware Image Transformer (GIT). The GIT combines visual texture and temporal correlation features to offer rich semantics and dynamics for better point cloud representation. During training, we employ multiple knowledge transfer techniques, including temporal consistency losses and masked self-attention, to strengthen the knowledge transfer between modalities. This leads to enhanced performance during inference using single-modal 4D point cloud inputs. Extensive experiments demonstrate the superior performance of our framework on various 4D point cloud video understanding tasks, including action recognition, action segmentation and semantic segmentation. The results achieve 1st places, i.e., 85.3% (+7.9%) accuracy and 47.3% (+5.0%) mIoU for 4D action segmentation and semantic segmentation, on the HOI4D challenge, outperforming previous state-of-the-art by a large margin. We release the code at https://github.com/jinglinglingling/X4D. Linglin Jing, Ying Xue 0003, Xu Yan 0005, Chaoda Zheng, Dong Wang 0028, Ruimao Zhang, Zhigang Wang 0002, Hui Fang 0003, Bin Zhao 0001, Zhen Li 0026 |
AAAI | 10 |
| 2024 | WeakPCSOD: Overcoming the Bias of Box Annotations for Weakly Supervised Point Cloud Salient Object DetectionabstractPoint cloud salient object detection (PCSOD) is a newly proposed task in 3D dense segmentation. However, the acquisition of accurate 3D dense annotations comes at a high cost, severely limiting the progress of PCSOD. To address this issue, we propose the first weakly supervised PCSOD (named WeakPCSOD) model, which relies solely on cheap 3D bounding box annotations. In WeakPCSOD, we extract noise-free supervision from coarse 3D bounding boxes while mitigating shape biases inherent in box annotations. To achieve this, we introduce a novel mask-to-box (M2B) transformation and a color consistency (CC) loss. The M2B transformation, from a shape perspective, disentangles predictions from labels, enabling the extraction of noiseless supervision from labels while preserving object shapes independently of the box bias. From an appearance perspective, we further introduce the CC loss to provide dense supervision, which mitigates the non-unique predictions stemming from weak supervision and substantially reduces prediction variability. Furthermore, we employ a self-training (ST) strategy to enhance performance by utilizing high-confidence pseudo labels. Notably, the M2B transformation, CC loss, and ST strategy are seamlessly integrated into any model and incur no computational costs for inference. Extensive experiments demonstrate the effectiveness of our WeakPCSOD model, even comparable to fully supervised models utilizing dense annotations. Jun Wei 0006, Shaohua Kevin Zhou, Shuguang Cui, Zhen Li 0026 |
AAAI | 4 |
| 2024 | RadOcc: Learning Cross-Modality Occupancy Knowledge through Rendering Assisted Distillationabstract3D occupancy prediction is an emerging task that aims to estimate the occupancy states and semantics of 3D scenes using multi-view images. However, image-based scene perception encounters significant challenges in achieving accurate prediction due to the absence of geometric priors. In this paper, we address this issue by exploring cross-modal knowledge distillation in this task, i.e., we leverage a stronger multi-modal model to guide the visual model during training. In practice, we observe that directly applying features or logits alignment, proposed and widely used in bird's-eye-view (BEV) perception, does not yield satisfactory results. To overcome this problem, we introduce RadOcc, a Rendering assisted distillation paradigm for 3D Occupancy prediction. By employing differentiable volume rendering, we generate depth and semantic maps in perspective views and propose two novel consistency criteria between the rendered outputs of teacher and student models. Specifically, the depth consistency loss aligns the termination distributions of the rendered rays, while the semantic consistency loss mimics the intra-segment similarity guided by vision foundation models (VLMs). Experimental results on the nuScenes dataset demonstrate the effectiveness of our proposed method in improving various 3D occupancy prediction approaches, e.g., our proposed methodology enhances our baseline by 2.2% in the metric of mIoU and achieves 50% in Occ3D benchmark. Haiming Zhang 0001, Xu Yan 0005, Dongfeng Bai, Jiantao Gao, Shuguang Cui, Zhen Li 0026 |
AAAI | 8 |
| 2024 | MixPolyp: Integrating Mask, Box and Scribble Supervision for Enhanced Polyp SegmentationabstractLimited by the expensive labeling, polyp segmentation models are plagued by data shortages. To tackle this, we propose the mixed supervised polyp segmentation paradigm (MixPolyp). Unlike traditional models relying on a single type of annotation, MixPolyp combines diverse annotation types (mask, box, and scribble) within a single model, thereby expanding the range of available data and reducing labeling costs. To achieve this, MixPolyp introduces three novel supervision losses to handle various annotations: Subspace Projection loss $\left({{{\mathcal{L}}_{{\mathcal{S}}{\mathcal{P}}}}}\right)$, Binary Minimum Entropy loss $\left({{{\mathcal{L}}_{{\mathcal{B}}{\mathcal{M}}{\mathcal{E}}}}}\right)$, and Linear Regularization loss $\left({{{\mathcal{L}}_{{\mathcal{L}}{\mathcal{R}}}}}\right)$. For box annotations, ${{\mathcal{L}}_{{\mathcal{S}}{\mathcal{P}}}}$ eliminates shape inconsistencies between the prediction and the supervision. For scribble annotations, ${{\mathcal{L}}_{{\mathcal{B}}{\mathcal{M}}{\mathcal{E}}}}$ provides supervision for unlabeled pixels through minimum entropy constraint, thereby alleviating supervision sparsity. Furthermore, ${{\mathcal{L}}_{{\mathcal{L}}{\mathcal{R}}}}$ provides dense supervision by enforcing consistency among the predictions, thus reducing the non-uniqueness. These losses are independent of the model structure, making them generally applicable. They are used only during training, adding no computational cost during inference. Extensive experiments on five datasets demonstrate MixPolyp’s effectiveness. Yiwen Hu 0001, Jun Wei 0006, Yuncheng Jiang 0002, Shuguang Cui, Zhen Li 0026 |
BIBM | 6 |
| 2024 | Let Video Teaches You More: Video-to-Image Knowledge Distillation using Detection TRansformer for Medical Video Lesion DetectionabstractAI-assisted lesion detection models play a crucial role in the early screening of cancer. However, previous image-based models ignore the inter-frame contextual information present in videos. On the other hand, video-based models capture the inter-frame context but are computationally expensive. To mitigate this contradiction, we delve into Video-to-Image knowledge distillation leveraging DEtection TRansformer (V2I-DETR) for the task of medical video lesion detection. V2I-DETR adopts a teacher-student network paradigm. The teacher network aims at extracting temporal contexts from multiple frames and transferring them to the student network, and the student network is an image-based model dedicated to fast prediction in inference. By distilling multi-frame contexts into a single frame, the proposed V2I-DETR combines the advantages of utilizing temporal contexts from video-based models and the inference speed of image-based models. Through extensive experiments, V2I-DETR outperforms previous state-of-the-art methods by a large margin while achieving the real-time inference speed (30 FPS) as the image-based model. Yuncheng Jiang 0002, Zixun Zhang, Jun Wei 0006, Chun-Mei Feng 0001, Guanbin Li, Shuguang Cui, Zhen Li 0026 |
BIBM | 8 |
| 2024 | Visual Programming for Zero-Shot Open-Vocabulary 3D Visual Groundingabstract3D Visual Grounding (3DVG) aims at localizing 3D object based on textual descriptions. Conventional supervised methods for 3DVG often necessitate extensive annotations and a predefined vocabulary, which can be restrictive. To address this issue, we propose a novel visual programming approach for zero-shot open-vocabulary 3DVG, leveraging the capabilities of large language models (LLMs). Our approach begins with a unique dialog-based method, engaging with LLMs to establish a foundational understanding of zero-shot 3DVG. Building on this, we design a visual program that consists of three types of modules, i.e., view-independent, view-dependent, and functional modules. These modules, specifically tailored for 3D scenarios, work collaboratively to perform complex reasoning and inference. Furthermore, we develop an innovative language-object correlation module to extend the scope of existing 3D object detectors into open-vocabulary scenarios. Extensive experiments demonstrate that our zero-shot approach can outperform some supervised baselines, marking a significant stride towards effective 3DVG. Code is available at https://curryyuan.github.io/Z5VG3D. Zhihao Yuan, Jinke Ren, Chun-Mei Feng 0001, Hengshuang Zhao, Shuguang Cui, Zhen Li 0026 |
CVPR | 6 |
| 2024 | Compositional Substitutivity of Visual Reasoning for Visual Question Answering
Chuanhao Li 0001, Zhen Li 0026, Chenchen Jing, Yuwei Wu 0001, Mingliang Zhai, Yunde Jia |
ECCV (48) | 2 |
| 2024 | MonoTTA: Fully Test-Time Adaptation for Monocular 3D Object Detection
Yifan Zhang 0004, Shuaicheng Niu, Shuguang Cui, Zhen Li 0026 |
ECCV (44) | 5 |
| 2024 | In-Context Compositional Generalization for Large Vision-Language ModelsabstractRecent work has revealed that in-context learning for large language models exhibits compositional generalization capacity, which can be enhanced by selecting in-context demonstrations similar to test cases to provide contextual information.However, how to exhibit in-context compositional generalization (ICCG) of large vision-language models (LVLMs) is non-trival.Due to the inherent asymmetry between visual and linguistic modalities, ICCG in LVLMs faces an inevitable challenge-redundant information on the visual modality.The redundant information affects in-context learning from two aspects: (1) Similarity calculation may be dominated by redundant information, resulting in sub-optimal demonstration selection.(2) Redundant information in in-context demonstrations brings misleading contextual information to in-context learning.To alleviate these problems, we propose a demonstration selection method to achieve ICCG for LVLMs, by considering two key factors of demonstrations: content and structure, from a multimodal perspective.Specifically, we design a diversity-coverage-based matching score to select demonstrations with maximum coverage, and avoid selecting demonstrations with redundant information via their content redundancy and structural complexity.We build a GQA-ICCG dataset to simulate the ICCG setting, and conduct experiments on GQA-ICCG and the VQA v2 dataset.Experimental results demonstrate the effectiveness of our method. Chuanhao Li 0001, Chenchen Jing, Zhen Li 0026, Mingliang Zhai, Yuwei Wu 0001, Yunde Jia |
EMNLP | 3 |
| 2024 | DV-3DLane: End-to-end Multi-modal 3D Lane Detection with Dual-view RepresentationabstractAccurate 3D lane estimation is crucial for ensuring safety in autonomous driving. However, prevailing monocular techniques suffer from depth loss and lighting variations, hampering accurate 3D lane detection. In contrast, LiDAR points offer geometric cues and enable precise localization. In this paper, we present DV-3DLane, a novel end-to-end **D**ual-**V**iew multi-modal **3D Lane** detection framework that synergizes the strengths of both images and LiDAR points. We propose to learn multi-modal features in dual-view spaces, *i.e.*, *perspective view* (PV) and *bird's-eye-view* (BEV), effectively leveraging the modal-specific information. To achieve this, we introduce three designs: 1) A bidirectional feature fusion strategy that integrates multi-modal features into each view space, exploiting their unique strengths. 2) A unified query generation approach that leverages lane-aware knowledge from both PV and BEV spaces to generate queries. 3) A 3D dual-view deformable attention mechanism, which aggregates discriminative features from both PV and BEV spaces into queries for accurate 3D lane detection. Extensive experiments on the public benchmark, OpenLane, demonstrate the efficacy and efficiency of DV-3DLane. It achieves state-of-the-art performance, with a remarkable 11.2 gain in F1 score and a substantial 53.5% reduction in errors. Code is available on [github](https://github.com/JMoonr/dv-3dlane). Yueru Luo, Shuguang Cui, Zhen Li 0026 |
ICLR | 3 |
| 2024 | Unified Generation, Reconstruction, and Representation: Generalized Diffusion with Adaptive Latent Encoding-DecodingabstractThe vast applications of deep generative models are anchored in three core capabilities---*generating* new instances, *reconstructing* inputs, and learning compact *representations*---across various data types, such as discrete text/protein sequences and continuous images. Existing model families, like variational autoencoders (VAEs), generative adversarial networks (GANs), autoregressive models, and (latent) diffusion models, generally excel in specific capabilities and data types but fall short in others. We introduce *Generalized* ***E****ncoding*-***D****ecoding ****D****iffusion ****P****robabilistic ****M****odels* (EDDPMs) which integrate the core capabilities for broad applicability and enhanced performance. EDDPMs generalize the Gaussian noising-denoising in standard diffusion by introducing parameterized encoding-decoding. Crucially, EDDPMs are compatible with the well-established diffusion model objective and training recipes, allowing effective learning of the encoder-decoder parameters *jointly* with diffusion. By choosing appropriate encoder/decoder (e.g., large language models), EDDPMs naturally apply to different data types. Extensive experiments on text, proteins, and images demonstrate the flexibility to handle diverse data and tasks and the strong improvement over various existing models. Code is available at https://github.com/guangyliu/EDDPM . Guangyi Liu 0005, Yu Wang 0170, Zeyu Feng, Qiyu Wu 0001, Zhen Li 0026, Shuguang Cui, Julian J. McAuley, Eric P. Xing, Zhiting Hu |
ICML | 7 |
| 2024 | Chained Flexible Capsule Endoscope: Unraveling the Conundrum of Size Limitations and Functional Integration for Gastrointestinal TransitivityabstractCapsule endoscopes, predominantly serving diagnostic functions, provide lucid internal imagery but are devoid of surgical or therapeutic capabilities. Consequently, despite lesion detection, physicians frequently resort to traditional endoscopic or open surgical procedures for treatment, resulting in more complex, potentially risky interventions. To surmount these limitations, this study introduces a chained flexible capsule endoscope (FCE) design concept, specifically conceived to navigate the inherent volume constraints of capsule endoscopes whilst augmenting their therapeutic functionalities. The FCE’s distinctive flexibility originates from a conventional rotating joint design and the incision pattern in the flexible material. In vitro experiments validated the passive navigation ability of the FCE in rugged intestinal tracts. Further, the FCE demonstrates consistent reptile-like peristalsis under the influence of an external magnetic field, and possesses the capability for film expansion and disintegration under high-frequency electromagnetic stimulation. These findings illuminate a promising path toward amplifying the therapeutic capacities of capsule endoscopes without necessitating a size compromise. Sishen Yuan, Baijia Liang, Lailu Li, Qingzhuo Zheng, Shuang Song 0002, Zhen Li 0026, Hongliang Ren 0001 |
ICRA | 7 |
| 2024 | Magnetic-Guided Flexible Origami Robot toward Long-Term Phototherapy of H. pylori in the StomachabstractHelicobacter pylori, a pervasive bacterial infection associated with gastrointestinal disorders such as gastritis, peptic ulcer disease, and gastric cancer, impacts approximately 50% of the global population. The efficacy of standard clinical eradication therapies is diminishing due to the rise of antibiotic-resistant strains, necessitating alternative treatment strategies. Photodynamic therapy (PDT) emerges as a promising prospect in this context. This study presents the development and implementation of a magnetically-guided origami robot, incorporating flexible printed circuit units for sustained and stable phototherapy of Helicobacter pylori. Each integrated unit is equipped with wireless charging capabilities, producing an optimal power output that can concurrently illuminate up to 15 LEDs at their maximum intensity. Crucially, these units can be remotely manipulated via a magnetic field, facilitating both translational and rotational movements. We propose an open-loop manual control sequence that allows the formation of a stable, compliant triangular structure through the interaction of internal magnets. This adaptable configuration is uniquely designed to withstand the dynamic squeezing environment prevalent in real-world gastric applications. The research herein represents a significant stride in leveraging technology for innovative medical solutions, particularly in the management of antibiotic-resistant Helicobacter pylori infections. Sishen Yuan, Baijia Liang, Po Wa Wong, Mingjing Xu, Chi Hsuan Li, Zhen Li 0026, Hongliang Ren 0001 |
ICRA | 6 |
| 2024 | EndoUIC: Promptable Diffusion Transformer for Unified Illumination Correction in Capsule Endoscopy
Long Bai 0008, Tong Chen 0011, Qiaozhi Tan, Wan Jun Nah, Yanheng Li 0002, Zhicheng He 0010, Sishen Yuan, Zhen Chen 0018, Jinlin Wu, Mobarakol Islam, Zhen Li 0026, Hongbin Liu 0001, Hongliang Ren 0001 |
MICCAI (7) | 11 |
| 2024 | Towards a Benchmark for Colorectal Cancer Segmentation in Endorectal Ultrasound Videos: Dataset and Model Development
Yuncheng Jiang 0002, Yiwen Hu 0001, Zixun Zhang, Jun Wei 0006, Chun-Mei Feng 0001, Xuemei Tang, Yong Liu 0026, Shuguang Cui, Zhen Li 0026 |
MICCAI (8) | 10 |
| 2024 | SearchLVLMs: A Plug-and-Play Framework for Augmenting Large Vision-Language Models by Searching Up-to-Date Internet KnowledgeabstractLarge vision-language models (LVLMs) are ignorant of the up-to-date knowledge, such as LLaVA series, because they cannot be updated frequently due to the large amount of resources required, and therefore fail in many cases. For example, if a LVLM was released on January 2024, and it wouldn't know the singer of the theme song for the new Detective Conan movie, which wasn't released until April 2024. To solve the problem, a promising solution motivated by retrieval-augmented generation (RAG) is to provide LVLMs with up-to-date knowledge via internet search during inference, i.e., internet-augmented generation (IAG), which is already integrated in some closed-source commercial LVLMs such as GPT-4V. However, the specific mechanics underpinning them remain a mystery. In this paper, we propose a plug-and-play framework, for augmenting existing LVLMs in handling visual question answering (VQA) about up-to-date knowledge, dubbed SearchLVLMs. A hierarchical filtering model is trained to effectively and efficiently find the most helpful content from the websites returned by a search engine to prompt LVLMs with up-to-date knowledge. To train the model and evaluate our framework's performance, we propose a pipeline to automatically generate news-related VQA samples to construct a dataset, dubbed UDK-VQA. A multi-model voting mechanism is introduced to label the usefulness of website/content for VQA samples to construct the training set. Experimental results demonstrate the effectiveness of our framework, outperforming GPT-4o by $\sim$30\% in accuracy. Chuanhao Li 0001, Zhen Li 0026, Chenchen Jing, Wenqi Shao, Yuwei Wu 0001, Ping Luo 0002, Yu Qiao 0001, Kaipeng Zhang |
NeurIPS | 2 |
| 2024 | Towards Flexible 3D Perception: Object-Centric Occupancy Completion Augments 3D Object DetectionabstractWhile 3D object bounding box (bbox) representation has been widely used in autonomous driving perception, it lacks the ability to capture the precise details of an object's intrinsic geometry. Recently, occupancy has emerged as a promising alternative for 3D scene perception. However, constructing a high-resolution occupancy map remains infeasible for large scenes due to computational constraints. Recognizing that foreground objects only occupy a small portion of the scene, we introduce object-centric occupancy as a supplement to object bboxes. This representation not only provides intricate details for detected objects but also enables higher voxel resolution in practical applications. We advance the development of object-centric occupancy perception from both data and algorithm perspectives. On the data side, we construct the first object-centric occupancy dataset from scratch using an automated pipeline. From the algorithmic standpoint, we introduce a novel object-centric occupancy completion network equipped with an implicit shape decoder that manages dynamic-size occupancy generation. This network accurately predicts the complete object-centric occupancy volume for inaccurate object proposals by leveraging temporal information from long sequences. Our method demonstrates robust performance in completing object shapes under noisy detection and tracking conditions. Additionally, we show that our occupancy features significantly enhance the detection results of state-of-the-art 3D object detectors, especially for incomplete or distant objects in the Waymo Open Dataset. Chaoda Zheng, Feng Wang 0018, Naiyan Wang, Shuguang Cui, Zhen Li 0026 |
NeurIPS | 5 |
| 2024 | Benchmarking the Robustness of LiDAR Semantic Segmentation Models
Xu Yan 0005, Chaoda Zheng, Ying Xue 0003, Zhen Li 0026, Shuguang Cui, Dengxin Dai |
Int. J. Comput. Vis. | 4 |
| 2024 | An Effective Motion-Centric Paradigm for 3D Single Object Tracking in Point Cloudsabstract3D single object tracking in LiDAR point clouds (LiDAR SOT) plays a crucial role in autonomous driving. Current approaches all follow the Siamese paradigm based on appearance matching. However, LiDAR point clouds are usually textureless and incomplete, which hinders effective appearance matching. Besides, previous methods greatly overlook the critical motion clues among targets. In this work, beyond 3D Siamese tracking, we introduce amotion-centric paradigmto handle LiDAR SOT from a new perspective. Following this paradigm, we propose a matching-free two-stage trackerM$^{2}$2-Track. At the 1st-stage,$M^{2}$-Track localizes the target within successive frames viamotion transformation. Then it refines the target box throughmotion-assisted shape completion at the 2nd-stage. Due to the motion-centric nature, our method shows its impressive generalizability with limited training labels and provides good differentiability for end-to-end cycle training. This inspires us to explore semi-supervised LiDAR SOT by incorporating a pseudo-label-based motion augmentation and a self-supervised loss term. Under the fully-supervised setting, extensive experiments confirm that$M^{2}$-Track significantly outperforms previous state-of-the-arts on three large-scale datasets while running at57FPS($\sim$∼3%,$\sim$∼11%and$\sim$∼22%precision gains on KITTI, NuScenes, and Waymo Open Dataset respectively). While under the semi-supervised setting, our method performs on par with or even surpasses its fully-supervised counterpart using fewer than half labels from KITTI. Further analysis verifies each component's effectiveness and shows the motion-centric paradigm's promising potential for auto-labeling and unsupervised domain adaptation. Chaoda Zheng, Xu Yan 0005, Haiming Zhang 0001, Baoyuan Wang, Shenghui Cheng, Shuguang Cui, Zhen Li 0026 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2024 | Self-reconstruction network for fine-grained few-shot classificationabstractMetric-based methods are one of the most common methods to solve the problem of few-shot image classification. However, traditional metric-based few-shot methods suffer from overfitting and local feature misalignment. The recently proposed feature reconstruction-based approach, which reconstructs query image features from the support set features of a given class and compares the distance between the original query features and the reconstructed query features as the classification criterion, effectively solves the feature misalignment problem. However, the issue of overfitting still has not been considered. To this end, we propose a self-reconstruction metric module for diversifying query features and a restrained cross-entropy loss for avoiding over-confident predictions. By introducing them, the proposed self-reconstruction network can effectively alleviate overfitting. Extensive experiments on five benchmark fine-grained datasets demonstrate that our proposed method achieves state-of-the-art performance on both 5-way 1-shot and 5-way 5-shot classification tasks. Code is available at https://github.com/liz-lut/SRM-main. Zhen Li 0026, Jiyang Xie 0001, Jing-Hao Xue, Zhanyu Ma |
Pattern Recognit. | 2 |
| 2024 | A Wearable, Reconfigurable, and Modular Magnetic Tracking System for Wireless Capsule RobotsabstractWearable magnetic tracking systems (MTSs) offer a promising technology for the long-term tracking of wireless-capsule robots within the digestive tract. However, existing wearable MTSs are fixed in size and cannot accommodate patients with diverse abdominal circumferences. To address this limitation, we propose a wearable and reconfigurable MTS. First, we design a reconfigurable sensor array inspired by the structure of bamboo slips, allowing it to conform to the abdominal surface and accommodate individuals with different abdominal circumferences. Next, we formulate a magnetic tracking optimization problem based on the magnetic dipole model and our established kinematic model of the reconfigurable sensor array. Solving the magnetic tracking problem, we achieved outstanding localization accuracy of 1.44$\pm$0.50 mm and 1.07$\pm 0.16^\circ$. Experimental validation demonstrates our proposed system's portability, reconfigurability, and adaptability to varying abdominal circumferences, offering valuable technological means for diagnosing and treating gastrointestinal disorders. Shijian Su, Sishen Yuan, Zhen Li 0026, Miaomiao Ma, Hongliang Ren 0001 |
IEEE Trans. Ind. Informatics | 3 |
| 2024 | ECC-PolypDet: Enhanced CenterNet With Contrastive Learning for Automatic Polyp DetectionabstractAccurate polyp detection is critical for early colorectal cancer diagnosis. Although remarkable progress has been achieved in recent years, the complex colon environment and concealed polyps with unclear boundaries still pose severe challenges in this area. Existing methods either involve computationally expensive context aggregation or lack prior modeling of polyps, resulting in poor performance in challenging cases. In this paper, we propose the Enhanced CenterNet with Contrastive Learning (ECC-PolypDet), a two-stage training & end-to-end inference framework that leverages images and bounding box annotations to train a general model and fine-tune it based on the inference score to obtain a final robust model. Specifically, we conduct Box-assisted Contrastive Learning (BCL) during training to minimize the intra-class difference and maximize the inter-class difference between foreground polyps and backgrounds, enabling our model to capture concealed polyps. Moreover, to enhance the recognition of small polyps, we design the Semantic Flow-guided Feature Pyramid Network (SFFPN) to aggregate multi-scale features and the Heatmap Propagation (HP) module to boost the model's attention on polyp targets. In the fine-tuning stage, we introduce the IoU-guided Sample Re-weighting (ISR) mechanism to prioritize hard samples by adaptively adjusting the loss weight for each sample during fine-tuning. Extensive experiments on six large-scale colonoscopy datasets demonstrate the superiority of our model compared with previous state-of-the-art detectors. Yuncheng Jiang 0002, Zixun Zhang, Yiwen Hu 0001, Guanbin Li, Shuguang Cui, Silin Huang, Zhen Li 0026 |
IEEE J. Biomed. Health Informatics | 9 |
| 2024 | Hierarchical Weight Averaging for Deep Neural NetworksabstractDespite simplicity, stochastic gradient descent (SGD)-like algorithms are successful in training deep neural networks (DNNs). Among various attempts to improve SGD, weight averaging (WA), which averages the weights of multiple models, has recently received much attention in the literature. Broadly, WA falls into two categories: 1) online WA, which averages the weights of multiple models trained in parallel, is designed for reducing the gradient communication overhead of parallel mini-batch SGD and 2) offline WA, which averages the weights of one model at different checkpoints, is typically used to improve the generalization ability of DNNs. Though online and offline WA are similar in form, they are seldom associated with each other. Besides, these methods typically perform either offline parameter averaging or online parameter averaging, but not both. In this work, we first attempt to incorporate online and offline WA into a general training framework termed hierarchical WA (HWA). By leveraging both the online and offline averaging manners, HWA is able to achieve both faster convergence speed and superior generalization performance without any fancy learning rate adjustment. Besides, we also analyze the issues faced by the existing WA methods, and how our HWA addresses them, empirically. Finally, extensive experiments verify that HWA outperforms the state-of-the-art methods significantly. Xiaozhe Gu, Zixun Zhang, Yuncheng Jiang 0002, Tao Luo 0014, Ruimao Zhang, Shuguang Cui, Zhen Li 0026 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2024 | Adversarial Sample Synthesis for Visual Question AnsweringabstractLanguage prior is a major block to improving the generalization of visual question answering (VQA) models. Recent work has revealed that synthesizing extra training samples to balance training sets is a promising way to alleviate language priors. However, most existing methods synthesize extra samples in a manner independent of training processes, which neglect the fact that the language priors memorized by VQA models are changing during training, resulting in insufficient synthesized samples. In this article, we propose an adversarial sample synthesis method, which synthesizes different adversarial samples by adversarial masking at different training epochs to cope with the changing memorized language priors. The basic idea behind our method is to use adversarial masking to synthesize adversarial samples that will cause the model to make wrong answers. To this end, we design a generative module to carry out adversarial masking by attacking the VQA model and introduce a bias-oriented objective to supervise the training of the generative module. We couple the sample synthesis with the training process of the VQA model, which ensures that the synthesized samples at different training epochs are beneficial to the VQA model. We incorporated the proposed method into three VQA models including UpDn, LMH, and LXMERT and conducted experiments on three datasets including VQA-CP v1, VQA-CP v2, and VQA v2. Experimental results demonstrate that a large improvement of our method, such as 16.22% gains on LXMERT in the overall accuracy of VQA-CP v2. Chuanhao Li 0001, Chenchen Jing, Zhen Li 0026, Yuwei Wu 0001, Yunde Jia |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Comprehensive Visual Question Answering on Point Clouds through Compositional Scene ManipulationabstractVisual Question Answering on 3D Point Cloud (VQA-3D) is an emerging yet challenging field that aims at answering various types of textual questions given an entire point cloud scene. To tackle this problem, we propose the CLEVR3D, a large-scale VQA-3D dataset consisting of 171K questions from 8,771 3D scenes. Specifically, we develop a question engine leveraging 3D scene graph structures to generate diverse reasoning questions, covering the questions of objects' attributes (i.e., size, color, and material) and their spatial relationships. Through such a manner, we initially generated 44K questions from 1,333 real-world scenes. Moreover, a more challenging setup is proposed to remove the confounding bias and adjust the context from a common-sense layout. Such a setup requires the network to achieve comprehensive visual understanding when the 3D scene is different from the general co-occurrence context (e.g., chairs always exist with tables). To this end, we further introduce the compositional scene manipulation strategy and generate 127K questions from 7,438 augmented 3D scenes, which can improve VQA-3D models for real-world comprehension. Built upon the proposed dataset, we baseline several VQA-3D models, where experimental results verify that the CLEVR3D can significantly boost other 3D scene understanding tasks. Xu Yan 0005, Zhihao Yuan, Yinghong Liao, Yao Guo 0002, Shuguang Cui, Zhen Li 0026 |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2023 | Geometry-Aware Network for Domain Adaptive Semantic SegmentationabstractMeasuring and alleviating the discrepancies between the synthetic (source) and real scene (target) data is the core issue for domain adaptive semantic segmentation. Though recent works have introduced depth information in the source domain to reinforce the geometric and semantic knowledge transfer, they cannot extract the intrinsic 3D information of objects, including positions and shapes, merely based on 2D estimated depth. In this work, we propose a novel Geometry-Aware Network for Domain Adaptation (GANDA), leveraging more compact 3D geometric point cloud representations to shrink the domain gaps. In particular, we first utilize the auxiliary depth supervision from the source domain to obtain the depth prediction in the target domain to accomplish structure-texture disentanglement. Beyond depth estimation, we explicitly exploit 3D topology on the point clouds generated from RGB-D images for further coordinate-color disentanglement and pseudo-labels refinement in the target domain. Moreover, to improve the 2D classifier in the target domain, we perform domain-invariant geometric adaptation from source to target and unify the 2D semantic and 3D geometric segmentation results in two domains. Note that our GANDA is plug-and-play in any existing UDA framework. Qualitative and quantitative results demonstrate that our model outperforms state-of-the-arts on GTA5->Cityscapes and SYNTHIA->Cityscapes. Yinghong Liao, Wending Zhou, Xu Yan 0005, Zhen Li 0026, Yizhou Yu, Shuguang Cui |
AAAI | 4 |
| 2023 | CowClip: Reducing CTR Prediction Model Training Time from 12 Hours to 10 Minutes on 1 GPUabstractThe click-through rate (CTR) prediction task is to predict whether a user will click on the recommended item. As mind-boggling amounts of data are produced online daily, accelerating CTR prediction model training is critical to ensuring an up-to-date model and reducing the training cost. One approach to increase the training speed is to apply large batch training. However, as shown in computer vision and natural language processing tasks, training with a large batch easily suffers from the loss of accuracy. Our experiments show that previous scaling rules fail in the training of CTR prediction neural networks. To tackle this problem, we first theoretically show that different frequencies of ids make it challenging to scale hyperparameters when scaling the batch size. To stabilize the training process in a large batch size setting, we develop the adaptive Column-wise Clipping (CowClip). It enables an easy and effective scaling rule for the embeddings, which keeps the learning rate unchanged and scales the L2 loss. We conduct extensive experiments with four CTR prediction networks on two real-world datasets and successfully scaled 128 times the original batch size without accuracy loss. In particular, for CTR prediction model DeepFM training on the Criteo dataset, our optimization framework enlarges the batch size from 1K to 128K with over 0.1% AUC improvement and reduces training time from 12 hours to 10 minutes on a single V100 GPU. Our code locates at github.com/bytedance/LargeBatchCTR. Zangwei Zheng, Pengtai Xu, Xuan Zou, Da Tang, Zhen Li 0026, Chenguang Xi, Leqi Zou, Xiangzhuo Ding, Fuzhao Xue, Ziheng Qin, Youlong Cheng, Yang You 0001 |
AAAI | 5 |
| 2023 | ScribblePolyp: Scribble-Supervised Polyp Segmentation through Dual Consistency AlignmentabstractAutomatic polyp segmentation models play a pivotal role in the clinical diagnosis of gastrointestinal diseases. In previous studies, most methods relied on fully supervised approaches, necessitating pixel-level annotations for model training. However, the creation of pixel-level annotations is both expensive and time-consuming, impeding the development of model generalization. In response to this challenge, we introduce ScribblePolyp, a novel scribble-supervised polyp segmentation framework. Unlike fully-supervised models, ScribblePolyp only requires the annotation of two lines (scribble labels) for each image, significantly reducing the labeling cost. Despite the coarse nature of scribble labels, which leave a substantial portion of pixels unlabeled, we propose a two-branch consistency alignment approach to provide supervision for these unlabeled pixels. The first branch employs transformation consistency alignment to narrow the gap between predictions under different transformations of the same input image. The second branch leverages affinity propagation to refine predictions into a soft version, extending additional supervision to unlabeled pixels. In summary, ScribblePolyp is an efficient model that does not rely on teacher models or moving average pseudo labels during training. Extensive experiments on the SUN-SEG dataset underscore the effectiveness of ScribblePolyp, achieving a Dice score of 0.8155, with the potential for a 1.8% improvement in the Dice score through a straightforward self-training strategy. Zixun Zhang, Yuncheng Jiang 0002, Jun Wei 0006, Hannah Cui, Zhen Li 0026 |
BIBM | 5 |
| 2023 | Exploring the Effect of Primitives for Compositional Generalization in Vision-and-LanguageabstractCompositionality is one of the fundamental properties of human cognition (Fodor & Pylyshyn, 1988). Compositional generalization is critical to simulate the compositional capability of humans, and has received much attention in the vision-and-language (V&L) community. It is essential to understand the effect of the primitives, including words, image regions, and video frames, to improve the compositional generalization capability. In this paper, we explore the effect of primitives for compositional generalization in V&L. Specifically, we present a self-supervised learning based framework that equips existing V&L methods with two characteristics: semantic equivariance and semantic invariance. With the two characteristics, the methods understand primitives by perceiving the effect of primitive changes on sample semantics and ground-truth. Experimental results on two tasks: temporal video grounding and visual question answering, demonstrate the effectiveness of our framework. Chuanhao Li 0001, Zhen Li 0026, Chenchen Jing, Yunde Jia, Yuwei Wu 0001 |
CVPR | 2 |
| 2023 | Semantic Human Parsing via Scalable Semantic Transfer Over Multiple Label DomainsabstractThis paper presents Scalable Semantic Transfer (SST), a novel training paradigm, to explore how to leverage the mutual benefits of the data from different label domains (i.e. various levels of label granularity) to train a powerful human parsing network. In practice, two common application scenarios are addressed, termed universal parsing and dedicated parsing, where the former aims to learn homogeneous human representations from multiple label domains and switch predictions by only using different segmentation heads, and the latter aims to learn a specific domain prediction while distilling the semantic knowledge from other domains. The proposed SST has the following appealing benefits: (1) it can capably serve as an effective training scheme to embed semantic associations of human body parts from multiple label domains into the human representation learning process; (2) it is an extensible semantic transfer framework without predetermining the overall relations of multiple label domains, which allows continuously adding human parsing datasets to promote the training. (3) the relevant modules are only used for auxiliary training and can be removed during inference, eliminating the extra reasoning cost. Experimental results demonstrate SST can effectively achieve promising universal human parsing performance as well as impressive improvements compared to its counterparts on three human parsing benchmarks (i.e., PASCAL-Person-Part, ATR, and CIHP). Code is available at https://github.com/yangjie-cv/SST. Chaoqun Wang 0012, Zhen Li 0026, Junle Wang, Ruimao Zhang |
CVPR | 3 |
| 2023 | BEV@DC: Bird's-Eye View Assisted Training for Depth CompletionabstractDepth completion plays a crucial role in autonomous driving, in which cameras and LiDARs are two complementary sensors. Recent approaches attempt to exploit spatial geometric constraints hidden in LiDARs to enhance image-guided depth completion. However, only low efficiency and poor generalization can be achieved. In this paper, we propose BEV@DC, a more efficient and powerful multi-modal training scheme, to boost the performance of image-guided depth completion. In practice, the proposed BEV@DC model comprehensively takes advantage of LiDARs with rich geometric details in training, employing an enhanced depth completion manner in inference, which takes only images (RGB and depth) as input. Specifically, the geometric-aware LiDAR features are projected onto a unified BEV space, combining with RGB features to perform BEV completion. By equipping a newly proposed point-voxel spatial propagation network (PV-SPN), this auxiliary branch introduces strong guidance to the original image branches via 3D dense supervision and feature consistency. As a result, our baseline model demonstrates significant improvements with the sole image inputs. Concretely, it achieves state-of-the-art on several benchmarks, e.g., ranking Top-1 on the challenging KITTI depth completion benchmark. Wending Zhou, Xu Yan 0005, Yinghong Liao, Yuankai Lin, Gangming Zhao, Shuguang Cui, Zhen Li 0026 |
CVPR | 8 |
| 2023 | Composable Text Controls in Latent Space with ODEsabstractGuangyi Liu, Zeyu Feng, Yuan Gao, Zichao Yang, Xiaodan Liang, Junwei Bao, Xiaodong He, Shuguang Cui, Zhen Li, Zhiting Hu. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Guangyi Liu 0005, Zeyu Feng, Xiaodan Liang, Junwei Bao 0001, Xiaodong He 0001, Shuguang Cui, Zhen Li 0026, Zhiting Hu |
EMNLP | 9 |
| 2023 | LATR: 3D Lane Detection from Monocular Images with Transformerabstract3D lane detection from monocular images is a fundamental yet challenging task in autonomous driving. Recent advances primarily rely on structural 3D surrogates (e.g., bird’s eye view) built from front-view image features and camera parameters. However, the depth ambiguity in monocular images inevitably causes misalignment between the constructed surrogate feature map and the original image, posing a great challenge for accurate lane detection. To address the above issue, we present a novel LATR model, an end-to-end 3D lane detector that uses 3D-aware front-view features without transformed view representation. Specifically, LATR detects 3D lanes via cross-attention based on query and key-value pairs, constructed using our lane-aware query generator and dynamic 3D ground positional embedding. On the one hand, each query is generated based on 2D lane-aware features and adopts a hybrid embedding to enhance the lane information. On the other hand, 3D space information is injected as positional embedding from an iteratively-updated 3D ground plane. LATR outperforms previous state-of-the-art methods on both synthetic Apollo and realistic OpenLane and ONCE-3DLanes by large margins (e.g., 11.4 gain in terms of F1 score on OpenLane). Code will be released at https://github.com/JMoonr/LATR. Yueru Luo, Chaoda Zheng, Xu Yan 0005, Tang Kun, Chao Zheng 0004, Shuguang Cui, Zhen Li 0026 |
ICCV | 7 |
| 2023 | SupFusion: Supervised LiDAR-Camera Fusion for 3D Object DetectionabstractLiDAR-Camera fusion-based 3D detection is a critical task for automatic driving. In recent years, many LiDAR-Camera fusion approaches sprung up and gained promising performances compared with single-modal detectors, but always lack carefully designed and effective supervision for the fusion process. In this paper, we propose a novel training strategy called SupFusion, which provides an auxiliary feature level supervision for effective LiDAR-Camera fusion and significantly boosts detection performance. Our strategy involves a data enhancement method named Polar Sampling, which densifies sparse objects and trains an assistant model to generate high-quality features as the supervision. These features are then used to train the LiDAR-Camera fusion model, where the fusion feature is optimized to simulate the generated high-quality features. Furthermore, we propose a simple yet effective deep fusion module, which contiguously gains superior performance compared with previous fusion methods with SupFusion strategy. In such a manner, our proposal shares the following advantages. Firstly, SupFusion introduces auxiliary feature-level supervision which could boost LiDAR-Camera detection performance without introducing extra inference costs. Secondly, the proposed deep fusion could continuously improve the detector’s abilities. Our proposed SupFusion and deep fusion module is plug-and-play, we make extensive experiments to demon-strate its effectiveness. Specifically, we gain around 2% 3D mAP improvements on KITTI benchmark based on multiple LiDAR-Camera 3D detectors. Our code is available at https://github.com/IranQin/SupFusion. Yiran Qin, Chaoqun Wang 0012, Zijian Kang, Ningning Ma, Zhen Li 0026, Ruimao Zhang |
ICCV | 5 |
| 2023 | SkeletonMAE: Graph-based Masked Autoencoder for Skeleton Sequence Pre-trainingabstractSkeleton sequence representation learning has shown great advantages for action recognition due to its promising ability to model human joints and topology. However, the current methods usually require sufficient labeled data for training computationally expensive models. Moreover, these methods ignore how to utilize the fine-grained dependencies among different skeleton joints to pre-train an efficient skeleton sequence learning model that can generalize well across different datasets. In this paper, we propose an efficient skeleton sequence learning framework, named Skeleton Sequence Learning (SSL). To comprehensively capture the human pose and obtain discriminative skeleton sequence representation, we build an asymmetric graph-based encoder-decoder pre-training architecture named SkeletonMAE, which embeds skeleton joint sequence into graph convolutional network and reconstructs the masked skeleton joints and edges based on the prior human topology knowledge. Then, the pre-trained SkeletonMAE encoder is integrated with the Spatial-Temporal Representation Learning (STRL) module to build the SSL framework. Extensive experimental results show that our SSL generalizes well across different datasets and outperforms the state-of-the-art self-supervised skeleton-based methods on FineGym, Diving48, NTU 60 and NTU 120 datasets. Moreover, we obtain comparable performance to some fully supervised methods. The code is avaliable at https://github.com/HongYan1123/SkeletonMAE. Hong Yan 0004, Yang Liu 0267, Yushen Wei, Zhen Li 0026, Guanbin Li, Liang Lin 0004 |
ICCV | 4 |
| 2023 | RankMatch: Fostering Confidence and Consistency in Learning with Noisy LabelsabstractLearning with noisy labels (LNL) is one of the most important and challenging problems in weakly-supervised learning. Recent advances adopt the sample selection strategy to mitigate the interference of noisy labels and use small-loss criteria to select clean samples. However, the one-dimensional loss is an over-simplified metric that fails to accommodate the complex feature landscape of various samples, and, hence, is prone to introduce classification errors during sample selection. In this paper, we propose RankMatch, a novel LNL framework that investigates additional dimensions of confidence and consistency in order to combat noisy labels. Confidence-wise, we propose a novel sample selection strategy based on confidence representation voting instead of the widely-used small-loss criterion. This new strategy is capable of increasing sample selection quantity without sacrificing labeling accuracy. Consistency-wise, instead of the widely adopted feature distance metric for measuring the consistency of inner-class samples, we advocate that the rank of principal features is a much more robust indicator. Based on this metric, we propose rank contrastive loss, which strengthens the consistency of similar samples regardless of their labels and facilitates feature representation learning. Experimental results on noisy versions of CIFAR-10, CIFAR-100, Clothing1M and WebVision have validated the superiority of our approach over existing state-of-the-art methods. Weikai Chen 0001, Chaowei Fang, Zhen Li 0026, Lechao Chen, Liang Lin 0004, Guanbin Li |
ICCV | 4 |
| 2023 | EasyGaze3D: Towards Effective and Flexible 3D Gaze Estimation from a Single RGB CameraabstractEye gaze can convey rich information of human intentions, which enables the social robots to comprehend the cognition and behavior of human targets. However, the existing 3D gaze estimation methods generally have high requirements either on the dedicated hardware or the quantity and quality of training databases, which largely limits their practical application values. This paper proposes EasyGaze3D, an effective 3D gaze estimation framework using a single RGB camera. First, the framework detects the 2D facial landmarks and recovers the 3D facial shape from the input image, and derives the required camera parameters with these features. Then, without loss of generality, the gaze direction can be regarded as the vector pointing from the eyeball center to the pupil center, which are derived respectively from the detected facial landmarks and the spherical fitting performed on the recovered 3D facial shape. Besides, we propose a flexible yet efficient calibration module, namely Easy-Cali, for deriving the subject-specific 3D facial shape and eyeball centers. The features calibrated by Easy-Cali can further boost the performance of EasyGaze3D. Experimental results show that our proposed method, being plug-and-play and without the need of training on large-scale dataset, can achieve superior performance against the existing methods based on deep models. Jianxin Yang, Yuxuan Liu 0013, Zhen Li 0026, Guang-Zhong Yang, Yao Guo 0002 |
IROS | 4 |
| 2023 | ArSDM: Colonoscopy Images Synthesis with Adaptive Refinement Semantic Diffusion Models
Yuncheng Jiang 0002, Shuangyi Tan, Xusheng Wu, Qi Dou 0001, Zhen Li 0026, Guanbin Li |
MICCAI (2) | 6 |
| 2023 | YONA: You Only Need One Adjacent Reference-Frame for Accurate and Fast Video Polyp Detection
Yuncheng Jiang 0002, Zixun Zhang, Ruimao Zhang, Guanbin Li, Shuguang Cui, Zhen Li 0026 |
MICCAI (5) | 6 |
| 2023 | WeakPolyp: You only Look Bounding Box for Polyp Segmentation
Jun Wei 0006, Yiwen Hu 0001, Shuguang Cui, Shaohua Kevin Zhou, Zhen Li 0026 |
MICCAI (3) | 5 |
| 2023 | CPU: Codebook Lookup Transformer with Knowledge Distillation for Point Cloud UpsamplingabstractPoint clouds produced by 3D scanning are typically sparse, non-uniform, and noisy. Existing upsampling techniques directly learn the mapping from a sparse point set to a dense point set, which is often under-determined and ill-posed. To reduce the uncertainty and ambiguity of the upsampling mapping, this paper proposes a generic three-stage vector-quantization framework, which incorporates a Codebook lookup Transformer and knowledge distillation for Point Cloud Upsampling, named CPU. The proposed CPU reformulates the upsampling task into a relatively determinate code prediction task within a small, discrete proxy space. Since the traditional vector-quantization methods cannot be directly applied to point cloud upsampling scenarios, we introduce a knowledge distillation training scheme that facilitates efficient codebook learning and ensures full utilization of codebook entries. Specifically, we adopt a teacher-student training paradigm to avoid model collapse during codebook learning. In the first stage, we pre-train a vanilla auto-encoder of the dense point set as the teacher model, which provides rich guidance features to ensure sufficient codebook learning. In the second stage, we train a vector-quantized auto-encoder as a student model to capture high-fidelity geometric priors into a learned codebook with the aid of distillation. In the third stage, we propose a Codebook Lookup Transformer to model the global context of the sparse point set and predict the code indices. Then the coarse features of the sparse point set can be quantized and substituted by looking up the indices in the learned codebook. Benefiting from the expressive codebook priors and the distillation training scheme, the proposed CPU outperforms state-of-the-art methods quantitatively and qualitatively. Weibing Zhao, Haiming Zhang 0001, Chaoda Zheng, Xu Yan 0005, Shuguang Cui, Zhen Li 0026 |
ACM Multimedia | 6 |
| 2023 | CholecTriplet2021: A benchmark challenge for surgical action triplet recognition
Chinedu Innocent Nwoye, Deepak Alapatt, Tong Yu 0009, Armine Vardazaryan, Fangfang Xia, Tong Xia, Fucang Jia, Yuxuan Yang 0007, Hao Wang 0081, Derong Yu, Guoyan Zheng, Xiaotian Duan, Neil Getty, Ricardo Sanchez-Matilla, Maria Robu, Li Zhang 0040, Huabin Chen, Jiacheng Wang 0002, Liansheng Wang 0002, Beerend G. A. Gerats, Sista Raviteja, Rachana Sathish, Rong Tao, Satoshi Kondo, Winnie Pang, Hongliang Ren 0001, Julian Ronald Abbing, Mohammad Hasan Sarhan, Sebastian Bodenstedt, Nithya Bhasker, Bruno Oliveira 0002, Helena R. Torres, Finn Gaida, Tobias Czempiel, João L. Vilaça, Pedro Morais, Jaime C. Fonseca 0001, Ruby Mae Egging, Inge Nicole Wijma, Chen Qian 0006, Guibin Bian, Zhen Li 0026, Velmurugan Balasubramanian, Debdoot Sheet, Imanol Luengo, Yuanbo Zhu, Shuai Ding 0001, Jakob-Anton Aschenbrenner, Nicolas Elini van der Kar, Mengya Xu, Mobarakol Islam, Seenivasan Lalithkumar, Alexander Jenke, Danail Stoyanov, Didier Mutter, Pietro Mascagni, Barbara Seeliger, Cristians Gonzalez, Nicolas Padoy |
Medical Image Anal. | 45 |
| 2023 | SAVAnet: Surgical Action-Driven Visual Attention Network for Autonomous Endoscope ControlabstractAn endoscope holder must understand the detailed surgical actions and the surgeons’ visual attention to keep important targets in the field of endoscopic view during operations. From an intensive analysis of the surgeons’ attention mechanism, we included that surgical actions, like cutting, suturing, etc., play an important role in determining the positions and weights of visual attention points during a dynamic surgical scene. To perform this process, this work proposes a Surgical Action-driven Visual Attention network (SAVAnet) and applies the network in autonomous endoscope control. Four scenarios are constructed in the da Vinci V-rep simulator: pick&place and needle exercise in a general laparoscopic training environment, needle driving with and without obstacle removal in an abdominal cavity, to create datasets for network training. The results show that the network has an outstanding performance in surgical action prediction with a high average accuracy of over 91%. Additionally, with surgical action guidance, the attention point prediction has higher accuracy and accords with surgeons’ visual attention. Finally, the acquired attention points are utilized to execute visual servoing in simulation. The results verify that the SAVAnet is feasible for autonomous endoscope control in real-time and lays a theoretical foundation for future sim-to-real execution. Note to Practitioners—This paper was motivated by the problem of endowing an endoscope with surgeons’ visual attention mechanism, which is affected by surgical actions, for autonomous endoscope control. An eye-tracking device has been utilized to detect surgeon’s visual attention in real-time and then control the endoscope to follow what the surgeon is looking at. However, this approach is susceptible to the surgical environment. Besides, many instrument detection and segmentation algorithms are developed for automatic surgical instrument tracking. However, surgeons’ visual attention does not always focus on the instruments during operations. In this work, we propose a novel SAVAnet to determine visual attention based on surgical actions. We prove from many qualitative and quantitative experiments that surgical actions play a significant role in determining visual attention. The designed SAVAnet can predict surgical actions correctly and then effectively guide the choice of visual attention. Finally, the simulation results show that the SAVAnet can endow endoscope with surgeons’ visual attention to perform self-control in real time. In future research, we will train the SAVAnet using real datasets and conduct more physical experiments on real surgical robots. Huxin Gao, Weichen Fan, Liang Qiu 0002, Xiaoxiao Yang, Zhen Li 0026, Xiuli Zuo, Max Q.-H. Meng, Hongliang Ren 0001 |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2022 | Contact-Distil: Boosting Low Homologous Protein Contact Map Prediction by Self-Supervised DistillationabstractAccurate protein contact map prediction (PCMP) is essential for precise protein structure estimation and further biological studies. Recent works achieve significant performance on this task with high quality multiple sequence alignment (MSA). However, the PCMP accuracy drops dramatically while only poor MSA (e.g., absolute MSA count less than 10) is available. Therefore, in this paper, we propose the Contact-Distil to improve the low homologous PCMP accuracy through knowledge distillation on a self-supervised model. Particularly, two pre-trained transformers are exploited to learn the high quality and low quality MSA representation in parallel for the teacher and student model correspondingly. Besides, the co-evolution information is further extracted from pure sequence through a pretrained ESM-1b model, which provides auxiliary knowledge to improve student performance. Extensive experiments show Contact-Distil outperforms previous state-of-the-arts by large margins on CAMEO-L dataset for low homologous PCMP, i.e., around 13.3% and 9.5% improvements against Alphafold2 and MSA Transformer respectively when MSA count less than 10. Qin Wang 0011, Jiayang Chen, Yu Li 0006, Liangzhen Zheng, Sheng Wang 0001, Zhen Li 0026, Shuguang Cui |
AAAI | 7 |
| 2022 | APAUNet: Axis Projection Attention UNet for Small Target in 3D Medical Segmentation
Yuncheng Jiang 0002, Zixun Zhang, Shixi Qin, Yao Guo 0002, Zhen Li 0026, Shuguang Cui |
ACCV (6) | 5 |
| 2022 | Graph Enhanced Contrastive Learning for Radiology Findings SummarizationabstractThe impression section of a radiology report summarizes the most prominent observation from the findings section and is the most important section for radiologists to communicate to physicians.Summarizing findings is timeconsuming and can be prone to error for inexperienced radiologists, and thus automatic impression generation has attracted substantial attention.With the encoder-decoder framework, most previous studies explore incorporating extra knowledge (e.g., static pre-defined clinical ontologies or extra background information).Yet, they encode such knowledge by a separate encoder to treat it as an extra input to their models, which is limited in leveraging their relations with the original findings.To address the limitation, we propose a unified framework for exploiting both extra knowledge and the original findings in an integrated way so that the critical information (i.e., key words and their relations) can be extracted in an appropriate way to facilitate impression generation.In detail, for each input findings, it is encoded by a text encoder, and a graph is constructed through its entities and dependency tree.Then, a graph encoder (e.g., graph neural networks (GNNs)) is adopted to model relation information in the constructed graph.Finally, to emphasize the key words in the findings, contrastive learning is introduced to map positive samples (constructed by masking non-key words) closer and push apart negative ones (constructed by masking key words).The experimental results on OpenI and MIMIC-CXR confirm the effectiveness of our proposed method. 1 Jinpeng Hu, Zhen Li 0026, Tsung-Hui Chang |
ACL (1) | 4 |
| 2022 | X -Trans2Cap: Cross-Modal Knowledge Transfer using Transformer for 3D Dense Captioningabstract3D dense captioning aims to describe individual objects in 3D scenes by natural language, where 3D scenes are usually represented as RGB-D scans or point clouds. However, only exploiting single modal information, e.g., point cloud, previous approaches fail to produce faithful descriptions. Though aggregating 2D features into point clouds may be beneficial, it introduces an extra computational burden, especially in the inference phase. In this study, we investigate a cross-modal knowledge transfer using Transformer for 3D dense captioning, namely X-Trans2Cap. Our proposed X-Trans2Cap effectively boost the performance of single-modal 3D captioning through the knowledge distillation enabled by a teacher-student framework. In practice, during the training phase, the teacher network exploits auxiliary 2D modality and guides the student network that only takes point clouds as input through the feature consistency constraints. Owing to the well-designed cross-modal feature fusion module and the feature alignment in the training phase, X-Trans2Cap acquires rich appearance information embedded in 2D images with ease. Thus, a more faithful caption can be generated only using point clouds during the inference. Qualitative and quantitative results confirm that X-Trans2Cap outperforms previous state-of-the-art by a large margin, i.e., about +21 and +16 CIDEr points on ScanRefer and Nr3D datasets, respectively. Zhihao Yuan, Xu Yan 0005, Yinghong Liao, Yao Guo 0002, Guanbin Li, Shuguang Cui, Zhen Li 0026 |
CVPR | 7 |
| 2022 | Beyond 3D Siamese Tracking: A Motion-Centric Paradigm for 3D Single Object Tracking in Point Cloudsabstract3D single object tracking (3D SOT) in LiDAR point clouds plays a crucial role in autonomous driving. Current approaches all follow the Siamese paradigm based on appearance matching. However, LiDAR point clouds are usually textureless and incomplete, which hinders effective appearance matching. Besides, previous methods greatly overlook the critical motion clues among targets. In this work, beyond 3D Siamese tracking, we introduce a motion-centric paradigm to handle 3D SOT from a new perspective. Following this paradigm, we propose a matching-free two-stage tracker M2-Track. At the 1st-stage, M2-Track localizes the target within successive frames via motion transformation. Then it refines the target box through motion-assisted shape completion at the 2nd-stage. Extensive experiments confirm that M2-Track significantly outperforms previous state-of-the-arts on three large-scale datasets while running at 57FPS (~ 8%, ~ 17% and ~ 22% precision gains on KITTI, NuScenes, and Waymo Open Dataset respectively). Further analysis verifies each component's effectiveness and shows the motioncentric paradigm's promising potential when combined with appearance matching. Code will be made available at https://github.com/Ghostish/Open3DSOT. Chaoda Zheng, Xu Yan 0005, Haiming Zhang 0001, Baoyuan Wang, Shenghui Cheng, Shuguang Cui, Zhen Li 0026 |
CVPR | 7 |
| 2022 | Weakly Supervised Object Localization Through Inter-class Feature Similarity and Intra-class Appearance Consistency
Jun Wei 0006, Sheng Wang 0001, Shaohua Kevin Zhou, Shuguang Cui, Zhen Li 0026 |
ECCV (30) | 5 |
| 2022 | 2DPASS: 2D Priors Assisted Semantic Segmentation on LiDAR Point Clouds
Xu Yan 0005, Jiantao Gao, Chaoda Zheng, Chao Zheng 0004, Ruimao Zhang, Shuguang Cui, Zhen Li 0026 |
ECCV (28) | 7 |
| 2022 | Toward Clinically Assisted Colorectal Polyp Recognition via Structured Cross-Modal Representation Consistency
Weijie Ma 0001, Ruimao Zhang, Yiwen Hu 0001, Zhen Li 0026 |
MICCAI (3) | 6 |
| 2022 | BoxPolyp: Boost Generalized Polyp Segmentation Using Extra Coarse Bounding Box Annotations
Jun Wei 0006, Yiwen Hu 0001, Guanbin Li, Shuguang Cui, Shaohua Kevin Zhou, Zhen Li 0026 |
MICCAI (3) | 6 |
| 2022 | Semi-supervised Spatial Temporal Attention Network for Video Polyp Segmentation
Xinkai Zhao, Shuangyi Tan, Zhen Li 0026, Guanbin Li |
MICCAI (4) | 5 |
| 2022 | Don't Take It Literally: An Edit-Invariant Sequence Loss for Text GenerationabstractGuangyi Liu, Zichao Yang, Tianhua Tao, Xiaodan Liang, Junwei Bao, Zhen Li, Xiaodong He, Shuguang Cui, Zhiting Hu. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Guangyi Liu 0005, Tianhua Tao, Xiaodan Liang, Junwei Bao 0001, Zhen Li 0026, Xiaodong He 0001, Shuguang Cui, Zhiting Hu |
NAACL-HLT | 6 |
| 2022 | Let Images Give You More: Point Cloud Cross-Modal Training for Shape AnalysisabstractAlthough recent point cloud analysis achieves impressive progress, the paradigm of representation learning from single modality gradually meets its bottleneck. In this work, we take a step towards more discriminative 3D point cloud representation using 2D images, which inherently contain richer appearance information, e.g., texture, color, and shade. Specifically, this paper introduces a simple but effective point cloud cross-modality training (PointCMT) strategy, which utilizes view-images, i.e., rendered or projected 2D images of the 3D object, to boost point cloud classification. In practice, to effectively acquire auxiliary knowledge from view-images, we develop a teacher-student framework and formulate the cross-modal learning as a knowledge distillation problem. Through novel feature and classifier enhancement criteria, PointCMT eliminates the distribution discrepancy between different modalities and avoid potential negative transfer effectively. Note that PointCMT efficiently improves the point-only representation without any architecture modification. Sufficient experiments verify significant gains on various datasets based on several backbones, i.e., equipped with PointCMT, PointNet++ and PointMLP achieve state-of-the-art performance on two benchmarks, i.e., 94.4% and 86.7% accuracy on ModelNet40 and ScanObjectNN, respectively. Xu Yan 0005, Heshen Zhan, Chaoda Zheng, Jiantao Gao, Ruimao Zhang, Shuguang Cui, Zhen Li 0026 |
NeurIPS | 7 |
| 2022 | AMOS: A Large-Scale Abdominal Multi-Organ Benchmark for Versatile Medical Image SegmentationabstractDespite the considerable progress in automatic abdominal multi-organ segmentation from CT/MRI scans in recent years, a comprehensive evaluation of the models' capabilities is hampered by the lack of a large-scale benchmark from diverse clinical scenarios. Constraint by the high cost of collecting and labeling 3D medical data, most of the deep learning models to date are driven by datasets with a limited number of organs of interest or samples, which still limits the power of modern deep models and makes it difficult to provide a fully comprehensive and fair estimate of various methods. To mitigate the limitations, we present AMOS, a large-scale, diverse, clinical dataset for abdominal organ segmentation. AMOS provides 500 CT and 100 MRI scans collected from multi-center, multi-vendor, multi-modality, multi-phase, multi-disease patients, each with voxel-level annotations of 15 abdominal organs, providing challenging examples and test-bed for studying robust segmentation algorithms under diverse targets and scenarios. We further benchmark several state-of-the-art medical segmentation models to evaluate the status of the existing methods on this new challenging dataset. We have made our datasets, benchmark servers, and baselines publicly available, and hope to inspire future research. Information can be found at https://amos22.grand-challenge.org. Yuanfeng Ji, Haotian Bai, Chongjian Ge, Ruimao Zhang, Zhen Li 0026, Wanling Ma, Ping Luo 0002 |
NeurIPS | 7 |
| 2022 | Divide and Contrast: Source-free Domain Adaptation via Adaptive Contrastive LearningabstractWe investigate a practical domain adaptation task, called source-free domain adaptation (SFUDA), where the source pretrained model is adapted to the target domain without access to the source data. Existing techniques mainly leverage self-supervised pseudo-labeling to achieve class-wise global alignment [1] or rely on local structure extraction that encourages the feature consistency among neighborhoods [2]. While impressive progress has been made, both lines of methods have their own drawbacks – the “global” approach is sensitive to noisy labels while the “local” counterpart suffers from the source bias. In this paper, we present Divide and Contrast (DaC), a new paradigm for SFUDA that strives to connect the good ends of both worlds while bypassing their limitations. Based on the prediction confidence of the source model, DaC divides the target data into source-like and target-specific samples, where either group of samples is treated with tailored goals under an adaptive contrastive learning framework. Specifically, the source-like samples are utilized for learning global class clustering thanks to their relatively clean labels. The more noisy target-specific data are harnessed at the instance level for learning the intrinsic local structures. We further align the source-like domain with the target-specific samples using a memory bank-based Maximum Mean Discrepancy (MMD) loss to reduce the distribution mismatch. Extensive experiments on VisDA, Office-Home, and the more challenging DomainNet have verified the superior performance of DaC over current state-of-the-art approaches. The code is available at https://github.com/ZyeZhang/DaC.git. Weikai Chen 0001, Zhen Li 0026, Liang Lin 0004, Guanbin Li |
NeurIPS | 4 |
| 2022 | Prior knowledge facilitates low homologous protein secondary structure prediction with DSM distillationabstractMOTIVATION: Protein secondary structure prediction (PSSP) is one of the fundamental and challenging problems in the field of computational biology. Accurate PSSP relies on sufficient homologous protein sequences to build the multiple sequence alignment (MSA). Unfortunately, many proteins lack homologous sequences, which results in the low quality of MSA and poor performance. In this article, we propose the novel dynamic scoring matrix (DSM)-Distil to tackle this issue, which takes advantage of the pretrained BERT and exploits the knowledge distillation on the newly designed DSM features. Specifically, we propose the DSM to replace the widely used profile and PSSM (position-specific scoring matrix) features. DSM could automatically dig for the suitable feature for each residue, based on the original profile. Namely, DSM-Distil not only could adapt to the low homologous proteins but also is compatible with high homologous ones. Thanks to the dynamic property, DSM could adapt to the input data much better and achieve higher performance. Moreover, to compensate for low-quality MSA, we propose to generate the pseudo-DSM from a pretrained BERT model and aggregate it with the original DSM by adaptive residue-wise fusion, which helps to build richer and more complete input features. In addition, we propose to supervise the learning of low-quality DSM features using high-quality ones. To achieve this, a novel teacher-student model is designed to distill the knowledge from proteins with high homologous sequences to that of low ones. Combining all the proposed methods, our model achieves the new state-of-the-art performance for low homologous proteins. RESULTS: Compared with the previous state-of-the-art method 'Bagging', DSM-Distil achieves an improvement about 5% and 7.3% improvement for proteins with MSA count ≤30 and extremely low homologous cases, respectively. We also compare DSM-Distil with Alphafold2 which is a state-of-the-art framework for protein structure prediction. DSM-Distil outperforms Alphafold2 by 4.1% on extremely low-quality MSA on 8-state secondary structure prediction. Moreover, we release a large-scale up-to-date test dataset BC40 for low-quality MSA structure prediction evaluation. AVAILABILITY AND IMPLEMENTATION: BC40 dataset: https://drive.google.com/drive/folders/15vwRoOjAkhhwfjDk6-YoKGf4JzZXIMC. HardCase dataset: https://drive.google.com/drive/folders/1BvduOr2b7cObUHy6GuEWk-aUkKJgzTUv. Code: https://github.com/qinwang-ai/DSM-Distil. Qin Wang 0011, Jun Wei 0006, Mingzhi Lin, Ruobing Ren, Sheng Wang 0001, Shuguang Cui, Zhen Li 0026 |
Bioinform. | 8 |
| 2022 | VQAMix: Conditional Triplet Mixup for Medical Visual Question AnsweringabstractMedical visual question answering (VQA) aims to correctly answer a clinical question related to a given medical image. Nevertheless, owing to the expensive manual annotations of medical data, the lack of labeled data limits the development of medical VQA. In this paper, we propose a simple yet effective data augmentation method, VQAMix, to mitigate the data limitation problem. Specifically, VQAMix generates more labeled training samples by linearly combining a pair of VQA samples, which can be easily embedded into any visual-language model to boost performance. However, mixing two VQA samples would construct new connections between images and questions from different samples, which will cause the answers for those new fabricated image-question pairs to be missing or meaningless. To solve the missing answer problem, we first develop the Learning with Missing Labels (LML) strategy, which roughly excludes the missing answers. To alleviate the meaningless answer issue, we design the Learning with Conditional-mixed Labels (LCL) strategy, which further utilizes language-type prior to forcing the mixed pairs to have reasonable answers that belong to the same category. Experimental results on the VQA-RAD and PathVQA benchmarks show that our proposed method significantly improves the performance of the baseline by about 7% and 5% on the averaging result of two backbones, respectively. More importantly, VQAMix could improve confidence calibration and model interpretability, which is significant for medical VQA models in practical applications. All code and models are available at https://github.com/haifangong/VQAMix. Haifan Gong, Guanqi Chen, Mingzhi Mao, Zhen Li 0026, Guanbin Li |
IEEE Trans. Medical Imaging | 4 |
| 2021 | PSSM-Distil: Protein Secondary Structure Prediction (PSSP) on Low-Quality PSSM by Knowledge Distillation with Contrastive LearningabstractProtein secondary structure prediction (PSSP) is an essential task in computational biology. To achieve the accurate PSSP, the general and vital feature engineering is to use multiple sequence alignment (MSA) for Position-Specific Scoring Matrix (PSSM) extraction. However, when only low-quality PSSM can be obtained due to poor sequence homology, previous PSSP accuracy (merely around 65%) is far from practical usage for subsequent tasks. In this paper, we propose a novel PSSM-Distil framework for PSSP on low-quality PSSM, which not only enhances the PSSM feature at a lower level but also aligns the feature distribution at a higher level. In practice, the PSSM-Distil first exploits the proteins with high-quality PSSM to achieve a teacher network for PSSP in a full-supervised way. Under the guidance of the teacher network, the low-quality PSSM and corresponding student network with low discriminating capacity are effectively resolved by feature enhancement through EnhanceNet and distribution alignment through knowledge distillation with contrastive learning. Further, our PSSM-Distil supports the input from a pre-trained protein sequence language BERT model to provide auxiliary information, which is designed to address the extremely low-quality PSSM cases, i.e., no homologous sequence. Extensive experiments demonstrate the proposed PSSM-Distil outperforms state-of-the-art models on PSSP by 6% on average and nearly 8% in extremely low-quality cases on public benchmarks, BC40 and CB513. Qin Wang 0011, Zhenlei Xu, Jiaxiang Wu 0001, Peilin Zhao, Zhen Li 0026, Sheng Wang 0001, Junzhou Huang, Shuguang Cui |
AAAI | 6 |
| 2021 | Sparse Single Sweep LiDAR Point Cloud Segmentation via Learning Contextual Shape Priors from Scene CompletionabstractLiDAR point cloud analysis is a core task for 3D computer vision, especially for autonomous driving. However, due to the severe sparsity and noise interference in the single sweep LiDAR point cloud, the accurate semantic segmentation is non-trivial to achieve. In this paper, we propose a novel sparse LiDAR point cloud semantic segmentation framework assisted by learned contextual shape priors. In practice, an initial semantic segmentation (SS) of a single sweep point cloud can be achieved by any appealing network and then flows into the semantic scene completion (SSC) module as the input. By merging multiple frames in the LiDAR sequence as supervision, the optimized SSC module has learned the contextual shape priors from sequential LiDAR data, completing the sparse single sweep point cloud to the dense one. Thus, it inherently improves SS optimization through fully end-to-end training. Besides, a Point-Voxel Interaction (PVI) module is proposed to further enhance the knowledge fusion between SS and SSC tasks, i.e., promoting the interaction of incomplete local geometry of point cloud and complete voxel-wise global structure. Furthermore, the auxiliary SSC and PVI modules can be discarded during inference without extra burden for SS. Extensive experiments confirm that our JS3C-Net achieves superior performance on both SemanticKITTI and SemanticPOSS benchmarks, i.e., 4% and 3% improvement correspondingly. Xu Yan 0005, Jiantao Gao, Jie Li 0002, Ruimao Zhang, Zhen Li 0026, Shuguang Cui |
AAAI | 5 |
| 2021 | Shallow Feature Matters for Weakly Supervised Object LocalizationabstractWeakly supervised object localization (WSOL) aims to localize objects by only utilizing image-level labels. Class activation maps (CAMs) are the commonly used features to achieve WSOL. However, previous CAM-based methods did not take full advantage of the shallow features, despite their importance for WSOL. Because shallow features are easily buried in background noise through conventional fusion. In this paper, we propose a simple but effective Shallow feature-aware Pseudo supervised Object Localization (SPOL) model for accurate WSOL, which makes the utmost of low-level features embedded in shallow layers. In practice, our SPOL model first generates the CAMs through a novel element-wise multiplication of shallow and deep feature maps, which filters the background noise and generates sharper boundaries robustly. Besides, we further propose a general class-agnostic segmentation model to achieve the accurate object mask, by only using the initial CAMs as the pseudo label without any extra annotation. Eventually, a bounding box extractor is applied to the object mask to locate the target. Experiments verify that our SPOL outperforms the state-of-the-art on both CUB- 200 and ImageNet-1K benchmarks, achieving 93.44% and 67.15% (i.e., 3.93% and 2.13% improvement) Top-5 localization accuracy, respectively. Jun Wei 0006, Qin Wang 0011, Zhen Li 0026, Sheng Wang 0001, Shaohua Kevin Zhou, Shuguang Cui |
CVPR | 3 |
| 2021 | InstanceRefer: Cooperative Holistic Understanding for Visual Grounding on Point Clouds through Instance Multi-level Contextual ReferringabstractCompared with the visual grounding on 2D images, the natural-language-guided 3D object localization on point clouds is more challenging. In this paper, we propose a new model, named InstanceRefer1, to achieve a superior 3D visual grounding through the grounding-by-matching strategy. In practice, our model first predicts the target category from the language descriptions using a simple language classification model. Then, based on the category, our model sifts out a small number of instance candidates (usually less than 20) from the panoptic segmentation on point clouds. Thus, the non-trivial 3D visual grounding task has been effectively re-formulated as a simplified instance-matching problem, considering that instance-level candidates are more rational than the redundant 3D object proposals. Subsequently, for each candidate, we perform the multi-level contextual inference, i.e., referring from instance attribute perception, instance-to-instance relation perception, and instance-to-background global localization perception, respectively. Eventually, the most relevant candidate is selected and localized by ranking confidence scores, which are obtained by the cooperative holistic visual-language feature matching. Experiments confirm that our method outperforms previous state-of-the-arts on ScanRefer online benchmark and Nr3D/Sr3D datasets. Zhihao Yuan, Xu Yan 0005, Yinghong Liao, Ruimao Zhang, Sheng Wang 0001, Zhen Li 0026, Shuguang Cui |
ICCV | 6 |
| 2021 | Box-Aware Feature Enhancement for Single Object Tracking on Point CloudsabstractCurrent 3D single object tracking approaches track the target based on a feature comparison between the target template and the search area. However, due to the common occlusion in LiDAR scans, it is non-trivial to conduct accurate feature comparisons on severe sparse and incomplete shapes. In this work, we exploit the ground truth bounding box given in the first frame as a strong cue to enhance the feature description of the target object, enabling a more accurate feature comparison in a simple yet effective way. In particular, we first propose the BoxCloud, an informative and robust representation, to depict an object using the point-to-box relation. We further design an efficient box-aware feature fusion module, which leverages the aforementioned BoxCloud for reliable feature matching and embedding. Integrating the proposed general components into an existing model P2B [27], we construct a superior box-aware tracker (BAT)1. Experiments confirm that our proposed BAT outperforms the previous state-of-the-art by a large margin on both KITTI and NuScenes benchmarks, achieving a 12.8% improvement in terms of precision while running ∼20% faster. Chaoda Zheng, Xu Yan 0005, Jiantao Gao, Weibing Zhao, Wei Zhang 0001, Zhen Li 0026, Shuguang Cui |
ICCV | 6 |
| 2021 | Adaptive Residue-wise Profile Fusion for Low Homologous Protein Secondary Structure Prediction Using External KnowledgeabstractProtein secondary structure prediction (PSSP) is essential for protein function analysis. However, for low homologous proteins, the PSSP suffers from insufficient input features. In this paper, we explicitly import external self-supervised knowledge for low homologous PSSP under the guidance of residue-wise (amino acid wise) profile fusion. In practice, we firstly demonstrate the superiority of profile over Position-Specific Scoring Matrix (PSSM) for low homologous PSSP. Based on this observation, we introduce the novel self-supervised BERT features as the pseudo profile, which implicitly involves the residue distribution in all native discovered sequences as the complementary features. Furthermore, a novel residue-wise attention is specially designed to adaptively fuse different features (i.e., original low-quality profile, BERT based pseudo profile), which not only takes full advantage of each feature but also avoids noise disturbance. Besides, the feature consistency loss is proposed to accelerate the model learning from multiple semantic levels. Extensive experiments confirm that our method outperforms state-of-the-arts (i.e., 4.7% for extremely low homologous cases on BC40 dataset). Qin Wang 0011, Jun Wei 0006, Zhen Li 0026, Sheng Wang 0001, Shuguang Cui |
IJCAI | 4 |
| 2021 | PointLIE: Locally Invertible Embedding for Point Cloud Sampling and RecoveryabstractPoint Cloud Sampling and Recovery (PCSR) is critical for massive real-time point cloud collection and processing since raw data usually requires large storage and computation. This paper addresses a fundamental problem in PCSR: How to downsample the dense point cloud with arbitrary scales while preserving the local topology of discarded points in a case-agnostic manner (i.e., without additional storage for point relationships)? We propose a novel Locally Invertible Embedding (PointLIE) framework to unify the point cloud sampling and upsampling into one single framework through bi-directional learning. Specifically, PointLIE decouples the local geometric relationships between discarded points from the sampled points by progressively encoding the neighboring offsets to a latent variable. Once the latent variable is forced to obey a pre-defined distribution in the forward sampling path, the recovery can be achieved effectively through inverse operations. Taking the recover-pleasing sampled points and a latent embedding randomly drawn from the specified distribution as inputs, PointLIE can theoretically guarantee the fidelity of reconstruction and outperform state-of-the-arts quantitatively and qualitatively. Weibing Zhao, Xu Yan 0005, Jiantao Gao, Ruimao Zhang, Jiayan Zhang, Zhen Li 0026, Shuguang Cui |
IJCAI | 6 |
| 2021 | Multi-compound Transformer for Accurate Biomedical Image Segmentation
Yuanfeng Ji, Ruimao Zhang, Huijie Wang, Zhen Li 0026, Lingyun Wu, Shaoting Zhang 0001, Ping Luo 0002 |
MICCAI (1) | 4 |
| 2021 | Colorectal Polyp Classification from White-Light Colonoscopy Images via Domain Alignment
Qin Wang 0011, Hui Che, Weizhen Ding, Guanbin Li, Zhen Li 0026, Shuguang Cui |
MICCAI (7) | 6 |
| 2021 | Shallow Attention Network for Polyp Segmentation
Jun Wei 0006, Yiwen Hu 0001, Ruimao Zhang, Zhen Li 0026, Shaohua Kevin Zhou, Shuguang Cui |
MICCAI (1) | 4 |
| 2021 | Medical-VLBERT: Medical Visual Language BERT for COVID-19 CT Report Generation With Alternate LearningabstractMedical imaging technologies, including computed tomography (CT) or chest X-Ray (CXR), are largely employed to facilitate the diagnosis of the COVID-19. Since manual report writing is usually too time-consuming, a more intelligent auxiliary medical system that could generate medical reports automatically and immediately is urgently needed. In this article, we propose to use the medical visual language BERT (Medical-VLBERT) model to identify the abnormality on the COVID-19 scans and generate the medical report automatically based on the detected lesion regions. To produce more accurate medical reports and minimize the visual-and-linguistic differences, this model adopts an alternate learning strategy with two procedures that are knowledge pretraining and transferring. To be more precise, the knowledge pretraining procedure is to memorize the knowledge from medical texts, while the transferring procedure is to utilize the acquired knowledge for professional medical sentences generations through observations of medical images. In practice, for automatic medical report generation on the COVID-19 cases, we constructed a dataset of 368 medical findings in Chinese and 1104 chest CT scans from The First Affiliated Hospital of Jinan University, Guangzhou, China, and The Fifth Affiliated Hospital of Sun Yat-sen University, Zhuhai, China. Besides, to alleviate the insufficiency of the COVID-19 training samples, our model was first trained on the large-scale Chinese CX-CHR dataset and then transferred to the COVID-19 CT dataset for further fine-tuning. The experimental results showed that Medical-VLBERT achieved state-of-the-art performances on terminology prediction and report generation with the Chinese COVID-19 CT dataset and the CX-CHR dataset. The Chinese COVID-19 CT dataset is available at https://covid19ct.github.io/. Guangyi Liu 0005, Yinghong Liao, Fuyu Wang 0001, Lu Zhang 0051, Xiaodan Liang, Shaolin Li, Zhen Li 0026, Shuixing Zhang, Shuguang Cui |
IEEE Trans. Neural Networks Learn. Syst. | 9 |
| 2020 | PointASNL: Robust Point Clouds Processing Using Nonlocal Neural Networks With Adaptive SamplingabstractRaw point clouds data inevitably contains outliers or noise through acquisition from 3D sensors or reconstruction algorithms. In this paper, we present a novel end-to-end network for robust point clouds processing, named PointASNL, which can deal with point clouds with noise effectively. The key component in our approach is the adaptive sampling (AS) module. It first re-weights the neighbors around the initial sampled points from farthest point sampling (FPS), and then adaptively adjusts the sampled points beyond the entire point cloud. Our AS module can not only benefit the feature learning of point clouds, but also ease the biased effect of outliers. To further capture the neighbor and long-range dependencies of the sampled point, we proposed a local-nonlocal (L-NL) module inspired by the nonlocal operation. Such L-NL module enables the learning process insensitive to noise. Extensive experiments verify the robustness and superiority of our approach in point clouds processing tasks regardless of synthesis data, indoor data, and outdoor data with or without noise. Specifically, PointASNL achieves state-of-the-art robust performance for classification and segmentation tasks on all datasets, and significantly outperforms previous methods on real-world outdoor SemanticKITTI dataset with considerate noise. Xu Yan 0005, Chaoda Zheng, Zhen Li 0026, Sheng Wang 0001, Shuguang Cui |
CVPR | 3 |
| 2020 | Exemplar Normalization for Learning Deep RepresentationabstractNormalization techniques are important in different advanced neural networks and different tasks. This work investigates a novel dynamic learning-to-normalize (L2N) problem by proposing Exemplar Normalization (EN), which is able to learn different normalization methods for different convolutional layers and image samples of a deep network. EN significantly improves the flexibility of the recently proposed switchable normalization (SN), which solves a static L2N problem by linearly combining several normalizers in each normalization layer (the combination is the same for all samples). Instead of directly employing a multi-layer perceptron (MLP) to learn data-dependent parameters as conditional batch normalization (cBN) did, the internal architecture of EN is carefully designed to stabilize its optimization, leading to many appealing benefits. (1) EN enables different convolutional layers, image samples, categories, benchmarks, and tasks to use different normalization methods, shedding light on analyzing them in a holistic view. (2) EN is effective for various network architectures and tasks. (3) It could replace any normalization layers in a deep network and still produce stable model training. Extensive experiments demonstrate the effectiveness of EN in a wide spectrum of tasks including image recognition, noisy label learning, and semantic segmentation. For example, by replacing BN in the ordinary ResNet50, improvement produced by EN is 300% more than that of SN on both ImageNet and the noisy WebVision dataset. The codes and models will be released. Ruimao Zhang, Zhanglin Peng, Lingyun Wu, Zhen Li 0026, Ping Luo 0002 |
CVPR | 4 |
| 2020 | MetaSelection: Metaheuristic Sub-Structure Selection for Neural Network Pruning Using Evolutionary AlgorithmabstractNeural network pruning is widely applied to various mobile applications. Previous pruning methods mainly leverage ad-hoc criteria to evaluate channel importance. In this paper, we propose an effective metaheuristic sub-structure selection (MetaSelection) method for neural network pruning. MetaSelection exploits evolutionary algorithm (EA) to search the proper sub-structure satisfying the resource constraints. In comparison with previous AutoML based methods, MetaSelection can automatically achieve the pruning rate and channel selection at the same time instead of hand-crafted criteria in a cascaded way. Regarding the tremendous search space of channel selection as a combinatorial optimization problem, we further utilize a coarse-to-fine strategy and the novel probability distribution crossover (PDC) to speed up the search procedure. Besides, MetaSelection prunes the network globally rather than in a layer-by-layer way. We evaluate MetaSelection on several appealing deep neural networks, achieving superior results with adaptive depth and width. Concretely, on ImageNet, MetaSelection achieves a top-1 accuracy of 71.5% on MobileNetV2 under 70% FLOPs constraint and a FLOPs reduction of 30% with 76.4% top-1 accuracy for ResNet50. Zixun Zhang, Zhen Li 0026, Lin Lin 0008, Na Lei, Guanbin Li, Shuguang Cui |
ECAI | 2 |
| 2020 | Towards Content-Independent Multi-Reference Super-Resolution: Adaptive Pattern Matching and Feature Aggregation
Xu Yan 0005, Weibing Zhao, Kun Yuan 0004, Ruimao Zhang, Zhen Li 0026, Shuguang Cui |
ECCV (25) | 5 |
| 2020 | UXNet: Searching Multi-level Feature Aggregation for 3D Medical Image Segmentation
Yuanfeng Ji, Ruimao Zhang, Zhen Li 0026, Jiamin Ren, Shaoting Zhang 0001, Ping Luo 0002 |
MICCAI (1) | 3 |
| 2020 | Characterizing Label Errors: Confident Learning for Noisy-Labeled Image Segmentation
Minqing Zhang, Jiantao Gao, Zhen Lyu, Weibing Zhao, Qin Wang 0011, Weizhen Ding, Sheng Wang 0001, Zhen Li 0026, Shuguang Cui |
MICCAI (1) | 8 |
| 2020 | Adaptive Context Selection for Polyp Segmentation
Ruifei Zhang, Guanbin Li, Zhen Li 0026, Shuguang Cui, Dahong Qian, Yizhou Yu |
MICCAI (6) | 3 |
| 2020 | JAFPro: Joint Appearance Fusion and Propagation for Human Video Motion Transfer from Multiple Reference ImagesabstractWe present a novel framework for human video motion transfer. Deviating from recent studies that use only single source image, we propose to allow users to supply multiple source images by simply imitating some poses in the desired target video. To aggregate the appearance from multiple input images, we propose a JAFPro framework that incorporates two modules: an appearance fusion module that adaptively fuses the information in the supplied images and an appearance propagation module that propagates textures through flow-based warping to further improve the result. An attractive feature of JAFPro is that the quality of its results progressively improves as more imitating images are supplied. Furthermore, we build a new dataset containing a large variety of dancing videos in the wild. Extensive experiments conducted on this dataset demonstrate JAFPro outperforms state-of-the-art methods both qualitatively and quantitatively. We will release our code and dataset upon publication of this work. Xianggang Yu, Haolin Liu 0004, Xiaoguang Han 0001, Zhen Li 0026, Zixiang Xiong, Shuguang Cui |
ACM Multimedia | 4 |
| 2019 | Semi-Supervised Video Salient Object Detection Using Pseudo-LabelsabstractDeep learning-based video salient object detection has recently achieved great success with its performance significantly outperforming any other unsupervised methods. However, existing data-driven approaches heavily rely on a large quantity of pixel-wise annotated video frames to deliver such promising results. In this paper, we address the semi-supervised video salient object detection task using pseudo-labels. Specifically, we present an effective video saliency detector that consists of a spatial refinement network and a spatiotemporal module. Based on the same refinement network and motion information in terms of optical flow, we further propose a novel method for generating pixel-level pseudo-labels from sparsely annotated frames. By utilizing the generated pseudo-labels together with a part of manual annotations, our video saliency detector learns spatial and temporal cues for both contrast inference and coherence enhancement, thus producing accurate saliency maps. Experimental results demonstrate that our proposed semi-supervised method even greatly outperforms all the state-of-the-art fully supervised methods across three public benchmarks of VOS, DAVIS, and FBMS. Pengxiang Yan, Guanbin Li, Yuan Xie 0004, Zhen Li 0026, Chuan Wang 0001, Tianshui Chen, Liang Lin 0004 |
ICCV | 4 |
| 2017 | High-Resolution Shape Completion Using Deep Neural Networks for Global Structure and Local Geometry InferenceabstractWe propose a data-driven method for recovering missing parts of 3D shapes. Our method is based on a new deep learning architecture consisting of two sub-networks: a global structure inference network and a local geometry refinement network. The global structure inference network incorporates a long short-term memorized context fusion module (LSTM-CF) that infers the global structure of the shape based on multi-view depth information provided as part of the input. It also includes a 3D fully convolutional (3DFCN) module that further enriches the global structure representation according to volumetric information in the input. Under the guidance of the global structure network, the local geometry refinement network takes as input local 3D patches around missing regions, and progressively produces a high-resolution, complete surface through a volumetric encoder-decoder architecture. Our method jointly trains the global structure inference and local geometry refinement networks in an end-to-end manner. We perform qualitative and quantitative evaluations on six object categories, demonstrating that our method outperforms existing state-of-the-art work on shape completion. Xiaoguang Han 0001, Zhen Li 0026, Evangelos Kalogerakis, Yizhou Yu |
ICCV | 2 |
| 2017 | Folding Membrane Proteins by Deep Transfer Learning
Zhen Li 0026, Sheng Wang 0001, Yizhou Yu, Jinbo Xu |
RECOMB | 1 |
| 2017 | Accurate De Novo Prediction of Protein Contact Map by Ultra-Deep Learning ModelabstractMOTIVATION: Protein contacts contain key information for the understanding of protein structure and function and thus, contact prediction from sequence is an important problem. Recently exciting progress has been made on this problem, but the predicted contacts for proteins without many sequence homologs is still of low quality and not very useful for de novo structure prediction. METHOD: This paper presents a new deep learning method that predicts contacts by integrating both evolutionary coupling (EC) and sequence conservation information through an ultra-deep neural network formed by two deep residual neural networks. The first residual network conducts a series of 1-dimensional convolutional transformation of sequential features; the second residual network conducts a series of 2-dimensional convolutional transformation of pairwise information including output of the first residual network, EC information and pairwise potential. By using very deep residual networks, we can accurately model contact occurrence patterns and complex sequence-structure relationship and thus, obtain higher-quality contact prediction regardless of how many sequence homologs are available for proteins in question. RESULTS: Our method greatly outperforms existing methods and leads to much more accurate contact-assisted folding. Tested on 105 CASP11 targets, 76 past CAMEO hard targets, and 398 membrane proteins, the average top L long-range prediction accuracy obtained by our method, one representative EC method CCMpred and the CASP11 winner MetaPSICOV is 0.47, 0.21 and 0.30, respectively; the average top L/10 long-range accuracy of our method, CCMpred and MetaPSICOV is 0.77, 0.47 and 0.59, respectively. Ab initio folding using our predicted contacts as restraints but without any force fields can yield correct folds (i.e., TMscore>0.6) for 203 of the 579 test proteins, while that using MetaPSICOV- and CCMpred-predicted contacts can do so for only 79 and 62 of them, respectively. Our contact-assisted models also have much better quality than template-based models especially for membrane proteins. The 3D models built from our contact prediction have TMscore>0.5 for 208 of the 398 membrane proteins, while those from homology modeling have TMscore>0.5 for only 10 of them. Further, even if trained mostly by soluble proteins, our deep learning method works very well on membrane proteins. In the recent blind CAMEO benchmark, our fully-automated web server implementing this method successfully folded 6 targets with a new fold and only 0.3L-2.3L effective sequence homologs, including one β protein of 182 residues, one α+β protein of 125 residues, one α protein of 140 residues, one α protein of 217 residues, one α/β of 260 residues and one α protein of 462 residues. Our method also achieved the highest F1 score on free-modeling targets in the latest CASP (Critical Assessment of Structure Prediction), although it was not fully implemented back then. AVAILABILITY: http://raptorx.uchicago.edu/ContactMap/. Sheng Wang 0001, Zhen Li 0026, Jinbo Xu |
PLoS Comput. Biol. | 3 |
| 2016 | LSTM-CF: Unifying Context Modeling and Fusion with LSTMs for RGB-D Scene Labeling
Zhen Li 0026, Yukang Gan, Xiaodan Liang, Yizhou Yu, Liang Lin 0004 |
ECCV (2) | 1 |
| 2016 | Protein Secondary Structure Prediction Using Cascaded Convolutional and Recurrent Neural Networks
Zhen Li 0026, Yizhou Yu |
IJCAI | 1 |