VLDB 2026 Research / reviewers in the wild / expert
Le Xue
dblp:304/2195
· DBLP profile ↗
16ranked-venue papers
3as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PET2Rep: Towards Vision-Language Model-Drived Automated Radiology Report Generation for Positron Emission TomographyabstractPositron emission tomography (PET) is a cornerstone of modern oncologic and neurologic imaging, distinguished by its unique ability to illuminate dynamic metabolic processes that transcend the anatomical focus of traditional imaging technologies. Radiology reports are essential for clinical decision making, yet their manual creation is labor-intensive and time-consuming. Recent advancements of vision-language models (VLMs) have shown strong potential in medical applications, presenting a promising avenue for automating report generation. However, existing applications of VLMs in the medical domain have predominantly focused on structural imaging modalities, while the unique characteristics of molecular PET imaging have largely been overlooked. To bridge the gap, we introduce PET2Rep, a large-scale comprehensive benchmark for evaluation of general and medical VLMs for radiology report generation for PET images. PET2Rep stands out as the first dedicated dataset for PET report generation with metabolic information, uniquely capturing whole-body image-report pairs that cover dozens of organs to fill the critical gap in existing benchmarks and mirror real-world clinical comprehensiveness. In addition to widely recognized natural language generation metrics, we introduce a series of clinical efficiency metrics to evaluate the quality of radiotracer uptake pattern description in key organs in generated reports. We conduct a head-to-head comparison of 30 cutting-edge general-purpose and medical-specialized VLMs. The results show that the current state-of-the-art VLMs perform poorly on PET report generation task, falling considerably short of fulfilling practical needs. Moreover, we identify several key insufficiency that need to be addressed to advance the development in medical applications. We believe PET2Rep will serve as a platform for the development and application of VLMs for PET imaging, accelerating the development of trustworthy reporting tools that can genuinely alleviate radiologist burden and enhance patient care. Yichi Zhang 0007, Zehui Ling, Sisi Peng, Deshu Chen, Lanlan Li, Limei Han, Zixin Hu, Yuan Qi 0001, Le Xue |
AAAI | 15 |
| 2025 | LLAVIDAL: A Large LAnguage VIsion Model for Daily Activities of LivingabstractCurrent Large Language Vision Models (LLVMs) trained on web videos perform well in general video understanding but struggle with fine-grained details, complex human-object interactions (HOI), and view-invariant representation learning essential for Activities of Daily Living (ADL). This limitation stems from a lack of specialized ADL video instruction-tuning datasets and insufficient modality integration to capture discriminative action representations. To address this, we propose a semi-automated framework for curating ADL datasets, creating ADL-X, a multiview, multimodal RGBS instruction-tuning dataset. Additionally, we introduce LLAVIDAL, an LLVM integrating videos, 3D skeletons, and HOIs to model ADL’s complex spatiotemporal relationships. For training LLAVIDAL a simple joint alignment of all modalities yields suboptimal results; thus, we propose a Multimodal Progressive (MMPro) training strategy, incorporating modalities in stages following a curriculum. We also establish ADL MCQ and video description benchmarks to assess LLVM performance in ADL tasks. Trained on ADL-X, LLAVIDAL achieves state-of-the-art performance across ADL benchmarks. Code and data will be made publicly available at https://adl-x.github.io/. Dominick Reilly, Rajatsubhra Chakraborty, Arkaprava Sinha, Manish Kumar Govind, Pu Wang 0001, François Brémond, Le Xue, Srijan Das |
CVPR | 7 |
| 2025 | Contra4: Evaluating Contrastive Cross-Modal Reasoning in Audio, Video, Image, and 3DabstractArtemis Panagopoulou, Le Xue, Honglu Zhou, Silvio Savarese, Ran Xu, Caiming Xiong, Chris Callison-Burch, Mark Yatskar, Juan Carlos Niebles. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Artemis Panagopoulou, Le Xue, Honglu Zhou, Silvio Savarese, Ran Xu 0001, Caiming Xiong, Chris Callison-Burch, Mark Yatskar, Juan Carlos Niebles |
EMNLP | 2 |
| 2025 | SegAnyPET: Universal Promptable Segmentation from Positron Emission Tomography ImagesabstractPositron Emission Tomography (PET) is a powerful molecular imaging tool that plays a crucial role in modern medical diagnostics by visualizing radio-tracer distribution to reveal physiological processes. Accurate organ segmentation from PET images is essential for comprehensive multi-systemic analysis of interactions between different organs and pathologies. Existing segmentation methods are limited by insufficient annotation data and varying levels of annotation, resulting in weak generalization ability and difficulty in clinical application. Recent developments in segmentation foundation models have shown superior versatility across diverse segmentation tasks. Despite the efforts of medical adaptations, these works primarily focus on structural medical images with detailed physiological structural information and exhibit limited generalization performance on molecular PET imaging. In this paper, we collect and construct PETS-5k, the largest PET segmentation dataset to date, comprising 5,731 three-dimensional whole-body PET images and encompassing over 1.3M 2D images. Based on the established dataset, we develop SegAnyPET, a modality-specific 3D foundation model for universal promptable segmentation from PET images. To issue the challenge of discrepant annotation quality, we adopt a cross prompting confident learning (CPCL) strategy with an uncertainty-guided self-rectification process to robustly learn segmentation from high-quality labeled data and low-quality noisy labeled data for promptable segmentation. Experimental results demonstrate that SegAnyPET can segment seen and unseen target organs using only one or a few prompt points, outperforming state-of-the-art foundation models and task-specific fully supervised models with higher accuracy and strong generalization ability for universal segmentation. Yichi Zhang 0007, Le Xue, Lanlan Li, Chen Jiang 0006, Yuan Qi 0001 |
ICCV | 2 |
| 2025 | Towards Multi-scenario Generalization: Text-Guided Unified Framework for Low-Dose CT and Total-Body PET Reconstruction
Yanyan Huang, Shunjie Dong, Le Xue, Kuangyu Shi, Yu Fu 0008 |
MICCAI (2) | 4 |
| 2025 | SemiSAM+: Rethinking semi-supervised medical image segmentation in the era of foundation models
Yichi Zhang 0007, Bohao Lv, Le Xue, Yuan Qi 0001 |
Medical Image Anal. | 3 |
| 2024 | ULIP-2: Towards Scalable Multimodal Pre-Training for 3D UnderstandingabstractRecent advancements in multimodal pretraining have shown promising efficacy in 3D representation learning by aligning multimodal features across 3D shapes, their 2D counterparts, and language descriptions. However, the methods used by existing frameworks to curate such multimodal data, in particular language descriptions for 3D shapes, are not scalable, and the collected language descriptions are not diverse. To address this, we introduce ULIP-2, a simple yet effective tri-modal pretraining framework that leverages large multimodal models to automatically generate holistic language descriptions for 3D shapes. It only needs 3D data as input, eliminating the need for any manual 3D annotations, and is therefore scalable to large datasets. ULIP-2 is also equipped with scaled-up backbones for better multimodal representation learning. We conduct experiments on two large-scale 3D datasets, Objaverse and ShapeNet, and augment them with tri-modal datasets of 3D point clouds, images, and language for training ULIP-2. Experiments show that ULIP-2 demonstrates substantial benefits in three downstream tasks: zero-shot 3D classification, standard 3D classification with fine-tuning, and 3D captioning (3D-to-language generation). It achieves a new SOTA of 50.6% (top-1) on Objaverse-LVIS and 84.7% (top-1) on ModelNet40 in zero-shot classification. In the ScanObjectNN benchmark for standard fine-tuning, ULIP-2 reaches an overall accuracy of 91.5% with a compact model of only 1.4 million parameters. ULIP-2 sheds light on a new paradigm for scalable multimodal 3D representation learning without human annotations and shows significant improvements over existing baselines. The code and datasets are released at https://github.com/salesforce/ULIP. Le Xue, Ning Yu 0006, Shu Zhang 0007, Artemis Panagopoulou, Junnan Li 0001, Roberto Martin Martin, Jiajun Wu 0001, Caiming Xiong, Ran Xu 0001, Juan Carlos Niebles, Silvio Savarese |
CVPR | 1 |
| 2024 | X-InstructBLIP: A Framework for Aligning Image, 3D, Audio, Video to LLMs and its Emergent Cross-Modal Reasoning
Artemis Panagopoulou, Le Xue, Ning Yu 0006, Junnan Li 0001, Dongxu Li 0003, Shafiq R. Joty, Ran Xu 0001, Silvio Savarese, Caiming Xiong, Juan Carlos Niebles |
ECCV (45) | 2 |
| 2024 | Retroformer: Retrospective Large Language Agents with Policy Gradient OptimizationabstractRecent months have seen the emergence of a powerful new trend in which large language models (LLMs) are augmented to become autonomous language agents capable of performing objective oriented multi-step tasks on their own, rather than merely responding to queries from human users. Most existing language agents, however, are not optimized using environment-specific rewards. Although some agents enable iterative refinement through verbal feedback, they do not reason and plan in ways that are compatible with gradient-based learning from rewards. This paper introduces a principled framework for reinforcing large language agents by learning a retrospective model, which automatically tunes the language agent prompts from environment feedback through policy gradient. Specifically, our proposed agent architecture learns from rewards across multiple environments and tasks, for fine-tuning a pre-trained language model which refines the language agent prompt by summarizing the root cause of prior failed attempts and proposing action plans. Experimental results on various tasks demonstrate that the language agents improve over time and that our approach considerably outperforms baselines that do not properly leverage gradients from the environment. Weiran Yao, Shelby Heinecke, Juan Carlos Niebles, Zhiwei Liu 0001, Yihao Feng, Le Xue, Rithesh R. N., Zeyuan Chen 0001, Jianguo Zhang 0005, Devansh Arpit, Ran Xu 0001, Phil Mui, Huan Wang 0016, Caiming Xiong, Silvio Savarese |
ICLR | 6 |
| 2024 | Hierarchical Point Attention for Indoor 3D Object Detectionabstract3D object detection is an essential vision technique for various robotic systems, such as augmented reality and domestic robots. Transformers as versatile network architectures have recently seen great success in 3D point cloud object detection. However, the lack of hierarchy in a plain transformer restrains its ability to learn features at different scales. Such limitation makes transformer detectors perform worse on smaller objects and affects their reliability in indoor environments where small objects are the majority. This work proposes two novel attention operations as generic hierarchical designs for point-based transformer detectors. First, we propose Aggregated Multi-Scale Attention (MS-A) that builds multi-scale tokens from a single-scale input feature to enable more fine-grained feature learning. Second, we propose Size-Adaptive Local Attention (Local-A) with adaptive attention regions for localized feature aggregation within bounding box proposals. Both attention operations are model-agnostic network modules that can be plugged into existing point cloud transformers for end-to-end training. We evaluate our method on two widely used indoor detection benchmarks. By plugging our proposed modules into the state-of-the-art transformer-based 3D detectors, we improve the previous best results on both benchmarks, with more significant improvements on smaller objects. Manli Shu, Le Xue, Ning Yu 0006, Roberto Martin Martin, Caiming Xiong, Tom Goldstein, Juan Carlos Niebles, Ran Xu 0001 |
ICRA | 2 |
| 2024 | MINT-1T: Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion TokensabstractMultimodal interleaved datasets featuring free-form interleaved sequences of images and text are crucial for training frontier large multimodal models (LMMs). Despite the rapid progression of open-source LMMs, there remains a pronounced scarcity of large-scale, open-source multimodal interleaved datasets.In response, we introduce MINT-1T, the most extensive and diverse open-source Multimodal INTerleaved dataset to date. MINT-1T comprises of one trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. As scaling multimodal interleaved datasets requires substantial engineering effort, sharing the data curation process and releasing the dataset greatly benefits the community. Our experiments show that LMMs trained on MINT-1T rival the performance of models trained on the previous leading dataset, OBELICS. We release our data at https://github.com/mlfoundations/MINT-1T. Anas Awadalla, Le Xue, Oscar Lo, Manli Shu, Hannah Lee, Etash Kumar Guha, Sheng Shen 0001, Mohamed Awadalla, Silvio Savarese, Caiming Xiong, Ran Xu 0001, Yejin Choi 0001, Ludwig Schmidt |
NeurIPS | 2 |
| 2023 | ULIP: Learning a Unified Representation of Language, Images, and Point Clouds for 3D UnderstandingabstractThe recognition capabilities of current state-of-the-art 3D models are limited by datasets with a small number of annotated data and a pre-defined set of categories. In its 2D counterpart, recent advances have shown that similar problems can be significantly alleviated by employing knowledge from other modalities, such as language. Inspired by this, leveraging multimodal information for 3D modality could be promising to improve 3D understanding under the restricted data regime, but this line of research is not well studied. Therefore, we introduce ULIP to learn a unified representation of image, text, and 3D point cloud by pre-training with object triplets from the three modalities. To overcome the shortage of training triplets, ULIP leverages a pre-trained vision-language model that has already learned a common visual and textual space by training with massive image-text pairs. Then, ULIP learns a 3D representation space aligned with the common image-text space, using a small number of automatically synthesized triplets. ULIP is agnostic to 3D backbone networks and can easily be integrated into any 3D architecture. Experiments show that ULIP effectively improves the performance of multiple recent 3D backbones by simply pre-training them on ShapeNet55 using our framework, achieving state-of-the-art performance in both standard 3D classification and zero-shot 3D classification on ModelNet40 and ScanObjectNN. ULIP also improves the performance of PointMLP by around 3% in 3D classification on ScanObjectNN, and outperforms PointCLIP by 28.8% on top-1 accuracy for zero-shot 3D classification on ModelNet40. Our code and pre-trained models will be released. Le Xue, Mingfei Gao, Chen Xing, Roberto Martin Martin, Jiajun Wu 0001, Caiming Xiong, Ran Xu 0001, Juan Carlos Niebles, Silvio Savarese |
CVPR | 1 |
| 2023 | Robustness Evaluation of Transformer-Based Form Field Extractors via Form Attacks
Le Xue, Mingfei Gao, Zeyuan Chen 0001, Caiming Xiong, Ran Xu 0001 |
ICDAR (2) | 1 |
| 2023 | AIGAN: Attention-encoding Integrated Generative Adversarial Network for the reconstruction of low-dose CT and low-dose PET images
Yu Fu 0008, Shunjie Dong, Meng Niu, Le Xue, Hanning Guo, Yanyan Huang, Yuanfan Xu, Tianbai Yu, Kuangyu Shi, Qianqian Yang 0002, Yiyu Shi 0001, Cheng Zhuo |
Medical Image Anal. | 4 |
| 2022 | DocQueryNet: Value Retrieval with Arbitrary Queries for Form-like DocumentsabstractWe propose, DocQueryNet, a value retrieval method with arbitrary queries for form-like documents to reduce human effort of processing forms. Unlike previous methods that only address a fixed set of field items, our method predicts target value for an arbitrary query based on the understanding of the layout and semantics of a form. To further boost model performance, we propose a simple document language modeling (SimpleDLM) strategy to improve document understanding on large-scale model pre-training. Experimental results show that DocQueryNet outperforms previous designs significantly and the SimpleDLM further improves our performance on value retrieval by around 17% F1 score compared with the state-of-the-art pre-training method. Code is available here, https://github.com/salesforce/QVR-SimpleDLM. Mingfei Gao, Le Xue, Chetan Ramaiah, Chen Xing, Ran Xu 0001, Caiming Xiong |
COLING | 2 |
| 2021 | Cross-Modality Generation of Amyloid PET from FDG PET for Alzheimer's Disease DiagnosisabstractPositron Emission Tomography (PET) has been widely used in the early diagnosis and treatment monitoring of Alzheimer’s Disease (AD). As two radiotracers of neurodegeneration, [18F]Fluorodeoxyglucose ([18F]FDG) and [18F]Florbetapir ([18F]AV45) PET have been used to measure cerebral glucose metabolism and $\beta$-amyloid $(A\beta)$ deposition, respectively. The combination of different modality PET images, such as FDG PET and AV45 PET, can provide complementary information for clinical diagnosis and evaluation. However, compared to the actual and available FDG PET data, AV45 PET data is always deficient due to the institution-specific tracers. In this paper, we propose a lightweight Generative Adversarial Network (GAN)-based model, which is termed “dual perceptual loss based generative adversarial network (DPGAN) for fast 2. 5D-based cross-modality generation of AV45 PET from FDG PET. This model provides a potential supplementary solution to those clinical situations that only the FDG PET image is acquired, but the AV45 PET is missing. Our experimental results showed that the DPGAN outperformed recent CycleGAN and pGAN, given its stronger ability in capturing the $A \beta$ deposition patterns on the whole-brain scale. All qualitative and quantitative metrics demonstrated the strong similarity between the generated AV45 PET images using DPGAN and the original AV45 PET images. Yu Fu 0008, Le Xue, Meng Niu, Cheng Zhuo |
BIBM | 2 |