Runjia Li

dblp:352/5264 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
9since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Vision and language · 42% Segmentation and scene understanding · 17% 3D vision · 13%
Computer graphics and multimedia
2 papers
Visual content generation and editing · 90% Multimedia analysis and retrieval · 10%

Topics — the 23 heaviest of 25, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
vision-language model
1.622025
Generalized Few-shot 3D Point Cloud Segmentation with Vision-Language Model · CVPR 2025
CLIP as RNN: Segment Countless Visual Concepts without Training Endeavor · CVPR 2024
Computer vision › Segmentation and scene understanding
3d point cloud segmentation
0.912025
Multimodality Helps Few-shot 3D Point Cloud Semantic Segmentation · ICLR 2025
Computer vision › Vision and language › vision-language model
3d vision-language model
0.912025
Generalized Few-shot 3D Point Cloud Segmentation with Vision-Language Model · CVPR 2025
Machine learning › Optimization for machine learning
convergence analysis
0.912025
Unified Convergence Analysis for Score-Based Diffusion Models with Deterministic Samplers · ICLR 2025
Machine learning › Generative modeling
diffusion model
0.912025
Unified Convergence Analysis for Score-Based Diffusion Models with Deterministic Samplers · ICLR 2025
Computer vision › 3D vision › point cloud segmentation
few-shot point cloud segmentation
0.912025
Multimodality Helps Few-shot 3D Point Cloud Semantic Segmentation · ICLR 2025
Computer vision › 3D vision › point cloud segmentation
point cloud semantic segmentation
0.912025
Generalized Few-shot 3D Point Cloud Segmentation with Vision-Language Model · CVPR 2025
Machine learning › Optimization for machine learning › convergence analysis
sampling convergence
0.912025
Unified Convergence Analysis for Score-Based Diffusion Models with Deterministic Samplers · ICLR 2025
Machine learning › Generative modeling › diffusion model
score-based generative model
0.912025
Unified Convergence Analysis for Score-Based Diffusion Models with Deterministic Samplers · ICLR 2025
Computer vision › Vision and language › visual grounding
text-guided segmentation
0.912025
Multimodality Helps Few-shot 3D Point Cloud Semantic Segmentation · ICLR 2025
Visual content generation and editing › scene synthesis
interactive scene synthesis
0.912025
VMem: Consistent Interactive Video Scene Generation with Surfel-Indexed View Memory · ICCV 2025
Visual content generation and editing
video generation
0.912025
VMem: Consistent Interactive Video Scene Generation with Surfel-Indexed View Memory · ICCV 2025
Computer vision › Vision and language › image captioning
cross-lingual image captioning
0.812024
No Culture Left Behind: ArtELingo-28, a Benchmark of WikiArt with Captions in 28 Languages · EMNLP 2024
Computer vision › Segmentation and scene understanding › semantic segmentation
open-vocabulary segmentation
0.812024
CLIP as RNN: Segment Countless Visual Concepts without Training Endeavor · CVPR 2024
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation
0.812024
Label Alignment Regularization for Distribution Shift · J. Mach. Learn. Res. 2024
Computer vision › Segmentation and scene understanding › semantic segmentation › open-vocabulary segmentation
zero-shot semantic segmentation
0.812024
CLIP as RNN: Segment Countless Visual Concepts without Training Endeavor · CVPR 2024
Computer vision › Vision and language › multimodal understanding
humor understanding
0.712023
OxfordTVG-HIC: Can Machine Make Humorous Captions from Images? · ICCV 2023
Computer vision › Vision and language
image captioning
0.712023
OxfordTVG-HIC: Can Machine Make Humorous Captions from Images? · ICCV 2023
Computer vision › Vision and language
multimodal understanding
0.712023
OxfordTVG-HIC: Can Machine Make Humorous Captions from Images? · ICCV 2023
Computer vision › 3D vision
novel view synthesis
0.312025
VMem: Consistent Interactive Video Scene Generation with Surfel-Indexed View Memory · ICCV 2025
Computer vision › Vision and language
art image understanding
0.212024
No Culture Left Behind: ArtELingo-28, a Benchmark of WikiArt with Captions in 28 Languages · EMNLP 2024
Computer vision › Segmentation and scene understanding
referring image segmentation
0.212024
CLIP as RNN: Segment Countless Visual Concepts without Training Endeavor · CVPR 2024
Multimedia analysis and retrieval › multimedia dataset construction
multimodal dataset
0.212023
OxfordTVG-HIC: Can Machine Make Humorous Captions from Images? · ICCV 2023

Methods — techniques the papers use, named apart from their topics

surfel-indexed view memory · 1.7test-time adaptive cross-modal calibration · 0.9pseudo-label selection · 0.9novel-base mix · 0.9multimodal correlation fusion · 0.9exponential integrator · 0.9adaptive infilling · 0.9DDIM · 0.9recurrent framework · 0.8frozen vision-language model · 0.8explainability analysis · 0.7
YearPublicationVenuePosition
2025 DreamBeast: Distilling 3D Fantastical Animals with Part-Aware Knowledge Transfer
abstract
We present DreamBeast, a novel method based on score distillation sampling (SDS) for generating fantastical 3D animal assets composed of distinct parts. Existing SDS methods often struggle with this generation task due to a limited understanding of part-level semantics in text-to-image diffusion models. While recent diffusion models, such as Stable Diffusion 3, demonstrate a better part-level understanding, they are prohibitively slow and exhibit other common problems associated with single-view diffusion models. DreamBeast overcomes this limitation through a novel part-aware knowledge transfer mechanism. For each generated asset, we efficiently extract part-level knowledge from the Stable Diffusion 3 model into a 3D Part-Affinity implicit representation. This enables us to instantly generate Part-Affinity maps from arbitrary camera views, which we then use to modulate the guidance of a multi-view diffusion model during SDS to create 3D assets of fantastical animals. DreamBeast significantly enhances the quality of generated 3D creatures with user-specified part compositions while reducing computational overhead, as demonstrated by extensive quantitative and qualitative evaluations.
Runjia Li, Junlin Han, Luke Melas-Kyriazi, Chunyi Sun, Zhaochong An, Zhongrui Gui, Shuyang Sun, Philip Torr 0001, Tomas Jakab
3DV1
2025 Generalized Few-shot 3D Point Cloud Segmentation with Vision-Language Model
abstract
Generalized few-shot 3D point cloud segmentation (GFS-PCS) adapts models to new classes with few support samples while retaining base class segmentation. Existing GFS-PCS methods enhance prototypes via interacting with support or query features but remain limited by sparse knowledge from few-shot samples. Meanwhile, 3D vision-language models (3D VLMs), generalizing across open-world novel classes, contain rich but noisy novel class knowledge. In this work, we introduce a GFS-PCS framework that synergizes dense but noisy pseudo-labels from 3D VLMs with precise yet sparse few-shot samples to maximize the strengths of both, named GFS-VL. Specifically, we present a prototype-guided pseudo-label selection to filter low-quality regions, followed by an adaptive infilling strategy that combines knowledge from pseudo-label contexts and few-shot samples to adaptively label the filtered, un-labeled areas. Additionally, we design a novel-base mix strategy to embed few-shot samples into training scenes, preserving essential context for improved novel class learning. Moreover, recognizing the limited diversity in current GFS-PCS benchmarks, we introduce two challenging benchmarks with diverse novel classes for comprehensive generalization evaluation. Experiments validate the effectiveness of our framework across models and datasets. Our approach and benchmarks provide a solid foundation for advancing GFS-PCS in the real world. The code is at here.
Zhaochong An, Guolei Sun, Yun Liu 0011, Runjia Li, Junlin Han, Ender Konukoglu, Serge J. Belongie
CVPR4
2025 VMem: Consistent Interactive Video Scene Generation with Surfel-Indexed View Memory
Runjia Li, Philip Torr 0001, Andrea Vedaldi, Tomas Jakab
ICCV1
2025 Multimodality Helps Few-shot 3D Point Cloud Semantic Segmentation
abstract
Few-shot 3D point cloud segmentation (FS-PCS) aims at generalizing models to segment novel categories with minimal annotated support samples. While existing FS-PCS methods have shown promise, they primarily focus on unimodal point cloud inputs, overlooking the potential benefits of leveraging multimodal information. In this paper, we address this gap by introducing a multimodal FS-PCS setup, utilizing textual labels and the potentially available 2D image modality. Under this easy-to-achieve setup, we present the MultiModal Few-Shot SegNet (MM-FSS), a model effectively harnessing complementary information from multiple modalities. MM-FSS employs a shared backbone with two heads to extract intermodal and unimodal visual features, and a pretrained text encoder to generate text embeddings. To fully exploit the multimodal information, we propose a Multimodal Correlation Fusion (MCF) module to generate multimodal correlations, and a Multimodal Semantic Fusion (MSF) module to refine the correlations using text-aware semantic guidance. Additionally, we propose a simple yet effective Test-time Adaptive Cross-modal Calibration (TACC) technique to mitigate training bias, further improving generalization. Experimental results on S3DIS and ScanNet datasets demonstrate significant performance improvements achieved by our method. The efficacy of our approach indicates the benefits of leveraging commonly-ignored free modalities for FS-PCS, providing valuable insights for future research. The code is available at github.com/ZhaochongAn/Multimodality-3D-Few-Shot.
Zhaochong An, Guolei Sun, Yun Liu 0011, Runjia Li, Min Wu 0008, Ming-Ming Cheng, Ender Konukoglu, Serge J. Belongie
ICLR4
2025 Unified Convergence Analysis for Score-Based Diffusion Models with Deterministic Samplers
abstract
Score-based diffusion models have emerged as powerful techniques for generating samples from high-dimensional data distributions. These models involve a two-phase process: first, injecting noise to transform the data distribution into a known prior distribution, and second, sampling to recover the original data distribution from noise. Among the various sampling methods, deterministic samplers stand out for their enhanced efficiency. However, analyzing these deterministic samplers presents unique challenges, as they preclude the use of established techniques such as Girsanov's theorem, which are only applicable to stochastic samplers. Furthermore, existing analysis for deterministic samplers usually focuses on specific examples, lacking a generalized approach for general forward processes and various deterministic samplers. Our paper addresses these limitations by introducing a unified convergence analysis framework. To demonstrate the power of our framework, we analyze the variance-preserving (VP) forward process with the exponential integrator (EI) scheme, achieving iteration complexity of $\tilde{O}(d^2/\epsilon)$. Additionally, we provide a detailed analysis of Denoising Diffusion Implicit Models (DDIM)-type samplers, which have been underexplored in previous research, achieving polynomial iteration complexity.
Runjia Li, Qiwei Di, Quanquan Gu
ICLR1
2024 CLIP as RNN: Segment Countless Visual Concepts without Training Endeavor
abstract
Existing open-vocabulary image segmentation methods re-quire a fine-tuning step on mask labels and/or image-text datasets. Mask labels are labor-intensive, which limits the number of categories in segmentation datasets. Con-sequently, the vocabulary capacity of pre-trained VLMs is severely reduced after fine-tuning. However, without fine-tuning, VLMs trained under weak image-text supervision tend to make suboptimal mask predictions. To alleviate these issues, we introduce a novel recurrent framework that progressively filters out irrelevant texts and enhances mask quality without training efforts. The recurrent unit is a two-stage segmenter built upon a frozen VLM. Thus, our model retains the VLM's broad vocabulary space and equips it with segmentation ability. Experiments show that our method outperforms not only the training-free counter-parts, but also those fine-tuned with millions of data sam-ples, and sets the new state-of-the-art records for both zero-shot semantic and referring segmentation. Concretely, we improve the current record by 28.8, 16.0, and 6.9 mloU on Pascal VOC, COCO Object, and Pascal Context.
Shuyang Sun, Runjia Li, Philip Torr 0001, Xiuye Gu, Siyang Li 0002
CVPR2
2024 No Culture Left Behind: ArtELingo-28, a Benchmark of WikiArt with Captions in 28 Languages
abstract
Youssef Mohamed, Runjia Li, Ibrahim Said Ahmad, Kilichbek Haydarov, Philip Torr, Kenneth Church, Mohamed Elhoseiny. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Youssef Mohamed, Runjia Li, Ibrahim Said Ahmad, Kilichbek Haydarov, Philip Torr 0001, Kenneth Church 0001, Mohamed Elhoseiny 0001
EMNLP2
2024 Label Alignment Regularization for Distribution Shift
abstract
Recent work has highlighted the label alignment property (LAP) in supervised learning, where the vector of all labels in the dataset is mostly in the span of the top few singular vectors of the data matrix. Drawing inspiration from this observation, we propose a regularization method for unsupervised domain adaptation that encourages alignment between the predictions in the target domain and its top singular vectors. Unlike conventional domain adaptation approaches that focus on regularizing representations, we instead regularize the classifier to align with the unsupervised target data, guided by the LAP in both the source and target domains. Theoretical analysis demonstrates that, under certain assumptions, our solution resides within the span of the top right singular vectors of the target domain data and aligns with the optimal solution. By removing the reliance on the commonly used optimal joint risk assumption found in classic domain adaptation theory, we showcase the effectiveness of our method on addressing problems where traditional domain adaptation methods often fall short due to high joint error. Additionally, we report improved performance over domain adaptation baselines in well-known tasks such as MNIST-USPS domain adaptation and cross-lingual sentiment analysis. An implementation is available at https://github.com/EhsanEI/lar/.
Ehsan Imani, Runjia Li, Jun Luo 0009, Pascal Poupart, Philip Torr 0001, Yangchen Pan
J. Mach. Learn. Res.3
2023 OxfordTVG-HIC: Can Machine Make Humorous Captions from Images?
abstract
This paper presents OxfordTVG-HIC (Humorous Image Captions), a large-scale dataset for humour generation and understanding. Humour is an abstract, subjective, and context-dependent cognitive construct involving several cognitive factors, making it a challenging task to generate and interpret. Hence, humour generation and understanding can serve as a new task for evaluating the ability of deep-learning methods to process abstract and subjective information. Due to the scarcity of data, humourrelated generation tasks such as captioning remain underexplored. To address this gap, OxfordTVG-HIC offers approximately 2.9M image-text pairs with humour scores to train a generalizable humour captioning model. Contrary to existing captioning datasets, OxfordTVG-HIC features a wide range of emotional and semantic diversity resulting in out-of-context examples that are particularly conducive to generating humour. Moreover, OxfordTVG-HIC is curated devoid of offensive content. We also show how OxfordTVG-HIC can be leveraged for evaluating the humour of a generated text. Through explainability analysis of the trained models, we identify the visual and linguistic cues influential for evoking humour prediction (and generation). We observe qualitatively that these cues are aligned with the benign violation theory of humour in cognitive psychology.
Runjia Li, Shuyang Sun, Mohamed Elhoseiny 0001, Philip Torr 0001
ICCV1