EDBT 2026 Demo / reviewers in the wild / expert
Ruichuan An
dblp:362/5891
· DBLP profile ↗
9ranked-venue papers
1as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Vision and language · 41% Generative modeling · 20% Transfer learning and domain adaptation · 13% | |
| Computer graphics and multimedia
1 paper |
Multimedia analysis and retrieval · 100% |
Topics — the 22 heaviest of 25, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language
vision-language model |
1.7 | 2 | 2025 | UniCTokens: Boosting Personalized Understanding and Generation via Unified Concept Tokens · NeurIPS 2025 MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders · CVPR 2025 |
Multimedia analysis and retrieval › cross-modal retrieval
text-video retrieval |
1.0 | 1 | 2026 | LoVR: A Benchmark for Long Video Retrieval in Multimodal Contexts · WWW 2026 |
Multimedia analysis and retrieval
video retrieval |
1.0 | 1 | 2026 | LoVR: A Benchmark for Long Video Retrieval in Multimodal Contexts · WWW 2026 |
Machine learning › Generative modeling › diffusion model › guided diffusion
classifier-free guidance |
0.9 | 1 | 2025 | Adaptive Classifier-Free Guidance via Dynamic Low-Confidence Masking · NeurIPS 2025 |
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | Adaptive Classifier-Free Guidance via Dynamic Low-Confidence Masking · NeurIPS 2025 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.9 | 1 | 2025 | MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders · CVPR 2025 |
Machine learning › Generative modeling › diffusion model › discrete diffusion model
masked diffusion language model |
0.9 | 1 | 2025 | Adaptive Classifier-Free Guidance via Dynamic Low-Confidence Masking · NeurIPS 2025 |
Machine learning › Efficient and distributed learning
model compression |
0.9 | 1 | 2025 | MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders · CVPR 2025 |
Computer vision › Vision and language › vision-language model › multimodal large language model
multimodal instruction following |
0.9 | 1 | 2025 | Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want · ICLR 2025 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.9 | 1 | 2025 | Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want · ICLR 2025 |
Computer vision › Segmentation and scene understanding
object segmentation |
0.9 | 1 | 2025 | Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos · NeurIPS 2025 |
Machine learning › Generative modeling › diffusion model
personalized image generation |
0.9 | 1 | 2025 | UniCTokens: Boosting Personalized Understanding and Generation via Unified Concept Tokens · NeurIPS 2025 |
Computer vision › Segmentation and scene understanding › image segmentation
region-based segmentation |
0.9 | 1 | 2025 | Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos · NeurIPS 2025 |
Computer vision › Vision and language › image captioning › grounded image captioning
region captioning |
0.9 | 1 | 2025 | Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos · NeurIPS 2025 |
Computer vision › Vision and language
region-level understanding |
0.9 | 1 | 2025 | Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos · NeurIPS 2025 |
Computer vision › Vision and language
visual prompting |
0.9 | 1 | 2025 | Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want · ICLR 2025 |
Computer vision › Vision and language › multimodal understanding
visual prompt understanding |
0.9 | 1 | 2025 | Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want · ICLR 2025 |
Machine learning › Transfer learning and domain adaptation › domain adaptation
graph domain adaptation |
0.8 | 1 | 2024 | Can Modifying Data Address Graph Domain Adaptation? · KDD 2024 |
Natural language and speech › Language models and text generation
language model analysis |
0.8 | 1 | 2024 | LLM as Dataset Analyst: Subpopulation Structure Discovery with Large Language Model · ECCV (33) 2024 |
Machine learning › Transfer learning and domain adaptation › domain adaptation › graph domain adaptation
unsupervised graph domain adaptation |
0.8 | 1 | 2024 | Can Modifying Data Address Graph Domain Adaptation? · KDD 2024 |
Computer vision › Vision and language
video captioning |
0.3 | 1 | 2026 | LoVR: A Benchmark for Long Video Retrieval in Multimodal Contexts · WWW 2026 |
Natural language and speech › Language models and text generation
text generation |
0.3 | 1 | 2025 | Adaptive Classifier-Free Guidance via Dynamic Low-Confidence Masking · NeurIPS 2025 |
Methods — techniques the papers use, named apart from their topics
large language model · 2.4vision-language model · 2.0semantic fusion · 2.0unified concept tokens · 0.9progressive training · 0.9multimodal large language model · 0.9mixture of experts · 0.9low-rank adaptation · 0.9instruction tuning · 0.9attention-based distillation · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LoVR: A Benchmark for Long Video Retrieval in Multimodal ContextsabstractLong videos contain a vast amount of information, making video-text retrieval an essential and challenging task in multimodal learning and web-scale search. On today's Web, where users increasingly expect to locate not only relevant pages but also specific long videos or fine-grained clips, existing benchmarks fall short due to limited video duration, low-quality captions, and coarse annotation granularity. To address these limitations, we introduce LoVR, a benchmark specifically designed for long video-text retrieval. LoVR contains 467 long videos and over 40,804 fine-grained clips with high-quality captions. To overcome the issue of poor machine-generated annotations, we propose an efficient caption generation framework that integrates VLM automatic generation, caption quality scoring, and dynamic refinement. This pipeline improves annotation accuracy while maintaining scalability. Furthermore, we introduce a semantic fusion method to generate coherent full-video captions without losing important contextual information. Our benchmark introduces longer videos, more detailed captions, and a larger-scale dataset, presenting new challenges for video understanding and retrieval. Extensive experiments on various advanced models demonstrate that LoVR is a challenging benchmark, revealing the limitations of current approaches and providing valuable insights for future research. We release the code link at https://github.com/TechNomad-ds/LoVR-benchmark/. Hao Liang 0017, Qifeng Cai, Hejun Dong, Meiyi Qiang, Ruichuan An, Quanqing Xu, Bin Cui 0001, Wentao Zhang 0001 |
WWW | 6 |
| 2025 | MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual EncodersabstractVisual encoders are fundamental components in vision-language models (VLMs), each showcasing unique strengths derived from various pre-trained visual foundation models. To leverage the various capabilities of these encoders, recent studies incorporate multiple encoders within a single VLM, leading to a considerable increase in computational cost. In this paper, we present Mixture-of-Visual-Encoder Knowledge Distillation (MoVEKD), a novel framework that distills the unique proficiencies of multiple vision encoders into a single, efficient encoder model. Specifically, to mitigate conflicts and retain the unique characteristics of each teacher encoder, we employ low-rank adaptation (LoRA) and mixture-of-experts (MoEs) to selectively activate specialized knowledge based on input features, enhancing both adaptability and efficiency. To regularize the KD process and enhance performance, we propose an attention-based distillation strategy that adaptively weighs the different encoders and emphasizes valuable visual tokens, reducing the burden of replicating comprehensive but distinct features from multiple teachers. Comprehensive experiments on popular VLMs, such as LLaVA and LLaVA-NeXT, validate the effectiveness of our method. Our code is available at: https://github.com/hey-cjj/MoVE-KD. Jiajun Cao, Yuan Zhang 0020, Tao Huang 0020, Ming Lu 0002, Qizhe Zhang, Ruichuan An, Ningning Ma, Shanghang Zhang |
CVPR | 6 |
| 2025 | Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You WantabstractIn this paper, we present the Draw-and-Understand framework, exploring how to integrate visual prompting understanding capabilities into Multimodal Large Language Models (MLLMs). Visual prompts allow users to interact through multi-modal instructions, enhancing the models' interactivity and fine-grained image comprehension. In this framework, we propose a general architecture adaptable to different pre-trained MLLMs, enabling it to recognize various types of visual prompts (such as points, bounding boxes, and free-form shapes) alongside language understanding. Additionally, we introduce MDVP-Instruct-Data, a multi-domain dataset featuring 1.2 million image-visual prompt-text triplets, including natural images, document images, scene text images, mobile/web screenshots, and remote sensing images. Building on this dataset, we introduce MDVP-Bench, a challenging benchmark designed to evaluate a model's ability to understand visual prompting instructions. The experimental results demonstrate that our framework can be easily and effectively applied to various MLLMs, such as SPHINX-X and LLaVA. After training with MDVP-Instruct-Data and image-level instruction datasets, our models exhibit impressive multimodal interaction capabilities and pixel-level understanding, while maintaining their image-level visual perception performance. Weifeng Lin, Ruichuan An, Peng Gao 0007, Bocheng Zou, Yulin Luo, Siyuan Huang 0004, Shanghang Zhang, Hongsheng Li 0001 |
ICLR | 3 |
| 2025 | UniCTokens: Boosting Personalized Understanding and Generation via Unified Concept TokensabstractPersonalized models have demonstrated remarkable success in understanding and generating concepts provided by users. However, existing methods use separate concept tokens for understanding and generation, treating these tasks in isolation. This may result in limitations for generating images with complex prompts. For example, given the concept $\langle bo\rangle$, generating "$\langle bo\rangle$ wearing its hat" without additional textual descriptions of its hat. We call this kind of generation \textit{\textbf{personalized attribute-reasoning generation}}. To address the limitation, we present UniCTokens, a novel framework that effectively integrates personalized information into a unified vision language model (VLM) for understanding and generation. UniCTokens trains a set of unified concept tokens to leverage complementary semantics, boosting two personalized tasks. Moreover, we propose a progressive training strategy with three stages: understanding warm-up, bootstrapping generation from understanding, and deepening understanding from generation to enhance mutual benefits between both tasks. To quantitatively evaluate the unified VLM personalization, we present UnifyBench, the first benchmark for assessing concept understanding, concept generation, and attribute-reasoning generation. Experimental results on UnifyBench indicate that UniCTokens shows competitive performance compared to leading methods in concept understanding, concept generation, and achieving state-of-the-art results in personalized attribute-reasoning generation. Our research demonstrates that enhanced understanding improves generation, and the generation process can yield valuable insights into understanding. Our code and dataset will be released at: \href{https://github.com/arctanxarc/UniCTokens}{https://github.com/arctanxarc/UniCTokens}. Ruichuan An, Renrui Zhang, Zijun Shen, Ming Lu 0002, Gaole Dai, Hao Liang 0017, Shilin Yan, Yulin Luo, Bocheng Zou, Wentao Zhang 0001 |
NeurIPS | 1 |
| 2025 | Adaptive Classifier-Free Guidance via Dynamic Low-Confidence MaskingabstractClassifier-Free Guidance (CFG) significantly enhances controllability in generative models by interpolating conditional and unconditional predictions. However, standard CFG often employs a static unconditional input, which can be suboptimal for iterative generation processes where model uncertainty varies dynamically. We introduce Adaptive Classifier-Free Guidance (A-CFG), a novel method that tailors the unconditional input by leveraging the model's instantaneous predictive confidence. At each step of an iterative (masked) diffusion language model, A-CFG identifies tokens in the currently generated sequence for which the model exhibits low confidence. These tokens are temporarily re-masked to create a dynamic, localized unconditional input. This focuses CFG's corrective influence precisely on areas of ambiguity, leading to more effective guidance. We integrate A-CFG into a state-of-the-art masked diffusion language model and demonstrate its efficacy. Experiments on diverse language generation benchmarks show that A-CFG yields substantial improvements over standard CFG, achieving, for instance, a 3.9 point gain on GPQA. Our work highlights the benefit of dynamically adapting guidance mechanisms to model uncertainty in iterative generation. Shilin Yan, Jiayin Cai, Renrui Zhang, Ruichuan An |
NeurIPS | 5 |
| 2025 | Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and VideosabstractWe present Perceive Anything Model (PAM), a conceptually straightforward and efficient framework for comprehensive region-level visual understanding in images and videos. Our approach extends the powerful segmentation model SAM 2 by integrating Large Language Models (LLMs), enabling simultaneous object segmentation with the generation of diverse, region-specific semantic outputs, including categories, label definition, functional explanations, and detailed captions. A key component, Semantic Perceiver, is introduced to efficiently transform SAM 2's rich visual features, which inherently carry general vision, localization, and semantic priors into multi-modal tokens for LLM comprehension. To support robust multi-granularity understanding, we also develop a dedicated data refinement and augmentation pipeline, yielding a high-quality dataset of 1.5M image and 0.6M video region-semantic annotations, including novel region-level streaming video caption data. PAM is designed for lightweightness and efficiency, while also demonstrates strong performance across a diverse range of region understanding tasks. It runs 1.2$-$2.4$\times$ faster and consumes less GPU memory than prior approaches, offering a practical solution for real-world applications. We believe that our effective approach will serve as a strong baseline for future research in region-level visual understanding. Weifeng Lin, Ruichuan An, Tianhe Ren, Renrui Zhang, Wentao Zhang 0001, Lei Zhang 0006, Hongsheng Li 0001 |
NeurIPS | 3 |
| 2025 | CodeRankEval: Benchmarking and Analyzing LLM Performance for Code Ranking
Li-Guo Chen, Yi-Jiang Xu, Ruichuan An, Yang-Ning Li, Ying-Hui Li, Yi-Dong Wang, Zheng-Ran Zeng, Shi-Kun Zhang |
J. Comput. Sci. Technol. | 4 |
| 2024 | LLM as Dataset Analyst: Subpopulation Structure Discovery with Large Language Model
Yulin Luo, Ruichuan An, Bocheng Zou, Jiaming Liu 0003, Shanghang Zhang |
ECCV (33) | 2 |
| 2024 | Can Modifying Data Address Graph Domain Adaptation?abstractGraph neural networks (GNNs) have demonstrated remarkable success in numerous graph analytical tasks. Yet, their effectiveness is often compromised in real-world scenarios due to distribution shifts, limiting their capacity for knowledge transfer across changing environments or domains. Recently, Unsupervised Graph Domain Adaptation (UGDA) has been introduced to resolve this issue. UGDA aims to facilitate knowledge transfer from a labeled source graph to an unlabeled target graph. Current UGDA efforts primarily focus on model-centric methods, such as employing domain invariant learning strategies and designing model architectures. However, our critical examination reveals the limitations inherent to these model-centric methods, while a data-centric method allowed to modify the source graph provably demonstrates considerable potential. This insight motivates us to explore UGDA from a data-centric perspective. By revisiting the theoretical generalization bound for UGDA, we identify two data-centric principles for UGDA: alignment principle and rescaling principle. Guided by these principles, we propose GraphAlign, a novel UGDA method that generates a small yet transferable graph. By exclusively training a GNN on this new graph with classic Empirical Risk Minimization (ERM), GraphAlign attains exceptional performance on the target graph. Extensive experiments under various transfer scenarios demonstrate the GraphAlign outperforms the best baselines by an average of 2.16%, training on the generated graph as small as 0.25~1% of the original training graph. Renhong Huang, Jiarong Xu, Xin Jiang 0015, Ruichuan An, Yang Yang 0009 |
KDD | 4 |