EDBT 2026 Demo / reviewers in the wild / expert
Yifan Yang 0004
dblp:83/89-4
· DBLP profile ↗
21ranked-venue papers
1as first author
21since 2021 · last 2026
0000-0002-5481-2851ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 1 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 12 since 2021Computer networks · 4 · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning Systems Expansion with Efficient Heterogeneity-aware Knowledge TransferabstractModern AI services must continually adapt to newly joined domains, yet delivering high-quality customized models is hampered by label sparsity, domain shifts, and tight budgets. We formulate this challenge as the learning system expansion problem and introduce HaT, an efficient heterogeneity-aware knowledge-transfer framework. HaT first selects a small set of high-quality source models with minimal overhead, and then fuses their imperfect predictions through a sample-wise attention mixer. Later, it adaptively distills the fused knowledge into target models via a knowledge dictionary. Extensive experiments on different tasks and modalities show that HaT outperforms state-of-the-art baselines by up to 16.5% accuracy, and saves 31.1% training time and up to 93.0% traffic. Gaole Dai, Huatao Xu, Yifan Yang 0004, Rui Tan 0001, Mo Li 0001 |
AAAI | 3 |
| 2026 | LLM2CLIP: Powerful Language Model Unlocks Richer Cross-Modality RepresentationabstractCLIP is a seminal multimodal model that maps images and text into a shared representation space by contrastive learning on billions of image–caption pairs. Inspired by the rapid progress of large language models (LLMs), we investigate how the superior linguistic understanding and broad world knowledge of LLMs can further strengthen CLIP—particularly in handling long, complex captions. We introduce an efficient fine-tuning framework that embeds an LLM into a pretrained CLIP while incurring almost the same training cost as regular CLIP fine-tuning. Our method first “embedding-izes” the LLM for the CLIP setting, then couples it to the pretrained CLIP vision encoder through a lightweight adaptor trained on only a few million image–caption pairs. With this strategy we achieve large performance gains—without large-scale retraining—over state-of-the-art CLIP variants such as EVA02 and SigLIP-2. The LLM-enhanced CLIP delivers consistent improvements across a wide spectrum of downstream tasks, including linear-probe classification, zero-shot image–text retrieval with both short and long captions (in English and other languages), zero-shot/supervised image segmentation, object detection, and used as tokenizer for multimodal large-model benchmarks. Weiquan Huang, Aoqi Wu, Yifan Yang 0004, Xufang Luo, Yuqing Yang 0001, Usman Naseem, Chunyu Wang 0001, Qi Dai 0001, Xiyang Dai, Dongdong Chen 0001, Chong Luo 0001, Lili Qiu, Liang Hu 0004 |
AAAI | 3 |
| 2026 | HiTVideo: Hierarchical Tokenizers for Enhancing Text-to-Video Generation with Autoregressive Large Language ModelsabstractText-to-video generation poses significant challenges due to the inherent complexity of video data, which spans both temporal and spatial dimensions. It introduces additional redundancy, abrupt variations, and a domain gap between language and vision tokens while generation. Addressing these challenges requires an effective video tokenizer that can efficiently encode video data while preserving essential semantic and spatiotemporal information, serving as a critical bridge between text and vision. Inspired by the observation in VQ-VAE-2, we propose HiTVideo, a novel approach for text-to-video generation with hierarchical tokenizers. It utilizes a 3D causal VAE with a multi-layer discrete token framework, encoding video content into hierarchically structured codebooks. Higher layers capture semantic information with higher compression, while lower layers focus on fine-grained spatiotemporal details, striking a balance between compression efficiency and reconstruction quality. Our approach efficiently encodes longer video sequences (e.g., 8 seconds, 64 frames), reducing bits per pixel (bpp) by approximately 70% compared to previous tokenizers, while maintaining competitive reconstruction quality. We explore the trade-offs between compression and reconstruction, while emphasizing the advantages of high-compressed semantic tokens in text-to-video tasks. HiTVideo aims to address the potential limitations of existing video tokenizers in text-to-video generation tasks, striving for higher compression ratios, improved token quality, and simplify LLMs modeling under language guidance, offering a scalable and promising framework for advancing text to video generation. Ziqin Zhou, Yifan Yang 0004, Yuqing Yang 0001, Tianyu He, Houwen Peng, Qi Dai 0001, Lili Qiu, Chong Luo 0001, Lingqiao Liu |
AAAI | 2 |
| 2026 | AVA: Towards Agentic Video Analytics with Vision Language Models
Yuxuan Yan, Shiqi Jiang 0002, Ting Cao 0003, Yifan Yang 0004, Qianqian Yang 0002, Yuanchao Shu, Yuqing Yang 0001, Lili Qiu |
NSDI | 4 |
| 2026 | MageBench: Bridging Large Multimodal Models to AgentsabstractRecent models like OpenAI’s O1 and DeepSeek’s R1, which utilize test-time scaling techniques, have demonstrated remarkable improvements in reasoning capabilities. We anticipate that in the near future, multimodal models will also experience significant breakthroughs in multimodal reasoning. This will require some highly challenging and specialized evaluations. As one of the most crucial real-world applications of multimodal models, visual agents require complex and comprehensive capabilities such as spatial planning and vision-in-the-chain type reasoning. These capabilities are currently lacking in existing multimodal benchmarks. In this paper, we introduce MageBench, a Multimodal reasoning benchmark built upon light-weight AGEnt environments that pose significant reasoning challenges and hold substantial practical value. The results show that only a few product-level models are better than random acting, and all of them are far inferior to human level. We analyze and summarize their errors and capability gaps in visual planning. Furthermore, we found that rule-based RL can significantly boost visual reasoning capabilities. This highlights that our benchmark could serve as a valuable testing ground for the emerging field of agentic RL research. Miaosen Zhang, Qi Dai 0001, Yifan Yang 0004, Jianmin Bao, Dongdong Chen 0001, Chong Luo 0001, Xin Geng 0001, Baining Guo |
WACV | 3 |
| 2026 | AMID: Model-Agnostic Dataset Distillation by Adversarial Mutual Information MinimizationabstractThe escalating energy consumption and carbon footprint of training large-scale Web AI models pose urgent challenges for sustainable development. Dataset Distillation (DD) offers a promising avenue for green AI by compressing large datasets into small synthetic ones for efficient training. However, most existing DD methods overfit to the inductive biases of specific source architectures (e.g., CNNs or ViTs), resulting in poor cross-model generalization. This limitation necessitates redundant re-distillation processes for different architectures, severely undermining the energy-saving potential of DD. To address this, we introduce Adversarial Mutual Information Distillation (AMID), a rigorous framework designed to create highly reusable and robust synthetic datasets. From an information-theoretic perspective, we cast model-agnosticism as minimizing the mutual information (MI) between the synthetic data and the specific identity of the distillation model. We convert this intractable objective into a tractable two-player adversarial game, which unifies knowledge preservation with adversarial unlearning of architectural bias. Extensive experiments on CIFAR-10 and Tiny ImageNet demonstrate that AMID achieves state-of-the-art cross-architecture generalization across diverse CNNs and ViTs. Crucially, our analysis confirms that AMID significantly reduces the computational overhead and CO2 emissions of downstream training while maintaining robust performance, paving the way for energy-efficient, transferable, and sustainable Web AI ecosystems. Aoqi Wu, Weiquan Huang, Liang Hu 0004, Yifan Yang 0004, Qi Zhang 0020, Jiaxing Miao, Yuhan Tang, Zhongyuan Lai |
WWW | 6 |
| 2026 | Joint Latency-Energy Optimization for Two-Tier Multiuser Multitask Offloading in AI-Agent Communication NetworksabstractArtificial intelligence-agent communication networks (ACNs) in the sixth-generation (6G) enable collaborative task execution among agents and butler. However, compared with traditional mobile edge computing (MEC), in ACNs, a large number of agents possess comparable computing capabilities and task proportions need to be jointly determined rather than being predefined, causing high optimization complexity in large-scale deployment scenarios. In this paper, we propose an effective framework to solve the large-scale coupled optimization problem under acceptable complexity. Specifically, we model the joint task allocation, resource allocation, task offloading and computation frequency adjustment problem as an NP-hard nonconvex mixed-integer nonlinear programming (MINLP) problem, and derive its lower bound through Lagrangian relaxation. We decompose the problem into resource allocation and task offloading subproblems, which are solved via proximal policy optimization (PPO) and minimum-cost models within a block coordinate descent (BCD) framework. Simulations demonstrate the tightness of the lower bound, achieving stable convergence for 30 agents while reducing latency from 150 ms to 50 ms and the energy consumption by 28.18%. Our algorithm has great potential for ACNs with many agents deployed in future 6G scenarios, such as autonomous vehicles, robotic swarms and precision telemedicine. Jie Zeng 0001, Yifan Yang 0004, Wei Feng 0001, Tiejun Lv |
IEEE Internet Things J. | 3 |
| 2025 | ProLongVid: A Simple but Strong Baseline for Long-context Video Instruction TuningabstractRui Wang, Bohao Li, Xiyang Dai, Jianwei Yang, Yi-Ling Chen, Zhen Xing, Yifan Yang, Dongdong Chen, Xipeng Qiu, Zuxuan Wu, Yu-Gang Jiang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Rui Wang 0095, Xiyang Dai, Yifan Yang 0004, Dongdong Chen 0001, Xipeng Qiu, Zuxuan Wu, Yu-Gang Jiang 0001 |
EMNLP | 7 |
| 2025 | StreamMind: Unlocking Full Frame Rate Streaming Video Dialogue through Event-Gated Cognition
Hao Wu 0067, Yifan Yang 0004, Shiqi Jiang 0002, Qianxi Zhang, Donglin Bai, Zhibo Chen 0001, Ting Cao 0003 |
ICCV | 3 |
| 2025 | REDUCIO! Generating 1K Video Within 16 Seconds Using Extremely Compressed Motion Latents
Qi Dai 0001, Jianmin Bao, Yifan Yang 0004, Chong Luo 0001, Zuxuan Wu, Yu-Gang Jiang 0001 |
ICCV | 5 |
| 2025 | DreamDistribution: Learning Prompt Distribution for Diverse In-distribution GenerationabstractThe popularization of Text-to-Image (T2I) diffusion models enables the generation of high-quality images from text descriptions. However, generating diverse customized images with reference visual attributes remains challenging. This work focuses on personalizing T2I diffusion models at a more abstract concept or category level, adapting commonalities from a set of reference images while creating new instances with sufficient variations. We introduce a solution that allows a pretrained T2I diffusion model to learn a set of soft prompts, enabling the generation of novel images by sampling prompts from the learned distribution. These prompts offer text-guided editing capabilities and additional flexibility in controlling variation and mixing between multiple distributions. We also show the adaptability of the learned prompt distribution to other tasks, such as text-to-3D. Finally we demonstrate effectiveness of our approach through quantitative analysis including automatic evaluation and human assessment. Brian Nlong Zhao, Xinyang Jiang, Yifan Yang 0004, Dongsheng Li 0002, Laurent Itti, Vibhav Vineet, Yunhao Ge |
ICLR | 5 |
| 2025 | Babel: A Scalable Pre-trained Model for Multi-Modal Sensing via Expandable Modality AlignmentabstractThis paper presents Babel, the expandable modality alignment model, specially designed for multi-modal sensing. While there has been considerable work on multi-modality alignment, they all struggle to effectively incorporate multiple sensing modalities due to the data scarcity constraints. How to utilize multi-modal data with partial pairings in sensing remains an unresolved challenge. Shenghong Dai, Shiqi Jiang 0002, Yifan Yang 0004, Ting Cao 0003, Mo Li 0001, Suman Banerjee 0001, Lili Qiu |
SenSys | 3 |
| 2025 | Efficient and Adaptive Diffusion Model Inference Through Lookup Table on Mobile DevicesabstractDiffusion models have revolutionized image synthesis applications. Many studies focus on using approximate computation such as model quantization to reduce inference costs on mobile devices. However, due to their extensive model parameters and autoregressive inference fashion, the overhead of diffusion models remains high, which is challenging for mobile devices to handle. To reduce the inference overhead of diffusion models on mobile devices, we proposeLUT-Diff, an algorithm-system co-design specifically tailored for mobile device diffusion model inference optimization.LUT-Diffoptimizes using lookup tables and can efficiently generate a series of lookup table candidates for diffusion models without end-to-end training. During inference,LUT-Diffadaptively selects the best inference strategy based on the application/user's latency budget. Additionally,LUT-Diffincludes a parallel inference engine that rapidly completes model inference through CPU-GPU co-scheduling. Extensive experiments demonstrate thatLUT-Diffcan generate images comparable to the original model, with an up to 0.012 MSE in generated images.LUT-Diffcan also achieve up to 9.1× inference acceleration and reduce the inference memory footprint by up to 70.9% compared to baseline methods. Moreover,LUT-Diffcan save at least 3281× the learning cost of lookup tables. Qipeng Wang 0001, Shiqi Jiang 0002, Yifan Yang 0004, Ruiqi Liu 0001, Yuanchun Li 0003, Ting Cao 0003, Xuanzhe Liu |
IEEE Trans. Mob. Comput. | 3 |
| 2024 | Unified Medical Image Pre-training in Language-Guided Common Semantic Space
Xiaoxuan He, Yifan Yang 0004, Xinyang Jiang, Xufang Luo, Haoji Hu, Siyun Zhao, Dongsheng Li 0002, Yuqing Yang 0001, Lili Qiu |
ECCV (81) | 2 |
| 2024 | Online Video Quality Enhancement with Spatial-Temporal Look-Up Tables
Zefan Qu, Xinyang Jiang, Yifan Yang 0004, Dongsheng Li 0002, Cairong Zhao |
ECCV (72) | 3 |
| 2024 | Understanding and Improving Training-free Loss-based Diffusion GuidanceabstractAdding additional guidance to pretrained diffusion models has become an increasingly popular research area, with extensive applications in computer vision, reinforcement learning, and AI for science. Recently, several studies have proposed training-free loss-based guidance by using off-the-shelf networks pretrained on clean images. This approach enables zero-shot conditional generation for universal control formats, which appears to offer a free lunch in diffusion guidance. In this paper, we aim to develop a deeper understanding of training-free guidance, as well as overcome its limitations. We offer a theoretical analysis that supports training-free guidance from the perspective of optimization, distinguishing it from classifier-based (or classifier-free) guidance. To elucidate their drawbacks, we theoretically demonstrate that training-free guidance is more susceptible to misaligned gradients and exhibits slower convergence rates compared to classifier guidance. We then introduce a collection of techniques designed to overcome the limitations, accompanied by theoretical rationale and empirical evidence. Our experiments in image and motion generation confirm the efficacy of these techniques. Yifei Shen 0004, Xinyang Jiang, Yifan Yang 0004, Yezhen Wang, Dongsheng Li 0002 |
NeurIPS | 3 |
| 2023 | Similarity Distribution Based Membership Inference Attack on Person Re-identificationabstractWhile person Re-identification (Re-ID) has progressed rapidly due to its wide real-world applications, it also causes severe risks of leaking personal information from training data. Thus, this paper focuses on quantifying this risk by membership inference (MI) attack. Most of the existing MI attack algorithms focus on classification models, while Re-ID follows a totally different training and inference paradigm. Re-ID is a fine-grained recognition task with complex feature embedding, and model outputs commonly used by existing MI like logits and losses are not accessible during inference. Since Re-ID focuses on modelling the relative relationship between image pairs instead of individual semantics, we conduct a formal and empirical analysis which validates that the distribution shift of the inter-sample similarity between training and test set is a critical criterion for Re-ID membership inference. As a result, we propose a novel membership inference attack method based on the inter-sample similarity distribution. Specifically, a set of anchor images are sampled to represent the similarity distribution conditioned on a target image, and a neural network with a novel anchor selection module is proposed to predict the membership of the target image. Our experiments validate the effectiveness of the proposed approach on both the Re-ID task and conventional classification task. Junyao Gao 0002, Xinyang Jiang, Huishuai Zhang, Yifan Yang 0004, Shuguang Dou, Dongsheng Li 0002, Duoqian Miao 0001, Cheng Deng 0002, Cairong Zhao |
AAAI | 4 |
| 2023 | Towards Inference Efficient Deep Ensemble LearningabstractEnsemble methods can deliver surprising performance gains but also bring significantly higher computational costs, e.g., can be up to 2048X in large-scale ensemble tasks. However, we found that the majority of computations in ensemble methods are redundant. For instance, over 77% of samples in CIFAR-100 dataset can be correctly classified with only a single ResNet-18 model, which indicates that only around 23% of the samples need an ensemble of extra models. To this end, we propose an inference efficient ensemble learning method, to simultaneously optimize for effectiveness and efficiency in ensemble learning. More specifically, we regard ensemble of models as a sequential inference process and learn the optimal halting event for inference on a specific sample. At each timestep of the inference process, a common selector judges if the current ensemble has reached ensemble effectiveness and halt further inference, otherwise filters this challenging sample for the subsequent models to conduct more powerful ensemble. Both the base models and common selector are jointly optimized to dynamically adjust ensemble inference for different samples with various hardness, through the novel optimization goals including sequential ensemble boosting and computation saving. The experiments with different backbones on real-world datasets illustrate our method can bring up to 56% inference cost reduction while maintaining comparable performance to full ensemble, achieving significantly better ensemble utility than other baselines. Code and supplemental materials are available at https://seqml.github.io/irene. Kan Ren, Yifan Yang 0004, Xinyang Jiang, Yuqing Yang 0001, Dongsheng Li 0002 |
AAAI | 3 |
| 2023 | Attentive Mask CLIPabstractIn vision-language modeling, image token removal is an efficient augmentation technique to reduce the cost of encoding image features. The CLIP-style models, however, have been found to be negatively impacted by this technique. We hypothesize that removing a large portion of image tokens may inadvertently destroy the semantic information associated to a given text description, resulting in misaligned paired data in CLIP training. To address this issue, we propose an attentive token removal approach, which retains a small number of tokens that have a strong semantic correlation to the corresponding text description. The correlation scores are dynamically evaluated through an EMA-updated vision encoder. Our method, termed attentive mask CLIP, outperforms original CLIP and CLIP variant with random token removal while saving the training time. In addition, our approach also enables efficient multi-view contrastive learning. Experimentally, by training ViT-B on YFCC-15M dataset, our approach achieves 43.9% top-1 accuracy on ImageNet-1K zero-shot classification, 62.7/42.1 and 38.0/23.2 I2T/T2I retrieval accuracy on Flickr30K and MS COCO, outperforming SLIP by +1.1%, +5.5/+0.9, and +4.4/+1.3, respectively, while being 2.30× faster. An efficient version of our approach runs 1.16× faster than the plain CLIP model, while achieving significant gains of +5.3%, +11.3/+8.0, and +9.5/+4.9 on these benchmarks, respectively. Code will be release in https://github.com/microsoft/A-CLIP. Yifan Yang 0004, Weiquan Huang, Yixuan Wei, Houwen Peng, Xinyang Jiang, Huiqiang Jiang, Fangyun Wei, Han Hu 0001, Lili Qiu, Yuqing Yang 0001 |
ICCV | 1 |
| 2023 | ImageBrush: Learning Visual In-Context Instructions for Exemplar-Based Image ManipulationabstractWhile language-guided image manipulation has made remarkable progress, the challenge of how to instruct the manipulation process faithfully reflecting human intentions persists. An accurate and comprehensive description of a manipulation task using natural language is laborious and sometimes even impossible, primarily due to the inherent uncertainty and ambiguity present in linguistic expressions.
Is it feasible to accomplish image manipulation without resorting to external cross-modal language information? If this possibility exists, the inherent modality gap would be effortlessly eliminated. In this paper, we propose a novel manipulation methodology, dubbed ImageBrush, that learns visual instructions for more accurate image editing.
Our key idea is to employ a pair of transformation images as visual instructions, which not only precisely captures human intention but also facilitates accessibility in real-world scenarios. Capturing visual instructions is particularly challenging because it involves extracting the underlying intentions solely from visual demonstrations and then applying this operation to a new image. To address this challenge, we formulate visual instruction learning as a diffusion-based inpainting problem, where the contextual information is fully exploited through an iterative process of generation. A visual prompting encoder is carefully devised to enhance the model's capacity in uncovering human intent behind the visual instructions. Extensive experiments show that our method generates engaging manipulation results conforming to the transformations entailed in demonstrations. Moreover, our model exhibits robust generalization capabilities on various downstream tasks such as pose transfer, image translation and video inpainting. Yasheng Sun, Yifan Yang 0004, Houwen Peng, Yifei Shen 0004, Yuqing Yang 0001, Han Hu 0001, Lili Qiu, Hideki Koike |
NeurIPS | 2 |
| 2023 | Online Video Super-Resolution With Convolutional Kernel Bypass GraftsabstractDeep learning-based models have achieved remarkable performance in video super-resolution (VSR) in recent years, but most of these models are less applicable to online video applications. These methods solely consider the distortion quality and ignore crucial requirements for online applications, e.g., low latency and low model complexity. In this paper, we focus on online video transmission in which VSR algorithms are required to generate high-resolution video sequences frame by frame in real time. To address such challenges, we propose an extremely low-latency VSR algorithm based on a novel kernel knowledge transfer method, named the convolutional kernel bypass graft (CKBG). First, we design a lightweight network structure that does not require future frames as inputs and saves extra time for caching these frames. Then, our proposed CKBG method enhances this lightweight base model by bypassing the original network with “kernel grafts”, which are extra convolutional kernels containing the prior knowledge of the external pretrained image SR models. During the testing phase, we further accelerate the grafted multibranch network by converting it into a simple single-path structure. The experimental results show that our proposed method can process online video sequences up to 110 FPS with very low model complexity and competitive SR performance. Jun Xiao 0010, Xinyang Jiang, Ningxin Zheng, Huan Yang 0005, Yifan Yang 0004, Yuqing Yang 0001, Dongsheng Li 0002, Kin-Man Lam 0001 |
IEEE Trans. Multim. | 5 |