EDBT 2026 Demo / reviewers in the wild / expert
Haokun Lin
dblp:338/6350
· DBLP profile ↗
15ranked-venue papers
2as first author
15since 2021 · last 2026
0009-0000-1084-7115ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 2 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EfficientLLM: Unified Pruning-Aware Pretraining for Auto-Designed Compact Language ModelsabstractXingrun Xing, Zheng Liu, Shitao Xiao, Boyan Gao, Yiming Liang, Haokun Lin, Xianlin Zeng, Guoqi Li, Jiajun Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xingrun Xing, Shitao Xiao, Boyan Gao, Yiming Liang, Haokun Lin, Xianlin Zeng, Guoqi Li 0002 |
ACL (1) | 6 |
| 2026 | DapQ-DiT: Distribution-Aware Post-Training Quantization for Efficient Generative Tasks in Diffusion TransformersabstractDiffusion Transformers (DiTs) have demonstrated remarkable performance in image and video generation tasks. However, their high computational and memory overheads severely restrict their practical deployment on resource-constrained devices. Post-training quantization (PTQ), an efficient and practical model compression technique, serves as a solution to alleviate this issue. Nevertheless, existing PTQ methods tailored for DiTs suffer from significant performance degradation when conducting low-bit weight–activation quantization. In this work, we identify two key factors responsible for such performance degradation. First, the weights of DiTs exhibit Gaussian-like distributions, which makes uniform quantization poorly matched to the actual weight density and introduces large quantization errors. Second, activation outliers with extremely large magnitudes, especially in specific linear layers, significantly widen the value range and severely reduce the suppression effectiveness of fixed Hadamard rotation. To address the above degradation issues, we propose DapQ-DiT, a novel distribution-aware post-training quantization framework tailored for DiTs. First, we introduce an arctan quantizer that explicitly adapts to Gaussian-like weight distributions, concentrating more quantization intervals in the high-density central region while preserving representation accuracy for critical weights. Second, we enhance the fixed Hadamard rotation by leveraging principal components derived from activation covariance, which allows the transformation to better align with real activation distributions and more effectively suppress diverse and extreme activation outliers. Extensive experiments conducted on text-to-image generation with PixArt and text-to-video generation with OpenSORA demonstrate that DapQ-DiT consistently outperforms existing PTQ methods across various prompt sets and diverse bit-width configurations. Lianwei Yang, Haokun Lin, Zhenan Sun, Qingyi Gu |
ICMR | 2 |
| 2026 | Reshape and rotate: Adaptive weight reshaping and fine-grained rotation for ultra-low-bit diffusion transformers quantization
Lianwei Yang, Haokun Lin, Caifeng Shan, Zhenan Sun, Qingyi Gu |
Neurocomputing | 2 |
| 2026 | Singular Value Fine-Tuning for Few-Shot Class-Incremental LearningabstractClass-Incremental Learning (CIL) aims to prevent catastrophic forgetting of previously learned classes while sequentially incorporating new ones. The more challenging Few-shot CIL (FSCIL) setting further complicates this by providing only a limited number of samples for each new class, increasing the risk of overfitting in addition to standard CIL challenges. While catastrophic forgetting has been extensively studied, overfitting in FSCIL, especially with large foundation models, has received less attention. To fill this gap, we propose the Singular Value Fine-tuning for FSCIL (SVFCL) and compared it with existing approaches for adapting foundation models to FSCIL, which primarily build on Parameter Efficient Fine-Tuning (PEFT) methods like prompt tuning and Low-Rank Adaptation (LoRA). Specifically, SVFCL applies singular value decomposition to the foundation model weights, keeping the singular vectors fixed while fine-tuning the singular values for each task, and then merging them. This simple yet effective approach not only alleviates the forgetting problem but also mitigates overfitting more effectively while significantly reducing trainable parameters. Extensive experiments on four benchmark datasets, along with visualizations and ablation studies, validate the effectiveness of SVFCL. The code will be made available. Zhiwu Wang, Renzhen Wang, Haokun Lin, Quanziang Wang, Qian Zhao 0002, Deyu Meng |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | DOGR: Towards Versatile Visual Document Grounding and Referring
Yinan Zhou, Haokun Lin, Shuyu Yang, Zhongang Qi, Chen Ma 0001, Li Zhu 0003 |
ICCV | 3 |
| 2025 | Image-level Memorization Detection via Inversion-based Inference PerturbationabstractRecent studies have discovered that widely used text-to-image diffusion models can replicate training samples during image generation, a phenomenon known as memorization. Existing detection methods primarily focus on identifying memorized prompts. However, in real-world scenarios, image owners may need to verify whether their proprietary or personal images have been memorized by the model, even in the absence of paired prompts or related metadata. We refer to this challenge as image-level memorization detection, where current methods relying on original prompts fall short. In this work, we uncover two characteristics of memorized images after perturbing the inference procedure: lower similarity of the original images and larger magnitudes of TCNP.
Building on these insights, we propose Inversion-based Inference Perturbation (IIP), a new framework for image-level memorization detection. Our approach uses unconditional DDIM inversion to derive latent codes that contain core semantic information of original images and optimizes random prompt embeddings to introduce effective perturbation. Memorized images exhibit distinct characteristics within the proposed pipeline, providing a robust basis for detection. To support this task, we construct a comprehensive setup for the image-level memorization detection, carefully curating datasets to simulate realistic memorization scenarios. Using this setup, we evaluate our IIP framework across three different memorization settings, demonstrating its state-of-the-art performance in identifying memorized images in various settings, even in the presence of data augmentation attacks. Haokun Lin, Bo Peng 0002, Zhili Liu, Yueming Lyu, Xing Zheng, Jing Dong 0003 |
ICLR | 2 |
| 2025 | PhyBlock: A Progressive Benchmark for Physical Understanding and Planning via 3D Block AssemblyabstractWhile vision-language models (VLMs) have demonstrated promising capabilities in reasoning and planning for embodied agents, their ability to comprehend physical phenomena, particularly within structured 3D environments, remains severely limited. To close this gap, we introduce PhyBlock, a progressive benchmark designed to assess VLMs on physical understanding and planning through robotic 3D block assembly tasks. PhyBlock integrates a novel four-level cognitive hierarchy assembly task alongside targeted Visual Question Answering (VQA) samples, collectively aimed at evaluating progressive spatial reasoning and fundamental physical comprehension, including object properties, spatial relationships, and holistic scene understanding. PhyBlock includes 2600 block tasks (400 assembly tasks, 2200 VQA tasks) and evaluates models across three key dimensions: partial completion, failure diagnosis, and planning robustness. We benchmark 23 state-of-the-art VLMs, highlighting their strengths and limitations in physically grounded, multi-step planning. Our empirical findings indicate that the performance of VLMs exhibits pronounced limitations in high-level planning and reasoning capabilities, leading to a notable decline in performance for the growing complexity of the tasks.Error analysis reveals persistent difficulties in spatial orientation and dependency reasoning.We position PhyBlock as a unified testbed to advance embodied reasoning, bridging vision-language understanding and real-world physical problem-solving. Jiajun Wen 0003, Rongtao Xu, Xiwen Liang, Bingqian Lin, Ziming Wei 0001, Haokun Lin, Mingfei Han 0002, Meng Cao 0002, Bokui Chen, Ivan Laptev, Xiaodan Liang |
NeurIPS | 10 |
| 2025 | Scale Up Composed Image Retrieval Learning via Modification Text GenerationabstractComposed Image Retrieval (CIR) aims to search an image of interest using a combination of a reference image and modification text as the query. Despite recent advancements, this task remains challenging due to limited training data and laborious triplet annotation processes. To address this issue, this paper proposes to synthesize the training triplets to augment the training resource for the CIR problem. Specifically, we commence by training a modification text generator exploiting large-scale multimodal models and scale up the CIR learning throughout both the pretraining and fine-tuning stages. During pretraining, we leverage the trained generator to directly create Modification Text-oriented Synthetic Triplets (MTST) conditioned on pairs of images. For fine-tuning, we first synthesize reverse modification text to connect the target image back to the reference image. Subsequently, we devise a two-hop alignment strategy to incrementally close the semantic gap between the multimodal pair and the target image. We initially learn an implicit prototype utilizing both the original triplet and its reversed version in a cycle manner, followed by combining the implicit prototype feature with the modification text to facilitate accurate alignment with the target image. Extensive experiments validate the efficacy of the generated triplets and confirm that our proposed methodology attains competitive recall on both the CIRR and FashionIQ benchmarks. Codes and datasets will be made publicly accessible. Yinan Zhou, Yaxiong Wang, Haokun Lin, Chen Ma 0001, Li Zhu 0003, Zhedong Zheng |
IEEE Trans. Multim. | 3 |
| 2024 | MoPE-CLIP: Structured Pruning for Efficient Vision-Language Models with Module-Wise Pruning Error MetricabstractVision-language pretrained models have achieved impressive performance on various downstream tasks. However, their large model sizes hinder their utilization on platforms with limited computational resources. We find that directly using smaller pretrained models and applying magnitude-based pruning on CLIP models leads to in-flexibility and inferior performance. Recent efforts for VLP compression either adopt uni-modal compression metrics resulting in limited performance or involve costly mask-search processes with learnable masks. In this paper, we first propose the Module-wise Pruning Error (MoPE) met-ric, accurately assessing CLIP module importance by performance decline on cross-modal tasks. Using the MoPE metric, we introduce a unified pruning framework applica-ble to both pretraining and task-specific fine-tuning compression stages. For pretraining, MoPE-CLIP effectively leverages knowledge from the teacher model, significantly reducing pretraining costs while maintaining strong zero-shot capabilities. For fine-tuning, consecutive pruning from width to depth yields highly competitive task-specific models. Extensive experiments in two stages demonstrate the effectiveness of the MoPE metric, and MoPE-CLIP outperforms previous state-of-the-art VLP compression methods. Haokun Lin, Haoli Bai, Zhili Liu, Lu Hou 0002, Muyi Sun, Linqi Song, Ying Wei 0001, Zhenan Sun |
CVPR | 1 |
| 2024 | Contrastive Learning with Counterfactual Explanations for Radiology Report Generation
Mingjie Li 0006, Haokun Lin, Xiaodan Liang, Ling Chen 0006, Abdulmotaleb El Saddik, Xiaojun Chang |
ECCV (43) | 2 |
| 2024 | MATHVERSE: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?
Renrui Zhang, Dongzhi Jiang, Haokun Lin, Pengshuo Qiu, Aojun Zhou, Pan Lu, Kai-Wei Chang 0001, Yu Qiao 0001, Peng Gao 0007, Hongsheng Li 0001 |
ECCV (8) | 4 |
| 2024 | Plug-and-Play: An Efficient Post-training Pruning Method for Large Language ModelsabstractWith the rapid growth of large language models (LLMs), there is increasing demand for memory and computation in LLMs. Recent efforts on post-training pruning of LLMs aim to reduce the model size and computation requirements, yet the performance is still sub-optimal.
In this paper, we present a plug-and-play solution for post-training pruning of LLMs.
The proposed solution has two innovative components: 1) **Relative Importance and Activations (RIA)**, a new pruning metric that jointly considers the weight and activations efficiently on LLMs, and 2) **Channel Permutation**, a new approach to maximally preserves important weights under N:M sparsity.
The two proposed components can be readily combined to further enhance the N:M semi-structured pruning of LLMs.
Our empirical experiments show that RIA alone can already surpass all existing post-training pruning methods on prevalent LLMs, e.g., LLaMA ranging from 7B to 65B. Furthermore, N:M semi-structured pruning with channel permutation can even outperform the original LLaMA2-70B on zero-shot tasks, together with practical speed-up on specific hardware.
Our code is available at: https://github.com/biomedical-cybernetics/Relative-importance-and-activation-pruning Yingtao Zhang, Haoli Bai, Haokun Lin, Jialin Zhao 0004, Lu Hou 0002, Carlo V. Cannistraci |
ICLR | 3 |
| 2024 | DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMsabstractQuantization of large language models (LLMs) faces significant challenges, particularly due to the presence of outlier activations that impede efficient low-bit representation. Traditional approaches predominantly address Normal Outliers, which are activations across all tokens with relatively large magnitudes. However, these methods struggle with smoothing Massive Outliers that display significantly larger values, which leads to significant performance degradation in low-bit quantization. In this paper, we introduce DuQuant, a novel approach that utilizes rotation and permutation transformations to more effectively mitigate both massive and normal outliers. First, DuQuant starts by constructing the rotation matrix, using specific outlier dimensions as prior knowledge, to redistribute outliers to adjacent channels by block-wise rotation. Second, We further employ a zigzag permutation to balance the distribution of outliers across blocks, thereby reducing block-wise variance. A subsequent rotation further smooths the activation landscape, enhancing model performance. DuQuant simplifies the quantization process and excels in managing outliers, outperforming the state-of-the-art baselines across various sizes and types of LLMs on multiple tasks, even with 4-bit weight-activation quantization. Our code is available at https://github.com/Hsu1023/DuQuant. Haokun Lin, Jingzhi Cui, Yingtao Zhang, Linzhan Mou, Linqi Song, Zhenan Sun, Ying Wei 0001 |
NeurIPS | 1 |
| 2024 | Web2Code: A Large-scale Webpage-to-Code Dataset and Evaluation Framework for Multimodal LLMsabstractMultimodal large language models (MLLMs) have shown impressive success across modalities such as image, video, and audio in a variety of understanding and generation tasks. However, current MLLMs are surprisingly poor at understanding webpage screenshots and generating their corresponding HTML code. To address this problem, we propose Web2Code, a benchmark consisting of a new large-scale webpage-to-code dataset for instruction tuning and an evaluation framework for the webpage understanding and HTML code translation abilities of MLLMs. For dataset construction, we leverage pretrained LLMs to enhance existing webpage-to-code datasets as well as generate a diverse pool of new webpages rendered into images. Specifically, the inputs are webpage images and instructions, while the responses are the webpage's HTML code. We further include diverse natural language QA pairs about the webpage content in the responses to enable a more comprehensive understanding of the web content. To evaluate model performance in these tasks, we develop an evaluation framework for testing MLLMs' abilities in webpage understanding and web-to-code generation. Extensive experiments show that our proposed dataset is beneficial not only to our proposed tasks but also in the general visual domain. We hope our work will contribute to the development of general MLLMs suitable for web-based content generation and task automation. Our data and code are available at https://github.com/MBZUAI-LLM/web2code. Sukmin Yun, Haokun Lin, Rusiru Thushara, Mohammad Qazim Bhat, Zutao Jiang, Mingkai Deng, Tianhua Tao, Haonan Li 0002, Preslav Nakov, Timothy Baldwin, Zhengzhong Liu 0001, Eric P. Xing, Xiaodan Liang |
NeurIPS | 2 |
| 2023 | Dynamic Graph Enhanced Contrastive Learning for Chest X-Ray Report GenerationabstractAutomatic radiology reporting has great clinical potential to relieve radiologists from heavy workloads and improve diagnosis interpretation. Recently, researchers have enhanced data-driven neural networks with medical knowledge graphs to eliminate the severe visual and textual bias in this task. The structures of such graphs are exploited by using the clinical dependencies formed by the disease topic tags via general knowledge and usually do not update during the training process. Consequently, the fixed graphs can not guarantee the most appropriate scope of knowledge and limit the effectiveness. To address the limitation, we propose a knowledge graph with Dynamic structure and nodes to facilitate chest X-ray report generation with Contrastive Learning, named DCL. In detail, the fundamental structure of our graph is pre-constructed from general knowledge. Then we explore specific knowledge extracted from the retrieved reports to add additional nodes or redefine their relations in a bottom-up manner. Each image feature is integrated with its very own updated graph before being fed into the decoder module for report generation. Finally, this paper introduces Image-Report Contrastive and Image-Report Matching losses to better represent visual features and textual information. Evaluated on IU-Xray and MIMIC-CXR datasets, our DCL outperforms previous state-of-the-art models on these two benchmarks. Mingjie Li 0006, Bingqian Lin, Zicong Chen, Haokun Lin, Xiaodan Liang, Xiaojun Chang |
CVPR | 4 |