VLDB 2026 Research / reviewers in the wild / expert
Haoqin Tu
dblp:309/7386
· DBLP profile ↗
16ranked-venue papers
6as first author
16since 2021 · last 2026
0000-0002-5627-249XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 6 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | STAR-1: Safer Alignment of Reasoning LLMs with 1K DataabstractThis paper introduces STAR-1, a high-quality, just-1k-scale safety dataset specifically designed for large reasoning models (LRMs) like DeepSeek-R1. Built on three core principles --- diversity, deliberative reasoning, and rigorous filtering --- STAR-1 aims to address the critical needs for safety alignment in LRMs. Specifically, we begin by integrating existing open-source safety datasets from diverse sources. Then, we curate safety policies to generate policy-grounded deliberative reasoning samples. Lastly, we apply a GPT-4o-based safety scoring system to select training examples aligned with best practices. Experimental results show that fine-tuning LRMs with STAR-1 leads to an average 40% improvement in safety performance across four benchmarks, while only incurring a marginal decrease (e.g., an average of 1.1%) in reasoning ability measured across five reasoning tasks. Extensive ablation studies further validate the importance of our design principles in constructing STAR-1 and analyze its efficacy across both LRMs and traditional LLMs. Haoqin Tu, Yuhan Wang 0001, Juncheng Wu, Jieru Mei, Brian R. Bartoldson, Bhavya Kailkhura, Cihang Xie |
AAAI | 2 |
| 2025 | ViLBench: A Suite for Vision-Language Process Reward ModelingabstractProcess-supervised reward models serve as a fine-grained function that provides detailed step-wise feedback to model responses, facilitating effective selection of reasoning trajectories for complex tasks.Despite its advantages, evaluation on PRMs remains less explored, especially in the multimodal domain.To address this gap, this paper first benchmarks current vision large language models (VLLMs) as two types of reward models: output reward models (ORMs) and process reward models (PRMs) on multiple vision-language benchmarks, which reveal that neither ORM nor PRM consistently outperforms across all tasks, and superior VLLMs do not necessarily yield better rewarding performance.To further advance evaluation, we introduce VILBENCH, a vision-language benchmark designed to require intensive process reward signals.Notably, Ope-nAI's GPT-4o with Chain-of-Thought (CoT) achieves only 27.3% accuracy, challenging current VLLMs.Lastly, we preliminarily showcase a promising pathway towards bridging the gap between general VLLMs and reward models-by collecting 73.6K vision-language process reward data using an enhanced treesearch algorithm, our 3B model is able to achieve an average improvement of 3.3% over standard CoT and up to 2.5% compared to its untrained counterpart on VILBENCH by selecting OpenAI o1's generations.We will release our code, model, and data at https: //ucsc-vlaa.github.io/ViLBench. Haoqin Tu, Hardy Chen, Hui Liu 0033, Xianfeng Tang, Cihang Xie |
EMNLP | 1 |
| 2025 | Language Models Can See Better: Visual Contrastive Decoding For LLM Multimodal ReasoningabstractAlthough Large Language Models (LLMs) excel in reasoning and generation for language tasks, they are not specifically designed for multimodal challenges. Training Multimodal Large Language Models (MLLMs), however, is resource-intensive and constrained by various training limitations. In this paper, we propose the Modular-based Visual Contrastive Decoding (MVCD) framework to move this obstacle. Our framework leverages LLMs’ In-Context Learning (ICL) capability and the proposed visual contrastive-example decoding (CED), specifically tailored for this framework, without requiring any additional training. By converting visual signals into text and focusing on contrastive output distributions during decoding, we can highlight the new information introduced by contextual examples, explore their connections, and avoid over-reliance on prior encoded knowledge. MVCD enhances LLMs’ visual perception to make it see and reason over the input visuals. To demonstrate MVCD’s effectiveness, we conduct experiments with four LLMs across five question answering datasets. Our results not only show consistent improvement in model accuracy but well explain the effective components inside our decoding strategy. Our code will be available at https://github.com/Pbhgit/MVCD. Yuqi Pang, Haoqin Tu, Yun Cao 0001 |
ICASSP | 3 |
| 2025 | OpenVision: A Fully-Open, Cost-Effective Family of Advanced Vision Encoders for Multimodal LearningabstractOpenAI's CLIP, released in early 2021, have long been the go-to choice of vision encoder for building multimodal foundation models. Although recent alternatives such as SigLIP have begun to challenge this status quo, to our knowledge none are fully open: their training data remains proprietary and/or their training recipes are not released. This paper fills this gap with OpenVision, a fully-open, cost-effective family of vision encoders that match or surpass the performance of OpenAI's CLIP when integrated into multimodal frameworks like LLaVA. OpenVision builds on existing works -- e.g., CLIPS for training framework and Recap-DataComp-1B for training data -- while revealing multiple key insights in enhancing encoder quality and showcasing practical benefits in advancing multimodal models. By releasing vision encoders spanning from 5.9M to 632.1M parameters, OpenVision offers practitioners a flexible trade-off between capacity and efficiency in building multimodal models: larger models deliver enhanced multimodal performance, while smaller versions enable lightweight, edge-ready multimodal deployments. Xianhang Li, Haoqin Tu, Cihang Xie |
ICCV | 3 |
| 2025 | Autoregressive Pretraining with Mamba in VisionabstractThe vision community has started to build with the recently developed state space model, Mamba, as the new backbone for a range of tasks. This paper shows that Mamba's visual capability can be significantly enhanced through autoregressive pretraining, a direction not previously explored. Efficiency-wise, the autoregressive nature can well capitalize on the Mamba's unidirectional recurrent structure, enabling faster overall training speed compared to other training strategies like mask modeling. Performance-wise, autoregressive pretraining equips the Mamba architecture with markedly higher accuracy over its supervised-trained counterparts and, more importantly, successfully unlocks its scaling potential to large and even huge model sizes. For example, with autoregressive pretraining, a base-size Mamba attains 83.2\% ImageNet accuracy, outperforming its supervised counterpart by 2.0\%; our huge-size Mamba, the largest Vision Mamba to date, attains 85.0\% ImageNet accuracy (85.5\% when finetuned with $384\times384$ inputs), notably surpassing all other Mamba variants in vision. The code is available at \url{https://github.com/OliverRensu/ARM}. Sucheng Ren, Xianhang Li, Haoqin Tu, Fangxun Shu, Jieru Mei, Alan L. Yuille, Cihang Xie |
ICLR | 3 |
| 2025 | What If We Recaption Billions of Web Images with LLaMA-3?abstractWeb-crawled image-text pairs are inherently noisy. Prior studies demonstrate that semantically aligning and enriching textual descriptions of these pairs can significantly enhance model training across various vision-language tasks, particularly text-to-image generation. However, large-scale investigations in this area remain predominantly closed-source. Our paper aims to bridge this community effort, leveraging the powerful and $\textit{open-sourced}$ LLaMA-3, a GPT-4 level LLM. Our recaptioning pipeline is simple: first, we fine-tune a LLaMA-3-8B powered LLaVA-1.5 and then employ it to recaption ~1.3 billion images from the DataComp-1B dataset. Our empirical results confirm that this enhanced dataset, Recap-DataComp-1B, offers substantial benefits in training advanced vision-language models. For discriminative models like CLIP, we observe an average of 3.1% enhanced zero-shot performance cross four cross-modal retrieval tasks using a mixed set of the original and our captions. For generative models like text-to-image Diffusion Transformers, the generated images exhibit a significant improvement in alignment with users' text instructions, especially in following complex queries. Our project page is https://www.haqtu.me/Recap-Datacomp-1B/. Xianhang Li, Haoqin Tu, Mude Hui, Zeyu Wang 0008, Bingchen Zhao, Junfei Xiao, Sucheng Ren, Jieru Mei, Qing Liu 0017, Huangjie Zheng, Yuyin Zhou, Cihang Xie |
ICML | 2 |
| 2025 | MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?abstractWhile text-to-image models like GPT-4o-Image and FLUX are rapidly proliferating, they often encounter challenges such as hallucination, bias, and the production of unsafe, low-quality output. To effectively address these issues, it is crucial to align these models with desired behaviors based on feedback from a multimodal judge. Despite their significance, current multimodal judges frequently undergo inadequate evaluation of their capabilities and limitations, potentially leading to misalignment and unsafe fine-tuning outcomes. To address this issue, we introduce MJ-Bench, a novel benchmark which incorporates a comprehensive preference dataset to evaluate multimodal judges in providing feedback for image generation models across six key perspectives: alignment, safety, image quality, bias, composition, and visualization. Specifically, we evaluate a large variety of multimodal judges including smaller-sized CLIP-based scoring models, open-source VLMs, and close-source VLMs on each decomposed subcategory of our preference dataset. Experiments reveal that close-source VLMs generally provide better feedback, with GPT-4o outperforming other judges in average. Compared with open-source VLMs, smaller-sized scoring models can provide better feedback regarding text-image alignment and image quality, while VLMs provide more accurate feedback regarding safety and generation bias due to their stronger reasoning capabilities. Further studies in feedback scale reveal that VLM judges can generally provide more accurate and stable feedback in natural language than numerical scales. Notably, human evaluations on end-to-end and fine-tuned models using separate feedback from these multimodal judges provide similar conclusions, further confirming the effectiveness of MJ-Bench. Zhaorun Chen, Zichen Wen, Yichao Du, Yiyang Zhou, Chenhang Cui, Siwei Han, Jen Weng, Chaoqi Wang, Zhengwei Tong, Leria Huang, Canyu Chen, Haoqin Tu, Qinghao Ye, Zhihong Zhu 0001, Zhuokai Zhao, Rafael Rafailov, Chelsea Finn, Huaxiu Yao |
NeurIPS | 12 |
| 2024 | How Many Are in This Image A Safety Evaluation Benchmark for Vision LLMs
Haoqin Tu, Chenhang Cui, Yiyang Zhou, Bingchen Zhao, Junlin Han, Wangchunshu Zhou, Huaxiu Yao, Cihang Xie |
ECCV (51) | 1 |
| 2024 | Tuning LayerNorm in Attention: Towards Efficient Multi-Modal LLM FinetuningabstractThis paper introduces an efficient strategy to transform Large Language Models (LLMs) into Multi-Modal Large Language Models.
By conceptualizing this transformation as a domain adaptation process, \ie, transitioning from text understanding to embracing multiple modalities, we intriguingly note that, within each attention block, tuning LayerNorm suffices to yield strong performance.
Moreover, when benchmarked against other tuning approaches like full parameter finetuning or LoRA, its benefits on efficiency are substantial.
For example, when compared to LoRA on a 13B model scale, performance can be enhanced by an average of over 20\% across five multi-modal tasks, and meanwhile,
results in a significant reduction of trainable parameters by 41.9\% and a decrease in GPU memory usage by 17.6\%. On top of this LayerNorm strategy, we showcase that selectively tuning only with conversational data can improve efficiency further.
Beyond these empirical outcomes, we provide a comprehensive analysis to explore the role of LayerNorm in adapting LLMs to the multi-modal domain and improving the expressive power of the model. Bingchen Zhao, Haoqin Tu, Chen Wei 0005, Jieru Mei, Cihang Xie |
ICLR | 2 |
| 2024 | VHELM: A Holistic Evaluation of Vision Language ModelsabstractCurrent benchmarks for assessing vision-language models (VLMs) often focus on their perception or problem-solving capabilities and neglect other critical aspects such as fairness, multilinguality, or toxicity. Furthermore, they differ in their evaluation procedures and the scope of the evaluation, making it difficult to compare models. To address these issues, we extend the HELM framework to VLMs to present the Holistic Evaluation of Vision Language Models (VHELM). VHELM aggregates various datasets to cover one or more of the 9 aspects: visual perception, knowledge, reasoning, bias, fairness, multilinguality, robustness, toxicity, and safety. In doing so, we produce a comprehensive, multi-dimensional view of the capabilities of the VLMs across these important factors. In addition, we standardize the standard inference parameters, methods of prompting, and evaluation metrics to enable fair comparisons across models. Our framework is designed to be lightweight and automatic so that evaluation runs are cheap and fast. Our initial run evaluates 22 VLMs on 21 existing datasets to provide a holistic snapshot of the models. We uncover new key findings, such as the fact that efficiency-focused models (e.g., Claude 3 Haiku or Gemini 1.5 Flash) perform significantly worse than their full models (e.g., Claude 3 Opus or Gemini 1.5 Pro) on the bias benchmark but not when evaluated on the other aspects. For transparency, we release the raw model generations and complete results on our website at https://crfm.stanford.edu/helm/vhelm/v2.0.1. VHELM is intended to be a living benchmark, and we hope to continue adding new datasets and models over time. Haoqin Tu, Chi Heem Wong, Yiyang Zhou, Yifan Mai 0001, Josselin Somerville Roberts, Michihiro Yasunaga, Huaxiu Yao, Cihang Xie, Percy Liang |
NeurIPS | 2 |
| 2024 | FET-LM: Flow-Enhanced Variational Autoencoder for Topic-Guided Language ModelingabstractVariational autoencoder (VAE) is widely used in tasks of unsupervised text generation due to its potential of deriving meaningful latent spaces, which, however, often assumes that the distribution of texts follows a common yet poor-expressed isotropic Gaussian. In real-life scenarios, sentences with different semantics may not follow simple isotropic Gaussian. Instead, they are very likely to follow a more intricate and diverse distribution due to the inconformity of different topics in texts. Considering this, we propose a flow-enhanced VAE for topic-guided language modeling (FET-LM). The proposed FET-LM models topic and sequence latent separately, and it adopts a normalized flow composed of householder transformations for sequence posterior modeling, which can better approximate complex text distributions. FET-LM further leverages a neural latent topic component by considering learned sequence knowledge, which not only eases the burden of learning topic without supervision but also guides the sequence component to coalesce topic information during training. To make the generated texts more correlative to topics, we additionally assign the topic encoder to play the role of a discriminator. Encouraging results on abundant automatic metrics and three generation tasks demonstrate that the FET-LM not only learns interpretable sequence and topic representations but also is fully capable of generating high-quality paragraphs that are semantically consistent. Haoqin Tu, Zhongliang Yang, Jinshuai Yang, Linna Zhou, Yongfeng Huang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | ReSee: Responding through Seeing Fine-grained Visual Knowledge in Open-domain DialogueabstractIncorporating visual knowledge into text-only dialogue systems has become a potential direction to imitate the way humans think, imagine, and communicate.However, existing multimodal dialogue systems are either confined by the scale and quality of available datasets or the coarse concept of visual knowledge.To address these issues, we provide a new paradigm of constructing multimodal dialogues as well as two datasets extended from text-only dialogues under such paradigm (RESEE-WoW, RESEE-DD).We propose to explicitly split the visual knowledge into finer granularity ("turn-level" and "entity-level").To further boost the accuracy and diversity of augmented visual information, we retrieve them from the Internet or a large image dataset.To demonstrate the superiority and universality of the provided visual knowledge, we propose a simple but effective framework RESEE to add visual representation into vanilla dialogue models by modality concatenations.We also conduct extensive experiments and ablations w.r.t.different model configurations and visual knowledge settings.Empirically, encouraging results not only demonstrate the effectiveness of introducing visual knowledge at both entity and turn level but also verify the proposed model RESEE outperforms several state-of-the-art methods on automatic and human evaluations.By leveraging text and vision knowledge, RESEE can produce informative responses with real-world visual concepts.Our code is available at https: //github.com/ImKeTT/ReSee. Haoqin Tu, Fei Mi, Zhongliang Yang |
EMNLP | 1 |
| 2023 | ZeroGen: Zero-Shot Multimodal Controllable Text Generation with Multiple Oracles
Haoqin Tu, Xianfeng Zhao |
NLPCC (2) | 1 |
| 2023 | Linguistic Steganalysis Toward Social NetworkabstractWith the rapid development of the internet and social media, linguistic steganography can be easily abused in social networks to make considerable damage to varied aspects like personal privacy, network virus and national defense. Currently, considerable linguistic steganalysis methods are proposed to detect harmful steganographic carriers. However, almost all the existing methods fail in real social networks, since they are only devoted to the linguistic features that are extreme insufficient owing to the extreme sparsity and extreme fragmentation challenges of real social networks. In this paper, we attempt to fill the long-standing gap that the datasets and effective methods are absent for hunting steganographic texts in social network scenarios. Concretely, we construct a dataset called Stego-Sandbox to simulate the real social network scenarios, which contains texts and their relation. And we propose an effective linguistic steganalysis framework integrating linguistic features contained in texts and context features represented by these connections. Extensive experimental results demonstrate owing to the captured context features, our proposed framework can effectively compensate for shortcomings of these existing methods and tremendously improve their detection ability in real social network scenarios. Jinshuai Yang, Zhongliang Yang, Haoqin Tu, Yongfeng Huang 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2022 | PCAE: A framework of plug-in conditional auto-encoder for controllable text generation
Haoqin Tu, Zhongliang Yang, Jinshuai Yang, Si-yu Zhang 0001, Yongfeng Huang 0001 |
Knowl. Based Syst. | 1 |
| 2022 | SeSy: Linguistic Steganalysis Framework Integrating Semantic and Syntactic FeaturesabstractWith the rapid development of natural language processing technology and linguistic steganography, linguistic steganalysis gains considerable interest in recent years. Current advanced methods dominantly focus on statistical features in semantic view yet ignore syntax structure of text, which leads to limited performance to some newly statistically indistinguishable steganography algorithms. To fill this gap, in this paper, we propose a novel linguistic steganalysis framework named SeSy to integrate bothsemantic andsyntactic features. Specifically, we propose to employ transformer-architecture language model as semantics extractor and leverage a graph attention network to retain syntactic features. Extensive experimental results show that owing to additional syntactic information, the SeSy framework effectively brings about remarkable improvement to current advanced linguistic steganalysis methods. Jinshuai Yang, Zhongliang Yang, Si-yu Zhang 0001, Haoqin Tu, Yongfeng Huang 0001 |
IEEE Signal Process. Lett. | 4 |