Jiuhai Chen

dblp:263/3943 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
10since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion
abstract
We present Florence-VL, a new family of multimodal large language models (MLLMs) with enriched visual representations produced by Florence-2 [45], a generative vision foundation model. Unlike the widely used CLIP-style vision transformer [35] trained by contrastive learning, Florence-2 can capture different levels and aspects of visual features, which are more versatile to be adapted to diverse downstream tasks. We propose a novel feature-fusion architecture and an innovative training recipe that effectively integrates Florence-2’s visual features into pre-trained LLMs, such as Phi 3.5 and LLama 3. In particular, we propose "depth-breath fusion (DBFusion)" to fuse the visual features extracted from different depths and under multiple prompts. Our model training is composed of end-to-end pretraining of the whole model followed by finetuning of the projection layer and the LLM, on a carefully designed recipe of diverse open-source datasets that include high-quality image captions and instruction-tuning pairs. Our quantitative analysis and visualization of Florence-VL’s visual features show its advantages over popular vision encoders on vision-language alignment, where the enriched depth and breath play important roles. Florence-VL achieves significant improvements over existing state-of-the-art MLLMs across various multi-modal and vision-centric benchmarks covering general VQA, perception, hallucination, OCR, Chart, knowledge-intensive understanding, etc. To facilitate future research, our models and the complete training recipe are open-sourced. https://github.com/JiuhaiChen/Florence-VL
Jiuhai Chen, Haiping Wu, Dianqi Li, Jianfeng Gao 0001, Tianyi Zhou 0001, Bin Xiao 0004
CVPR1
2025 ColorBench: Can VLMs See and Understand the Colorful World? A Comprehensive Benchmark for Color Perception, Reasoning, and Robustness
abstract
Color plays an important role in human perception and usually provides critical clues in visual reasoning. However, it is unclear whether and how vision-language models (VLMs) can perceive, understand, and leverage color as humans.This paper introduces ColorBench, an innovative benchmark meticulously crafted to assess the capabilities of VLMs in color understanding, including color perception, reasoning, and robustness. By curating a suite of diverse test scenarios, with grounding in real applications, ColorBench evaluates how these models perceive colors, infer meanings from color-based cues, and maintain consistent performance under varying color transformations. Through an extensive evaluation of 32 VLMs with varying language models and vision encoders, our paper reveals some undiscovered findings: (i) The scaling law (larger models are better) still holds on ColorBench, while the language model plays a more important role than the vision encoder. (ii) However, the performance gaps across models are relatively small, indicating that color understanding has been largely neglected by existing VLMs. (iii) CoT reasoning improves color understanding accuracies and robustness, though they are vision-centric tasks. (iv) Color clues are indeed leveraged by VLMs on ColorBench but they can also mislead models in some tasks.These findings highlight the critical limitations of current VLMs and underscore the need to enhance color comprehension. Our ColorBench can serve as a foundational tool for advancing the study of human-level color understanding of multimodal AI.
Yijun Liang, Ming Li 0010, Chenrui Fan, Kwesi Cobbina, Shweta Bhardwaj, Jiuhai Chen, Fuxiao Liu, Tianyi Zhou 0001
NeurIPS8
2024 Quantifying Uncertainty in Answers from any Language Model and Enhancing their Trustworthiness
abstract
We introduce BSDETECTOR, a method for detecting bad and speculative answers from a pretrained Large Language Model by estimating a numeric confidence score for any output it generated.Our uncertainty quantification technique works for any LLM accessible only via a black-box API, whose training data remains unknown.By expending a bit of extra computation, users of any LLM API can now get the same response as they would ordinarily, as well as a confidence estimate that cautions when not to trust this response.Experiments on both closed and open-form Question-Answer benchmarks reveal that BSDETECTOR more accurately identifies incorrect LLM responses than alternative uncertainty estimation procedures (for both GPT-3 and ChatGPT).By sampling multiple responses from the LLM and considering the one with the highest confidence score, we can additionally obtain more accurate responses from the same LLM, without any extra training steps.In applications involving automated evaluation with LLMs, accounting for our confidence scores leads to more reliable evaluation in both human-in-the-loop and fullyautomated settings (across both GPT 3.5 and 4).... Observed Consistency + Ans...
Jiuhai Chen, Jonas Mueller 0001
ACL (1)1
2024 ODIN: Disentangled Reward Mitigates Hacking in RLHF
abstract
In this work, we study the issue of reward hacking on the response length, a challenge emerging in Reinforcement Learning from Human Feedback (RLHF) on LLMs. A well-formatted, verbose but less helpful response from the LLMs can often deceive LLMs or even human evaluators and achieve high scores. The same issue also holds for some reward models in RL. To address the challenges in both training and evaluation, we establish a more reliable evaluation protocol for comparing different training configurations, which inspects the trade-off between LLM evaluation score and response length obtained by varying training hyperparameters. Based on this evaluation, we conduct large-scale studies, where the results shed insights into the efficacy of hyperparameters and tricks used in RL on mitigating length bias. We further propose to improve the reward model by jointly training two linear heads to predict the preference, one trained to correlate with length and the other trained to decorrelate with length and therefore focusing more on the actual content. We then discard the length head in RL to ignore the spurious length reward. Experiments demonstrate that our approach eliminates the reward correlation with length, and improves the obtained policy by a significant margin.
Lichang Chen, Chen Zhu 0001, Jiuhai Chen, Davit Soselia, Tianyi Zhou 0001, Tom Goldstein, Heng Huang 0001, Mohammad Shoeybi, Bryan Catanzaro
ICML3
2024 InstructZero: Efficient Instruction Optimization for Black-Box Large Language Models
abstract
Large language models (LLMs) are instruction followers but the performance varies under different instructions. It is challenging to create the best instruction, especially for black-box LLMs on which backpropagation is forbidden. Instead of directly optimizing the discrete instruction, we optimize a low-dimensional soft prompt applied to an open-source LLM to generate the instruction for the black-box LLM. In each optimization step of the proposed method InstructZero, a soft prompt is converted into an instruction by the open-source LLM, which is then submitted to the black-box LLM for zero-shot evaluation, whose result is sent to Bayesian optimization to produce new soft prompts improving the zero-shot performance. We evaluate InstructZero on different combinations of open-source LLMs and APIs including Vicuna and ChatGPT. InstructZero outperforms SOTA auto-instruction methods across a variety of downstream tasks.
Lichang Chen, Jiuhai Chen, Tom Goldstein, Heng Huang 0001, Tianyi Zhou 0001
ICML2
2024 From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning
abstract
Ming Li, Yong Zhang, Zhitao Li, Jiuhai Chen, Lichang Chen, Ning Cheng, Jianzong Wang, Tianyi Zhou, Jing Xiao. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Ming Li 0010, Yong Zhang 0058, Zhitao Li 0002, Jiuhai Chen, Lichang Chen, Ning Cheng 0001, Jianzong Wang, Tianyi Zhou 0001, Jing Xiao 0006
NAACL-HLT4
2023 PTP: Boosting Stability and Performance of Prompt Tuning with Perturbation-Based Regularizer
abstract
Recent studies show that prompt tuning can better leverage the power of large language models than fine-tuning on downstream natural language understanding tasks.Nonetheless, current prompt tuning methods encounter instability during training, marked by a high variance in scores given different random seeds.In addressing this crucial issue, we uncover that the loss landscape of standard prompt tuning, when visualized, is remarkably steep, i.e., minor alterations in the input data can trigger substantial fluctuations in the loss landscape, which is an essential factor that leads to the training instability.In light of this finding, we incorporate perturbation-based regularizers to temper the loss landscape within the prompt tuning process.We thus present a novel algorithm, called Prompt Tuning with Perturbation-based regularizer (PTP), that can significantly reduce training instability and concurrently enhance the performance of prompt tuning.Specifically, we design two variants of perturbation-based regularizers: one that employs random noise, and another that uses an adversarial approach.Importantly, our proposed perturbations display flexibility in both the text and embedding spaces.Extensive experiments show the effectiveness of our proposed methods in stabilizing the training.Our new algorithms improve the state-of-the-art prompt tuning methods by 1.94% and 2.34% on SuperGLUE and FewGLUE benchmarks, respectively.
Lichang Chen, Jiuhai Chen, Heng Huang 0001, Minhao Cheng
EMNLP2
2023 GOAT: A Global Transformer on Large-scale Graphs
abstract
Graph transformers have been competitive on graph classification tasks, but they fail to outperform Graph Neural Networks (GNNs) on node classification, which is a common task performed on large-scale graphs for industrial applications. Meanwhile, existing GNN architectures are limited in their ability to perform equally well on both homophilious and heterophilious graphs as their inductive biases are generally tailored to only one setting. To address these issues, we propose GOAT, a scalable global graph transformer. In GOAT, each node conceptually attends to all the nodes in the graph and homophily/heterophily relationships can be learnt adaptively from the data. We provide theoretical justification for our approximate global self-attention scheme, and show it to be scalable to large-scale graphs. We demonstrate the competitiveness of GOAT on both heterophilious and homophilious graphs with millions of nodes.
Kezhi Kong, Jiuhai Chen, John Kirchenbauer, Renkun Ni, C. Bayan Bruss, Tom Goldstein
ICML2
2022 Does your graph need a confidence boost? Convergent boosted smoothing on graphs with tabular node features
Jiuhai Chen, Jonas Mueller 0001, Vassilis N. Ioannidis, Soji Adeshina, Yangkun Wang, Tom Goldstein, David P. Wipf
ICLR1
2022 Why Propagate Alone? Parallel Use of Labels and Features on Graphs
Yangkun Wang, Jiarui Jin, Weinan Zhang 0001, Yongyi Yang, Jiuhai Chen, Yong Yu 0001, Zheng Zhang 0001, Zengfeng Huang, David P. Wipf
ICLR5