VLDB 2026 Research / reviewers in the wild / expert
Jinyi Hu
dblp:261/9291
· DBLP profile ↗
14ranked-venue papers
3as first author
12since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 3 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GUICourse: From General Vision Language Model to Versatile GUI AgentabstractUtilizing Graphic User Interfaces (GUIs) for human-computer interaction is essential for accessing various digital tools. Recent advancements in Vision Language Models (VLMs) reveal significant potential for developing versatile agents that assist humans in navigating GUIs. However, current VLMs face challenges related to fundamental abilities, such as OCR and grounding, as well as a lack of knowledge about GUI elements functionalities and control methods. These limitations hinder their effectiveness as practical GUI agents. To address these challenges, we introduce GUICourse, a series of datasets for training visual-based GUI agents using general VLMs. First, we enhance the OCR and grounding capabilities of VLMs using the GUIEnv dataset. Next, we enrich the GUI knowledge of VLMs using the GUIAct and GUIChat datasets. Our experiments demonstrate that even a small-sized GUI agent (with 3.1 billion parameters) performs effectively on both single-step and multi-step GUI tasks. We further finetune our GUI agents on other GUI tasks with different action spaces (AITW and Mind2Web), and the results show that our agents are better than their baseline VLMs. Additionally, we analyze the impact of OCR and grounding capabilities through an ablation study, revealing a positive correlation with GUI navigation ability. Wentong Chen, Junbo Cui, Jinyi Hu, Yujia Qin, Junjie Fang, Chongyi Wang, Guirong Chen, Yupeng Huo, Yuan Yao 0013, Yankai Lin 0001, Zhiyuan Liu 0001, Maosong Sun 0001 |
ACL (1) | 3 |
| 2025 | NVILA: Efficient Frontier Visual Language ModelsabstractVisual language models (VLMs) have made significant advances in accuracy in recent years. However, their efficiency has received much less attention. This paper introduces NVILA, a family of open VLMs designed to optimize both efficiency and accuracy. Building on top of VILA, we improve its model architecture by first scaling up the spatial and temporal resolutions, and then compressing visual tokens. This "scale-then-compress" approach enables NVILA to efficiently process high-resolution images and long videos. We also conduct a systematic investigation to enhance the efficiency of NVILA throughout its entire lifecycle, from training to deployment. NVILA matches or surpasses the accuracy of many leading open and proprietary VLMs across a wide range of image and video benchmarks. At the same time, it reduces training costs by 1.9-5.1×, prefilling latency by 1.6-2.2×, and decoding latency by 1.2-2.8×. Ligeng Zhu, Baifeng Shi, Zhuoyang Zhang, Yuming Lou, Shang Yang, Haocheng Xi, Shiyi Cao, Yuxian Gu, Dacheng Li, Xiuyu Li, Haotian Tang, Yunhao Fang, Yukang Chen, Cheng-Yu Hsieh, De-An Huang, An-Chieh Cheng, Jinyi Hu, Sifei Liu, Ranjay Krishna, Pavlo Molchanov 0001, Jan Kautz, Hongxu Yin, Song Han 0003, Yao Lu 0006 |
CVPR | 18 |
| 2025 | DC-AR: Efficient Masked Autoregressive Image Generation with Deep Compression Hybrid TokenizerabstractWe introduce DC-AR, a novel masked autoregressive (AR) text-to-image generation framework that delivers superior image generation quality with exceptional computational efficiency. Due to the tokenizers' limitations, prior masked AR models have lagged behind diffusion models in terms of quality or efficiency. We overcome this limitation by introducing DC-HT - a deep compression hybrid tokenizer for AR models that achieves a 32x spatial compression ratio while maintaining high reconstruction fidelity and cross-resolution generalization ability. Building upon DC-HT, we extend MaskGIT and create a new hybrid masked autoregressive image generation framework that first produces the structural elements through discrete tokens and then applies refinements via residual tokens. DC-AR achieves state-of-the-art results with a gFID of 5.49 on MJHQ-30K and an overall score of 0.69 on GenEval, while offering 1.5-7.9x higher throughput and 2.0-3.5x lower latency compared to prior leading diffusion and autoregressive models. Yecheng Wu, Junyu Chen 0003, Zhuoyang Zhang, Enze Xie, Junsong Chen, Jinyi Hu, Yao Lu 0006, Song Han 0003, Han Cai |
ICCV | 7 |
| 2024 | OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific ProblemsabstractChaoqun He, Renjie Luo, Yuzhuo Bai, Shengding Hu, Zhen Thai, Junhao Shen, Jinyi Hu, Xu Han, Yujie Huang, Yuxiang Zhang, Jie Liu, Lei Qi, Zhiyuan Liu, Maosong Sun. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Chaoqun He, Renjie Luo, Yuzhuo Bai, Shengding Hu, Zhen Leng Thai, Jinyi Hu, Xu Han 0007, Zhiyuan Liu 0001, Maosong Sun 0001 |
ACL (1) | 7 |
| 2024 | Revisiting Non-Autoregressive Transformers for Efficient Image SynthesisabstractThe field of image synthesis is currently flourishing due to the advancements in diffusion models. While diffusion models have been successful, their computational inten-sity has prompted the pursuit of more efficient alternatives. As a representative work, non-autoregressive Transformers (NATs) have been recognized for their rapid generation. However, a major drawback of these models is their in-ferior performance compared to diffusion models. In this paper, we aim to re-evaluate the full potential of NATs by revisiting the design of their training and inference strategies. Specifically, we identify the complexities in properly configuring these strategies and indicate the possible sub-optimality in existing heuristic-driven designs. Recognizing this, we propose to go beyond existing methods by directly solving the optimal strategies in an automatic framework. The resulting method, named AutoNAT, advances the performance boundaries of NATs notably, and is able to perform comparably with the latest diffusion models with a significantly reduced inference cost. The effectiveness of AutoNAT is comprehensively validated on four benchmark datasets, i.e., ImageNet-256 & 512, MS-COCO, and CC3M. Code and pretrained models will be available at htt P s: / /gi thub. com/LeapLabTHU/ImprovedNAT. Zanlin Ni, Yulin Wang 0002, Renping Zhou, Jinyi Hu, Zhiyuan Liu 0001, Shiji Song, Yuan Yao 0013, Gao Huang 0001 |
CVPR | 5 |
| 2024 | RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-Grained Correctional Human FeedbackabstractMultimodal Large Language Models (MLLMs) have recently demonstrated impressive capabilities in multimodal understanding, reasoning, and interaction. However, existing MLLMs prevalently suffer from serious hallucination problems, generating text that is not factually grounded in associated images. The problem makes existing MLLMs untrustworthy and thus impractical in real-world (especially high-stakes) applications. To address the challenge, we present RLHF-V, which enhances MLLM trustworthiness via behavior alignment from fine-grained correctional human feedback. Specifically, RLHF-V collects human preference in the form of segment-level corrections on hallucinations, and performs dense direct preference optimization over the human feedback. Comprehensive experiments on five benchmarks in both automatic and human evaluation show that, RLHF-V can enable substantially more trustworthy MLLM behaviors with promising data and computation efficiency. Remarkably, using 1.4k annotated data samples, RLHF-V significantly reduces the hallucination rate of the base MLLM by 34.8%, outperforming the concurrent LLaVA-RLHF trained on 10k annotated data. The final model achieves state-of-the-art performance in trustwor-thiness among open-source MLLMs, and shows better ro-bustness than GPT-4V in preventing hallucinations aroused from over-generalization. Tianyu Yu 0002, Yuan Yao 0013, Haoye Zhang, Taiwen He, Yifeng Han, Ganqu Cui, Jinyi Hu, Zhiyuan Liu 0001, Hai-Tao Zheng 0002, Maosong Sun 0001 |
CVPR | 7 |
| 2024 | AdaNAT: Exploring Adaptive Policy for Token-Based Image Generation
Zanlin Ni, Yulin Wang 0002, Renping Zhou, Rui Lu 0001, Jinyi Hu, Zhiyuan Liu 0001, Yuan Yao 0013, Gao Huang 0001 |
ECCV (16) | 6 |
| 2024 | Large Multilingual Models Pivot Zero-Shot Multimodal Learning across LanguagesabstractRecently there has been a significant surge in multimodal learning in terms of both image-to-text and text-to-image generation. However, the success is typically limited to English, leaving other languages largely behind. Building a competitive counterpart in other languages is highly challenging due to the low-resource nature of non-English multimodal data (i.e., lack of large-scale, high-quality image-text data). In this work, we propose MPM, an effective training paradigm for training large multimodal models in low-resource languages. MPM demonstrates that Multilingual language models can Pivot zero-shot Multimodal learning across languages. Specifically, based on a strong multilingual large language model, multimodal models pretrained on English-only image-text data can well generalize to other languages in a (quasi)-zero-shot manner, even surpassing models trained on image-text data in native languages. Taking Chinese as a practice of MPM, we build large multimodal models VisCPM in image-to-text and text-to-image generation, which achieve state-of-the-art (open-source) performance in Chinese. To facilitate future research, we open-source codes and model weights at https://github.com/OpenBMB/VisCPM. Jinyi Hu, Yuan Yao 0013, Chongyi Wang, Shan Wang 0015, Yinxu Pan, Tianyu Yu 0002, Hanghao Wu, Haoye Zhang, Xu Han 0007, Yankai Lin 0001, Jiao Xue, Dahai Li, Zhiyuan Liu 0001, Maosong Sun 0001 |
ICLR | 1 |
| 2024 | scMulan: A Multitask Generative Pre-Trained Language Model for Single-Cell Analysis
Haiyang Bian, Xiaomin Dong, Chen Li 0001, Minsheng Hao, Jinyi Hu, Maosong Sun 0001, Lei Wei 0009, Xuegong Zhang |
RECOMB | 7 |
| 2022 | Evade the Trap of Mediocrity: Promoting Diversity and Novelty in Text Generation via Concentrating AttentionabstractRecently, powerful Transformer architectures have proven superior in generating high-quality sentences.Nevertheless, these models tend to produce dull high-frequency phrases, severely hurting the diversity and novelty of generated text.In this work, we dig into the intrinsic mechanism of this problem and found that sparser attention values in Transformer could improve diversity.To understand such a phenomenon, we first conduct both empirical and theoretical analysis and then attribute it to representation degeneration caused by the attentive mixture of the hidden states during training.We term this process the Trap of Mediocrity.To escape from such a trap, we introduce a novel attention regularization loss to control the sharpness of the attention distribution, which is transparent to model structures and can be easily implemented within 20 lines of python code.We prove that this method could be mathematically regarded as learning a Bayesian approximation of posterior attention.Experiments show that our method improved the diversity and novelty of the generated text while maintaining comparable quality on a variety of conditional and unconditional generation tasks. Wenhao Li 0003, Xiaoyuan Yi, Jinyi Hu, Maosong Sun 0001, Xing Xie 0001 |
EMNLP | 3 |
| 2022 | Fuse It More Deeply! A Variational Transformer with Layer-Wise Latent Variable Inference for Text GenerationabstractJinyi Hu, Xiaoyuan Yi, Wenhao Li, Maosong Sun, Xing Xie. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Jinyi Hu, Xiaoyuan Yi, Wenhao Li 0003, Maosong Sun 0001, Xing Xie 0001 |
NAACL-HLT | 1 |
| 2021 | Aspect-Level Sentiment-Controllable Review Generation with Mutual Learning FrameworkabstractReview generation, aiming to automatically generate review text according to the given information, is proposed to assist in the unappealing review writing. However, most of existing methods only consider the overall sentiments of reviews and cannot achieve aspect-level sentiment control. Even though some previous studies attempt to generate aspect-level sentiment-controllable reviews, they usually require large-scale human annotations which are unavailable in the real world. To address this issue, we propose a mutual learning framework to take advantage of unlabeled data to assist the aspect-level sentiment-controllable review generation. The framework consists of a generator and a classifier which utilize confidence mechanism and reconstruction reward to enhance each other. Experimental results show our model can achieve aspect-sentiment control accuracy up to 88% without losing generation quality. Yankai Lin 0001, Fanchao Qi, Jinyi Hu, Peng Li 0030, Jie Zhou 0016, Maosong Sun 0001 |
AAAI | 4 |
| 2020 | Generating Major Types of Chinese Classical Poetry in a Uniformed FrameworkabstractPoetry generation is an interesting research topic in the field of text generation. As one of the most valuable literary and cultural heritages of China, Chinese classical poetry is very familiar and loved by Chinese people from generation to generation. It has many particular characteristics in its language structure, ranging from form, sound to meaning, thus is regarded as an ideal testing task for text generation. In this paper, we propose a GPT-2 based uniformed framework for generating major types of Chinese classical poems. We define a unified format for formulating all types of training samples by integrating detailed form information, then present a simple form- stressed weighting method in GPT-2 to strengthen the control to the form of the generated poems, with special emphasis on those forms with longer body length. Preliminary experimental results show this enhanced model can generate Chinese classical poems of major types with high quality in both form and content, validating the effectiveness of the proposed strategy. The model has been incorporated into Jiuge, the most influential Chinese classical poetry generation system developed by Tsinghua University. Jinyi Hu, Maosong Sun 0001 |
LREC | 1 |
| 2020 | Towards Interpretable Natural Language Understanding with Explanations as Latent VariablesabstractRecently generating natural language explanations has shown very promising results in not only offering interpretable explanations but also providing additional information and supervision for prediction. However, existing approaches usually require a large set of human annotated explanations for training while collecting a large set of explanations is not only time consuming but also expensive. In this paper, we develop a general framework for interpretable natural language understanding that requires only a small set of human annotated explanations for training. Our framework treats natural language explanations as latent variables that model the underlying reasoning process of a neural model. We develop a variational EM framework for optimization where an explanation generation module and an explanation-augmented prediction module are alternatively optimized and mutually enhance each other. Moreover, we further propose an explanation-based self-training method under this framework for semi-supervised learning. It alternates between assigning pseudo-labels to unlabeled data and generating new explanations to iteratively improve each other. Experiments on two natural language understanding tasks demonstrate that our framework can not only make effective predictions in both supervised and semi-supervised settings, but is also able to generate good natural language explanations. Wangchunshu Zhou, Jinyi Hu, Hanlin Zhang 0002, Xiaodan Liang, Maosong Sun 0001, Chenyan Xiong, Jian Tang 0005 |
NeurIPS | 2 |