Zhengzhong Liu 0001

dblp:166/0352 · also Hector Liu, Hector Zhengzhong Liu · DBLP profile ↗
← Back
20ranked-venue papers
3as first author
10since 2021 · last 2026
0009-0007-6686-2577ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 3 first-author · 10 since 2021Databases, data management, data science and information retrieval · 3Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Vision-G1: Towards General Reasoning Vision-Language Models via Reinforcement Learning
abstract
Recent vision-language models (VLMs) show strong reasoning capabilities through training with reinforcement learning from verifiable rewards (RLVR). Despite their impressive capabilities, current VLMs focus on a limited range of reasoning tasks, such as mathematical and logical reasoning, due to the lack of readily available verifiable reward data in broader domains. As a result, these models struggle to generalize their reasoning abilities to the wide variety of challenges encountered in real-world environments. To address this limitation, we collect and assemble a comprehensive RL-ready visual reasoning training dataset encompassing 46 datasets across 13 dimensions of 5 domains, covering a wide range of realistic scenarios such as infographic reasoning, mathematical reasoning, spatial reasoning, and general science reasoning. Based on this dataset, we propose an influence function-based data filtering strategy and a multi-round data curriculum method to iteratively strengthen general visual reasoning abilities. Using this approach, we train a general reasoning VLM, namely Vision-G1. Our 7B model achieves state-of-the-art performance across nine visual reasoning benchmarks, surpassing previous similar-sized VLMs and even GPT-4o and Gemini-1.5 Flash.
Yuheng Zha, Kun Zhou 0002, Yujia Wu, Yushu Wang, Shibo Hao, Zhengzhong Liu 0001, Eric P. Xing, Zhiting Hu
AAAI8
2026 Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models
abstract
Yanbin Yin, Kun Zhou, Zhen Wang, Xiangdong Zhang, Yifei Shao, Shibo Hao, Yi Gu, Jieyuan Liu, Somanshu Singla, Tianyang Liu, Eric P. Xing, Zhengzhong Liu, Haojian Jin, Zhiting Hu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yanbin Yin, Kun Zhou 0002, Zhen Wang 0041, Yifei Shao, Shibo Hao, Yi Gu 0002, Jieyuan Liu, Somanshu Singla, Tianyang Liu 0003, Eric P. Xing, Zhengzhong Liu 0001, Haojian Jin, Zhiting Hu
ACL (1)12
2025 Scaling Long Context Training Data by Long-Distance Referrals
abstract
Training large language models for long context understanding faces the challenge of data shortage. Previous data engineering approaches mechanically concatenate short documents, which may create many pseudo long documents but raise concerns about data quality. In this paper, we study the core attribute of high quality data for long context training, and provide a data pipeline, LongPack, to scale such data. We found that long distance referrals, which occur in natural long documents, are crucial for long-context training. However, simply concatenating short documents does not reliably generate these relations. We further show that the density of long-distance referrals, which is higher in longer documents, has a key role in training efficiency, making previous upsampling methods suboptimal. To enrich long documents, we propose LongPack, a data pipeline that constructs long documents by packing shorter ones based on referral relationships. Specifically, for web pages, which are the primary source for language model training, we found hyper-link a native signal for such a relation. By packing web pages through their hyper-link connection, we can create longer, high-quality documents. Our experiments demonstrate that LongPackis highly scalable, generating a corpus of long documents equivalent in size to an entire pretraining dataset using just 0.5% root documents. Furthermore, the constructed documents have a ‘near-natural’ quality as innate long documents for long context training, reaching a 32.7% higher score than previous state-of-the-art methods.
Yonghao Zhuang 0001, Lanxiang Hu, Longfei Yun, Souvik Kundu 0009, Zhengzhong Liu 0001, Eric P. Xing, Hao Zhang 0025
ICLR5
2025 Data Mixing Optimization for Supervised Fine-Tuning of Large Language Models
abstract
Optimizing data mixtures for supervised fine-tuning (SFT) of large language models (LLMs) is critical for developing general-purpose models, yet this area remains underexplored. In this paper, we frame data mixing as an optimization problem and introduce a novel method designed to minimize validation loss. Our approach parametrizes the loss by modeling effective data transferred and leveraging scaling laws for fine-tuning. By experimenting with various small-scale data mixtures, we fit these parameters and derive the optimal weights. We provide both mathematical proofs and empirical results demonstrating that our algorithm achieves excellent overall and individual performance across all domains. Through controlled experiments, we show that models trained with our optimized weights perform on par with those using optimal weights determined via grid search, with per-domain loss only 0.66% higher than the best domain loss from grid search on average. Additionally, we show that reweighting popular SFT datasets using our method improves both validation loss and downstream performance. Finally, we discuss how our method can generalize to guide data selection for domain-specific models and provide insights into SFT.
Yuan Li 0032, Zhengzhong Liu 0001, Eric P. Xing
ICML2
2025 Fast Video Generation with Sliding Tile Attention
abstract
Diffusion Transformers (DiTs) with 3D full attention power state-of-the-art video generation, but suffer from prohibitive compute cost -- when generating just a 5-second 720P video, attention alone takes 800 out of 950 seconds of total inference time. This paper introduces sliding tile attention (STA) to address this challenge. STA leverages the observation that attention scores in pretrained video diffusion models predominantly concentrate within localized 3D windows. By sliding and attending over local spatial-temporal region, STA eliminates redundancy from full attention. Unlike traditional token-wise sliding window attention (SWA), STA operates tile-by-tile with a novel hardware-aware sliding window design, preserving expressiveness while being \emph{hardware-efficient}. With careful kernel-level optimizations, STA offers the first efficient 2D/3D sliding-window-like attention implementation, achieving 58.79\% MFU -- 7.17× faster than prior art methods. On the leading video DiT model, Hunyuan, it accelerates attention by 1.6–10x over FlashAttention-3, yielding a 1.36–3.53× end-to-end speedup with no or minimum quality loss.
Peiyuan Zhang, Runlong Su, Hangliang Ding, Ion Stoica, Zhengzhong Liu 0001, Hao Zhang 0025
ICML6
2025 Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective
abstract
Reinforcement learning (RL) has shown promise in enhancing large language model (LLM) reasoning, yet progress towards broader capabilities is limited by the availability of high-quality, multi-domain datasets. This work introduces \ours, a 92K RL-for-reasoning dataset designed to address this gap, covering six reasoning domains: Math, Code, Science, Logic, Simulation, and Tabular, each with corresponding verifiers. We build \ours via a careful data-curation pipeline, including sourcing, deduplication, reward design, and domain-specific and difficulty-based filtering, to facilitate the systematic investigation of cross-domain RL generalization. Our study using \ours suggests the efficacy of a simple mixed-domain RL training approach and reveals several key aspects affecting cross-domain transferability. We further train two models {\ours}-7B and {\ours}-32B purely with RL on our curated data and observe largely improved performance over leading open RL reasoning model baselines, with gains of 7.3\% and 7.8\% respectively on an extensive 17-task, six-domain evaluation suite. We are releasing our dataset, code, and evaluation suite to the community, aiming to support further research and development of more general RL-enhanced reasoning models.
Jorge (Zhoujun) Cheng, Shibo Hao, Tianyang Liu 0003, Yuexin Bian, Nilabjo Dey, Yonghao Zhuang 0001, Yuheng Zha, Yi Gu 0002, Kun Zhou 0002, Yuan Li 0032, Richard Fan, Jianshu She, Chengqian Gao, Abulhair Saparov, Taylor W. Killian, Haonan Li 0002, Mikhail Yurochkin, Eric P. Xing, Zhengzhong Liu 0001, Zhiting Hu
NeurIPS23
2025 Faster Video Diffusion with Trainable Sparse Attention
abstract
Scaling video diffusion transformers (DiTs) is limited by their quadratic 3D attention, even though most of the attention mass concentrates on a small subset of positions. We turn this observation into VSA, a trainable, hardware-efficient sparse attention that replaces full attention at both training and inference. In VSA, a lightweight coarse stage pools tokens into tiles and identifies high-weight critical tokens; a fine stage computes token-level attention only inside those tiles subjecting to block computing layout to ensure hard efficiency. This leads to a single differentiable kernel that trains end-to-end, requires no post-hoc profiling, and sustains 85\% of FlashAttention3 MFU. We perform a large sweep of ablation studies and scaling-law experiments by pretraining DiTs from 60M to 1.4B parameters. VSA reaches a Pareto point that cuts training FLOPS by 2.53$\times$ with no drop in diffusion loss. Retrofitting the open-source Wan2.1-1.3B model speeds up attention time by 6$\times$ and lowers end-to-end generation time from 31s to 18s with comparable quality, while for the 14B model, end-to-end generation time is reduced from 1274s to 576s. Furthermore, we introduce a preliminary study of Sparse-Distill, the first method to enable sparse attention and distillation concurrently, achieving 50.9x speed up for Wan-1.3B while maintaining quality. These results establish trainable sparse attention as a practical alternative to full attention and a key enabler for further scaling of video diffusion models. Code is available at https://github.com/hao-ai-lab/FastVideo.
Peiyuan Zhang, Haofeng Huang, Will Lin, Zhengzhong Liu 0001, Ion Stoica, Eric P. Xing, Hao Zhang 0025
NeurIPS5
2024 Web2Code: A Large-scale Webpage-to-Code Dataset and Evaluation Framework for Multimodal LLMs
abstract
Multimodal large language models (MLLMs) have shown impressive success across modalities such as image, video, and audio in a variety of understanding and generation tasks. However, current MLLMs are surprisingly poor at understanding webpage screenshots and generating their corresponding HTML code. To address this problem, we propose Web2Code, a benchmark consisting of a new large-scale webpage-to-code dataset for instruction tuning and an evaluation framework for the webpage understanding and HTML code translation abilities of MLLMs. For dataset construction, we leverage pretrained LLMs to enhance existing webpage-to-code datasets as well as generate a diverse pool of new webpages rendered into images. Specifically, the inputs are webpage images and instructions, while the responses are the webpage's HTML code. We further include diverse natural language QA pairs about the webpage content in the responses to enable a more comprehensive understanding of the web content. To evaluate model performance in these tasks, we develop an evaluation framework for testing MLLMs' abilities in webpage understanding and web-to-code generation. Extensive experiments show that our proposed dataset is beneficial not only to our proposed tasks but also in the general visual domain. We hope our work will contribute to the development of general MLLMs suitable for web-based content generation and task automation. Our data and code are available at https://github.com/MBZUAI-LLM/web2code.
Sukmin Yun, Haokun Lin, Rusiru Thushara, Mohammad Qazim Bhat, Zutao Jiang, Mingkai Deng, Tianhua Tao, Haonan Li 0002, Preslav Nakov, Timothy Baldwin, Zhengzhong Liu 0001, Eric P. Xing, Xiaodan Liang
NeurIPS14
2021 Cross-document Event Identity via Dense Annotation
abstract
In this paper, we study the identity of textual events from different documents.While the complex nature of event identity is previously studied (Hovy et al., 2013), the case of events across documents is unclear.Prior work on cross-document event coreference has two main drawbacks.First, they restrict the annotations to a limited set of event types.Second, they insufficiently tackle the concept of event identity.Such annotation setup reduces the pool of event mentions and prevents one from considering the possibility of quasiidentity relations.We propose a dense annotation approach for cross-document event coreference, comprising a rich source of event mentions and a dense annotation effort between related document pairs.To this end, we design a new annotation workflow with careful quality control and an easy-to-use annotation interface.In addition to the links, we further collect overlapping event contexts, including time, location, and participants, to shed some light on the relation between identity decisions and context.We present an open-access dataset for cross-document event coreference, CDEC-WN, collected from English Wikinews and open-source our annotation toolkit to encourage further research on cross-document tasks. 1
Adithya Pratapa, Zhengzhong Liu 0001, Kimihiro Hasegawa, Yukari Yamakawa, Shikun Zhang, Teruko Mitamura
CoNLL2
2021 Compression, Transduction, and Creation: A Unified Framework for Evaluating Natural Language Generation
abstract
Natural language generation (NLG) spans a broad range of tasks, each of which serves for specific objectives and desires different properties of generated text.The complexity makes automatic evaluation of NLG particularly challenging.Previous work has typically focused on a single task and developed individual evaluation metrics based on specific intuitions.In this paper, we propose a unifying perspective based on the nature of information change in NLG tasks, including compression (e.g., summarization), transduction (e.g., text rewriting), and creation (e.g., dialog).Information alignment between input, context, and output text plays a common central role in characterizing the generation.With automatic alignment prediction models, we develop a family of interpretable metrics that are suitable for evaluating key aspects of different NLG tasks, often without need of gold reference data.Experiments show the uniformly designed metrics achieve stronger or comparable correlations with human judgement compared to state-of-the-art metrics in each of diverse tasks, including text summarization, style transfer, and knowledgegrounded dialog. 1
Mingkai Deng, Bowen Tan, Zhengzhong Liu 0001, Eric P. Xing, Zhiting Hu
EMNLP (1)3
2020 A Two-Step Approach for Implicit Event Argument Detection
abstract
In this work, we explore the implicit event argument detection task, which studies event arguments beyond sentence boundaries.The addition of cross-sentence argument candidates imposes great challenges for modeling.To reduce the number of candidates, we adopt a two-step approach, decomposing the problem into two sub-problems: argument head-word detection and head-to-span expansion.Evaluated on the recent RAMS dataset (Ebner et al., 2020), our model achieves overall better performance than a strong sequence labeling baseline.We further provide detailed error analysis, presenting where the model mainly makes errors and indicating directions for future improvements.It remains a challenge to detect implicit arguments, calling for more future work of document-level modeling for this task.
Zhisong Zhang, Xiang Kong, Zhengzhong Liu 0001, Xuezhe Ma, Eduard H. Hovy
ACL3
2018 Graph Based Decoding for Event Sequencing and Coreference Resolution
abstract
Events in text documents are interrelated in complex ways. In this paper, we study two types of relation: Event Coreference and Event Sequencing. We show that the popular tree-like decoding structure for automated Event Coreference is not suitable for Event Sequencing. To this end, we propose a graph-based decoding algorithm that is applicable to both tasks. The new decoding algorithm supports flexible feature sets for both tasks. Empirically, our event coreference system has achieved state-of-the-art performance on the TAC-KBP 2015 event coreference task and our event sequencing system beats a strong temporal-based, oracle-informed baseline. We discuss the challenges of studying these event relations.
Zhengzhong Liu 0001, Teruko Mitamura, Eduard H. Hovy
COLING1
2018 Automatic Event Salience Identification
abstract
Identifying the salience (i.e.importance) of discourse units is an important task in language understanding.While events play important roles in text documents, little research exists on analyzing their saliency status.This paper empirically studies the Event Salience task and proposes two salience detection models based on content similarities and discourse relations.The first is a feature based salience model that incorporates similarities among discourse units.The second is a neural model that captures more complex relations between discourse units.Tested on our new largescale event salience corpus, both methods significantly outperform the strong frequency baseline, while our neural model further improves the feature based one by a large margin.Our analyses demonstrate that our neural model captures interesting connections between salience and discourse unit relations (e.g., scripts and frame structures).
Zhengzhong Liu 0001, Chenyan Xiong, Teruko Mitamura, Eduard H. Hovy
EMNLP1
2018 Towards Better Text Understanding and Retrieval through Kernel Entity Salience Modeling
abstract
This paper presents a Kernel Entity Salience Model (KESM) that improves text understanding and retrieval by better estimating entity salience (importance) in documents. KESM represents entities by knowledge enriched distributed representations, models the interactions between entities and words by kernels, and combines the kernel scores to estimate entity salience. The whole model is learned end-to-end using entity salience labels. The salience model also improves ad hoc search accuracy, providing effective ranking features by modeling the salience of query entities in candidate documents. Our experiments on two entity salience corpora and two TREC ad hoc search datasets demonstrate the effectiveness of KESM over frequency-based and feature-based methods. We also provide examples showing how KESM conveys its text understanding ability learned from entity salience to search.
Chenyan Xiong, Zhengzhong Liu 0001, Jamie Callan, Tie-Yan Liu
SIGIR2
2017 JointSem: Combining Query Entity Linking and Entity based Document Ranking
abstract
Entity-based ranking systems often employ entity linking systems to align entities to query and documents. Previously, entity linking systems were not designed specifically for search engines and were mostly used as a preprocessing step. This work presents JointSem, a joint semantic ranking system that combines query entity linking and entity-based document ranking. In JointSem, the spotting and linking signals are used to describe the importance of candidate entities in the query, and the linked entities are utilized to provide additional ranking features for the documents. The linking signals and the ranking signals are combined by a joint learning-to-rank model, and the whole system is fully optimized towards end-to-end ranking performance. Experiments on TREC Web Track datasets demonstrate the effectiveness of joint learning of entity linking and entity-based ranking.
Chenyan Xiong, Zhengzhong Liu 0001, Jamie Callan, Eduard H. Hovy
CIKM2
2017 De-duping URLs with Sequence-to-Sequence Neural Networks
abstract
Many URLs on the Internet point to identical contents, which increase the burden of web crawlers. Techniques that detect such URLs (known as URL de-duping) can greatly save resources such as bandwidth and storage for crawlers. Traditional de-duping methods are usually limited to heavily engineered rule matching strategies.In this work, we propose a novel URL de-duping framework based on sequence-to-sequence (Seq2Seq) neural networks. A single concise translation model can take the place of thousands of explicit rules. Experiments indicate that a vanilla Seq2Seq architecture yields robust and accurate results in detecting duplicate URLs. Furthermore, we demonstrate the efficiency of this framework in the real large-scale web environment.
Keyang Xu, Zhengzhong Liu 0001, Jamie Callan
SIGIR2
2016 Harnessing Deep Neural Networks with Logic Rules
abstract
Combining deep neural networks with structured logic rules is desirable to harness flexibility and reduce uninterpretability of the neural models.We propose a general framework capable of enhancing various types of neural networks (e.g., CNNs and RNNs) with declarative first-order logic rules.Specifically, we develop an iterative distillation method that transfers the structured information of logic rules into the weights of neural networks.We deploy the framework on a CNN for sentiment analysis, and an RNN for named entity recognition.With a few highly intuitive rules, we obtain substantial improvements and achieve state-of-the-art or comparable results to previous best-performing systems.
Zhiting Hu, Xuezhe Ma, Zhengzhong Liu 0001, Eduard H. Hovy, Eric P. Xing
ACL (1)3
2016 Unsupervised Ranking Model for Entity Coreference Resolution
abstract
Coreference resolution is one of the first stages in deep language understanding and its importance has been well recognized in the natural language processing community. In this paper, we propose a generative, unsupervised ranking model for entity coreference resolution by introducing resolution mode variables. Our unsupervised system achieves 58.44% F1 score of the CoNLL metric on the English data from the CoNLL-2012 shared task (Pradhan et al., 2012), outperforming the Stanford deterministic system (Lee et al., 2013) by 3.01%.
Xuezhe Ma, Zhengzhong Liu 0001, Eduard H. Hovy
HLT-NAACL2
2014 Detecting Subevent Structure for Event Coreference Resolution
Jun Araki, Zhengzhong Liu 0001, Eduard H. Hovy, Teruko Mitamura
LREC2
2014 Supervised Within-Document Event Coreference using Information Propagation
Zhengzhong Liu 0001, Jun Araki, Eduard H. Hovy, Teruko Mitamura
LREC1