Quzhe Huang

dblp:278/1884 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
14since 2021 · last 2025
0000-0001-8391-8260ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 4 first-author · 14 since 2021
YearPublicationVenuePosition
2025 Automating Legal Interpretation with LLMs: Retrieval, Generation, and Evaluation
abstract
Interpreting the law is always essential for the law to adapt to the ever-changing society.It is a critical and challenging task even for legal practitioners, as it requires meticulous and professional annotations and summarizations by legal experts, which are admittedly timeconsuming and expensive to collect at scale.To alleviate the burden on legal experts, we propose a method for automated legal interpretation.Specifically, by emulating doctrinal legal research, we introduce a novel framework, ATRIE, to address Legal Concept Interpretation, a typical task in legal interpretation.ATRIE utilizes large language models (LLMs) to AuTomatically Retrieve conceptrelated information, Interpret legal concepts, and Evaluate generated interpretations, eliminating dependence on legal experts.ATRIE comprises a legal concept interpreter and a legal concept interpretation evaluator.The interpreter uses LLMs to retrieve relevant information from previous cases and interpret legal concepts.The evaluator uses performance changes on Legal Concept Entailment, a downstream task we propose, as a proxy of interpretation quality.Automated and multifaceted human evaluations indicate that the quality of our interpretations is comparable to those written by legal experts, with superior comprehensiveness and readability.Although there remains a slight gap in accuracy, it can already assist legal practitioners in improving the efficiency of legal interpretation.1
Kangcheng Luo, Quzhe Huang, Yansong Feng 0002
ACL (1)2
2025 JUREX-4E: Juridical Expert-Annotated Four-Element Knowledge Base for Legal Reasoning
abstract
In recent years, Large Language Models (LLMs) have been widely applied to legal tasks.To enhance their understanding of legal texts and improve reasoning accuracy, a promising approach is to incorporate legal theories.One of the most widely adopted theories is the Four-Element Theory (FET), which defines the crime constitution through four elements: Subject, Object, Subjective Aspect, and Objective Aspect.While recent work has explored prompting LLMs to follow FET, our evaluation demonstrates that LLM-generated four-elements are often incomplete and less representative, limiting their effectiveness in legal reasoning.To address these issues, we present JUREX-4E, an expert-annotated fourelement knowledge base covering 155 criminal charges.The annotations follow a progressive hierarchical framework grounded in legal source validity and incorporate diverse interpretive methods to ensure precision and authority.We evaluate JUREX-4E on the Similar Charge Disambiguation task and apply it to Legal Case Retrieval.Experimental results validate the high quality of JUREX-4E and its substantial impact on downstream legal tasks, underscoring its potential for advancing legal AI applications.
Huanghai Liu, Quzhe Huang, Qingjing Chen, Yiran Hu, Jiayu Ma, Weixing Shen, Yansong Feng 0002
EMNLP2
2025 Pyramidal Flow Matching for Efficient Video Generative Modeling
abstract
Video generation requires modeling a vast spatiotemporal space, which demands significant computational resources and data usage. To reduce the complexity, the prevailing approaches employ a cascaded architecture to avoid direct training with full resolution latent. Despite reducing computational demands, the separate optimization of each sub-stage hinders knowledge sharing and sacrifices flexibility. This work introduces a unified pyramidal flow matching algorithm. It reinterprets the original denoising trajectory as a series of pyramid stages, where only the final stage operates at the full resolution, thereby enabling more efficient video generative modeling. Through our sophisticated design, the flows of different pyramid stages can be interlinked to maintain continuity. Moreover, we craft autoregressive video generation with a temporal pyramid to compress the full-resolution history. The entire framework can be optimized in an end-to-end manner and with a single unified Diffusion Transformer (DiT). Extensive experiments demonstrate that our method supports generating high-quality 5-second (up to 10-second) videos at 768p resolution and 24 FPS within 20.7k A100 GPU training hours. All code and models are open-sourced at https://pyramid-flow.github.io.
Zhicheng Sun 0001, Ningyuan Li 0002, Kun Xu 0005, Hao Jiang 0032, Nan Zhuang, Quzhe Huang, Yang Song 0008, Yadong Mu, Zhouchen Lin
ICLR7
2024 Harder Task Needs More Experts: Dynamic Routing in MoE Models
abstract
Quzhe Huang, Zhenwei An, Nan Zhuang, Mingxu Tao, Chen Zhang, Yang Jin, Kun Xu, Kun Xu, Liwei Chen, Songfang Huang, Yansong Feng. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Quzhe Huang, Zhenwei An, Nan Zhuang, Mingxu Tao, Chen Zhang 0019, Kun Xu 0005, Songfang Huang, Yansong Feng 0002
ACL (1)1
2024 MC²: Towards Transparent and Culturally-Aware NLP for Minority Languages in China
abstract
Current large language models demonstrate deficiencies in understanding low-resource languages, particularly the minority languages in China.This limitation stems from the scarcity of available pre-training data.To address this accessibility challenge, we present MC 2 , a Multilingual Corpus of Minority Languages in China, which is the largest open-source corpus of its kind so far.MC 2 includes four underrepresented languages: Tibetan, Uyghur, Kazakh, and Mongolian.Notably, we focus on the less common writing systems of Kazakh and Mongolian, i.e., Kazakh Arabic script and traditional Mongolian script, respectively, which have been long neglected in previous corpus construction efforts.Recognizing the prevalence of language contamination within existing corpora, we adopt a quality-centric solution for collecting MC 2 , prioritizing accuracy while enhancing diversity.Furthermore, we underscore the importance of attending to the multiplicity of writing systems, which is closely related to the cultural awareness of the resulting models.The MC 2 corpus and related models are made public to the community 1 .
Chen Zhang 0019, Mingxu Tao, Quzhe Huang, Jiuheng Lin, Zhibin Chen 0003, Yansong Feng 0002
ACL (1)3
2024 Probing Multimodal Large Language Models for Global and Local Semantic Representations
abstract
The advancement of Multimodal Large Language Models (MLLMs) has greatly accelerated the development of applications in understanding integrated texts and images. Recent works leverage image-caption datasets to train MLLMs, achieving state-of-the-art performance on image-to-text tasks. However, there are few studies exploring which layers of MLLMs make the most effort to the global image information, which plays vital roles in multimodal comprehension and generation. In this study, we find that the intermediate layers of models can encode more global semantic information, whose representation vectors perform better on visual-language entailment tasks, rather than the topmost layers. We further probe models regarding local semantic representations through object recognition tasks. We find that the topmost layers may excessively focus on local information, leading to a diminished ability to encode global information. Our code and data are released via https://github.com/kobayashikanna01/probing_MLLM_rep.
Mingxu Tao, Quzhe Huang, Kun Xu 0005, Yansong Feng 0002, Dongyan Zhao 0001
LREC/COLING2
2024 Unified Language-Vision Pretraining in LLM with Dynamic Discrete Visual Tokenization
abstract
Recently, the remarkable advance of the Large Language Model (LLM) has inspired researchers to transfer its extraordinary reasoning capability to both vision and language data. However, the prevailing approaches primarily regard the visual input as a prompt and focus exclusively on optimizing the text generation process conditioned upon vision content by a frozen LLM. Such an inequitable treatment of vision and language heavily constrains the model's potential. In this paper, we break through this limitation by representing both vision and language in a unified form. Specifically, we introduce a well-designed visual tokenizer to translate the non-linguistic image into a sequence of discrete tokens like a foreign language that LLM can read. The resulting visual tokens encompass high-level semantics worthy of a word and also support dynamic sequence length varying from the image. Coped with this tokenizer, the presented foundation model called LaVIT can handle both image and text indiscriminately under the same generative learning paradigm. This unification empowers LaVIT to serve as an impressive generalist interface to understand and generate multi-modal content simultaneously. Extensive experiments further showcase that it outperforms the existing models by a large margin on massive vision-language tasks. Our code and models are available at https://github.com/jy0205/LaVIT.
Kun Xu 0005, Chao Liao, Jianchao Tan, Quzhe Huang, Chengru Song, Dai Meng, Di Zhang 0026, Wenwu Ou, Kun Gai, Yadong Mu
ICLR6
2024 Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization
abstract
In light of recent advances in multimodal Large Language Models (LLMs), there is increasing attention to scaling them from image-text data to more informative real-world videos. Compared to static images, video poses unique challenges for effective large-scale pre-training due to the modeling of its spatiotemporal dynamics. In this paper, we address such limitations in video-language pre-training with an efficient video decomposition that represents each video as keyframes and temporal motions. These are then adapted to an LLM using well-designed tokenizers that discretize visual and temporal information as a few tokens, thus enabling unified generative pre-training of videos, images, and text. At inference, the generated tokens from the LLM are carefully recovered to the original continuous pixel space to create various video content. Our proposed framework is both capable of comprehending and generating image and video content, as demonstrated by its competitive performance across 13 multimodal benchmarks in image and video understanding and generation. Our code and models are available at https://video-lavit.github.io.
Zhicheng Sun 0001, Kun Xu 0005, Hao Jiang 0032, Quzhe Huang, Chengru Song, Di Zhang 0026, Yang Song 0008, Kun Gai, Yadong Mu
ICML7
2024 Only One Relation Possible? Modeling the Ambiguity in Temporal Relation Extraction
Yutong Hu 0002, Quzhe Huang, Yansong Feng 0002
NLPCC (1)2
2023 More than Classification: A Unified Framework for Event Temporal Relation Extraction
abstract
Event temporal relation extraction (ETRE) is usually formulated as a multi-label classification task, where each type of relation is simply treated as a one-hot label.This formulation ignores the meaning of relations and wipes out their intrinsic dependency.After examining the relation definitions in various ETRE tasks, we observe that all relations can be interpreted using the start and end time points of events.For example, relation Includes could be interpreted as event 1 starting no later than event 2 and ending no earlier than event 2. In this paper, we propose a unified event temporal relation extraction framework, which transforms temporal relations into logical expressions of time points and completes the ETRE by predicting the relations between certain time point pairs.Experiments on TB-Dense and MATRES show significant improvements over a strong baseline and outperform the state-of-the-art model by 0.3% on both datasets.By representing all relations in a unified framework, we can leverage the relations with sufficient data to assist the learning of other relations, thus achieving stable improvement in low-data scenarios.When the relation definitions are changed, our method can quickly adapt to the new ones by simply modifying the logic expressions that map time points to new event relations.The code is released at https://github.com/AndrewZhe/ A-Unified-Framework-for-ETRE.
Quzhe Huang, Yutong Hu 0002, Shengqi Zhu 0002, Yansong Feng 0002, Chang Liu 0076, Dongyan Zhao 0001
ACL (1)1
2022 Does Recommend-Revise Produce Reliable Annotations? An Analysis on Missing Instances in DocRED
abstract
DocRED is a widely used dataset for documentlevel relation extraction.In the large-scale annotation, a recommend-revise scheme is adopted to reduce the workload.Within this scheme, annotators are provided with candidate relation instances from distant supervision, and they then manually supplement and remove relational facts based on the recommendations.However, when comparing Do-cRED with a subset relabeled from scratch, we find that this scheme results in a considerable amount of false negative samples and an obvious bias towards popular entities and relations.Furthermore, we observe that the models trained on DocRED have low recall on our relabeled dataset and inherit the same bias in the training data.Through the analysis of annotators' behaviors, we figure out the underlying reason for the problems above: the scheme actually discourages annotators from supplementing adequate instances in the revision phase.We appeal to future research to take into consideration the issues with the recommend-revise scheme when designing new models and annotation schemes.The relabeled dataset is released at https://github. com/AndrewZhe/Revisit-DocRED, to serve as a more reliable test set of document RE models.
Quzhe Huang, Shibo Hao, Yuan Ye 0001, Shengqi Zhu 0002, Yansong Feng 0002, Dongyan Zhao 0001
ACL (1)1
2022 Rethinking Task-Specific Knowledge Distillation: Contextualized Corpus as Better Textbook
abstract
Knowledge distillation has been proven effective when customizing small language models for specific tasks.Here, a corpus as 'textbook' plays an indispensable role, only through which the teacher can teach the student.Prevailing methods adopt a two-stage distillation paradigm: general distillation first with taskagnostic general corpus and task-specific distillation next with augmented task-specific corpus.We argue that such a paradigm may not be optimal.In general distillation, it's extravagant to let the diverse but desultory general knowledge overwhelms the limited model capacity of the student.While in task-specific distillation, the task corpus is usually limited and narrow, preventing the student from learning enough knowledge.To mitigate the issues in the two gapped corpora, we present a better textbook for the student to learn: contextualized corpus that contextualizes task corpus with large-scale general corpus through relevance-based text retrieval.Experimental results on GLUE benchmark demonstrate that contextualized corpus is the better textbook compared with jointly using general corpus and augmented task-specific corpus.Surprisingly, it enables task-specific distillation from scratch without general distillation while maintaining comparable performance, making it more flexible to customize the student model with desired model size under various computation constraints.
Chang Liu 0076, Chongyang Tao, Jianxin Liang, Tao Shen 0001, Jiazhan Feng, Quzhe Huang, Dongyan Zhao 0001
EMNLP6
2022 Knowledge-Enhanced Iterative Instruction Generation and Reasoning for Knowledge Base Question Answering
Haowei Du, Quzhe Huang, Chen Zhang 0019, Dongyan Zhao 0001
NLPCC (1)2
2021 Exploring Distantly-Labeled Rationales in Neural Network Models
abstract
Quzhe Huang, Shengqi Zhu, Yansong Feng, Dongyan Zhao. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Quzhe Huang, Shengqi Zhu 0002, Yansong Feng 0002, Dongyan Zhao 0001
ACL/IJCNLP (1)1