Zuchao Li

dblp:198/9339 · DBLP profile ↗
← Back
91ranked-venue papers
20as first author
75since 2021 · last 2026
0000-0003-0436-8446ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 80 · 19 first-author · 64 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 6 first-author · 19 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 End-to-End Contrastive Language-Speech Pretraining Model for Long-Form Spoken Question Answering
abstract
Significant progress has been made in spoken question answering (SQA) in recent years. However, many existing methods, including large audio language models, struggle with processing long audio. Follow the success of retrieval augmented generation, a speech-related retriever shows promising in help preprocessing long-form speech. But the performance of existing speech-related retrievers is lacking. To address this challenge, we propose CLSR, an end-to-end contrastive language-speech retriever that efficiently extracts question-relevant segments from long audio recordings for downstream SQA task. Unlike conventional speech-text contrastive models, CLSR incorporates an intermediate step that converts acoustic features into text-like representations prior to alignment, thereby more effectively bridging the gap between modalities. Experimental results across four cross-modal retrieval datasets demonstrate that CLSR surpasses both end-to-end speech related retrievers and pipeline approaches combining speech recognition with text retrieval, providing a robust foundation for advancing practical long-form SQA applications.
Jiliang Hu 0001, Zuchao Li, Baoyuan Qi, Guoming Liu, Ping Wang 0028
AAAI2
2026 Scaling LLM Speculative Decoding: Non-Autoregressive Forecasting in Large-Batch Scenarios
abstract
Speculative decoding accelerates LLM inference by utilizing otherwise idle computational resources during memory-to-chip data transfer. Current speculative decoding methods typically assume a considerable amount of available computing power, then generate a complex and massive draft tree using a small autoregressive language model to improve overall prediction accuracy. However, methods like batching have been widely applied in mainstream model inference systems as a superior alternative to speculative decoding, as they compress the available idle computing power. Therefore, performing speculative decoding with low verification resources and low scheduling costs has become an important research problem. We believe that more capable models that allow for parallel generation on draft sequences are what we truly need. Recognizing the fundamental nature of draft models to only generate sequences of limited length, we propose SpecFormer, a novel architecture that integrates unidirectional and bidirectional attention mechanisms. SpecFormer combines the autoregressive model’s ability to extract information from the entire input sequence with the parallel generation benefits of non-autoregressive models. This design eliminates the reliance on large prefix trees and achieves consistent acceleration, even in large-batch scenarios. Through lossless speculative decoding experiments across models of various scales, we demonstrate that SpecFormer sets a new standard for scaling LLM inference with lower training demands and reduced computational costs.
Luohe Shi, Zuchao Li, Lefei Zhang, Baoyuan Qi, Guoming Liu, Hai Zhao 0001
AAAI2
2026 Ghost in the Transformer: Detecting Model Reuse with Invariant Spectral Signatures
abstract
Large Language Models (LLMs) are widely adopted, but their high training cost leads many developers to fine-tune existing open-source models. While most adhere to open-source licenses, some falsely claim original training despite clear derivation from public models, raising pressing concerns about intellectual property protection and the need to verify model provenance. In this paper, we propose GhostSpec, a lightweight yet effective method for verifying LLM lineage without access to training data or modification of model behavior. Our approach constructs compact and robust fingerprints by applying singular value decomposition (SVD) to invariant products of internal attention weight matrices. Unlike watermarking or output-based methods, GhostSpec is fully data-free, non-invasive, and computationally efficient. Extensive experiments show it is robust to fine-tuning, pruning, expansion, and adversarial transformations, reliably tracing lineage with minimal overhead. By offering a practical solution for model verification, our method contributes to intellectual property protection and fosters a transparent, trustworthy LLM ecosystem.
Suqing Wang, Zuchao Li
AAAI4
2026 TRACE: Traversal Retrieval-Augmented Chain of Evidence for Document Understanding
abstract
Early Long-context Document Visual Question Answering (DocVQA) methods struggle with preserving visual semantics or handling finite context windows.Conversely, recent RAG-based approaches suffer from "semantic gaps" and "structural disconnections" due to passive retrieval mechanisms that ignore logical dependencies.To address these challenges, we introduce TRACE (Traversal Retrieval-Augmented Chain of Evidence).By navigating a Bi-Layered Graph that encodes both physical adjacency and semantic relevance, TRACE transforms retrieval from static matching into adaptive evidence chain construction.Furthermore, we propose M5BookVQA, a benchmark designed to assess deep, multihop reasoning in books, addressing the limitations of existing datasets.Extensive experiments show that TRACE achieves an average accuracy improvement of 14.07% on M5BookVQA and exhibits robust generalization with a 13.38% gain across four established benchmarks.
Liqi He, Zuchao Li, Ping Wang 0028
ACL (1)2
2026 Vista-LLM: Decoupled Query-Guided Visual Token Pruning for Efficient Long-Video Large Language Models
abstract
Long-video understanding is bottlenecked by the high cost of processing massive visual tokens.Current reduction strategies often rely on static allocation or inefficient in-network selection that disrupts optimized attention kernels.In this paper, we introduce Vista-LLM, a decoupled framework for query-guided visual token pruning.By filtering redundancy prior to inference with minimal overhead, Vista-LLM ensures full compatibility with Flash Attention.Our method employs a coarse-tofine pipeline: (1) Query-Guided Dynamic Budgeting for adaptive temporal allocation; (2) a lightweight Semantic Scout for fine-grained, query-specific selection; and (3) Structure-Aware Compensation to preserve global context.Extensive experiments on benchmarks like Video-MME and MLVU demonstrate a significantly improved Pareto frontier.Notably, on LLaVA-OneVision, Vista-LLM reduces visual tokens by 90% and accelerates inference while retaining over 98% of baseline performance on average, effectively filtering visual noise.Our code is available at https: //github.com/lizhenyu-123/Vista-LLM.
Zuchao Li, Ping Wang 0028, Lefei Zhang, Haojun Ai
ACL (1)2
2026 BoYaEval: Evaluating Multimodal Large Language Models on Understanding Ancient Chinese Musical Scores
abstract
Jiajia Li, Weizhi Xue, Yao Yao, Qiwei Li, Chenchong, Zuchao Li, Ping Wang, Hai Zhao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jiajia Li 0005, Weizhi Xue, Yao Yao 0008, Qiwei Li 0002, Zuchao Li, Ping Wang 0028, Hai Zhao 0001
ACL (1)6
2026 From AR to Diffusion: Efficiently Adapting Large Language Models with Strictly Causal and Elastic Horizons
abstract
Diffusion models promise efficient parallel text generation but rely on bidirectional attention, creating a structural mismatch with pre-trained Autoregressive (AR) models.This incompatibility precludes reusing robust AR priors, necessitating prohibitive pre-training from scratch.To bridge this gap, we propose FLUID, a framework that efficiently adapts AR backbones to the diffusion paradigm.By enforcing Strictly Causal Alignment, FLUID enables seamless initialization from standard GPT-style checkpoints, circumventing the need for massive pre-training.Furthermore, we introduce Elastic Horizons, an entropy-driven mechanism that dynamically modulates denoising strides based on local information density rather than fixed schedules.Experiments demonstrate that FLUID achieves state-of-theart performance while reducing training costs by orders of magnitude, effectively reconciling established AR foundations with efficient parallel generation.
Teng Xiao, Zuchao Li, Lefei Zhang
ACL (1)3
2026 PAR: Training-Free Positional Perturbation and Attention Recycling for Faithful OCR
abstract
In high-precision scenarios, vision language models suffer from Linguistic Priors Hallucination.When processing familiar text, models tend to over-rely on internal parametric knowledge, effectively "reciting" the content rather than "reading" the image.In this paper, we first systematically investigate this phenomenon by constructing the GlitchText Probing Dataset.We discover that the model's reliance on visual grounding diminishes significantly as the generation length increases.To mitigate this, we propose PAR (Positional Perturbation and Attention Recycling), a training-free, inferencetime intervention framework.PAR consists of two parts: (1) Positional Perturbation (PP) injects structured phase noise into the rotary positional embeddings; (2) Foveal Attention Recycling (FAR) detects over-confident linguistic priors and dynamically redistributes attention mass back to important visual regions.Extensive experiments across state-of-the-art models, demonstrate that PAR significantly reduces hallucination rates (reducing CER by 12%), particularly in long-context scenarios, while maintaining robust generalization on standard benchmarks.Our code is publicly available at https://github.com/Zoeyyao27/PAR-for- Faithful-OCR.
Yao Yao 0008, Manwen Liao, Weitian Zhang, Zuchao Li, Hai Zhao 0001
ACL (1)4
2026 Generalized and group spherical linear interpolation for token-level context compression
Jinhao Tian, Zuchao Li, Meng-Jia Shen, Lefei Zhang
Neural Networks2
2025 Imitate Before Detect: Aligning Machine Stylistic Preference for Machine-Revised Text Detection
abstract
Large Language Models (LLMs) have revolutionized text generation, making detecting machine-generated text increasingly challenging. Although past methods have achieved good performance on detecting pure machine-generated text, those detectors have poor performance on distinguishing machine-revised text (rewriting, expansion, and polishing), which can have only minor changes from its original human prompt. As the content of text may originate from human prompts, detecting machine-revised text often involves identifying distinctive machine styles, e.g., worded favored by LLMs. However, existing methods struggle to detect machine-style phrasing hidden within the content contributed by humans. We propose the “Imitate Before Detect” (ImBD) approach, which first imitates the machine-style token distribution, and then compares the distribution of the text to be tested with the machine-style distribution to determine whether the text has been machine-revised. To this end, we introduce Style Preference Optimization (SPO), which aligns a scoring LLM model to the preference of text styles generated by machines. The aligned scoring model is then used to calculate the style-conditional probability curvature (Style-CPC), quantifying the log probability difference between the original and conditionally sampled texts for effective detection. We conduct extensive comparisons across various scenarios, encompassing text revisions by six LLMs, four distinct text domains, and three machine revision types. Compared to existing state-of-the-art methods, our method yields a 13% increase in AUC for detecting text revised by open-source LLMs, and improves performance by 5% and 19% for detecting GPT-3.5 and GPT-4o revised text, respectively. Notably, our method surpasses the commercially trained GPT-Zero with just 1,000 samples and five minutes of SPO, demonstrating its efficiency and effectiveness.
Xiaoye Zhu, Yiwen Yuan, Chak Tou Leong, Zuchao Li, Tang Long, Chenyu Yan, Guanghao Mei, Lefei Zhang
AAAI8
2025 SongSong: A Time Phonograph for Chinese SongCi Music from Thousand of Years Away
abstract
Recently, there have been significant advancements in music generation. However, existing models primarily focus on creating modern pop songs, making it challenging to produce ancient music with distinct rhythms and styles, such as ancient Chinese SongCi. In this paper, we introduce SongSong, the first music generation model capable of restoring Chinese SongCi to our knowledge. Our model first predicts the melody from the input SongCi, then separately generates the singing voice and accompaniment based on that melody, and finally combines all elements to create the final piece of music. Additionally, to address the lack of ancient music datasets, we create OpenSongSong, a comprehensive dataset of ancient Chinese SongCi music, featuring 29.9 hours of compositions by various renowned SongCi music masters. To assess SongSong's proficiency in performing SongCi, we randomly select 85 SongCi sentences that were not part of the training set for evaluation against SongSong and music generation platforms such as Suno and SkyMusic. The subjective and objective outcomes indicate that our proposed model achieves leading performance in generating high-quality SongCi music.
Jiliang Hu 0001, Jiajia Li 0005, Ziyi Pan, Zuchao Li, Ping Wang 0028, Lefei Zhang
AAAI5
2025 Dialogue-RAG: Enhancing Retrieval for LLMs via Node-Linking Utterance Rewriting
abstract
Large Language Models (LLMs) and Retrieval Augmented Generation (RAG) methods have demonstrated significant potential on tasks across multiple domains. However, ellipses and coreferences, as common phenomena in dialogue scenes, pose challenges to LLMs’ understanding and RAG’s retrieval accuracy. The previous works ignore the negative impact of this fuzzy data on RAG system.We explore the capabilities of LLMs and RAG systems in dialogue scenarios and use Incomplete Utterance Rewriting (IUR) to complete the key information in dialogue to enhance retrieval.Besides, we propose a lightweight IUR model for query rewriting. It is an end-to-end framework for node linking and iterative inference, incorporating two newly proposed probing semantic features derived from generative pre-training. This framework treats IUR as a series of link decisions on the input sequence and the incrementally constructed rewriting outputs.To test the performance of RAG system in the model multi-round dialogue scenario, we construct an RAG dialogue dataset on English and Chinese, Dialogue-RAG-MULTI-v1.0.Experiment results show that utterance rewriting can effectively improve the retrieval and generation ability of RAG system in dialogue scenes. Experiments on IUR tasks demonstrate the excellent performance of our lightweight IUR method.
Qiwei Li 0002, Teng Xiao, Zuchao Li, Ping Wang 0028, Mengjia Shen, Hai Zhao 0001
ACL (1)3
2025 KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding
abstract
Large language models (LLMs) based on Transformer Decoders have become the preferred choice for conversational generative AI.Despite the overall superiority of the Decoder architecture, the gradually increasing Key-Value (KV) cache during inference has emerged as a primary efficiency bottleneck, both in aspects of memory consumption and data transfer bandwidth limitations.To address these challenges, we propose a paradigm called KV-Latent.By down-sampling the Key-Value vector dimensions into a latent space, we can significantly reduce the KV Cache footprint and improve inference speed, only with a small amount of extra training, less than 1% of pretraining takes.Besides, we enhanced the stability of Rotary Positional Embedding applied on lower-dimensional vectors by modifying its frequency sampling mechanism, avoiding noise introduced by higher frequencies while retaining position attenuation.Our experiments, including both models with Grouped Query Attention and those without, have yielded satisfactory results.Finally, we conducted comparative experiments to study the impact of separately reducing Key and Value components on model's performance.Our approach allows for the construction of more efficient language model systems, and opens the new possibility on KV Cache saving and efficient LLMs.Our code is available at https://github.com/ShiLuohe/KV- Latent.
Luohe Shi, Zuchao Li, Lefei Zhang, Baoyuan Qi, Guoming Liu, Hai Zhao 0001
ACL (1)2
2025 SpindleKV: A Novel KV Cache Reduction Method Balancing Both Shallow and Deep Layers
abstract
Zicong Tang, Shi Luohe, Zuchao Li, Baoyuan Qi, Liu Guoming, Lefei Zhang, Ping Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Zicong Tang, Luohe Shi, Zuchao Li, Baoyuan Qi, Guoming Liu, Lefei Zhang, Ping Wang 0028
ACL (1)3
2025 Segment First or Comprehend First? Explore the Limit of Unsupervised Word Segmentation with Large Language Models
abstract
Word segmentation stands as a cornerstone of Natural Language Processing (NLP).Based on the concept of "comprehend first, segment later", we propose a new framework to explore the limit of unsupervised word segmentation with Large Language Models (LLMs) and evaluate the semantic understanding capabilities of LLMs based on word segmentation.We employ current mainstream LLMs to perform word segmentation across multiple languages to assess LLMs' "comprehension".Our findings reveal that LLMs are capable of following simple prompts to segment raw text into words.There is a trend suggesting that models with more parameters tend to perform better on multiple languages.Additionally, we introduce a novel unsupervised method, termed LLACA (Large Language Model-Inspired Aho-Corasick Automaton).Leveraging the advanced pattern recognition capabilities of Aho-Corasick automata, LLACA innovatively combines these with the deep insights of well-pretrained LLMs.This approach not only enables the construction of a dynamic n-gram model that adjusts based on contextual information but also integrates the nuanced understanding of LLMs, offering significant improvements over traditional methods.
Zihong Zhang, Liqi He, Zuchao Li, Lefei Zhang, Hai Zhao 0001, Bo Du 0001
ACL (1)3
2025 IAM: Efficient Inference through Attention Mapping between Different-scale LLMs
abstract
LLMs encounter significant challenges in resource consumption nowadays, especially with long contexts. Despite extensive efforts dedicate to enhancing inference efficiency, these methods primarily exploit internal sparsity within the models, without leveraging external information for optimization. We identify the high similarity of attention matrices across different-scale LLMs, which offers a novel perspective for optimization. We first conduct a comprehensive analysis of how to measure similarity, how to select mapping Layers and whether mapping is consistency. Based on these insights, we introduce the IAM framework, which achieves dual benefits of accelerated attention computation and reduced KV cache usage by performing attention mapping between small and large LLMs. Our experimental results demonstrate that IAM can accelerate prefill by 15% and reduce KV cache usage by 22.1% without appreciably sacrificing performance. Experiments on different series of models show the generalizability of IAM. Importantly, it is also orthogonal to many existing KV cache optimization methods, making it a versatile addition to the current toolkit for enhancing LLM efficiency.
Zuchao Li, Hai Zhao 0001
ACL (1)2
2025 DAC: A Dynamic Attention-aware Approach for Task-Agnostic Prompt Compression
abstract
Task-agnostic prompt compression leverages the redundancy in natural language to reduce computational overhead and enhance information density within prompts, especially in longcontext scenarios.Existing methods predominantly rely on information entropy as the metric to compress lexical units, aiming to achieve minimal information loss.However, these approaches overlook two critical aspects: (i) the importance of attention-critical tokens at the algorithmic level, and (ii) shifts in information entropy during the compression process.Motivated by these challenges, we propose a dynamic attention-aware approach for taskagnostic prompt compression (DAC).This approach effectively integrates entropy and attention information, dynamically sensing entropy shifts during compression to achieve fine-grained prompt compression.Extensive experiments across various domains, including LongBench, GSM8K, and BBH, show that DAC consistently yields robust and substantial improvements across a diverse range of tasks and LLMs, offering compelling evidence of its efficacy.
Zuchao Li, Hai Zhao 0001, Baoyuan Qi, Guoming Liu
ACL (1)2
2025 Faster In-Context Learning for LLMs via N-Gram Trie Speculative Decoding
abstract
As a crucial method in prompt engineering, In-Context Learning (ICL) enhances the generalization and knowledge utilization capabilities of Large Language Models (LLMs) (Dong et al., 2024).However, the lengthy retrieved contexts and limited token throughput in autoregressive models significantly constrain reasoning speed.To address this challenge, we propose N-Gram Trie Speculative Decoding, a novel approach that leverages the overlap between context and model output.This method constructs an n-gram trie from the context to generate drafts, accelerating token generation for LLMs.We evaluate our approach on summarization, Retrieval-Augmented Generation (RAG), and contextbased Question Answering (QA) tasks.Experimental results on Vicuna-7B, Llama2-7B-Chat, and Llama3-8B-Instruct demonstrate substantial speed improvements without compromising accuracy.Compared with various strong baselines, our method achieves the highest mean speedup, showcasing its effectiveness and efficiency.Our implement code is available here: https://github.com/mrlife219/Ngram-Trie.
Jinglin Chen, Qiwei Li 0002, Zuchao Li, Baoyuan Qi, Guoming Liu, Haojun Ai, Hai Zhao 0001, Ping Wang 0028
EMNLP3
2025 ToM: Leveraging Tree-oriented MapReduce for Long-Context Reasoning in Large Language Models
abstract
Large Language Models (LLMs), constrained by limited context windows, often face significant performance degradation when reasoning over long contexts.To address this, Retrieval-Augmented Generation (RAG) retrieves and reasons over chunks but frequently sacrifices logical coherence due to its reliance on similarity-based rankings.Similarly, divideand-conquer frameworks (DCF) split documents into small chunks for independent reasoning and aggregation.While effective for local reasoning, DCF struggles to capture longrange dependencies and risks inducing conflicts by processing chunks in isolation.To overcome these limitations, we propose ToM, a novel Tree-oriented MapReduce framework for long-context reasoning.ToM leverages the inherent hierarchical structure of long documents (e.g., main headings and subheadings) by constructing a DocTree through hierarchical semantic parsing and performing bottom-up aggregation.Using a Tree MapReduce approach, ToM enables recursive reasoning: in the Map step, rationales are generated at child nodes; in the Reduce step, these rationales are aggregated across sibling nodes to resolve conflicts or reach consensus at parent nodes.Experimental results on 70B+ LLMs show that ToM significantly outperforms existing divide-andconquer frameworks and retrieval-augmented generation methods, achieving better logical coherence and long-context reasoning.Our code is available at https://github.com/gjn12- 31/ToM.
Jiani Guo, Zuchao Li, Jie Wu 0001, Qianren Wang, Yun Li 0011, Lefei Zhang, Hai Zhao 0001, Yujiu Yang 0001
EMNLP2
2025 From Parameters to Performance: A Data-Driven Study on LLM Structure and Development
abstract
Large language models (LLMs) have achieved remarkable success across various domains, driving significant technological advancements and innovations.Despite the rapid growth in model scale and capability, systematic, data-driven research on how structural configurations affect performance remains scarce.To address this gap, we present a large-scale dataset encompassing diverse open-source LLM structures and their performance across multiple benchmarks.Leveraging this dataset, we conduct a systematic, data mining-driven analysis to validate and quantify the relationship between structural configurations and performance.Our study begins with a review of the historical development of LLMs and an exploration of potential future trends.We then analyze how various structural choices impact performance across benchmarks and further corroborate our findings using mechanistic interpretability techniques.By providing data-driven insights into LLM optimization, our work aims to guide the targeted development and application of future models.We release our dataset at
Suqing Wang, Zuchao Li, Luohe Shi, Bo Du 0001, Hai Zhao 0001, Yun Li 0011, Qianren Wang
EMNLP2
2025 Teaching Your Models to Understand Code via Focal Preference Alignment
abstract
Jie Wu, Haoling Li, Xin Zhang, Xiao Liu, Yangyu Huang, Jianwen Luo, Yizhen Zhang, Zuchao Li, Ruihang Chu, Yujiu Yang, Scarlett Li. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Jie Wu 0001, Haoling Li, Xin Zhang 0099, Xiao Liu 0029, Yangyu Huang, Zuchao Li, Ruihang Chu, Yujiu Yang 0001, Scarlett Li
EMNLP8
2025 Can Large Language Models Be Good Language Teachers?
abstract
Large language models (LLMs) have achieved remarkable success across diverse domains.However, their potential as effective language teachers-particularly in complex pedagogical scenarios like teaching Chinese as a second language-remains inadequately assessed.To address this gap, we propose the first pedagogical competence benchmark for LLMs, rigorously evaluating their performance against international standards for Chinese language teachers.Our framework spans three core dimensions: (1) basic knowledge evaluation, covering 32 subtopics across five major categories;(2) international teacher examination, based on data collected from international Chinese teacher certification exams; and (3) teaching practice evaluation, where target LLMs summarize knowledge points and design instructional content for student models, followed by testing the student models to assess the LLM's ability to distill and teach key concepts.We conduct a comprehensive evaluation of 13 latest multilingual and Chinese LLMs.While most models demonstrate promising pedagogical potential, there remains substantial room for improvement in their teaching capabilities.This study contributes to the development of AI-assisted language education tools capable of rivaling human teaching excellence.
LiQing Xu, Qiwei Li 0002, Tianshuo Peng, Zuchao Li, Hai Zhao 0001, Ping Wang 0028
EMNLP4
2025 XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer Compression
abstract
Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse natural language processing tasks.However, their extensive memory requirements, particularly due to KV cache growth during long-text understanding and generation, present significant challenges for deployment in resourceconstrained environments.Quantization has emerged as a promising solution to reduce memory consumption while preserving historical information.We propose XQuant, a training-free and plug-and-play framework that achieves ultra-low equivalent bit-width KV cache quantization.XQuant introduces two key innovations: a computationally negligible data-free calibration method and cross-layer KV cache compression, enabling quantization to sub-1.4 bits.Extensive experiments on TruthfulQA and LongBench demonstrate that XQuant outperforms state-of-the-art methods (e.g., KIVI-2bit and AsymKV-1.5bit)by achieving lower bit-width while maintaining superior performance, establishing a better trade-off between memory efficiency and model accuracy.The source code is available at https: //github.com/brinenick511/XQuant.KeyCache[l][0], 14 KeyCache[l][1], 15 KeyCache[l][2] 16 else 17 DequantizedKey ← Dequantize 18 KeyCache[l -1][0], 19 KeyCache[l -1][1], 20 KeyCache[l][2] 21 if l < vm or l mod 2 == 0 then 22 DequantizedValue ← Dequantize 23 ValueCache[l][0], 24 ValueCache[l][1], 25 ValueCache[l][2] 26 else 27 DequantizedValue ← Dequantize 28 ValueCache[l -1][0], 29 ValueCache[l -1][1], 30 ValueCache[l][2]
Haoqi Yang 0001, Yao Yao 0008, Zuchao Li, Baoyuan Qi, Guoming Liu, Hai Zhao 0001
EMNLP3
2025 Joint Automatic Speech Recognition And Structure Learning For Better Speech Understanding
abstract
Spoken language understanding (SLU) is a structure prediction task in the field of speech. Recently, many works on SLU that treat it as a sequence-to-sequence task have achieved great success. However, This method is not suitable for simultaneous speech recognition and understanding. In this paper, we propose a joint speech recognition and structure learning framework (JSRSL), an end-to-end SLU model based on span, which can accurately transcribe speech and extract structured content simultaneously. We conduct experiments on name entity recognition and intent classification using the Chinese dataset AISHELL-NER and the English dataset SLURP. The results show that our proposed method not only outperforms the traditional sequence-to-sequence method in both transcription and extraction capabilities but also achieves state-of-the-art performance on the two datasets.
Jiliang Hu 0001, Zuchao Li, Mengjia Shen, Haojun Ai, Sheng Li 0010
ICASSP2
2025 What Limits Bidirectional Model's Generative Capabilities? A Uni-Bi-Directional Mixture-of-Expert Method For Bidirectional Fine-tuning
abstract
Large Language Models (LLMs) excel in generation tasks, yet their causal attention mechanisms limit performance in embedding tasks. While bidirectional modeling may enhance embeddings, naively fine-tuning unidirectional models bidirectionally severely degrades generative performance. To investigate this trade-off, we analyze attention weights as dependence indicators and find that bidirectional fine-tuning increases subsequent dependence, impairing unidirectional generation. Through systematic Transformer module evaluations, we discover the FFN layer is least affected by such dependence. Leveraging this discovery, we propose UBMoE-LLM, a novel Uni-Bi-directional Mixture-of-Experts LLM, which integrates the original unidirectional FFN with a bidirectionally fine-tuned FFN via unsupervised contrastive learning. This MoE-based approach enhances embedding performance while preserving robust generation. Extensive experiments across diverse datasets and model scales validate our attention dependence metric and demonstrate UBMoE-LLM’s superior generative quality and reduced hallucination. Code is available at: https://github.com/heiyonghua/ubmoe_llm.
Zuchao Li, Yonghua Hei, Qiwei Li 0002, Lefei Zhang, Ping Wang 0028, Hai Zhao 0001, Baoyuan Qi, Guoming Liu
ICML1
2025 Label Drop for Multi-Aspect Relation Modeling in Universal Information Extraction
abstract
Lu Yang, Jiajia Li, En Ci, Lefei Zhang, Zuchao Li, Ping Wang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Lu Yang 0008, Jiajia Li 0005, En Ci, Lefei Zhang, Zuchao Li, Ping Wang 0028
NAACL (Long Papers)5
2025 SmallKV: Small Model Assisted Compensation of KV Cache Compression for Efficient LLM Inference
abstract
KV cache eviction has emerged as an effective solution to alleviate resource constraints faced by LLMs in long-context scenarios. However, existing token-level eviction methods often overlook two critical aspects: (1) their irreversible eviction strategy fails to adapt to dynamic attention patterns during decoding (the saliency shift problem), and (2) they treat both marginally important tokens and truly unimportant tokens uniformly, despite the collective significance of marginal tokens to model performance (the marginal information over-compression problem). To address these issues, we design two compensation mechanisms based on the high similarity of attention matrices between LLMs with different scales. We propose SmallKV, a small model assisted compensation method for KV cache compression. SmallKV can maintain attention matching between different-scale LLMs to: 1) assist the larger model in perceiving globally important information of attention; and 2) use the smaller model’s attention scores to approximate those of marginal tokens in the larger model. Extensive experiments on benchmarks including GSM8K, BBH, MT-Bench, and LongBench demonstrate the effectiveness of SmallKV. Moreover, efficiency evaluations show that SmallKV achieves 1.75 - 2.56 times higher throughput than baseline methods, highlighting its potential for efficient and performant LLM inference in resource constrained environments.
Yajuan Peng, Cam-Tu Nguyen, Zuchao Li, Xiaoliang Wang 0001, Hai Zhao 0001, Xiaoming Fu 0001
NeurIPS4
2025 Enhancing Perception of Key Changes in Remote Sensing Image Change Captioning
abstract
Recently, while significant progress has been made in remote sensing image change captioning, existing methods fail to filter out areas unrelated to actual changes, making models susceptible to irrelevant features. In this article, we propose a novel multimodal model for remote sensing image change captioning, guided by Key Change Features and Instruction-tuned (KCFI). This model aims to fully leverage the intrinsic knowledge of large language models through visual instructions and enhance the effectiveness and accuracy of change features using pixel-level change detection tasks. Specifically, KCFI includes a ViTs encoder for extracting bi-temporal remote sensing image features, a key feature perceiver for identifying critical change areas, a pixel-level change detection decoder to constrain key change features, and an instruction-tuned decoder based on a large language model. Moreover, to ensure that change captioning and change detection tasks are jointly optimized, we employ a dynamic weight-averaging strategy to balance the losses between the two tasks. We also explore various feature combinations for visual fine-tuning instructions and demonstrate that using only key change features to guide the large language model is the optimal choice. To validate the effectiveness of our approach, we compare it against several state-of-the-art change captioning methods on the LEVIR-CC dataset, achieving the best performance. Our code will be available at https://github.com/yangcong356/KCFI.git.
Zuchao Li, Hongzan Jiao, Zhi Gao 0005, Lefei Zhang
IEEE Trans. Image Process.2
2024 Multi-Modal Latent Space Learning for Chain-of-Thought Reasoning in Language Models
abstract
Chain-of-thought (CoT) reasoning has exhibited impressive performance in language models for solving complex tasks and answering questions. However, many real-world questions require multi-modal information, such as text and images. Previous research on multi-modal CoT has primarily focused on extracting fixed image features from off-the-shelf vision models and then fusing them with text using attention mechanisms. This approach has limitations because these vision models were not designed for complex reasoning tasks and do not align well with language thoughts. To overcome this limitation, we introduce a novel approach for multi-modal CoT reasoning that utilizes latent space learning via diffusion processes to generate effective image features that align with language thoughts. Our method fuses image features and text representations at a deep level and improves the complex reasoning ability of multi-modal CoT. We demonstrate the efficacy of our proposed method on multi-modal ScienceQA and machine translation benchmarks, achieving state-of-the-art performance on ScienceQA. Overall, our approach offers a more robust and effective solution for multi-modal reasoning in language models, enhancing their ability to tackle complex real-world problems.
Liqi He, Zuchao Li, Xiantao Cai, Ping Wang 0028
AAAI2
2024 A Novel Energy Based Model Mechanism for Multi-Modal Aspect-Based Sentiment Analysis
abstract
Multi-modal aspect-based sentiment analysis (MABSA) has recently attracted increasing attention. The span-based extraction methods, such as FSUIE, demonstrate strong performance in sentiment analysis due to their joint modeling of input sequences and target labels. However, previous methods still have certain limitations: (i) They ignore the difference in the focus of visual information between different analysis targets (aspect or sentiment). (ii) Combining features from uni-modal encoders directly may not be sufficient to eliminate the modal gap and can cause difficulties in capturing the image-text pairwise relevance. (iii) Existing span-based methods for MABSA ignore the pairwise relevance of target span boundaries. To tackle these limitations, we propose a novel framework called DQPSA. Specifically, our model contains a Prompt as Dual Query (PDQ) module that uses the prompt as both a visual query and a language query to extract prompt-aware visual information and strengthen the pairwise relevance between visual information and the analysis target. Additionally, we introduce an Energy-based Pairwise Expert (EPE) module that models the boundaries pairing of the analysis target from the perspective of an Energy-based Model. This expert predicts aspect or sentiment span based on pairwise stability. Experiments on three widely used benchmarks demonstrate that DQPSA outperforms previous approaches and achieves a new state-of-the-art performance. The code will be released at https://github.com/pengts/DQPSA.
Tianshuo Peng, Zuchao Li, Ping Wang 0028, Lefei Zhang, Hai Zhao 0001
AAAI2
2024 N-gram Unsupervised Compoundation and Feature Injection for Better Symbolic Music Understanding
abstract
The first step to apply deep learning techniques for symbolic music understanding is to transform musical pieces (mainly in MIDI format) into sequences of predefined tokens like note pitch, note velocity, and chords. Subsequently, the sequences are fed into a neural sequence model to accomplish specific tasks. Music sequences exhibit strong correlations between adjacent elements, making them prime candidates for N-gram techniques from Natural Language Processing (NLP). Consider classical piano music: specific melodies might recur throughout a piece, with subtle variations each time. In this paper, we propose a novel method, NG-Midiformer, for understanding symbolic music sequences that leverages the N-gram approach. Our method involves first processing music pieces into word-like sequences with our proposed unsupervised compoundation, followed by using our N-gram Transformer encoder, which can effectively incorporate N-gram information to enhance the primary encoder part for better understanding of music sequences. The pre-training process on large-scale music datasets enables the model to thoroughly learn the N-gram information contained within music sequences, and subsequently apply this information for making inferences during the fine-tuning stage. Experiment on various datasets demonstrate the effectiveness of our method and achieved state-of-the-art performance on a series of music understanding downstream tasks. The code and model weights will be released at https://github.com/CinqueOrigin/NG-Midiformer.
Jinhao Tian, Zuchao Li, Jiajia Li 0005, Ping Wang 0028
AAAI2
2024 Hypergraph based Understanding for Document Semantic Entity Recognition
abstract
Semantic entity recognition is an important task in the field of visually-rich document understanding.It distinguishes the semantic types of text by analyzing the position relationship between text nodes and the relation between text content.The existing document understanding models mainly focus on entity categories while ignoring the extraction of entity boundaries.We build a novel hypergraph attention document semantic entity recognition framework, HGA, which uses hypergraph attention to focus on entity boundaries and entity categories at the same time.It can conduct a more detailed analysis of the document text representation analyzed by the upstream model and achieves a better performance of semantic information.We apply this method on the basis of GraphLayoutLM to construct a new semantic entity recognition model HGALayoutLM.Our experiment results on FUNSD, CORD, XFUND and SROIE show that our method can effectively improve the performance of semantic entity recognition tasks based on the original model.The results of HGALayoutLM on FUNSD and XFUND reach the new state-ofthe-art results.
Qiwei Li 0002, Zuchao Li, Ping Wang 0028, Haojun Ai, Hai Zhao 0001
ACL (1)2
2024 SirLLM: Streaming Infinite Retentive LLM
abstract
As Large Language Models (LLMs) become increasingly prevalent in various domains, their ability to process inputs of any length and maintain a degree of memory becomes essential.However, the one-off input of overly long texts is limited, as studies have shown that when input lengths exceed the LLMs' pre-trained text length, there is a dramatic decline in text generation capabilities.Moreover, simply extending the length of pre-training texts is impractical due to the difficulty in obtaining long text data and the substantial memory consumption costs this would entail for LLMs.Recent efforts have employed streaming inputs to alleviate the pressure of excessively long text inputs, but this approach can significantly impair the model's long-term memory capabilities.Motivated by this challenge, we introduce Streaming Infinite Retentive LLM (SirLLM), which allows LLMs to maintain longer memory during infinite-length dialogues without the need for fine-tuning.SirLLM utilizes the Token Entropy metric and a memory decay mechanism to filter key phrases, endowing LLMs with both long-lasting and flexible memory.We designed three distinct tasks and constructed three datasets to measure the effectiveness of SirLLM from various angles: (1) DailyDialog; (2) Grocery Shopping; (3) Rock-Paper-Scissors.Our experimental results robustly demonstrate that SirLLM can achieve stable and significant improvements across different LLMs and tasks, compellingly proving its effectiveness.When having a coversation, "A sir could forget himself," but SirLLM never does!Our
Yao Yao 0008, Zuchao Li, Hai Zhao 0001
ACL (1)2
2024 Soft-Prompting with Graph-of-Thought for Multi-modal Representation Learning
abstract
The chain-of-thought technique has been received well in multi-modal tasks. It is a step-by-step linear reasoning process that adjusts the length of the chain to improve the performance of generated prompts. However, human thought processes are predominantly non-linear, as they encompass multiple aspects simultaneously and employ dynamic adjustment and updating mechanisms. Therefore, we propose a novel Aggregation-Graph-of-Thought (AGoT) mechanism for soft-prompt tuning in multi-modal representation learning. The proposed AGoT models the human thought process not only as a chain but also models each step as a reasoning aggregation graph to cope with the overlooked multiple aspects of thinking in single-step reasoning. This turns the entire reasoning process into prompt aggregation and prompt flow operations. Experiments show that our multi-modal model enhanced with AGoT soft-prompting achieves good results in several tasks such as text-image retrieval, visual question answering, and image recognition. In addition, we demonstrate that it has good domain generalization performance due to better reasoning.
Juncheng Yang, Zuchao Li, Shuai Xie, Wei Yu 0009, Shijun Li 0001, Bo Du 0001
LREC/COLING2
2024 A Coin Has Two Sides: A Novel Detector-Corrector Framework for Chinese Spelling Correction
abstract
Chinese Spelling Correction (CSC) stands as a foundational Natural Language Processing (NLP) task, which primarily focuses on the correction of erroneous characters in Chinese texts. Certain existing methodologies opt to disentangle the error correction process, employing an additional error detector to pinpoint error positions. However, owing to the inherent performance limitations of error detector, precision and recall are like two sides of the coin which can not be both facing up simultaneously. Furthermore, it is also worth investigating how the error position information can be judiciously applied to assist the error correction. In this paper, we introduce a novel approach based on error detector-corrector framework. Our detector is designed to yield two error detection results, each characterized by high precision and recall. Given that the occurrence of errors is context-dependent and detection outcomes may be less precise, we incorporate the error detection results into the CSC task using an innovative feature fusion strategy and a selective masking strategy. Empirical experiments conducted on mainstream CSC datasets substantiate the efficacy of our proposed method.
Xiangke Zeng, Zuchao Li, Lefei Zhang, Ping Wang 0028, Hongqiu Wu, Hai Zhao 0001
ECAI2
2024 VHASR: A Multimodal Speech Recognition System With Vision Hotwords
abstract
The image-based multimodal automatic speech recognition (ASR) model enhances speech recognition performance by incorporating audio-related image.However, some works suggest that introducing image information to model does not help improving ASR performance.In this paper, we propose a novel approach effectively utilizing audio-related image information and set up VHASR, a multimodal speech recognition system that uses vision as hotwords to strengthen the model's speech recognition capability.Our system utilizes a dual-stream architecture, which firstly transcribes the text on the two streams separately, and then combines the outputs.We evaluate the proposed model on four datasets: Flickr8k, ADE20k, COCO, and OpenImages.The experimental results show that VHASR can effectively utilize key information in images to enhance the model's speech recognition ability.Its performance not only surpasses unimodal ASR, but also achieves SOTA among existing image-based multimodal ASR. 1
Jiliang Hu 0001, Zuchao Li, Ping Wang 0028, Haojun Ai, Lefei Zhang, Hai Zhao 0001
EMNLP2
2024 Semantics-Preserved Distortion for Personal Privacy Protection in Information Management
Jiajia Li 0005, Lu Yang 0008, Letian Peng, Shitou Zhang, Ping Wang 0028, Zuchao Li, Hai Zhao 0001
ICANN (5)6
2024 Cross-Modal Adapter: Parameter-Efficient Transfer Learning Approach for Vision-Language Models
abstract
Adapter-based parameter-efficient transfer learning has achieved exciting results in vision-language models. Traditional adapter methods often require training or fine-tuning, facing challenges such as insufficient samples or resource limitations. While some methods overcome the need for training by leveraging image modality cache and retrieval, they overlook the text modality’s importance and cross-modal cues for the efficient adaptation of parameters in visual-language models. This work introduces a cross-modal parameter-efficient approach named XMAdapter. XMAdapter establishes cache models for both text and image modalities. It then leverages retrieval through visual-language bimodal information to gather clues for inference. By dynamically adjusting the affinity ratio, it achieves cross-modal fusion, decoupling different modal similarities to assess their respective contributions. Additionally, it explores hard samples based on differences in cross-modal affinity and enhances model performance through adaptive adjustment of sample learning intensity. Extensive experimental results on benchmark datasets demonstrate that XMAdapter outperforms previous adapter-based methods significantly regarding accuracy, generalization, and efficiency.
Juncheng Yang, Zuchao Li, Shuai Xie, Weiping Zhu 0004, Wei Yu 0009, Shijun Li 0001
ICME2
2024 Sparse is Enough in Fine-tuning Pre-trained Large Language Models
abstract
With the prevalence of pre-training-fine-tuning paradigm, how to efficiently adapt the pre-trained model to the downstream tasks has been an intriguing issue. $\textbf{P}$arameter-$\textbf{E}$fficient $\textbf{F}$ine-$\textbf{T}$uning(PEFT) methods have been proposed for low-cost adaptation. Although PEFT has demonstrated effectiveness and been widely applied, the underlying principles are still unclear. In this paper, we adopt the PAC-Bayesian generalization error bound, viewing pre-training as a shift of prior distribution which leads to a tighter bound for generalization error. We validate this shift from the perspectives of oscillations in the loss landscape and the quasi-sparsity in gradient distribution. Based on this, we propose a gradient-based sparse fine-tuning algorithm, named $\textbf{S}$parse $\textbf{I}$ncrement $\textbf{F}$ine-$\textbf{T}$uning(SIFT), and validate its effectiveness on a range of tasks including the GLUE Benchmark and Instruction-tuning. The code is accessible at https://github.com/song-wx/SIFT/.
Weixi Song, Zuchao Li, Lefei Zhang, Hai Zhao 0001, Bo Du 0001
ICML2
2024 Multi-modal Auto-regressive Modeling via Visual Tokens
abstract
Large Language Models (LLMs), benefiting from the auto-regressive modelling approach performed on massive unannotated texts corpora, demonstrates powerful perceptual and reasoning capabilities. However, as for extending auto-regressive modelling to multi-modal scenarios to build Large Multi-modal Models (LMMs), there lies a great difficulty that the image information is processed in the LMM as continuous visual embeddings, which cannot obtain discrete supervised labels for classification. In this paper, we successfully perform multi-modal auto-regressive modeling with a unified objective for the first time. Specifically, we propose the concept of visual tokens, which maps the visual features to probability distributions over LLM's vocabulary, providing supervision information for visual modelling. We further explore the distribution of visual features in the semantic space within LMM and the possibility of using text embeddings to represent visual information. Experimental results and ablation studies on 5 VQA tasks and 4 benchmark toolkits validate the powerful performance of our proposed approach.
Tianshuo Peng, Zuchao Li, Lefei Zhang, Hai Zhao 0001, Ping Wang 0028, Bo Du 0001
ACM Multimedia2
2024 Reference Trustable Decoding: A Training-Free Augmentation Paradigm for Large Language Models
abstract
Large language models (LLMs) have rapidly advanced and demonstrated impressive capabilities. In-Context Learning (ICL) and Parameter-Efficient Fine-Tuning (PEFT) are currently two mainstream methods for augmenting LLMs to downstream tasks. ICL typically constructs a few-shot learning scenario, either manually or by setting up a Retrieval-Augmented Generation (RAG) system, helping models quickly grasp domain knowledge or question-answering patterns without changing model parameters. However, this approach involves trade-offs, such as slower inference speed and increased space occupancy. PEFT assists the model in adapting to tasks through minimal parameter modifications, but the training process still demands high hardware requirements, even with a small number of parameters involved. To address these challenges, we propose Reference Trustable Decoding (RTD), a paradigm that allows models to quickly adapt to new tasks without fine-tuning, maintaining low inference costs. RTD constructs a reference datastore from the provided training examples and optimizes the LLM's final vocabulary distribution by flexibly selecting suitable references based on the input, resulting in more trustable responses and enabling the model to adapt to downstream tasks at a low cost. Experimental evaluations on various LLMs using different benchmarks demonstrate that RTD establishes a new paradigm for augmenting models to downstream tasks. Furthermore, our method exhibits strong orthogonality with traditional methods, allowing for concurrent usage. Our code can be found at https://github.com/ShiLuohe/ReferenceTrustableDecoding.
Luohe Shi, Yao Yao 0008, Zuchao Li, Lefei Zhang, Hai Zhao 0001
NeurIPS3
2024 Centroid-Centered Modeling for Efficient Vision Transformer Pre-Training
Xin Yan 0008, Zuchao Li, Lefei Zhang
PRCV (4)2
2024 Confidence-based Syntax encoding network for better ancient Chinese understanding
Shitou Zhang, Ping Wang 0028, Zuchao Li, Jingrui Hou, Qibiao Hu
Inf. Process. Manag.3
2024 Generalizing to unseen domains via PatchMix
Juncheng Yang, Zuchao Li, Shuai Xie, Wei Yu 0009, Shijun Li 0001
Multim. Syst.2
2024 Bidirectional correlation-driven inter-frame interaction Transformer for referring video object segmentation
Meng Lan, Fu Rong, Zuchao Li, Wei Yu 0009, Lefei Zhang
Pattern Recognit.3
2024 Enhancing Lyrics Rewriting with Weak Supervision from Grammatical Error Correction Pre-training and Reference Knowledge Fusion
abstract
Lyric rewriting involves taking the original lyrics of a song and creatively rephrasing them while preserving their core meaning and emotional essence. Sequence-to-sequence methods often face the problem of lack of annotated corpus and difficulty in understanding lyrics when dealing with the lyric rewriting task. Inspired by the language rewriting technique, grammatical error correction (GEC) and sequence-to-sequence generation techniques, and neural machine translation (NMT) methods, we propose novel self-supervised learning methods that can effectively solve the problem of the lack of a lyric rewriting corpus. In addition, we also propose a new pretrained DAE Transformer model with data prior knowledge fusion to enhance the lyric rewriting ability. The reference-as-context model (RaC-Large) constructed by us based on these two methods achieves the best results in comparison with the baseline including large language models, fully verifying the effectiveness of the new method. We also validate the effectiveness of our approach on GEC and NMT tasks, further demonstrating the potential of our approach on a broad range of sequence-to-sequence tasks.
Jiajia Li 0005, Ping Wang 0028, Zuchao Li, Kevin Parnow, Hai Zhao 0001, Weiping Ding 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2024 Entity-Relation Extraction as Full Shallow Semantic Dependency Parsing
abstract
Entity-relation extraction is the essential information extraction task and can be decomposed into Named Entity Recognition (NER) and Relation Extraction (RE) subtasks. This paper proposes a novel joint entity-relation extraction method that models the entity-relation extraction task as full shallow semantic dependency graph parsing. Specifically, it jointly and simultaneously converts the entities and relation mentions as the edges of the semantic dependency graph to be parsed and their types as the labels. This model also integrates the advantages of multiple feature tagging methods and enriches the token representation. Furthermore, second-order scoring is introduced to exploit the relationships between entities and relations, which improves the model performance. Our work is the first time to fully model entities and relations into a graph and uses higher-order modules to address their interaction problems. Compared with state-of-the-art scores on five benchmarks (ACE04, ACE05, CoNLL04, ADE, and SciERC), empirical results show that our proposed model makes significant improvements and demonstrates its effectiveness and practicability.
Zuchao Li, Hai Zhao 0001, Weiping Ding 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2024 Bootstrapping Interactive Image-Text Alignment for Remote Sensing Image Captioning
abstract
Recently, remote sensing image captioning has gained significant attention in the remote sensing community. Due to the significant differences in spatial resolution of remote sensing images, existing methods in this field have predominantly concentrated on the fine-grained extraction of remote sensing image features, but they cannot effectively handle the semantic consistency between visual features and textual features. To efficiently align the image-text, we propose a novel two-stage vision-language pre-training-based approach to bootstrap interactive image-text alignment for remote sensing image captioning, called BITA, which relies on the design of a lightweight interactive Fourier Transformer to better align remote sensing image-text features. The Fourier layer in the interactive Fourier Transformer is capable of extracting multi-scale features of remote sensing images in the frequency domain, thereby reducing the redundancy of remote sensing visual features. Specifically, the first stage involves preliminary alignment through image-text contrastive learning, which aligns the learned multi-scale remote sensing features from the interactive Fourier Transformer with textual features. In the second stage, the interactive Fourier Transformer connects the frozen image encoder with a large language model. Then, prefix causal language modeling is utilized to guide the text generation process using visual features. Ultimately, across the UCM-caption, RSICD, and NWPU-caption datasets, the experimental results clearly demonstrate that BITA outperforms other advanced comparative approaches. The code is available at https://github.com/yangcong356/BITA.
Zuchao Li, Lefei Zhang
IEEE Trans. Geosci. Remote. Sens.2
2024 MGIMM: Multi-Granularity Instruction Multimodal Model for Attribute-Guided Remote Sensing Image Detailed Description
abstract
Generating detailed textual descriptions of remote sensing images is challenging because it requires capturing both global and local visual information. The complexity of backgrounds and the scale variations among targets make it difficult to align visual regions with corresponding textual attributes. Furthermore, large multimodal models, while effective in general scenarios, struggle in remote sensing due to their lack of specialized knowledge and regional awareness. To address these issues, this article proposes an attribute-guided multi-granularity instruction multimodal model (MGIMM) for remote sensing image detailed description. MGIMM guides the multimodal model to learn the consistency between visual regions and corresponding text attributes (such as object names, colors, and shapes) through region-level instruction tuning. Then, with the multimodal model aligned on region attribute, guided by multigrain visual features, MGIMM fully perceives both region-level and global image information, utilizing large language models for comprehensive descriptions of remote sensing images. Due to the lack of a standard benchmark for generating detailed descriptions of remote sensing images, we construct a dataset featuring 38320 region-attribute pairs and 23463 image-detailed description pairs. Compared with various advanced methods on this dataset, the results demonstrate the effectiveness of MGIMM’s region-attribute-guided learning approach. The code is available athttps://github.com/yangcong356/MGIMM.git.
Zuchao Li, Lefei Zhang
IEEE Trans. Geosci. Remote. Sens.2
2023 Fine-Grained Position Helps Memorizing More, a Novel Music Compound Transformer Model with Feature Interaction Fusion
abstract
Due to the particularity of the simultaneous occurrence of multiple events in music sequences, compound Transformer is proposed to deal with the challenge of long sequences. However, there are two deficiencies in the compound Transformer. First, since the order of events is more important for music than natural language, the information provided by the original absolute position embedding is not precise enough. Second, there is an important correlation between the tokens in the compound word, which is ignored by the current compound Transformer. Therefore, in this work, we propose an improved compound Transformer model for music understanding. Specifically, we propose an attribute embedding fusion module and a novel position encoding scheme with absolute-relative consideration. In the attribute embedding fusion module, different attributes are fused through feature permutation by using a multi-head self-attention mechanism in order to capture rich interactions between attributes. In the novel position encoding scheme, we propose RoAR position encoding, which realizes rotational absolute position encoding, relative position encoding, and absolute-relative position interactive encoding, providing clear and rich orders for musical events. Empirical study on four typical music understanding tasks shows that our attribute fusion approach and RoAR position encoding brings large performance gains. In addition, we further investigate the impact of masked language modeling and casual language modeling pre-training on music understanding.
Zuchao Li, Ruhan Gong, Yineng Chen, Kehua Su
AAAI1
2023 FSUIE: A Novel Fuzzy Span Mechanism for Universal Information Extraction
abstract
Universal Information Extraction (UIE) has been introduced as a unified framework for various Information Extraction (IE) tasks and has achieved widespread success.Despite this, UIE models have limitations.For example, they rely heavily on span boundaries in the data during training, which does not reflect the reality of span annotation challenges.Slight adjustments to positions can also meet requirements.Additionally, UIE models lack attention to the limited span length feature in IE.To address these deficiencies, we propose the Fuzzy Span Universal Information Extraction (FSUIE) framework.Specifically, our contribution consists of two concepts: fuzzy span loss and fuzzy span attention.Our experimental results on a series of main IE tasks show significant improvement compared to the baseline, especially in terms of fast convergence and strong performance with small amounts of data and training epochs.These results demonstrate the effectiveness and generalization of FSUIE in different tasks, settings, and scenarios.
Tianshuo Peng, Zuchao Li, Lefei Zhang, Bo Du 0001, Hai Zhao 0001
ACL (1)2
2023 Bidirectional Looking with A Novel Double Exponential Moving Average to Adaptive and Non-adaptive Momentum Optimizers
abstract
Optimizer is an essential component for the success of deep learning, which guides the neural network to update the parameters according to the loss on the training set. SGD and Adam are two classical and effective optimizers on which researchers have proposed many variants, such as SGDM and RAdam. In this paper, we innovatively combine the backward-looking and forward-looking aspects of the optimizer algorithm and propose a novel Admeta (**A** **D**ouble exponential **M**oving averag**E** **T**o **A**daptive and non-adaptive momentum) optimizer framework. For backward-looking part, we propose a DEMA variant scheme, which is motivated by a metric in the stock market, to replace the common exponential moving average scheme. While in the forward-looking part, we present a dynamic lookahead strategy which asymptotically approaches a set value, maintaining its speed at early stage and high convergence performance at final stage. Based on this idea, we provide two optimizer implementations, AdmetaR and AdmetaS, the former based on RAdam and the latter based on SGDM. Through extensive experiments on diverse tasks, we find that the proposed Admeta optimizer outperforms our base optimizers and shows advantages over recently proposed competitive optimizers. We also provide theoretical proof of these two algorithms, which verifies the convergence of our proposed Admeta.
Yineng Chen, Zuchao Li, Lefei Zhang, Bo Du 0001, Hai Zhao 0001
ICML2
2023 iRe2f: Rethinking Effective Refinement in Language Structure Prediction via Efficient Iterative Retrospecting and Reasoning
abstract
Refinement plays a critical role in language structure prediction, a process that deals with complex situations such as structural edge interdependencies. Since language structure prediction usually modeled as graph parsing, typical refinement methods involve taking an initial parsing graph as input and refining it using language input and other relevant information. Intuitively, a refinement component, i.e., refiner, should be lightweight and efficient, as it is only responsible for correcting faults in the initial graph. However, current refiners add a significant burden to the parsing process due to their reliance on time-consuming encoding-decoding procedure on the language input and graph. To make the refiner more practical for real-world applications, this paper proposes a lightweight but effective iterative refinement framework, iRe^2f, based on iterative retrospecting and reasoning without involving the re-encoding process on the graph. iRe^2f iteratively refine the parsing graph based on interaction between graph and sequence and efficiently learns the shortcut to update the sequence and graph representations in each iteration. The shortcut is calculated based on the graph representation in the latest iteration. iRe^2f reduces the number of refinement parameters by 90% compared to the previous smallest refiner. Experiments on a variety of language structure prediction tasks show that iRe^2f performs comparably or better than current state-of-the-art refiners, with a significant increase in efficiency.
Zuchao Li, Xingyi Guo, Letian Peng, Lefei Zhang, Hai Zhao 0001
IJCAI1
2023 MAPLE: Semi-Supervised Learning with Multi-Alignment and Pseudo-Learning
abstract
Data augmentation has undoubtedly enabled a significant leap forward in training a high-accuracy deep network. Besides the commonly used augmentation to target data, e.g., random cropping, flipping, and rotation, recent works have been dedicated to mining generalized knowledge by using multiple sources. However, along with plentiful data comes the huge data distribution gap between the target and different sources (hybrid shift). To mitigate this problem, existing methods tend to manually annotate more data. Unlike previous methods, this paper focuses on the study of learning deep models by gathering knowledge from multiple sources in a labor-free fashion and further proposes the "Multi-Alignment and Pseudo-Learning'' method, dubbed MAPLE. MAPLE constructs the multi-alignment module, which consists of multiple discriminators to align different data distributions via an adversarial process. In addition, a novel semi-supervised learning (SSL) manner is introduced to further facilitate the utility of our MAPLE. Extensive evaluations conducted on four benchmarks show the effectiveness of the proposed MAPLE, which achieves state-of-the-art performance outperforming existing methods by an obvious margin.
Juncheng Yang, Zuchao Li, Wei Yu 0009, Bo Du 0001, Shijun Li 0001
KDD3
2023 Enhancing Visually-Rich Document Understanding via Layout Structure Modeling
abstract
In recent years, the use of multi-modal pre-trained Transformers has led to significant advancements in visually-rich document understanding. However, existing models have mainly focused on features such as text and vision while neglecting the importance of layout relationship between text nodes. In this paper, we propose GraphLayoutLM, a novel document understanding model that leverages the modeling of layout structure graph to inject document layout knowledge into the model. GraphLayoutLM utilizes a graph reordering algorithm to adjust the text sequence based on the graph structure. Additionally, our model uses a layout-aware multi-head self-attention layer to learn document layout knowledge. The proposed model enables the understanding of the spatial arrangement of text elements, improving document comprehension. We evaluate our model on various benchmarks, including FUNSD, XFUND and CORD and it achieves state-of-the-art results among these datasets. Our experiment results demonstrate that our proposed method provides a significant improvement over existing approaches and showcases the importance of incorporating layout information into document understanding models. We also conduct an ablation study to investigate the contribution of each component of our model. The results show that both the graph reordering algorithm and the layout-aware multi-head self-attention layer play a crucial role in achieving the best performance.
Qiwei Li 0002, Zuchao Li, Xiantao Cai, Bo Du 0001, Hai Zhao 0001
ACM Multimedia2
2023 Cross-Lingual Universal Dependency Parsing Only From One Monolingual Treebank
abstract
Syntactic parsing is a highly linguistic processing task whose parser requires training on treebanks from the expensive human annotation. As it is unlikely to obtain a treebank for every human language, in this work, we propose an effective cross-lingual UD parsing framework for transferring parser from only one source monolingual treebank to any other target languages without treebank available. To reach satisfactory parsing accuracy among quite different languages, we introduce two language modeling tasks into the training process of dependency parsing as multi-tasking. Assuming only unlabeled data from target languages plus the source treebank can be exploited together, we adopt a self-training strategy for further performance improvement in terms of our multi-task framework. Our proposed cross-lingual parsers are implemented for English, Chinese, and 29 UD treebanks. The empirical study shows that our cross-lingual parsers yield promising results for all target languages, approaching the parser performance which is trained in its own target treebank.
Kailai Sun, Zuchao Li, Hai Zhao 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Universal Multimodal Representation for Language Understanding
abstract
Representation learning is the foundation of natural language processing (NLP). This work presents new methods to employ visual information as assistant signals to general NLP tasks. For each sentence, we first retrieve a flexible number of images either from a light topic-image lookup table extracted over the existing sentence-image pairs or a shared cross-modal embedding space that is pre-trained on out-of-shelf text-image pairs. Then, the text and images are encoded by a Transformer encoder and convolutional neural network, respectively. The two sequences of representations are further fused by an attention layer for the interaction of the two modalities. In this study, the retrieval process is controllable and flexible. The universal visual representation overcomes the lack of large-scale bilingual sentence-image pairs. Our method can be easily applied to text-only tasks without manually annotated multimodal parallel corpora. We apply the proposed method to a wide range of natural language generation and understanding tasks, including neural machine translation, natural language inference, and semantic similarity. Experimental results show that our method is generally effective for different tasks and languages. Analysis indicates that the visual signals enrich textual representations of content words, provide fine-grained grounding information about the relationship between concepts and events, and potentially conduce to disambiguation.
Zhuosheng Zhang 0001, Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita, Zuchao Li, Hai Zhao 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2023 Semi-Supervised Semantic Role Labeling with Bidirectional Language Models
abstract
The recent success of neural networks in NLP applications has provided a strong impetus to develop supervised models for semantic role labeling (SRL) that forego the requirement for extensive feature engineering. Recent state-of-the-art approaches require high-quality annotated datasets that are costly to obtain and almost unavailable for low-resource languages. We present a semi-supervised approach that utilizes both labeled and unlabeled data to provide performance improvement over a mere supervised SRL model. We show that our proposed semi-supervised SRL model provides larger improvement over a supervised model in the scenario where labeled training data size is small. Our SRL system leverages unlabeled data under the language modeling paradigm. We demonstrate that the incorporation of a self pre-trained bidirectional language model (S-PrLM) into a SRL system can help in SRL performance improvement by learning composition functions from the unlabeled data. Previous researches have concluded that syntax information is very useful for high-performing SRL systems, so we incorporate syntax information by employing an unsupervised approach to leverage dependency path information to connect argument candidates in vector space, which helps in distinguishing arguments with similar contexts but different syntactic functions. The basic idea is to connect predicate ( w p ) with argument candidate ( w a ) with the dependency path ( r ) between them in the embedding space. Experiments on the CoNLL-2008 and CoNLL-2009 datasets confirm that our full SRL model outperforms previous best models in terms of F 1 score.
Kashif Munir, Hai Zhao 0001, Zuchao Li
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2022 BiBL: AMR Parsing and Generation with Bidirectional Bayesian Learning
abstract
Abstract Meaning Representation (AMR) offers a unified semantic representation for natural language sentences. Thus transformation between AMR and text yields two transition tasks in opposite directions, i.e., Text-to-AMR parsing and AMR-to-Text generation. Existing AMR studies only focus on one-side improvements despite the duality of the two tasks, and their improvements are greatly attributed to the inclusion of large extra training data or complex structure modifications which harm the inference speed. Instead, we propose data-efficient Bidirectional Bayesian learning (BiBL) to facilitate bidirectional information transition by adopting a single-stage multitasking strategy so that the resulting model may enjoy much lighter training at the same time. Evaluation on benchmark datasets shows that our proposed BiBL outperforms strong previous seq2seq refinements without the help of extra data which is indispensable in existing counterpart models. We release the codes of BiBL at: https://github.com/KHAKhazeus/BiBL.
Ziming Cheng, Zuchao Li, Hai Zhao 0001
COLING2
2022 Nested Named Entity Recognition as Corpus Aware Holistic Structure Parsing
abstract
As a fundamental natural language processing task and one of core knowledge extraction techniques, named entity recognition (NER) is widely used to extract information from texts for downstream tasks. Nested NER is a branch of NER in which the named entities (NEs) are nested with each other. However, most of the previous studies on nested NER usually apply linear structure to model the nested NEs which are actually accommodated in a hierarchical structure. Thus in order to address this mismatch, this work models the full nested NEs in a sentence as a holistic structure, then we propose a holistic structure parsing algorithm to disclose the entire NEs once for all. Besides, there is no research on applying corpus-level information to NER currently. To make up for the loss of this information, we introduce Point-wise Mutual Information (PMI) and other frequency features from corpus-aware statistics for even better performance by holistic modeling from sentence-level to corpus-level. Experiments show that our model yields promising results on widely-used benchmarks which approach or even achieve state-of-the-art. Further empirical studies show that our proposed corpus-aware features can substantially improve NER domain adaptation, which demonstrates the surprising advantage of our proposed corpus-level holistic structure modeling.
Zuchao Li, Hai Zhao 0001
COLING2
2022 Reorder and then Parse, Fast and Accurate Discontinuous Constituency Parsing
abstract
Discontinuous constituency parsing is still kept developing for its efficiency and accuracy are far behind its continuous counterparts.Motivated by the observation that a discontinuous constituent tree can be simply transformed into a pseudo-continuous one by artificially reordering words in the sentence, we propose a novel reordering method, thereby construct fast and accurate discontinuous constituency parsing systems working in continuous way.Specifically, we model the relative position changes of words as a list of actions.By parsing and performing this actions, the corresponding pseudocontinuous sequence is derived.Discontinuous parse tree can be further inferred via integrating a high-performance pseudo-continuous constituency parser.Our systems are evaluated on three classical discontinuous constituency treebanks, achieving new state-of-the-art on two treebanks and showing a distinct advantage in speed.
Kailai Sun, Zuchao Li, Hai Zhao 0001
EMNLP2
2022 Explicit Alignment Learning for Neural Machine Translation
abstract
Even though neural machine translation (NMT) has become the state-of-the-art solution for end-to-end translation, it still suffers from a lack of translation interpretability, which may be conveniently enhanced by explicit alignment learning (EAL), as performed in traditional statistical machine translation (SMT). To provide the benefits of both NMT and SMT, this paper presents a novel model design that enhances NMT with an additional training process for EAL, in addition to the end-to-end translation training. Thus, we propose two approaches an explicit alignment learning approach, in which we further remove the need for the additional alignment model, and perform embedding mixup with the alignment based on encoder--decoder attention weights in the NMT model. We conducted experiments on both small-scale (IWSLT14 De->En and IWSLT13 Fr->En) and large-scale (WMT14 En->De, En->Fr, WMT17 Zh->En) benchmarks. Evaluation results show that our EAL methods significantly outperformed strong baseline methods, which shows the effectiveness of EAL. Further explorations show that the translation improvements are due to a better spatial alignment of the source and target language embeddings. Our method improves translation performance without the need to increase model parameters and training data, which verifies that the idea of incorporating techniques of SMT into NMT is worthwhile.
Zuchao Li, Hai Zhao 0001, Fengshun Xiao, Masao Utiyama, Eiichiro Sumita
IJCAI1
2022 Incorporating rich syntax information in Grammatical Error Correction
Zuchao Li, Kevin Parnow, Hai Zhao 0001
Inf. Process. Manag.1
2022 Neural Character-Level Syntactic Parsing for Chinese
abstract
In this work, we explore character-level neural syntactic parsing for Chinese with two typical syntactic formalisms: the constituent formalism and a dependency formalism based on a newly released character-level dependency treebank. Prior works in Chinese parsing have struggled with whether to de ne words when modeling character interactions. We choose to integrate full character-level syntactic dependency relationships using neural representations from character embeddings and richer linguistic syntactic information from human-annotated character-level Parts-Of-Speech and dependency labels. This has the potential to better understand the deeper structure of Chinese sentences and provides a better structural formalism for avoiding unnecessary structural ambiguities. Specifically, we first compare two different character-level syntax annotation styles: constituency and dependency. Then, we discuss two key problems for character-level parsing: (1) how to combine constituent and dependency syntactic structure in full character-level trees and (2) how to convert from character-level to word-level for both constituent and dependency trees. In addition, we also explore several other key parsing aspects, including di erent character-level dependency annotations and joint learning of Parts-Of-Speech and syntactic parsing. Finally, we evaluate our models on the Chinese Penn Treebank (CTB) and our published Shanghai Jiao Tong University Chinese Character Dependency Treebank (SCDT). The results show the e effectiveness of our model on both constituent and dependency parsing. We further provide empirical analysis and suggest several directions for future study.
Zuchao Li, Junru Zhou, Hai Zhao 0001, Zhisong Zhang, Haonan Li 0002, Yuqi Ju
J. Artif. Intell. Res.1
2022 Text Compression-Aided Transformer Encoding
abstract
Text encoding is one of the most important steps in Natural Language Processing (NLP). It has been done well by the self-attention mechanism in the current state-of-the-art Transformer encoder, which has brought about significant improvements in the performance of many NLP tasks. Though the Transformer encoder may effectively capture general information in its resulting representations, the backbone information, meaning the gist of the input text, is not specifically focused on. In this paper, we propose explicit and implicit text compression approaches to enhance the Transformer encoding and evaluate models using this approach on several typical downstream tasks that rely on the encoding heavily. Our explicit text compression approaches use dedicated models to compress text, while our implicit text compression approach simply adds an additional module to the main model to handle text compression. We propose three ways of integration, namely backbone source-side fusion, target-side fusion, and both-side fusion, to integrate the backbone information into Transformer-based models for various downstream tasks. Our evaluation on benchmark datasets shows that the proposed explicit and implicit text compression approaches improve results in comparison to strong baselines. We therefore conclude, when comparing the encodings to the baseline models, text compression helps the encoders to learn better language representations.
Zuchao Li, Zhuosheng Zhang 0001, Hai Zhao 0001, Rui Wang 0015, Kehai Chen, Masao Utiyama, Eiichiro Sumita
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Tri-training for Dependency Parsing Domain Adaptation
abstract
In recent years, the research on dependency parsing focuses on improving the accuracy of the domain-specific (in-domain) test datasets and has made remarkable progress. However, there are innumerable scenarios in the real world that are not covered by the dataset, namely, the out-of-domain dataset. As a result, parsers that perform well on the in-domain data usually suffer from significant performance degradation on the out-of-domain data. Therefore, to adapt the existing in-domain parsers with high performance to a new domain scenario, cross-domain transfer learning methods are essential to solve the domain problem in parsing. This paper examines two scenarios for cross-domain transfer learning: semi-supervised and unsupervised cross-domain transfer learning. Specifically, we adopt a pre-trained language model BERT for training on the source domain (in-domain) data at the subword level and introduce self-training methods varied from tri-training for these two scenarios. The evaluation results on the NLPCC-2019 shared task and universal dependency parsing task indicate the effectiveness of the adopted approaches on cross-domain transfer learning and show the potential of self-learning to cross-lingual transfer learning.
Zuchao Li, Hai Zhao 0001, Bao-Liang Lu, Rui Wang 0015
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2022 Dependency and Span, Cross-Style Semantic Role Labeling on PropBank and NomBank
abstract
The latest developments in neural semantic role labeling (SRL) have shown great performance improvements with both the dependency and span formalism/styles. Although the two styles share many similarities in linguistic meaning and computation, most previous studies focus on a single style. In this article, we define a new cross-style semantic role label convention and propose a new cross-style joint optimization model designed around the most basic linguistic meaning of a semantic role. Our work provides a solution to make the results of the two styles more comparable and allowing both formalisms of SRL to benefit from their natural connections in both linguistics and computation. Our model learns a general semantic argument structure and is capable of outputting in either style. Additionally, we propose a syntax-aided method to uniformly enhance the learning of both dependency and span representations. Experiments show that the proposed methods are effective on both span and dependency SRL benchmarks.
Zuchao Li, Hai Zhao 0001, Junru Zhou, Kevin Parnow, Shexia He
ACM Trans. Asian Low Resour. Lang. Inf. Process.1
2022 HPSG-Inspired Joint Neural Constituent and Dependency Parsing in O($n^3$) Time Complexity
abstract
Constituent and dependency parsing, the two classic forms of syntactic parsing, have been found to benefit from joint training and decoding under a uniform formalism, inspired by Head-driven Phrase Structure Grammar (HPSG). We thus refer to this joint parsing of constituency and dependency as HPSG-like parsing. However, in HPSG-like parsing, decoding this unified grammar has a higher time complexity ($O(n^5)$) than decoding either form individually ($O(n^3)$) since more factors have to be considered during decoding. We thus propose an improved head scorer that helps achieve a novel performance-preserved parser in$O(n^3$) time complexity. Furthermore, on the basis of this proposed practical HPSG-like parser, we investigated the strengths of HPSG-like parsing and explored the general method of training an HPSG-like parser from only a constituent or dependency annotations in a multilingual scenario. We thus present a more effective, more in-depth, and general work on HPSG-like parsing.
Zuchao Li, Junru Zhou, Hai Zhao 0001, Kevin Parnow
IEEE ACM Trans. Audio Speech Lang. Process.1
2021 Seeking Common but Distinguishing Difference, A Joint Aspect-based Sentiment Analysis Model
abstract
Aspect-based sentiment analysis (ABSA) task consists of three typical subtasks: aspect term extraction, opinion term extraction, and sentiment polarity classification.These three subtasks are usually performed jointly to save resources and reduce the error propagation in the pipeline.However, most of the existing joint models only focus on the benefits of encoder sharing between subtasks but ignore the difference.Therefore, we propose a joint ABSA model, which not only enjoys the benefits of encoder sharing but also focuses on the difference to improve the effectiveness of the model.In detail, we introduce a dual-encoder design, in which a pair encoder especially focuses on candidate aspect-opinion pair classification, and the original encoder keeps attention on sequence labeling.Empirical results show that our proposed model shows robustness and significantly outperforms the previous state-ofthe-art on four benchmark datasets.
Hongjiang Jing, Zuchao Li, Hai Zhao 0001
EMNLP (1)2
2021 Unsupervised Neural Machine Translation with Universal Grammar
abstract
Machine translation usually relies on parallel corpora to provide parallel signals for training.The advent of unsupervised machine translation has brought machine translation away from this reliance, though performance still lags behind traditional supervised machine translation.In unsupervised machine translation, the model seeks symmetric language similarities as a source of weak parallel signal to achieve translation.Chomsky's Universal Grammar theory postulates that grammar is an innate form of knowledge to humans and is governed by universal principles and constraints.Therefore, in this paper, we seek to leverage such shared grammar clues to provide more explicit language parallel signals to enhance the training of unsupervised machine translation models.Through experiments on multiple typical language pairs, we demonstrate the effectiveness of our proposed approaches.
Zuchao Li, Masao Utiyama, Eiichiro Sumita, Hai Zhao 0001
EMNLP (1)1
2021 Multilingual Pre-training with Universal Dependency Learning
abstract
The pre-trained language model (PrLM) demonstrates domination in downstream natural language processing tasks, in which multilingual PrLM takes advantage of language universality to alleviate the issue of limited resources for low-resource languages. Despite its successes, the performance of multilingual PrLM is still unsatisfactory, when multilingual PrLMs only focus on plain text and ignore obvious universal linguistic structure clues. Existing PrLMs have shown that monolingual linguistic structure knowledge may bring about better performance. Thus we propose a novel multilingual PrLM that supports both explicit universal dependency parsing and implicit language modeling. Syntax in terms of universal dependency parse serves as not only pre-training objective but also learned representation in our model, which brings unprecedented PrLM interpretability and convenience in downstream task use. Our model outperforms two popular multilingual PrLM, multilingual-BERT and XLM-R, on cross-lingual natural language understanding (NLU) benchmarks and linguistic structure parsing datasets, demonstrating the effectiveness and stronger cross-lingual modeling capabilities of our approach.
Kailai Sun, Zuchao Li, Hai Zhao 0001
NeurIPS2
2021 Syntax Role for Neural Semantic Role Labeling
abstract
Semantic role labeling (SRL) is dedicated to recognizing the semantic predicate-argument structure of a sentence. Previous studies in terms of traditional models have shown syntactic information can make remarkable contributions to SRL performance; however, the necessity of syntactic information was challenged by a few recent neural SRL studies that demonstrate impressive performance without syntactic backbones and suggest that syntax information becomes much less important for neural semantic role labeling, especially when paired with recent deep neural network and large-scale pre-trained language models. Despite this notion, the neural SRL field still lacks a systematic and full investigation on the relevance of syntactic information in SRL, for both dependency and both monolingual and multilingual settings. This paper intends to quantify the importance of syntactic information for neural SRL in the deep learning framework. We introduce three typical SRL frameworks (baselines), sequence-based, tree-based, and graph-based, which are accompanied by two categories of exploiting syntactic information: syntax pruning-based and syntax feature-based. Experiments are conducted on the CoNLL-2005, -2009, and -2012 benchmarks for all languages available, and results show that neural SRL models can still benefit from syntactic information under certain conditions. Furthermore, we show the quantitative significance of syntax to neural SRL models together with a thorough empirical survey using existing models.
Zuchao Li, Hai Zhao 0001, Shexia He, Jiaxun Cai
Comput. Linguistics1
2021 Neural Unsupervised Semantic Role Labeling
abstract
The task of semantic role labeling ( SRL ) is dedicated to finding the predicate-argument structure. Previous works on SRL are mostly supervised and do not consider the difficulty in labeling each example which can be very expensive and time-consuming. In this article, we present the first neural unsupervised model for SRL. To decompose the task as two argument related subtasks, identification and clustering, we propose a pipeline that correspondingly consists of two neural modules. First, we train a neural model on two syntax-aware statistically developed rules. The neural model gets the relevance signal for each token in a sentence, to feed into a BiLSTM, and then an adversarial layer for noise-adding and classifying simultaneously, thus enabling the model to learn the semantic structure of a sentence. Then we propose another neural model for argument role clustering, which is done through clustering the learned argument embeddings biased toward their dependency relations. Experiments on the CoNLL-2009 English dataset demonstrate that our model outperforms the previous state-of-the-art baseline in terms of non-neural models for argument identification and classification.
Kashif Munir, Hai Zhao 0001, Zuchao Li
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2021 Adaptive Convolution for Semantic Role Labeling
abstract
Semantic role labeling (SRL) aims at elaborating the meaning of a sentence by forming a predicate-argument structure. Recent researches depicted that the effective use of syntax can improve SRL performance. However, syntax is a complicated linguistic clue and is hard to be effectively applied in a downstream task like SRL. This work effectively encodes syntax using adaptive convolution which endows strong flexibility to existing convolutional networks. The existing CNNs may help in encoding a complicated structure like syntax for SRL, but it still has shortcomings. Contrary to traditional convolutional networks that use same filters for different inputs, adaptive convolution uses adaptively generated filters conditioned on syntactically-informed inputs. We achieve this with the integration of a filter generation network which generates the input specific filters. This helps the model to focus on important syntactic features present inside the input, thus enlarging the gap between syntax-aware and syntax-agnostic SRL systems. We further study a hashing technique to compress the size of the filter generation network for SRL in terms of trainable parameters. Experiments on CoNLL-2009 dataset confirm that the proposed model substantially outperforms most previous SRL systems for both English and Chinese languages.
Kashif Munir, Hai Zhao 0001, Zuchao Li
IEEE ACM Trans. Audio Speech Lang. Process.3
2021 Learning Context-Aware Convolutional Filters for Implicit Discourse Relation Classification
abstract
Implicit discourse relation classification (IDRC) is considered the most difficult component of shallow discourse parsing as the relation prediction in the absence of necessary clues requires a deep understanding of the context information of the sentences. Convolutional neural networks (CNNs) have emerged as an important encoding block for sentences in natural language processing (NLP). CNNs use a specific set of filters for the inputs which may lead to the partial coverage of contextual clues. Furthermore, conventional CNNs may not allow the initial communication between the sentences which is a crucial step for IDRC. We present an adaptive convolution approach for IDRC that utilizes context aware filters for the convolution operation. The goal is to abstract the context of sentences in the filters and let them interact with sentence representations, i.e. learning the representations through learned filters. Our model acts as a cross questioning agent by generating filters from one argument and convolving them with the other for the IDRC task. This process is analogous to the attention mechanism because both methods aim at abstracting contextual information. Different from the attention mechanism, our approach directly encodes the contextual representations in the form of filters and allows the initial communication between arguments during encoding. Furthermore, the adaptive convolution can also work alongside the attention mechanism to enhance the representational ability of the adaptive CNN encoder. Experiments on PDTB 2.0 and CDTB datasets show that our approach outperforms all the baselines by a fair margin and achieves excellent results.
Kashif Munir, Hai Zhao 0001, Zuchao Li
IEEE ACM Trans. Audio Speech Lang. Process.3
2020 Semantics-Aware BERT for Language Understanding
abstract
The latest work on language representations carefully integrates contextualized features into language model training, which enables a series of success especially in various machine reading comprehension and natural language inference tasks. However, the existing language representation models including ELMo, GPT and BERT only exploit plain context-sensitive features such as character or word embeddings. They rarely consider incorporating structured semantic information which can provide rich semantics for language representation. To promote natural language understanding, we propose to incorporate explicit contextual semantics from pre-trained semantic role labeling, and introduce an improved language representation model, Semantics-aware BERT (SemBERT), which is capable of explicitly absorbing contextual semantics over a BERT backbone. SemBERT keeps the convenient usability of its BERT precursor in a light fine-tuning way without substantial task-specific modifications. Compared with BERT, semantics-aware BERT is as simple in concept but more powerful. It obtains new state-of-the-art or substantially improves results on ten reading comprehension and language inference tasks.
Zhuosheng Zhang 0001, Yuwei Wu 0003, Hai Zhao 0001, Zuchao Li, Shuailiang Zhang, Xiang Zhou 0007
AAAI4
2020 Explicit Sentence Compression for Neural Machine Translation
abstract
State-of-the-art Transformer-based neural machine translation (NMT) systems still follow a standard encoder-decoder framework, in which source sentence representation can be well done by an encoder with self-attention mechanism. Though Transformer-based encoder may effectively capture general information in its resulting source sentence representation, the backbone information, which stands for the gist of a sentence, is not specifically focused on. In this paper, we propose an explicit sentence compression method to enhance the source sentence representation for NMT. In practice, an explicit sentence compression goal used to learn the backbone information in a sentence. We propose three ways, including backbone source-side fusion, target-side fusion, and both-side fusion, to integrate the compressed sentence into NMT. Our empirical tests on the WMT English-to-French and English-to-German translation tasks show that the proposed sentence compression method significantly improves the translation performances over strong baselines.
Zuchao Li, Rui Wang 0015, Kehai Chen, Masao Utiyama, Eiichiro Sumita, Zhuosheng Zhang 0001, Hai Zhao 0001
AAAI1
2020 Global Greedy Dependency Parsing
abstract
Most syntactic dependency parsing models may fall into one of two categories: transition- and graph-based models. The former models enjoy high inference efficiency with linear time complexity, but they rely on the stacking or re-ranking of partially-built parse trees to build a complete parse tree and are stuck with slower training for the necessity of dynamic oracle training. The latter, graph-based models, may boast better performance but are unfortunately marred by polynomial time inference. In this paper, we propose a novel parsing order objective, resulting in a novel dependency parsing model capable of both global (in sentence scope) feature extraction as in graph models and linear time inference as in transitional models. The proposed global greedy parser only uses two arc-building actions, left and right arcs, for projective parsing. When equipped with two extra non-projective arc-building actions, the proposed parser may also smoothly support non-projective parsing. Using multiple benchmark treebanks, including the Penn Treebank (PTB), the CoNLL-X treebanks, and the Universal Dependency Treebanks, we evaluate our parser and demonstrate that the proposed novel parser achieves good performance with faster training and decoding.
Zuchao Li, Hai Zhao 0001, Kevin Parnow
AAAI1
2020 Neural Machine Translation with Universal Visual Representation
Zhuosheng Zhang 0001, Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita, Zuchao Li, Hai Zhao 0001
ICLR6
2020 Data-dependent Gaussian Prior Objective for Language Generation
Zuchao Li, Rui Wang 0015, Kehai Chen, Masao Utiyama, Eiichiro Sumita, Zhuosheng Zhang 0001, Hai Zhao 0001
ICLR1
2020 Memory Network for Linguistic Structure Parsing
abstract
Memory-based learning can be characterized as a lazy learning method in machine learning terminology because it delays the processing of input by storing the input until needed. Linguistic structure parsing, which has been in a performance improvement bottleneck since the latest series of works was presented, determines the syntactic or semantic structure of a sentence. In this article, we construct a memory component and use it to augment a linguistic structure parser which allows the parser to directly extract patterns from the known training treebank to form memory. The experimental results show that existing state-of-the-art parsers reach new heights of performance on the main benchmarks for dependency parsing and semantic role labeling with this memory network.
Zuchao Li, Chaoyu Guan, Hai Zhao 0001, Rui Wang 0015, Kevin Parnow, Zhuosheng Zhang 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2019 Dependency or Span, End-to-End Uniform Semantic Role Labeling
abstract
Semantic role labeling (SRL) aims to discover the predicateargument structure of a sentence. End-to-end SRL without syntactic input has received great attention. However, most of them focus on either span-based or dependency-based semantic representation form and only show specific model optimization respectively. Meanwhile, handling these two SRL tasks uniformly was less successful. This paper presents an end-to-end model for both dependency and span SRL with a unified argument representation to deal with two different types of argument annotations in a uniform fashion. Furthermore, we jointly predict all predicates and arguments, especially including long-term ignored predicate identification subtask. Our single model achieves new state-of-the-art results on both span (CoNLL 2005, 2012) and dependency (CoNLL 2008, 2009) SRL benchmarks.
Zuchao Li, Shexia He, Hai Zhao 0001, Yiqing Zhang 0002, Zhuosheng Zhang 0001, Xiang Zhou 0007
AAAI1
2019 Syntax-aware Multilingual Semantic Role Labeling
abstract
Shexia He, Zuchao Li, Hai Zhao. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Shexia He, Zuchao Li, Hai Zhao 0001
EMNLP/IJCNLP (1)2
2019 Cross-Domain Transfer Learning for Dependency Parsing
Zuchao Li, Junru Zhou, Hai Zhao 0001, Rui Wang 0015
NLPCC (2)1
2019 Effective Representation for Easy-First Dependency Parsing
Zuchao Li, Jiaxun Cai, Hai Zhao 0001
PRICAI (1)1
2019 Effective Subword Segmentation for Text Comprehension
abstract
Representation learning is the foundation of machine reading comprehension and inference. In state-of-the-art models, character-level representations have been broadly adopted to alleviate the problem of effectively representing rare or complex words. However, character itself is not a natural minimal linguistic unit for representation or word embedding composing due to ignoring the linguistic coherence of consecutive characters inside word. This paper presents a general subword-augmented embedding framework for learning and composing computationally derived subword-level representations. We survey a series of unsupervised segmentation methods for subword acquisition and different subword-augmented strategies for text understanding, showing that subword-augmented embedding significantly improves our baselines in various types of text understanding tasks on both English and Chinese benchmarks.
Zhuosheng Zhang 0001, Hai Zhao 0001, Kangwei Ling, Jiangtong Li, Zuchao Li, Shexia He, Guohong Fu
IEEE ACM Trans. Audio Speech Lang. Process.5
2018 Syntax for Semantic Role Labeling, To Be, Or Not To Be
abstract
Semantic role labeling (SRL) is dedicated to recognizing the predicate-argument structure of a sentence.Previous studies have shown syntactic information has a remarkable contribution to SRL performance.However, such perception was challenged by a few recent neural SRL models which give impressive performance without a syntactic backbone.This paper intends to quantify the importance of syntactic information to dependency SRL in deep learning framework.We propose an enhanced argument labeling model companying with an extended korder argument pruning algorithm for effectively exploiting syntactic information.Our model achieves state-of-the-art results on the CoNLL-2008, 2009 benchmarks for both English and Chinese, showing the quantitative significance of syntax to neural SRL together with a thorough empirical survey over existing models.
Shexia He, Zuchao Li, Hai Zhao 0001, Hongxiao Bai
ACL (1)2
2018 A Full End-to-End Semantic Role Labeler, Syntactic-agnostic Over Syntactic-aware?
abstract
Semantic role labeling (SRL) is to recognize the predicate-argument structure of a sentence, including subtasks of predicate disambiguation and argument labeling. Previous studies usually formulate the entire SRL problem into two or more subtasks. For the first time, this paper introduces an end-to-end neural model which unifiedly tackles the predicate disambiguation and the argument labeling in one shot. Using a biaffine scorer, our model directly predicts all semantic role labels for all given word pairs in the sentence without relying on any syntactic parse information. Specifically, we augment the BiLSTM encoder with a non-linear transformation to further distinguish the predicate and the argument in a given sentence, and model the semantic role labeling process as a word pair classification task by employing the biaffine attentional mechanism. Though the proposed model is syntax-agnostic with local decoder, it outperforms the state-of-the-art syntax-aware SRL systems on the CoNLL-2008, 2009 benchmarks for both English and Chinese. To our best knowledge, we report the first syntax-agnostic SRL model that surpasses all known syntax-aware models.
Jiaxun Cai, Shexia He, Zuchao Li, Hai Zhao 0001
COLING3
2018 Seq2seq Dependency Parsing
abstract
This paper presents a sequence to sequence (seq2seq) dependency parser by directly predicting the relative position of head for each given word, which therefore results in a truly end-to-end seq2seq dependency parser for the first time. Enjoying the advantage of seq2seq modeling, we enrich a series of embedding enhancement, including firstly introduced subword and node2vec augmentation. Meanwhile, we propose a beam search decoder with tree constraint and subroot decomposition over the sequence to furthermore enhance our seq2seq parser. Our parser is evaluated on benchmark treebanks, being on par with the state-of-the-art parsers by achieving 94.11% UAS on PTB and 88.78% UAS on CTB, respectively.
Zuchao Li, Jiaxun Cai, Shexia He, Hai Zhao 0001
COLING1
2018 A Unified Syntax-aware Framework for Semantic Role Labeling
abstract
Semantic role labeling (SRL) aims to recognize the predicate-argument structure of a sentence.Syntactic information has been paid a great attention over the role of enhancing SRL.However, the latest advance shows that syntax would not be so important for SRL with the emerging much smaller gap between syntax-aware and syntax-agnostic SRL.To comprehensively explore the role of syntax for SRL task, we extend existing models and propose a unified framework to investigate more effective and more diverse ways of incorporating syntax into sequential neural networks.Exploring the effect of syntactic input quality on SRL performance, we confirm that high-quality syntactic parse could still effectively enhance syntactically-driven SRL.Using empirically optimized integration strategy, we even enlarge the gap between syntax-aware and syntax-agnostic SRL.Our framework achieves state-of-the-art results on CoNLL-2009 benchmarks both for English and Chinese, substantially outperforming all previous models.
Zuchao Li, Shexia He, Jiaxun Cai, Zhuosheng Zhang 0001, Hai Zhao 0001, Gongshen Liu, Linlin Li 0001, Luo Si
EMNLP1
2016 High precision gesture sensing via quantitative characterization of the Doppler effect
abstract
This paper presents a high precision gesture recognition system that leverages the Doppler effect of ultrasound to sense in-air hand gestures. The system can precisely identify a wider variety of gestures than other systems without any modification to consumer laptops. The system recognizes quantitatively detailed and complex movements from the signals reflected by a moving body. A Hidden Markov Model is used to construct a library of independent, discrete gestures. The gestures can be mapped to diverse application actions. Our method can distinguish among similar gestures with slight difference by extracting fewer, more effective features. Our proposed system reduces false positives caused by unintended motions and is versatile and adaptable to multiple device. We implemented a proof-of-concept prototype on a laptop and extensively evaluated the system. Our results show that the system recognizes six gestures with an average accuracy of 98.6% and 18 gestures including similar ones with 95% accuracy. The flexibility and robustness on multiple devices highlights its ability to enable future ubiquitous non-contact gesture-based interaction with computing devices.
Haojun Ai, Yifang Men, Liangliang Han, Zuchao Li
ICPR4