Zhouhan Lin

dblp:121/7919 · DBLP profile ↗
← Back
44ranked-venue papers
1as first author
32since 2021 · last 2025
0009-0009-7204-0689ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 33 · 1 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 9 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4
YearPublicationVenuePosition
2025 Gumbel Reranking: Differentiable End-to-End Reranker Optimization
abstract
RAG systems rely on rerankers to identify relevant documents. However, fine-tuning these models remains challenging due to the scarcity of annotated query-document pairs. Existing distillation-based approaches suffer from training-inference misalignment and fail to capture interdependencies among candidate documents. To overcome these limitations, we reframe the reranking process as an attention-mask problem and propose Gumbel Reranking, an end-to-end training framework for rerankers aimed at minimizing the training-inference gap. In our approach, reranker optimization is reformulated as learning a stochastic, document-wise Top-k attention mask using the Gumbel Trick and Relaxed Top-k Sampling. This formulation enables end-to-end optimization by minimizing the overall language loss. Experiments across various settings consistently demonstrate performance gains, including a 10.4% improvement in recall on HotpotQA for distinguishing indirectly relevant documents.
Siyuan Huang 0003, Jintao Du, Changhua Meng, Weiqiang Wang 0002, Jingwen Leng, Minyi Guo, Zhouhan Lin
ACL (1)8
2025 Training LLMs to be Better Text Embedders through Bidirectional Reconstruction
abstract
Large language models (LLMs) have increasingly been explored as powerful text embedders.Existing LLM-based text embedding approaches often leverage the embedding of the final token, typically a reserved special token such as [EOS].However, these tokens have not been intentionally trained to capture the semantics of the whole context, limiting their capacity as text embeddings, especially for retrieval and re-ranking tasks.We propose to add a new training stage before contrastive learning to enrich the semantics of the final token embedding.This stage employs bidirectional generative reconstruction tasks, namely EBQ2D (Embedding-Based Query-to-Document) and EBD2Q (Embedding-Based Document-to-Query), which interleave to anchor the [EOS] embedding and reconstruct either side of Query-Document pairs.Experimental results demonstrate that our additional training stage significantly improves LLM performance on the Massive Text Embedding Benchmark (MTEB), achieving new state-ofthe-art results across different LLM base models and scales. 1
Dengliang Shi, Siyuan Huang 0003, Jintao Du, Changhua Meng, Yu Cheng 0005, Weiqiang Wang 0002, Zhouhan Lin
EMNLP8
2025 Efficient Long Document Ranking via Adaptive Token Pruning with Query-Document Alignment
abstract
Transformer-based models have achieved great success in document ranking, yet they suffer from substantial computational costs due to the quadratic complexity of attention, particularly for Long Document Ranking (LDR). Token pruning is a promising approach to reducing computational costs, while existing methods have largely overlooked the interaction and alignment between document and query for guiding the token pruning process, which may result in mistakenly pruned tokens that are critical for the query but not for the document. Additionally, these methods often lack the flexibility required for varying input samples. To this end, we propose a novel framework, Adaptive Token Pruning with Query-Document Alignment (QD-ATP) for accelerating LDR. Specifically, we first introduce a well-designed Query-Document Alignment Guidance (QDAG) module to effectively align the semantic and matching information between query and document, ensuring that the pruned tokens are less important for both two fields. Furthermore, we design a novel Adaptive Token Pruning (ATP) module, which can dynamically adjust pruning ratios based on different input samples. Experimental results on three benchmark datasets demonstrate that QD-ATP can achieve up to a 7.3× latency speedup while preserving competitive performance.
Beiya Dai, Meilin Chen, Xinbing Wang, Chenghu Zhou, Zhouhan Lin
ICASSP6
2025 Training-free LLM-generated Text Detection by Mining Token Probability Sequences
abstract
Large language models (LLMs) have demonstrated remarkable capabilities in generating high-quality texts across diverse domains. However, the potential misuse of LLMs has raised significant concerns, underscoring the urgent need for reliable detection of LLM-generated texts. Conventional training-based detectors often struggle with generalization, particularly in cross-domain and cross-model scenarios. In contrast, training-free methods, which focus on inherent discrepancies through carefully designed statistical features, offer improved generalization and interpretability. Despite this, existing training-free detection methods typically rely on global text sequence statistics, neglecting the modeling of local discriminative features, thereby limiting their detection efficacy. In this work, we introduce a novel training-free detector, termed \textbf{Lastde}\footnote{The code and data are released at \url{https://github.com/TrustMedia-zju/Lastde_Detector}.} that synergizes local and global statistics for enhanced detection. For the first time, we introduce time series analysis to LLM-generated text detection, capturing the temporal dynamics of token probability sequences. By integrating these local statistics with global ones, our detector reveals significant disparities between human and LLM-generated texts. We also propose an efficient alternative, \textbf{Lastde++} to enable real-time detection. Extensive experiments on six datasets involving cross-domain, cross-model, and cross-lingual detection scenarios, under both white-box and black-box settings, demonstrated that our method consistently achieves state-of-the-art performance. Furthermore, our approach exhibits greater robustness against paraphrasing attacks compared to existing baseline methods.
Yihuai Xu, Yifei Bi, Huangsen Cao, Zhouhan Lin, Fei Wu 0001
ICLR5
2025 AdaptiveStep: Automatically Dividing Reasoning Step through Model Confidence
abstract
Current approaches for training Process Reward Models (PRMs) often involve deconposing responses into multiple reasoning steps using rule-based techniques, such as using predefined placeholder tokens or setting the reasoning step’s length to a fixed size. These approaches overlook the fact that certain words don’t usually indicate true decision points. To address this, we propose AdaptiveStep, a method that divides reasoning steps based on the model’s confidence in predicting the next word, offering more information on decision-making at each step, improving downstream tasks like reward model training. Moreover, our method requires no manual annotation. Experiments with AdaptiveStep-trained PRMs in mathematical reasoning and code generation show that the outcome PRM achieves state-of-the-art Best-of-N performance, surpassing greedy search strategy with token-level value-guided decoding, while also reducing construction costs by over 30% compared to existing open-source PRMs. We also provide a thorough analysis and case study on its performance, transferability, and generalization capabilities. We provide our code on https://github.com/Lux0926/ASPRM.
Chaofeng Qu, Zhaoling Chen, Zefan Cai, Jason Klein Liu, Chonghan Liu, Yunhui Xia, Li Zhao 0007, Jiang Bian 0002, Chuheng Zhang, Wei Shen 0005, Zhouhan Lin
ICML13
2025 How to Synthesize Text Data without Model Collapse?
abstract
Model collapse in synthetic data indicates that iterative training on self-generated data leads to a gradual decline in performance. With the proliferation of AI models, synthetic data will fundamentally reshape the web data ecosystem. Future GPT-$\{n\}$ models will inevitably be trained on a blend of synthetic and human-produced data. In this paper, we focus on two questions: what is the impact of synthetic data on language model training, and how to synthesize data without model collapse? We first pre-train language models across different proportions of synthetic data, revealing a negative correlation between the proportion of synthetic data and model performance. We further conduct statistical analysis on synthetic data to uncover distributional shift phenomenon and over-concentration of n-gram features. Inspired by the above findings, we propose token editing on human-produced data to obtain semi-synthetic data. As a proof of concept, we theoretically demonstrate that token-level editing can prevent model collapse, as the test error is constrained by a finite upper bound. We conduct extensive experiments on pre-training from scratch, continual pre-training, and supervised fine-tuning. The results validate our theoretical proof that token-level editing improves data quality and enhances model performance.
Xuekai Zhu, Daixuan Cheng, Hengli Li, Ermo Hua, Xingtai Lv, Ning Ding 0002, Zhouhan Lin, Zilong Zheng, Bowen Zhou 0002
ICML8
2025 Visual-informed Silent Video Identity Conversion
abstract
Conventional voice conversion modifies voice characteristics from a source speaker to a target speaker, relying on audio input from both sides. However, this process becomes infeasible when clean audio is unavailable, such as in silent videos or noisy environments. In this work, we focus on the task of Silent Face-based Voice Conversion (SFVC), which does voice conversion entirely from visual inputs. i.e., given images of a target speaker and a silent video of a source speaker containing lip motion, SFVC generates speech aligning the identity of the target speaker while preserving the speech content in the source silent video. As this task requires generating intelligible speech and converting identity using only visual cues, it is particularly challenging. To address this, we introduce MuteSwap, a novel framework that employs contrastive learning to align cross-modality identities and minimize mutual information to separate shared visual features. Experimental results show that MuteSwap achieves impressive performance in both speech synthesis and identity conversion, especially under noisy conditions where methods dependent on audio input fail to produce intelligible results, demonstrating both the effectiveness of our training approach and the feasibility of SFVC. Demo page is available at https://pussycat0700.github.io/MuteSwap-Demo/.
Yu Fang 0008, Zhouhan Lin
ACM Multimedia3
2025 Memory Decoder: A Pretrained, Plug-and-Play Memory for Large Language Models
abstract
Large Language Models (LLMs) have shown strong abilities in general language tasks, yet adapting them to specific domains remains a challenge. Current method like Domain Adaptive Pretraining (DAPT) requires costly full-parameter training and suffers from catastrophic forgetting. Meanwhile, Retrieval-Augmented Generation (RAG) introduces substantial inference latency due to expensive nearest-neighbor searches and longer context. This paper introduces \textit{Memory Decoder}, a plug-and-play pretrained memory that enables efficient domain adaptation without changing the original model's parameters. Memory Decoder employs a small transformer decoder that learns to imitate the behavior of an external non-parametric retriever. Once trained, Memory Decoder can be seamlessly integrated with any pretrained language model that shares the same tokenizer, requiring no model-specific modifications. Experimental results demonstrate that Memory Decoder enables effective adaptation of various Qwen and Llama models to three distinct specialized domains: biomedicine, finance, and law, reducing perplexity by an average of 6.17 points. Overall, Memory Decoder introduces a novel paradigm centered on a specially pretrained memory component designed for domain-specific adaptation. This memory architecture can be integrated in a plug-and-play manner, consistently enhancing performance across multiple models within the target domain.
Jiaqi Cao 0002, Rubin Wei, Qipeng Guo, Kai Chen 0026, Bowen Zhou 0002, Zhouhan Lin
NeurIPS7
2025 Alleviating LLM-based Generative Retrieval Hallucination in Alipay Search
abstract
Generative retrieval (GR) has revolutionized document retrieval with the advent of large language models (LLMs), and LLM-based GR is gradually being adopted by the industry. Despite its remarkable advantages and potential, LLM-based GR suffers from hallucination and generates documents that are irrelevant to the query in some instances, severely challenging its credibility in practical applications. We thereby propose an optimized GR framework designed to alleviate retrieval hallucination, which integrates knowledge distillation reasoning in model training and incorporate decision agent to further improve retrieval precision. Specifically, we employ LLMs to assess and reason GR retrieved query-document (q-d) pairs, and then distill the reasoning data as transferred knowledge to the GR model. Moreover, we utilize a decision agent as post-processing to extend the GR retrieved documents through retrieval model and select the most relevant ones from multi perspectives as the final generative retrieval result. Extensive offline experiments on real-world datasets and online A/B tests on Fund Search and Insurance Search in Alipay demonstrate our framework's superiority and effectiveness in improving search quality and conversion gains.
Yedan Shen, Kaixin Wu, Yuechen Ding, Jingyuan Wen, Mingjie Zhong, Zhouhan Lin, Jia Xu 0013, Linjian Mo
SIGIR7
2024 Breaking the Bottleneck on Graphs with Structured State Spaces
abstract
The majority of GNNs are based on message-passing mechanisms. However, Message Passing Neural Networks (MPNNs) have inherent limitations in capturing long-range interactions. The exponentially growing node information is compressed into fixed-size representations through multiple rounds of message passing, leading to the over-squashing problem. This issue severely hinders the flow of information across the graph and creates a bottleneck in graph learning. The natural idea of introducing global attention to point-to-point communication, as adopted in Graph Transformers (GTs), lacks inductive biases on graph structures and relies on complex positional encodings to enhance their performance in practical tasks. In this paper, we observe that the sensitivity between nodes in MPNNs decreases exponentially with the shortest path distance. In contrast, GTs have constant sensitivity, which leads to a loss of inductive bias. To address these issues, we introduce structured state spaces to capture the hierarchy of rooted trees, achieving linear sensitivity with theoretical guarantees. We further propose a novel state-space model-based graph convolution, resulting in a new paradigm that retains both the strong inductive biases from MPNNs and the long-range modeling capabilities from GTs. Extensive experimental results on long-range and general graph benchmarks demonstrate the superiority of our approach.
Yunchong Song, Siyuan Huang 0003, Jiacheng Cai, Xinbing Wang, Chenghu Zhou, Zhouhan Lin
CIKM6
2024 Enhancing SPARQL Generation by Triplet-order-sensitive Pre-training
abstract
Semantic parsing that translates natural language queries to SPARQL is of great importance for Knowledge Graph Question Answering (KGQA) systems. Although pre-trained language models like T5 have achieved significant success in the Text-to-SPARQL task, their generated outputs still exhibit notable errors specific to the SPARQL language, such as triplet flips. To address this challenge and further improve the performance, we propose an additional pre-training stage with a new objective, Triplet Order Correction (TOC), along with the commonly used Masked Language Modeling (MLM), to collectively enhance the model's sensitivity to triplet order and SPARQL syntax. We also propose to verbalize the Internationalized Resource Identifiers (IRIs) during training. Our method achieves state-of-the-art performances on three widely-used benchmarks.
Jiexing Qi, Zhouhan Lin
CIKM5
2024 Extracting Financial Events from Raw Texts via Matrix Chunking
abstract
Event Extraction (EE) is widely used in the Chinese financial field to provide valuable structured information. However, there are two key challenges for Chinese financial EE in application scenarios. First, events need to be extracted from raw texts, which sets it apart from previous works like the Automatic Content Extraction (ACE) EE task, where EE is treated as a classification problem given the entity spans. Second, recognizing financial entities can be laborious, as they may involve multiple elements. In this paper, we introduce CFTE, a novel task for Chinese Financial Text-to-Event extraction, which directly extracts financial events from raw texts. We further present FINEED, a Chinese FINancial Event Extraction Dataset, and an efficient MAtrix-ChunKing method called MACK, designed for the extraction of financial events from raw texts. Specifically, FINEED is manually annotated with rich linguistic features. We propose a novel two-dimensional annotation method for FINEED, which can visualize the interactions among text components. Our MACK method is fault-tolerant by preserving the tag frequency distribution when identifying financial entities. We conduct extensive experiments and the results verify the effectiveness of our MACK method.
Kunping Li, Zhouhan Lin
LREC/COLING5
2024 Towards Controlled Table-to-Text Generation with Scientific Reasoning
abstract
The sheer volume of scientific experimental results and complex technical statements, often presented in tabular formats, presents a formidable barrier to individuals acquiring preferred information. The realms of scientific reasoning and content generation that adhere to user preferences encounter distinct challenges. In this work, we present a new task for generating fluent and logical descriptions that match user preferences over scientific tabular data, aiming to automate scientific document analysis. To facilitate research in this direction, we construct a new challenging dataset CTRLSciTab consisting of table-description pairs extracted from the scientific literature, with highlighted cells and corresponding domain-specific knowledge base. We evaluated popular pre-trained language models to establish a baseline and proposed a novel architecture outperforming competing approaches. The results showed that large models struggle to produce accurate content that aligns with user preferences. As the first of its kind, our work should motivate further research in scientific domains.1
Zhixin Guo, Jianping Zhou 0004, Jiexing Qi, Mingxuan Yan, Ziwei He, Guanjie Zheng, Zhouhan Lin, Xinbing Wang, Chenghu Zhou
ICASSP7
2024 Graph Parsing Networks
abstract
Graph pooling compresses graph information into a compact representation. State-of-the-art graph pooling methods follow a hierarchical approach, which reduces the graph size step-by-step. These methods must balance memory efficiency with preserving node information, depending on whether they use node dropping or node clustering. Additionally, fixed pooling ratios or numbers of pooling layers are predefined for all graphs, which prevents personalized pooling structures from being captured for each individual graph. In this work, inspired by bottom-up grammar induction, we propose an efficient graph parsing algorithm to infer the pooling structure, which then drives graph pooling. The resulting Graph Parsing Network (GPN) adaptively learns personalized pooling structure for each individual graph. GPN benefits from the discrete assignments generated by the graph parsing algorithm, allowing good memory efficiency while preserving node information intact. Experimental results on standard benchmarks demonstrate that GPN outperforms state-of-the-art graph pooling methods in graph classification tasks while being able to achieve competitive performance in node classification tasks. We also conduct a graph reconstruction task to show GPN's ability to preserve node information and measure both memory and time efficiency through relevant tests.
Yunchong Song, Siyuan Huang 0003, Xinbing Wang, Chenghu Zhou, Zhouhan Lin
ICLR5
2024 SWave: Improving Vocoder Efficiency by Straightening the Waveform Generation Path
Jianping Zhou 0004, Xiaohua Tian, Zhouhan Lin
ICPR (6)4
2024 PaD: Program-aided Distillation Can Teach Small Models Reasoning Better than Chain-of-thought Fine-tuning
abstract
Xuekai Zhu, Biqing Qi, Kaiyan Zhang, Xinwei Long, Zhouhan Lin, Bowen Zhou. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Xuekai Zhu, Biqing Qi, Xinwei Long, Zhouhan Lin, Bowen Zhou 0002
NAACL-HLT5
2024 Cluster-wise Graph Transformer with Dual-granularity Kernelized Attention
abstract
In the realm of graph learning, there is a category of methods that conceptualize graphs as hierarchical structures, utilizing node clustering to capture broader structural information. While generally effective, these methods often rely on a fixed graph coarsening routine, leading to overly homogeneous cluster representations and loss of node-level information. In this paper, we envision the graph as a network of interconnected node sets without compressing each cluster into a single embedding. To enable effective information transfer among these node sets, we propose the Node-to-Cluster Attention (N2C-Attn) mechanism. N2C-Attn incorporates techniques from Multiple Kernel Learning into the kernelized attention framework, effectively capturing information at both node and cluster levels. We then devise an efficient form for N2C-Attn using the cluster-wise message-passing framework, achieving linear time complexity. We further analyze how N2C-Attn combines bi-level feature maps of queries and keys, demonstrating its capability to merge dual-granularity information. The resulting architecture, Cluster-wise Graph Transformer (Cluster-GT), which uses node clusters as tokens and employs our proposed N2C-Attn module, shows superior performance on various graph-level tasks. Code is available at https://github.com/LUMIA-Group/Cluster-wise-Graph-Transformer.
Siyuan Huang 0003, Yunchong Song, Jiayue Zhou, Zhouhan Lin
NeurIPS4
2024 HuRef: HUman-REadable Fingerprint for Large Language Models
abstract
Protecting the copyright of large language models (LLMs) has become crucial due to their resource-intensive training and accompanying carefully designed licenses. However, identifying the original base model of an LLM is challenging due to potential parameter alterations. In this study, we introduce HuRef, a human-readable fingerprint for LLMs that uniquely identifies the base model without interfering with training or exposing model parameters to the public. We first observe that the vector direction of LLM parameters remains stable after the model has converged during pretraining, with negligible perturbations through subsequent training steps, including continued pretraining, supervised fine-tuning, and RLHF, which makes it a sufficient condition to identify the base model. The necessity is validated by continuing to train an LLM with an extra term to drive away the model parameters' direction and the model becomes damaged. However, this direction is vulnerable to simple attacks like dimension permutation or matrix rotation, which significantly change it without affecting performance. To address this, leveraging the Transformer structure, we systematically analyze potential attacks and define three invariant terms that identify an LLM's base model. Due to the potential risk of information leakage, we cannot publish invariant terms directly. Instead, we map them to a Gaussian vector using an encoder, then convert it into a natural image using StyleGAN2, and finally publish the image. In our black-box setting, all fingerprinting steps are internally conducted by the LLMs owners. To ensure the published fingerprints are honestly generated, we introduced Zero-Knowledge Proof (ZKP). Experimental results across various LLMs demonstrate the effectiveness of our method. The code is available at https://github.com/LUMIA-Group/HuRef.
Boyi Zeng, Yuncong Hu, Yi Xu 0004, Chenghu Zhou, Xinbing Wang, Zhouhan Lin
NeurIPS8
2024 K2: A Foundation Language Model for Geoscience Knowledge Understanding and Utilization
abstract
Large language models (LLMs) have achieved great success in general domains of natural language processing. In this paper, we bring LLMs to the realm of geoscience with the objective of advancing research and applications in this field. To this end, we present the first-ever LLM in geoscience, K2, alongside a suite of resources developed to further promote LLM research within geoscience. For instance, we have curated the first geoscience instruction tuning dataset, GeoSignal, which aims to align LLM responses to geoscience-related user queries. Additionally, we have established the first geoscience benchmark, GeoBench, to evaluate LLMs in the context of geoscience. In this work, we experiment with a complete recipe to adapt a pre-trained general-domain LLM to the geoscience domain. Specifically, we further train the LLaMA-7B model on 5.5B tokens of geoscience text corpus, including over 1 million pieces of geoscience literature, and utilize GeoSignal's supervised data to fine-tune the model. Moreover, we share a protocol that can efficiently gather domain-specific data and construct domain-supervised data, even in situations where manpower is scarce. Meanwhile, we equip K2 with the abilities of using tools to be a naive geoscience aide. Experiments conducted on the GeoBench demonstrate the effectiveness of our approach and datasets on geoscience knowledge understanding and utilization.We open-source all the training data and K2 model checkpoints at https://github.com/davendw49/k2
Cheng Deng 0001, Tianhang Zhang, Zhongmou He, Qiyuan Chen 0002, Yi Xu 0004, Luoyi Fu, Weinan Zhang 0001, Xinbing Wang, Chenghu Zhou, Zhouhan Lin, Junxian He
WSDM11
2023 Unsupervised Graph-Text Mutual Conversion with a Unified Pretrained Language Model
abstract
Yi Xu, Shuqian Sheng, Jiexing Qi, Luoyi Fu, Zhouhan Lin, Xinbing Wang, Chenghu Zhou. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Yi Xu 0004, Shuqian Sheng, Jiexing Qi, Luoyi Fu, Zhouhan Lin, Xinbing Wang, Chenghu Zhou
ACL (1)5
2023 Asymmetric Polynomial Loss for Multi-Label Classification
abstract
Various tasks are reformulated as multi-label classification problems, in which the binary cross-entropy (BCE) loss is frequently utilized for optimizing well-designed models. However, the vanilla BCE loss cannot be tailored for diverse tasks, resulting in a suboptimal performance for different models. Besides, the imbalance between redundant negative samples and rare positive samples could degrade the model performance. In this paper, we propose an effective Asymmetric Polynomial Loss (APL) to mitigate the above issues. Specifically, we first perform Taylor expansion on BCE loss. Then we ameliorate the coefficients of polynomial functions. We further employ the asymmetric focusing mechanism to decouple the gradient contribution from the negative and positive samples. Moreover, we validate that the polynomial coefficients can recalibrate the asymmetric focusing hyperparameters. Experiments on relation extraction, text classification, and image classification show that our APL loss can consistently improve performance without extra training burden.1
Jiexing Qi, Xinbing Wang, Zhouhan Lin
ICASSP4
2023 Text Classification In The Wild: A Large-Scale Long-Tailed Name Normalization Dataset
abstract
Real-world data usually exhibits a long-tailed distribution, with a few frequent labels and a lot of few-shot labels. The study of institution name normalization is a perfect application case showing this phenomenon: there are many institutions worldwide, with enormous variations of their names in the publicly available literature. In this work, we first collect a large-scale institution name normalization dataset containing over 25k classes whose frequencies are naturally long-tail distributed. We construct our test set from four different subsets: many-, medium-, and few-shot sets, as well as a zero-shot open set, which are meant to isolate the few-shot and zero-shot learning scenarios from the massive many-shot classes. We also replicate several important benchmarks on our data, covering a wide range from search-based methods to neural network methods. Further, we propose our specially pretrained, BERT-based model that shows better out-of-distribution generalization on few-shot and zero-shot test sets. Compared to other datasets focusing on the long-tailed phenomenon, our dataset has one order of magnitude more training data than the largest existing long-tailed datasets and is naturally long-tailed rather than manually synthesized. We believe it provides an important and different scenario to study this problem. To our best knowledge, this is the first natural language dataset that focuses on this long-tailed and open-set classification problem.1
Jiexing Qi, Zhixin Guo, Chenghu Zhou, Weinan Zhang 0001, Xinbing Wang, Zhouhan Lin
ICASSP8
2023 Ordered GNN: Ordering Message Passing to Deal with Heterophily and Over-smoothing
Yunchong Song, Chenghu Zhou, Xinbing Wang, Zhouhan Lin
ICLR4
2023 Document-Level Relation Extraction with Relation Correlation Enhancement
Zhouhan Lin
ICONIP (13)2
2023 I2SRM: Intra- and Inter-Sample Relationship Modeling for Multimodal Information Extraction
abstract
Multimodal information extraction is attracting research attention nowadays, which requires aggregating representations from different modalities. In this paper, we present the Intra- and Inter-Sample Relationship Modeling (I2SRM) method for this task, which contains two modules. Firstly, the intra-sample relationship modeling module operates on a single sample and aims to learn effective representations. Embeddings from textual and visual modalities are shifted to bridge the modality gap caused by distinct pre-trained language and image models. Secondly, the inter-sample relationship modeling module considers relationships among multiple samples and focuses on capturing the interactions. An AttnMixup strategy is proposed, which not only enables collaboration among samples but also augments data to improve generalization. We conduct extensive experiments on the multimodal named entity recognition datasets Twitter-2015 and Twitter-2017, and the multimodal relation extraction dataset MNRE. Our proposed method I2SRM achieves competitive results, 77.12% F1-score on Twitter-2015, 88.40% F1-score on Twitter-2017, and 84.12% F1-score on MNRE. 1
Zhouhan Lin
MMAsia2
2023 Tailoring Self-Attention for Graph via Rooted Subtrees
abstract
Attention mechanisms have made significant strides in graph learning, yet they still exhibit notable limitations: local attention faces challenges in capturing long-range information due to the inherent problems of the message-passing scheme, while global attention cannot reflect the hierarchical neighborhood structure and fails to capture fine-grained local information. In this paper, we propose a novel multi-hop graph attention mechanism, named Subtree Attention (STA), to address the aforementioned issues. STA seamlessly bridges the fully-attentional structure and the rooted subtree, with theoretical proof that STA approximates the global attention under extreme settings. By allowing direct computation of attention weights among multi-hop neighbors, STA mitigates the inherent problems in existing graph attention mechanisms. Further we devise an efficient form for STA by employing kernelized softmax, which yields a linear time complexity. Our resulting GNN architecture, the STAGNN, presents a simple yet performant STA-based graph neural network leveraging a hop-aware attention strategy. Comprehensive evaluations on ten node classification datasets demonstrate that STA-based models outperform existing graph transformers and mainstream GNNs. The code is available at https://github.com/LUMIA-Group/SubTree-Attention.
Siyuan Huang 0003, Yunchong Song, Jiayue Zhou, Zhouhan Lin
NeurIPS4
2022 Block-Skim: Efficient Question Answering for Transformer
abstract
Transformer models have achieved promising results on natural language processing (NLP) tasks including extractive question answering (QA). Common Transformer encoders used in NLP tasks process the hidden states of all input tokens in the context paragraph throughout all layers. However, different from other tasks such as sequence classification, answering the raised question does not necessarily need all the tokens in the context paragraph. Following this motivation, we propose Block-skim, which learns to skim unnecessary context in higher hidden layers to improve and accelerate the Transformer performance. The key idea of Block-Skim is to identify the context that must be further processed and those that could be safely discarded early on during inference. Critically, we find that such information could be sufficiently derived from the self-attention weights inside the Transformer model. We further prune the hidden states corresponding to the unnecessary positions early in lower layers, achieving significant inference-time speedup. To our surprise, we observe that models pruned in this way outperform their full-size counterparts. Block-Skim improves QA models' accuracy on different datasets and achieves 3 times speedup on BERT-base model.
Yue Guan 0003, Zhengyi Li 0002, Zhouhan Lin, Yuhao Zhu 0001, Jingwen Leng, Minyi Guo
AAAI3
2022 Transkimmer: Transformer Learns to Layer-wise Skim
abstract
Transformer architecture has become the defacto model for many machine learning tasks from natural language processing and computer vision.As such, improving its computational efficiency becomes paramount.One of the major computational inefficiency of Transformer-based models is that they spend the identical amount of computation throughout all layers.Prior works have proposed to augment the Transformer model with the capability of skimming tokens to improve its computational efficiency.However, they suffer from not having effectual and end-to-end optimization of the discrete skimming predictor.To address the above limitations, we propose the Transkimmer architecture, which learns to identify hidden state tokens that are not required by each layer.The skimmed tokens are then forwarded directly to the final output, thus reducing the computation of the successive layers.The key idea in Transkimmer is to add a parameterized predictor before each layer that learns to make the skimming decision.We also propose to adopt reparameterization trick and add skim loss for the end-to-end training of Transkimmer.Transkimmer achieves 10.97× average speedup on GLUE benchmark compared with vanilla BERT base baseline with less than 1% accuracy degradation.
Yue Guan 0003, Zhengyi Li 0002, Jingwen Leng, Zhouhan Lin, Minyi Guo
ACL (1)4
2022 Leveraging Unimodal Self-Supervised Learning for Multimodal Audio-Visual Speech Recognition
abstract
Training Transformer-based models demands a large amount of data, while obtaining aligned and labelled data in multimodality is rather cost-demanding, especially for audio-visual speech recognition (AVSR).Thus it makes a lot of sense to make use of unlabelled unimodal data.On the other side, although the effectiveness of large-scale self-supervised learning is well established in both audio and visual modalities, how to integrate those pretrained models into a multimodal scenario remains underexplored.In this work, we successfully leverage unimodal self-supervised learning to promote the multimodal AVSR.In particular, audio and visual front-ends are trained on large-scale unimodal datasets, then we integrate components of both front-ends into a larger multimodal framework which learns to recognize parallel audio-visual data into characters through a combination of CTC and seq2seq decoding.We show that both components inherited from unimodal selfsupervised learning cooperate well, resulting in that the multimodal framework yields competitive results through fine-tuning.Our model is experimentally validated on both word-level and sentence-level tasks.Especially, even without an external language model, our proposed model raises the state-of-the-art performances on the widely accepted Lip Reading Sentences 2 (LRS2) dataset by a large margin, with a relative improvement of 30%.
Xichen Pan, Yichen Gong, Helong Zhou, Xinbing Wang, Zhouhan Lin
ACL (1)6
2022 RASAT: Integrating Relational Structures into Pretrained Seq2Seq Model for Text-to-SQL
abstract
Jiexing Qi, Jingyao Tang, Ziwei He, Xiangpeng Wan, Yu Cheng, Chenghu Zhou, Xinbing Wang, Quanshi Zhang, Zhouhan Lin. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Jiexing Qi, Ziwei He, Xiangpeng Wan, Yu Cheng 0003, Chenghu Zhou, Xinbing Wang, Quanshi Zhang, Zhouhan Lin
EMNLP9
2021 Incorporating External POS Tagger for Punctuation Restoration
abstract
Punctuation restoration is an important post-processing step in automatic speech recognition. Among other kinds of external information, part-of-speech (POS) taggers provide informative tags, suggesting each input token's syntactic role, which has been shown to be beneficial for the punctuation restoration task. In this work, we incorporate an external POS tagger and fuse its predicted labels into the existing language model to provide syntactic information. Besides, we propose sequence boundary sampling (SBS) to learn punctuation positions more efficiently as a sequence tagging task. Experimental results show that our methods can consistently obtain performance gains and achieve a new state-of-the-art on the common IWSLT benchmark. Further ablation studies illustrate that both large pre-trained language models and the external POS tagger take essential parts to improve the model's performance.
Ning Shi, Boxin Wang, Zhouhan Lin
Interspeech6
2021 Annotation Inconsistency and Entity Bias in MultiWOZ
abstract
MultiWOZ (Budzianowski et al., 2018) is one of the most popular multi-domain taskoriented dialog datasets, containing 10K+ annotated dialogs covering eight domains.It has been widely accepted as a benchmark for various dialog tasks, e.g., dialog state tracking (DST), natural language generation (NLG) and end-to-end (E2E) dialog modeling.In this work, we identify an overlooked issue with dialog state annotation inconsistencies in the dataset, where a slot type is tagged inconsistently across similar dialogs leading to confusion for DST modeling.We propose an automated correction for this issue, which is present in 70% of the dialogs.Additionally, we notice that there is significant entity bias in the dataset (e.g., "cambridge" appears in 50% of the destination cities in the train domain).The entity bias can potentially lead to named entity memorization in generative models, which may go unnoticed as the test set suffers from a similar entity bias as well.We release a new test set with all entities replaced with unseen entities.Finally, we benchmark joint goal accuracy (JGA) of the state-of-theart DST baselines on these modified versions of the data.Our experiments show that the annotation inconsistency corrections lead to 7-10% improvement in JGA.On the other hand, we observe a 29% drop in JGA when models are evaluated on the new test set with unseen entities.* The work of KQ and ZY was done as a research intern and a visiting research scientist at Facebook AI.
Kun Qian 0016, Ahmad Beirami, Zhouhan Lin, Ankita De, Alborz Geramifard, Zhou Yu 0005, Chinnadhurai Sankar
SIGDIAL3
2020 Exploiting Syntactic Structure for Better Language Modeling: A Syntactic Distance Approach
abstract
It is commonly believed that knowledge of syntactic structure should improve language modeling.However, effectively and computationally efficiently incorporating syntactic structure into neural language models has been a challenging topic.In this paper, we make use of a multi-task objective, i.e., the models simultaneously predict words as well as ground truth parse trees in a form called "syntactic distances", where information between these two separate objectives shares the same intermediate representation.Experimental results on the Penn Treebank and Chinese Treebank datasets show that when ground truth parse trees are provided as additional training signals, the model is able to achieve lower perplexity and induce trees with better quality.
Wenyu Du, Zhouhan Lin, Yikang Shen, Timothy J. O'Donnell, Yoshua Bengio, Yue Zhang 0004
ACL2
2019 Interactive Language Learning by Question Answering
abstract
Xingdi Yuan, Marc-Alexandre Côté, Jie Fu, Zhouhan Lin, Chris Pal, Yoshua Bengio, Adam Trischler. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Xingdi Yuan, Marc-Alexandre Côté, Jie Fu 0001, Zhouhan Lin, Christopher Joseph Pal, Yoshua Bengio, Adam Trischler
EMNLP/IJCNLP (1)4
2019 Ordered Memory
abstract
Stack-augmented recurrent neural networks (RNNs) have been of interest to the deep learning community for some time. However, the difficulty of training memory models remains a problem obstructing the widespread use of such models. In this paper, we propose the Ordered Memory architecture. Inspired by Ordered Neurons (Shen et al., 2018), we introduce a new attention-based mechanism and use its cumulative probability to control the writing and erasing operation of the memory. We also introduce a new Gated Recursive Cell to compose lower-level representations into higher-level representation. We demonstrate that our model achieves strong performance on the logical inference task (Bowman et al., 2015) and the ListOps (Nangia and Bowman, 2018) task. We can also interpret the model to retrieve the induced tree structure, and find that these induced structures align with the ground truth. Finally, we evaluate our model on the Stanford Sentiment Treebank tasks (Socher et al., 2013), and find that it performs comparatively with the state-of-the-art methods in the literature.
Yikang Shen, Shawn Tan, Seyed Arian Hosseini, Zhouhan Lin, Alessandro Sordoni, Aaron C. Courville
NeurIPS4
2018 Straight to the Tree: Constituency Parsing with Neural Syntactic Distance
abstract
Yikang Shen, Zhouhan Lin, Athul Paul Jacob, Alessandro Sordoni, Aaron Courville, Yoshua Bengio. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018.
Yikang Shen, Zhouhan Lin, Athul Paul Jacob, Alessandro Sordoni, Aaron C. Courville, Yoshua Bengio
ACL (1)2
2018 Neural Language Modeling by Jointly Learning Syntax and Lexicon
Yikang Shen, Zhouhan Lin, Chin-Wei Huang, Aaron C. Courville
ICLR (Poster)2
2018 Focused Hierarchical RNNs for Conditional Sequence Processing
abstract
Recurrent Neural Networks (RNNs) with attention mechanisms have obtained state-of-the-art results for many sequence processing tasks. Most of these models use a simple form of encoder with attention that looks over the entire sequence and assigns a weight to each token independently. We present a mechanism for focusing RNN encoders for sequence modelling tasks which allows them to attend to key parts of the input as needed. We formulate this using a multi-layer conditional hierarchical sequence encoder that reads in one token at a time and makes a discrete decision on whether the token is relevant to the context or question being asked. The discrete gating mechanism takes in the context embedding and the current hidden state as inputs and controls information flow into the layer above. We train it using policy gradient methods. We evaluate this method on several types of tasks with different attributes. First, we evaluate the method on synthetic tasks which allow us to evaluate the model for its generalization ability and probe the behavior of the gates in more controlled settings. We then evaluate this approach on large scale Question Answering tasks including the challenging MS MARCO and SearchQA tasks. Our models shows consistent improvements for both tasks over prior work and our baselines. It has also shown to generalize significantly better on synthetic tasks as compared to the baselines.
Nan Rosemary Ke, Konrad Zolna, Alessandro Sordoni, Zhouhan Lin, Adam Trischler, Yoshua Bengio, Joelle Pineau, Laurent Charlin, Christopher Joseph Pal
ICML4
2017 A Structured Self-Attentive Sentence Embedding
Zhouhan Lin, Minwei Feng, Cícero Nogueira dos Santos, Mo Yu, Bing Xiang, Bowen Zhou 0002, Yoshua Bengio
ICLR (Poster)1
2016 Architectural Complexity Measures of Recurrent Neural Networks
abstract
In this paper, we systematically analyze the connecting architectures of recurrent neural networks (RNNs). Our main contribution is twofold: first, we present a rigorous graph-theoretic framework describing the connecting architectures of RNNs in general. Second, we propose three architecture complexity measures of RNNs: (a) the recurrent depth, which captures the RNN’s over-time nonlinear complexity, (b) the feedforward depth, which captures the local input-output nonlinearity (similar to the “depth” in feedforward neural networks (FNNs)), and (c) the recurrent skip coefficient which captures how rapidly the information propagates over time. We rigorously prove each measure’s existence and computability. Our experimental results show that RNNs might benefit from larger recurrent depth and feedforward depth. We further demonstrate that increasing recurrent skip coefficient offers performance boosts on long term dependency problems.
Saizheng Zhang, Yuhuai Wu, Tong Che, Zhouhan Lin, Roland Memisevic, Ruslan Salakhutdinov, Yoshua Bengio
NIPS4
2014 Joint Adaboost and multifeature based ensemble for hyperspectral image classification
abstract
The paper presents a novel ensemble system which unites Adaboost with multifeature to increase diversity among individual classifiers. Adaboost gives rise to convenience for hyperspectral data classification. To improve the method further, we propose joint Adaboost and multifeature based ensemble (JAME), which assigns different multifeature sets to individual classifiers in Adaboost. Diverse spectral and spatial feature sets are integrated to form multifeature sets. As a result, compared with Adaboost the method has increased the diversity of ensemble system, and better overall accuracies are present. Experiments on hyperspectral data sets reveal that the proposed JAME obtains sound performances comparing with original Adaboost and single classifier.
Yushi Chen 0002, Xing Zhao 0005, Zhouhan Lin
IGARSS3
2013 Riemannian manifold learning based k-nearest-neighbor for hyperspectral image classification
abstract
The existence of nonlinear characteristics in hyperspectral data is considered as an influential factor curtailing the classification accuracy of canonical linear classifier like k-nearest neighbor (k-NN). To deal with the problem, we investigated approaches to combine manifold learning methods and the k-NN classifier to preserve nonlinear characteristics contained in hyperspectral imagery. Then we proposed a Riemannian manifold learning (RML) based k-NN classifier for hyperspectral image classification, which substitutes the Euclidean distances used in canonical kNN by geodesic distances yielded by RML. The experimental results on AVIRIS data show that in most cases, the RML-kNN Classifier accesses higher classification accuracies than canonical k-NN.
Yushi Chen 0002, Zhouhan Lin, Xing Zhao 0005
IGARSS2
2013 Supervised Locally Linear Embedding based dimension reduction for hyperspectral image classification
abstract
The nonlinear characteristics in hyperspectral data is considered as an influential factor curtailing the classification accuracy. To deal with the problem, a new method for classification is developed, especially for hyperspectral imagery (HSI). It is a supervised method based on Locally Linear Embedding (LLE) and k-Nearest Neighbor (KNN), named with KNN based supervised LLE (S-LLE KNN). We use two real HIS dataset of AVIRIS in experiment section and compare overall classification accuracy and accuracy of each class in different methods, the results shows that the supervised nonlinear feature extraction method contributes more to classification accuracies methods.
Changbo Qu, Zhouhan Lin
IGARSS3
2012 Parallel implementation for SAM algorithm based on GPU and distributed computing
abstract
Advances in sensor and computer technology are revolutionizing the way that remote sensing data with hundreds or even thousands of channels for the same area on the surface of the earth is collected, managed and analyzed. In this paper, the classical Spectral Angle Mapper (SAM) algorithm, which is fit for parallel and distributed computing, is implemented by using Graphic Processing Units (GPU) and distributed cluster respectively to accelerate the computations. A quantitative performance comparison between Compute Unified Device Architecture (CUDA) and Matlab platform is given by analyzing result of different parallel architectures' implementation of the same SAM algorithm.
Haicheng Qu, Junping Zhang, Yushi Chen 0002, Hao Chen 0014, Zhouhan Lin
IGARSS5