VLDB 2026 Research / reviewers in the wild / expert
Zhangyue Yin
dblp:314/5418
· DBLP profile ↗
26ranked-venue papers
5as first author
26since 2021 · last 2026
0009-0008-0620-8271ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 5 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Demystifying GNN-to-MLP Knowledge Transfer: Theoretical Grounding and Dual-Stream Distillation MethodabstractGraph Neural Networks (GNNs) have shown remarkable effectiveness across various applications, but their computational complexity poses significant scalability challenges. To this end, GNN-to-MLP Knowledge Distillation (KD) methods transfer relational inductive biases from GNNs to MLPs, equipping MLPs with graph-aware capabilities that rival or even surpass those of their teacher GNNs. However, a theoretical foundation for understanding GNN-to-MLP KD is still missing. In this paper, we provide a theoretical analysis of how knowledge distillation unlocks the potential of MLPs for graph tasks from the perspective of training dynamics. We demonstrate that label alignment in KD fundamentally reshapes the Neural Tangent Kernel (NTK) matrix of student MLPs, enabling them to learn the teacher model’s implicit graph bias. We further investigate finer-grained distillation paradigms and reveal that conventional layer-wise output alignment fails to effectively align the deep-layer graph propagation outcomes. To address this, we propose Dual-Stream Aligned MLP (DA-MLP), which incorporates complementary graph filters in a dual-stream architecture. This approach simultaneously enhances feature space dimensionality for improved representation alignment and preserves graph signals across different frequency bands. Comprehensive experiments on seven benchmark datasets validate that DA-MLP can be seamlessly integrated into existing knowledge distillation frameworks for performance enhancements in both transductive and inductive settings. Mingkai Lin, Zhangyue Yin, Shijian Xiao, Sanglu Lu |
AAAI | 4 |
| 2026 | OS-Sentinel: Towards Safety-Enhanced Mobile GUI Agents via Hybrid Validation in Realistic WorkflowsabstractQiushi Sun, Mukai Li, Zhoumianze Liu, Zhihui Xie, Fangzhi Xu, Zhangyue Yin, Kanzhi Cheng, Zehao Li, Zichen Ding, Qi Liu, Zhiyong Wu, Zhuosheng Zhang, Ben Kao, Lingpeng Kong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Qiushi Sun, Mukai Li, Zhoumianze Liu, Zhihui Xie 0002, Fangzhi Xu, Zhangyue Yin, Kanzhi Cheng, Zichen Ding 0002, Qi Liu 0049, Zhiyong Wu 0003, Zhuosheng Zhang 0001, Ben Kao, Lingpeng Kong |
ACL (1) | 6 |
| 2026 | Efficient KL Divergence Estimation via Truncated Top-K Integration for Large Language ModelsabstractKullback-Leibler (KL) divergence regularization is essential for stabilizing reinforcement learning from human feedback (RLHF) in large language models (LLMs), yet its exact computation requires summing over vocabularies of all tokens, incurring prohibitive memory costs during training.Existing stochastic estimators circumvent this bottleneck by estimating KL divergence using only the sampled token from the trajectory, but suffer from high variance (k 1 ) or systematic bias (k 2 ).We propose TIKE (Top-k Importance-weighted KL Estimator), which exploits the Zipfian structure of language model distributions: by deterministically integrating over only the top-k tokens, TIKE captures most of the probability mass while effectively reducing memory cost.To ensure correctness in off-policy settings characteristic of Group Relative Policy Optimization (GRPO), we incorporate importance sampling weights that correct for distribution shift between rollout and optimization policies.Experiments on models across diverse benchmarks demonstrate that TIKE consistently outperforms stochastic baselines, while exhibiting substantially lower gradient variance.Our analysis reveals that TIKE closely tracks the exact Rao-Blackwellized estimator with nearzero variance, offering a practical path toward stable, memory-efficient KL regularization for reasoning-intensive LLMs training.Code: Luozhijie Jin, Bo Wang 0084, Zhangyue Yin, Xipeng Qiu |
ACL (1) | 5 |
| 2026 | Thermometer of Thoughts: Enhancing LLM's Exploration via Attention Temperature ModulationabstractZhiyuan Yu, Shijian Xiao, Cam-Tu Nguyen, Zhangyue Yin, Lekai Xing, Wenzhong Li, Sanglu Lu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Shijian Xiao, Cam-Tu Nguyen, Zhangyue Yin, Lekai Xing, Sanglu Lu |
ACL (1) | 4 |
| 2025 | Contextual Structure Knowledge Transfer for Graph Neural NetworksabstractGraph transfer learning endeavors to develop a Graph Neural Network (GNN) model in a fully-labeled source domain, with the intention of deploying it on a target domain that has limited labeled data for inference. We reveal that prevalent graph transfer learning methods are susceptible to the homophily shift problem. This issue arises from the divergence in homophily structures between the source and target graphs, leading to a notable deterioration in the performance of GNN models. In this paper, we introduce a novel Contextual Structural Graph Neural Network (CS-GNN) method, leveraging a tailored attention mechanism to apprehend a variety of local structural cues, facilitating structural knowledge transfer across domains. It features an ego-network module to distill local structural diversity and a moment-based approach to gauge structural patterns without needing ground-truth labels. CS-GNN crafts a feature smoothness matrix from node attributes, guiding a customized attention mechanism for feature aggregation. A group-wise fairness loss is employed to balance learning across various structural patterns, enhancing the model's ability to transfer knowledge across domains. Comprehensive experiments conducted on six benchmark datasets substantiate the superiority of CS-GNN over the state-of-the-art methods, demonstrating significant improvements in accuracy and robustness against homophily shifts. Zhangyue Yin, Xiaobin Hong 0002, Shijian Xiao, Sanglu Lu |
AAAI | 3 |
| 2025 | Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities?abstractThe advent of test-time scaling in large language models (LLMs), exemplified by Ope-nAI's o1 series, has advanced reasoning capabilities by scaling computational resource allocation during inference.While successors like QwQ, Deepseek-R1 (R1) and LIMO replicate these advancements, whether these models truly possess test-time scaling capabilities remains underexplored.This study found that longer CoTs of these o1-like models do not consistently enhance accuracy; in fact, correct solutions are often shorter than incorrect ones for the same questions.Further investigation shows this phenomenon is closely related to models' self-revision capabilities -longer CoTs contain more self-revisions, which often lead to performance degradation.We then compare sequential and parallel scaling strategies on QwQ, R1 and LIMO, finding that parallel scaling achieves better coverage and scalability.Based on these insights, we propose Shortest Majority Vote, a method that combines parallel scaling strategies with CoT length characteristics, significantly improving models' test-time scalability compared to conventional majority voting approaches. Zhiyuan Zeng 0004, Qinyuan Cheng, Zhangyue Yin, Yunhua Zhou, Xipeng Qiu |
ACL (1) | 3 |
| 2025 | Dynamic and Generalizable Process Reward ModelingabstractZhangyue Yin, Qiushi Sun, Zhiyuan Zeng, Qinyuan Cheng, Xipeng Qiu, Xuanjing Huang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Zhangyue Yin, Qiushi Sun, Zhiyuan Zeng 0004, Qinyuan Cheng, Xipeng Qiu, Xuanjing Huang 0001 |
ACL (1) | 1 |
| 2025 | VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning TasksabstractGeneral-purposed embodied agents are designed to understand the users' natural instructions or intentions and act precisely to complete universal tasks. Recently, methods based on foundation models especially Vision-Language-Action models (VLAs) have shown a substantial potential to solve language-conditioned manipulation (LCM) tasks well. However, existing benchmarks do not adequately meet the needs of VLAs and relative algorithms. To better define such general-purpose tasks in the context of LLMs and advance the research in VLAs, we present VLABench, an open-source benchmark for evaluating universal LCM task learning. VLABench provides 100 carefully designed categories of tasks, with strong randomization in each category of task and a total of 2000+ objects. VLABench stands out from previous benchmarks in four key aspects: 1) tasks requiring world knowledge and common sense transfer, 2) natural language instructions with implicit human intentions rather than templates, 3) long-horizon tasks demanding multi-step reasoning, and 4) evaluation of both action policies and language model capabilities. The benchmark assesses multiple competencies including understanding of mesh\&texture, spatial relationship, semantic instruction, physical laws, knowledge transfer and reasoning, etc. To support the downstream finetuning, we provide high-quality training data collected via an automated framework incorporating heuristic skills and prior information. The experimental results indicate that both the current state-of-the-art pretrained VLAs and the workflow based on VLMs face challenges in our tasks. Shiduo Zhang, Peiju Liu, Qinghui Gao, Zhaoye Fei, Zhangyue Yin, Zuxuan Wu, Yu-Gang Jiang 0001, Xipeng Qiu |
ICCV | 8 |
| 2025 | The dual-edged sword: artificial intelligence's evolving role in academic peer review
Xuanjing Huang 0001, Shihan Dou, Zhangyue Yin |
Sci. China Inf. Sci. | 3 |
| 2025 | The rise and potential of large language model based agents: a survey
Zhiheng Xi, Wenxiang Chen, Wei He 0024, Yiwen Ding, Boyang Hong, Ming Zhang 0030, Junzhe Wang 0001, Senjie Jin, Enyu Zhou, Xiaoran Fan, Xiao Wang 0001, Limao Xiong, Yuhao Zhou 0005, Weiran Wang 0003, Changhao Jiang, Yicheng Zou, Zhangyue Yin, Shihan Dou, Rongxiang Weng, Wenjuan Qin, Yongyan Zheng, Xipeng Qiu, Xuanjing Huang 0001, Qi Zhang 0001, Tao Gui |
Sci. China Inf. Sci. | 20 |
| 2024 | Reasoning in Flux: Enhancing Large Language Models Reasoning through Uncertainty-aware Adaptive GuidanceabstractZhangyue Yin, Qiushi Sun, Qipeng Guo, Zhiyuan Zeng, Xiaonan Li, Junqi Dai, Qinyuan Cheng, Xuanjing Huang, Xipeng Qiu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Zhangyue Yin, Qiushi Sun, Qipeng Guo, Zhiyuan Zeng 0004, Junqi Dai, Qinyuan Cheng, Xuanjing Huang 0001, Xipeng Qiu |
ACL (1) | 1 |
| 2024 | Benchmarking Hallucination in Large Language Models Based on Unanswerable Math Word ProblemabstractLarge language models (LLMs) are highly effective in various natural language processing (NLP) tasks. However, they are susceptible to producing unreliable conjectures in ambiguous contexts called hallucination. This paper presents a new method for evaluating LLM hallucination in Question Answering (QA) based on the unanswerable math word problem (MWP). To support this approach, we innovatively develop a dataset called Unanswerable Math Word Problem (UMWP) which comprises 5200 questions across five categories. We developed an evaluation methodology combining text similarity and mathematical expression detection to determine whether LLM considers the question unanswerable. The results of extensive experiments conducted on 31 LLMs, including GPT-3, InstructGPT, LLaMA, and Claude, demonstrate that in-context learning and reinforcement learning with human feedback (RLHF) training significantly enhance the model’s ability to avoid hallucination. We show that utilizing MWP is a reliable and effective approach to assess hallucination. Our code and data are available at https://github.com/Yuki-Asuuna/UMWP. Yuhong Sun, Zhangyue Yin, Qipeng Guo, Jiawen Wu 0002, Xipeng Qiu |
LREC/COLING | 2 |
| 2024 | Aggregation of Reasoning: A Hierarchical Framework for Enhancing Answer Selection in Large Language ModelsabstractRecent advancements in Chain-of-Thought prompting have facilitated significant breakthroughs for Large Language Models (LLMs) in complex reasoning tasks. Current research enhances the reasoning performance of LLMs by sampling multiple reasoning chains and ensembling based on the answer frequency. However, this approach fails in scenarios where the correct answers are in the minority. We identify this as a primary factor constraining the reasoning capabilities of LLMs, a limitation that cannot be resolved solely based on the predicted answers. To address this shortcoming, we introduce a hierarchical reasoning aggregation framework AoR (Aggregation of Reasoning), which selects answers based on the evaluation of reasoning chains. Additionally, AoR incorporates dynamic sampling, adjusting the number of reasoning chains in accordance with the complexity of the task. Experimental results on a series of complex reasoning tasks show that AoR outperforms prominent ensemble methods. Further analysis reveals that AoR not only adapts various LLMs but also achieves a superior performance ceiling when compared to current methods. Zhangyue Yin, Qiushi Sun, Qipeng Guo, Zhiyuan Zeng 0004, Tianxiang Sun, Qinyuan Cheng, Xiaofeng Mou, Xipeng Qiu, Xuanjing Huang 0001 |
LREC/COLING | 1 |
| 2024 | Pixel-Level Semantic Correspondence Through Layout-Aware Representation Learning and Multi-Scale Matching IntegrationabstractEstablishing precise semantic correspondence across object instances in different images is a fundamental and challenging task in computer vision. In this task, difficulty arises often due to three challenges: confusing regions with similar appearance, inconsistent object scale, and indistinguishable nearby pixels. Recognizing these challenges, our paper proposes a novel semantic matching pipeline named LPMFlow toward extracting fine-grained semantics and geometry layouts for building pixel-level semantic correspondences. LPMFlow consists of three modules, each addressing one of the aforementioned challenges. The layout-aware representation learning module uniformly encodes source and target tokens to distinguish pixels or regions with similar appearances but different geometry semantics. The progressive feature superresolution module outputs four sets of 4D correlation tensors to generate accurate semantic flow between objects in different scales. Finally, the matching flow integration and refinement module is exploited to fuse matching flow in different scales to give the final flow predictions. The whole pipeline can be trained end-to-end, with a balance of computational cost and correspondence details. Extensive experiments based on benchmarks such as SPair-71K, PF-PASCAL, and PF-WILLOW have proved that the proposed method can well tackle the three challenges and outperform the previous methods, es-pecially in more stringent settings. Code is available at https://github.com/YXSUNMADMAX/LPMFlow. Yixuan Sun, Zhangyue Yin, Haibo Wang 0006, Yan Wang 0068, Xipeng Qiu, Weifeng Ge |
CVPR | 2 |
| 2024 | Explicit Memory Learning with Expectation MaximizationabstractLarge Language Models (LLMs) have revolutionized the landscape of natural language processing, demonstrating remarkable abilities across various complex tasks.However, their stateless nature limits the capability to retain information across interactions, hindering performance in scenarios requiring historical context recall.To mitigate this, current approaches primarily use explicit memory to allow LLMs to store useful information, which is accessible, readable, and interpretable.Nevertheless, explicit memory lacks the reliable learning mechanisms of implicit memory, which can be optimized end-to-end.To harness the benefits of both, we introduce EM 2 , a novel framework enhancing explicit memory updates via the Expectation-Maximization (EM) algorithm.EM 2 treats memory as a latent variable, ensuring continual learning and improvement during updates.Experimental results on streaming inference tasks demonstrate that EM 2 outperforms existing methods without memory or with static external memory.Our in-depth analysis highlights that EM 2 significantly enhances performance across various backbones and memory strategies, providing a robust solution for advancing LLM memory management and enabling explicit memory to learn and improve similarly to implicit memory. Zhangyue Yin, Qiushi Sun, Qipeng Guo, Zhiyuan Zeng 0004, Qinyuan Cheng, Xipeng Qiu, Xuanjing Huang 0001 |
EMNLP | 1 |
| 2024 | Turn Waste into Worth: Rectifying Top-k Router of MoEabstractZhiyuan Zeng, Qipeng Guo, Zhaoye Fei, Zhangyue Yin, Yunhua Zhou, Linyang Li, Tianxiang Sun, Hang Yan, Dahua Lin, Xipeng Qiu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Zhiyuan Zeng 0004, Qipeng Guo, Zhaoye Fei, Zhangyue Yin, Yunhua Zhou, Linyang Li, Tianxiang Sun, Hang Yan 0001, Dahua Lin, Xipeng Qiu |
EMNLP | 4 |
| 2024 | Memorize Step by Step: Efficient Long-Context Prefilling with Incremental Memory and Decremental ChunkabstractZhiyuan Zeng, Qipeng Guo, Xiaoran Liu, Zhangyue Yin, Wentao Shu, Mianqiu Huang, Bo Wang, Yunhua Zhou, Linlin Li, Qun Liu, Xipeng Qiu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Zhiyuan Zeng 0004, Qipeng Guo, Zhangyue Yin, Wentao Shu, Mianqiu Huang, Bo Wang 0084, Yunhua Zhou, Linlin Li 0001, Qun Liu 0001, Xipeng Qiu |
EMNLP | 4 |
| 2024 | Can AI Assistants Know What They Don't Know?abstractAI assistants powered by Large Language Models (LLMs) have demonstrated impressive performance in various tasks. However, LLMs still make factual errors in knowledge-intensive tasks such as open-domain question answering. These untruthful responses from AI assistants can pose significant risks in practical applications. Therefore, in this paper, we ask the question Can AI assistants know what they don’t know and express this awareness through natural language? To investigate this, we construct a model-specific "I don’t know" (Idk) dataset. This dataset includes Supervised Fine-tuning data and preference data, categorizing questions based on whether the assistant knows or does not know the answers. Then, we align the assistant with its corresponding Idk dataset using different alignment methods, including Supervised Fine-tuning and preference optimization. Experimental results show that, after alignment with the Idk dataset, the assistant is more capable of declining to answer questions outside its knowledge scope. The assistant aligned with the Idk dataset shows significantly higher truthfulness than the original assistant. Qinyuan Cheng, Tianxiang Sun, Zhangyue Yin, Linyang Li, Zhengfu He, Kai Chen 0026, Xipeng Qiu |
ICML | 5 |
| 2024 | LLatrieval: LLM-Verified Retrieval for Verifiable GenerationabstractXiaonan Li, Changtai Zhu, Linyang Li, Zhangyue Yin, Tianxiang Sun, Xipeng Qiu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Changtai Zhu, Linyang Li, Zhangyue Yin, Tianxiang Sun, Xipeng Qiu |
NAACL-HLT | 4 |
| 2023 | Graph Structure Learning via Lottery Hypothesis at Scale
Yuxin Wang 0005, Xiannian Hu, Jiaqing Xie, Zhangyue Yin, Yunhua Zhou, Xipeng Qiu, Xuanjing Huang 0001 |
ACML | 4 |
| 2023 | Correspondence Transformers with Asymmetric Feature Learning and Matching Flow Super-ResolutionabstractThis paper solves the problem of learning dense visual correspondences between different object instances of the same category with only sparse annotations. We decompose this pixel-level semantic matching problem into two easier ones: (i) First, local feature descriptors of source and target images need to be mapped into shared semantic spaces to get coarse matching flows. (ii) Second, matching flows in low resolution should be refined to generate accurate point-to-point matching results. We propose asymmetric feature learning and matching flow super-resolution based on vision transformers to solve the above problems. The asymmetric feature learning module exploits a biased cross-attention mechanism to encode token features of source images with their target counterparts. Then matching flow in low resolutions is enhanced by a super-resolution network to get accurate correspondences. Our pipeline is built upon vision transformers and can be trained in an end-to-end manner. Extensive experimental results on several popular benchmarks, such as PF-PASCAL, PF-WILLOW, and SPair-71 K, demonstrate that the proposed method can catch subtle semantic differences in pixels efficiently. Code is available on https://github.com/YXSUNMADMAX/ACTR. Yixuan Sun, Dongyang Zhao, Zhangyue Yin, Tao Gui, Weifeng Ge |
CVPR | 3 |
| 2023 | Exchange-of-Thought: Enhancing Large Language Model Capabilities through Cross-Model CommunicationabstractLarge Language Models (LLMs) have recently made significant strides in complex reasoning tasks through the Chain-of-Thought technique.Despite this progress, their reasoning is often constrained by their intrinsic understanding, lacking external insights.To address this, we propose Exchange-of-Thought (EoT), a novel framework that enables cross-model communication during problem-solving.Drawing inspiration from network topology, EoT integrates four unique communication paradigms: Memory, Report, Relay, and Debate.This paper delves into the communication dynamics and volume associated with each paradigm.To counterbalance the risks of incorrect reasoning chains, we implement a robust confidence evaluation mechanism within these communications.Our experiments across diverse complex reasoning tasks demonstrate that EoT significantly surpasses established baselines, underscoring the value of external insights in enhancing LLM performance.Furthermore, we show that EoT achieves these superior results in a cost-effective manner, marking a promising advancement for efficient and collaborative AI problem-solving."Two heads are better than one. Zhangyue Yin, Qiushi Sun, Qipeng Guo, Junqi Dai, Xuanjing Huang 0001, Xipeng Qiu |
EMNLP | 1 |
| 2023 | An anchor-guided sequence labeling model for event detection in both data-abundant and data-scarce scenarios
Zhigang Kan, Yanqi Shi, Zhangyue Yin, Liwen Peng, Linbo Qiao, Xipeng Qiu, Dongsheng Li 0001 |
Inf. Sci. | 3 |
| 2023 | A Composable Generative Framework Based on Prompt Learning for Various Information Extraction TasksabstractPrompt learning is an effective paradigm that bridges gaps between the pre-training tasks and the corresponding downstream applications. Approaches based on this paradigm have achieved great transcendent results in various applications. However, it still needs to be answered how to design a general-purpose framework based on the prompt learning paradigm for various information extraction tasks. In this article, we propose a novel composable prompt-based generative framework, which could be applied to a wide range of tasks in the field of information extraction. Specifically, we reformulate information extraction tasks into the form of filling slots in pre-designed type-specific prompts, which consist of one or multiple sub-prompts. A strategy of constructing composable prompts is proposed to enhance the generalization ability in data-scarce scenarios. Furthermore, to fit this framework, we transform relation extraction into the task of determining semantic consistency in prompts. The experimental results demonstrate that our approach surpasses compared baselines on real-world datasets in data-abundant and data-scarce scenarios. Further analysis of the proposed framework is presented, as well as numerical experiments conducted to investigate impact factors of performance on various tasks. Zhigang Kan, Linhui Feng, Zhangyue Yin, Linbo Qiao, Xipeng Qiu, Dongsheng Li 0001 |
IEEE Trans. Big Data | 3 |
| 2022 | Improving Abstractive Dialogue Summarization with Speaker-Aware Supervised Contrastive LearningabstractPre-trained models have brought remarkable success on the text summarization task. For dialogue summarization, the subdomain of text summarization, utterances are concatenated to flat text before being processed. As a result, existing summarization systems based on pre-trained models are unable to recognize the unique format of the speaker-utterance pair well in the dialogue. To investigate this issue, we conduct probing tests and manual analysis, and find that the powerful pre-trained model can not identify different speakers well in the conversation, which leads to various factual errors. Moreover, we propose three speaker-aware supervised contrastive learning (SCL) tasks: Token-level SCL, Turn-level SCL, and Global-level SCL. Comprehensive experiments demonstrate that our methods achieve significant performance improvement on two mainstream dialogue summarization datasets. According to detailed human evaluations, pre-trained models equipped with SCL tasks effectively generate summaries with better factual consistency. Zhichao Geng, Ming Zhong 0005, Zhangyue Yin, Xipeng Qiu, Xuanjing Huang 0001 |
COLING | 3 |
| 2022 | What Dense Graph Do You Need for Self-Attention?abstractTransformers have made progress in miscellaneous tasks, but suffer from quadratic computational and memory complexities. Recent works propose sparse transformers with attention on sparse graphs to reduce complexity and remain strong performance. While effective, the crucial parts of how dense a graph needs to be to perform well are not fully explored. In this paper, we propose Normalized Information Payload (NIP), a graph scoring function measuring information transfer on graph, which provides an analysis tool for trade-offs between performance and complexity. Guided by this theoretical analysis, we present Hypercube Transformer, a sparse transformer that models token interactions in a hypercube and shows comparable or even better results with vanilla transformer while yielding $O(N\log N)$ complexity with sequence length $N$. Experiments on tasks requiring various sequence lengths lay validation for our graph function well. Yuxin Wang 0005, Chu-Tak Lee, Qipeng Guo, Zhangyue Yin, Yunhua Zhou, Xuanjing Huang 0001, Xipeng Qiu |
ICML | 4 |