VLDB 2026 Research / reviewers in the wild / expert
Yu Zhang 0092
dblp:50/671-92
· DBLP profile ↗
8ranked-venue papers
5as first author
5since 2021 · last 2024
0000-0002-8345-3835ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 5 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Deep learning architectures and training · 49% Information extraction and text analysis · 26% Language models and text generation · 16% |
Topics — the 12 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis
syntactic parsing |
0.9 | 2 | 2020 | Fast and Accurate Neural CRF Constituency Parsing · IJCAI 2020 Efficient Second-Order TreeCRF for Neural Dependency Parsing · ACL 2020 |
Machine learning › Deep learning architectures and training
attention mechanism |
0.8 | 1 | 2024 | Gated Slot Attention for Efficient Linear-Time Sequence Modeling · NeurIPS 2024 |
Machine learning › Deep learning architectures and training › sequence modeling
efficient sequence modeling |
0.8 | 1 | 2024 | Gated Slot Attention for Efficient Linear-Time Sequence Modeling · NeurIPS 2024 |
Machine learning › Deep learning architectures and training › attention mechanism › efficient attention
linear attention |
0.8 | 1 | 2024 | Gated Slot Attention for Efficient Linear-Time Sequence Modeling · NeurIPS 2024 |
Machine learning › Deep learning architectures and training
sequence modeling |
0.8 | 1 | 2024 | Gated Slot Attention for Efficient Linear-Time Sequence Modeling · NeurIPS 2024 |
Natural language and speech › Language models and text generation › controllable text generation
text editing |
0.7 | 1 | 2023 | Non-autoregressive Text Editing with Copy-aware Latent Alignments · EMNLP 2023 |
Machine learning › Probabilistic and Bayesian machine learning
structured prediction |
0.6 | 2 | 2020 | Efficient Second-Order TreeCRF for Neural Dependency Parsing · ACL 2020 Fast and Accurate Neural CRF Constituency Parsing · IJCAI 2020 |
Natural language and speech › Information extraction and text analysis › syntactic parsing
constituency parsing |
0.4 | 1 | 2020 | Fast and Accurate Neural CRF Constituency Parsing · IJCAI 2020 |
Natural language and speech › Information extraction and text analysis › syntactic parsing
dependency parsing |
0.4 | 1 | 2020 | Efficient Second-Order TreeCRF for Neural Dependency Parsing · ACL 2020 |
Machine learning › Deep learning architectures and training
transformer |
0.2 | 1 | 2024 | Gated Slot Attention for Efficient Linear-Time Sequence Modeling · NeurIPS 2024 |
Natural language and speech › Language models and text generation › text generation
grammatical error correction |
0.2 | 1 | 2023 | Non-autoregressive Text Editing with Copy-aware Latent Alignments · EMNLP 2023 |
Natural language and speech › Language models and text generation › text summarization
sentence fusion |
0.2 | 1 | 2023 | Non-autoregressive Text Editing with Copy-aware Latent Alignments · EMNLP 2023 |
Methods — techniques the papers use, named apart from their topics
inside-outside algorithm · 0.9softmax attention · 0.8gating mechanism · 0.8latent alignment · 0.7copy mechanism · 0.7CTC alignment · 0.7viterbi algorithm · 0.4conditional random field · 0.4biaffine parser · 0.4biaffine attention · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Gated Slot Attention for Efficient Linear-Time Sequence ModelingabstractLinear attention Transformers and their gated variants, celebrated for enabling parallel training and efficient recurrent inference, still fall short in recall-intensive tasks compared to traditional Transformers and demand significant resources for training from scratch.
This paper introduces Gated Slot Attention (GSA), which enhances Attention with Bounded-memory-Control (ABC) by incorporating a gating mechanism inspired by Gated Linear Attention (GLA).
Essentially, GSA comprises a two-layer GLA linked via $\operatorname{softmax}$, utilizing context-aware memory reading and adaptive forgetting to improve memory capacity while maintaining compact recurrent state size.
This design greatly enhances both training and inference efficiency through GLA's hardware-efficient training algorithm and reduced state size.
Additionally, retaining the $\operatorname{softmax}$ operation is particularly beneficial in ``finetuning pretrained Transformers to RNNs'' (T2R) settings, reducing the need for extensive training from scratch.
Extensive experiments confirm GSA's superior performance in scenarios requiring in-context recall and in T2R settings. Yu Zhang 0092, Rui-Jie Zhu 0003, Yue Zhang 0004, Leyang Cui, Yiqiao Wang 0005, Bolun Wang, Freda Shi, Bailin Wang, Wei Bi, Peng Zhou 0017, Guohong Fu |
NeurIPS | 1 |
| 2023 | Non-autoregressive Text Editing with Copy-aware Latent AlignmentsabstractRecent work has witnessed a paradigm shift from Seq2Seq to Seq2Edit in the field of text editing, with the aim of addressing the slow autoregressive inference problem posed by the former.Despite promising results, Seq2Edit approaches still face several challenges such as inflexibility in generation and difficulty in generalizing to other languages.In this work, we propose a novel non-autoregressive text editing method to circumvent the above issues, by modeling the edit process with latent CTC alignments.We make a crucial extension to CTC by introducing the copy operation into the edit space, thus enabling more efficient management of textual overlap in editing.We conduct extensive experiments on GEC and sentence fusion tasks, showing that our proposed method significantly outperforms existing Seq2Edit models and achieves similar or even better results than Seq2Seq with over 4ˆ speedup.Moreover, it demonstrates good generalizability on German and Russian.In-depth analyses reveal the strengths of our method in terms of the robustness under various scenarios and generating fluent and flexible outputs. Yu Zhang 0092, Yue Zhang 0004, Leyang Cui, Guohong Fu |
EMNLP | 1 |
| 2022 | Semantic Role Labeling as Dependency Parsing: Exploring Latent Tree Structures inside ArgumentsabstractSemantic role labeling (SRL) is a fundamental yet challenging task in the NLP community. Recent works of SRL mainly fall into two lines: 1) BIO-based; 2) span-based. Despite ubiquity, they share some intrinsic drawbacks of not considering internal argument structures, potentially hindering the model’s expressiveness. The key challenge is arguments are flat structures, and there are no determined subtree realizations for words inside arguments. To remedy this, in this paper, we propose to regard flat argument spans as latent subtrees, accordingly reducing SRL to a tree parsing task. In particular, we equip our formulation with a novel span-constrained TreeCRF to make tree structures span-aware and further extend it to the second-order case. We conduct extensive experiments on CoNLL05 and CoNLL12 benchmarks. Results reveal that our methods perform favorably better than all previous syntax-agnostic works, achieving new state-of-the-art under both end-to-end and w/ gold predicates settings. Yu Zhang 0092, Qingrong Xia, Shilin Zhou 0002, Guohong Fu, Min Zhang 0005 |
COLING | 1 |
| 2022 | Fast and Accurate End-to-End Span-based Semantic Role Labeling as Word-based Graph ParsingabstractThis paper proposes to cast end-to-end span-based SRL as a word-based graph parsing task. The major challenge is how to represent spans at the word level. Borrowing ideas from research on Chinese word segmentation and named entity recognition, we propose and compare four different schemata of graph representation, i.e., BES, BE, BIES, and BII, among which we find that the BES schema performs the best. We further gain interesting insights through detailed analysis. Moreover, we propose a simple constrained Viterbi procedure to ensure the legality of the output graph according to the constraints of the SRL structure. We conduct experiments on two widely used benchmark datasets, i.e., CoNLL05 and CoNLL12. Results show that our word-based graph parsing approach achieves consistently better performance than previous results, under all settings of end-to-end and predicate-given, without and with pre-trained language models (PLMs). More importantly, our model can parse 669/252 sentences per second, without and with PLMs respectively. Shilin Zhou 0002, Qingrong Xia, Zhenghua Li, Yu Zhang 0092, Yu Hong 0001, Min Zhang 0005 |
COLING | 4 |
| 2021 | A Coarse-to-Fine Labeling Framework for Joint Word Segmentation, POS Tagging, and Constituent ParsingabstractThe most straightforward approach to joint word segmentation (WS), part-of-speech (POS) tagging, and constituent parsing (PAR) is converting a word-level tree into a char-level tree, which, however, leads to two severe challenges.First, a larger label set (e.g., ≥ 600) and longer inputs both increase computational cost.Second, it is difficult to rule out illegal trees containing conflicting production rules, which is important for reliable model evaluation.If a POS tag (like VV) is above a phrase tag (like VP) in the output tree, it becomes quite complex to decide word boundaries.To deal with both challenges, this work proposes a two-stage coarse-to-fine labeling framework for joint WS-POS-PAR.In the coarse labeling stage, the joint model outputs a bracketed tree, in which each node corresponds to one of four labels (i.e., phrase, subphrase, word, subword).The tree is guaranteed to be legal via constrained CKY decoding.In the fine labeling stage, the model expands each coarse label into a final label (such as VP, VP * , VV, VV * ).Experiments on Chinese Penn Treebank 5.1 and 7.0 show that our joint model consistently outperforms the pipeline approach on both settings of without and with BERT, and achieves new state-of-the-art performance. Yang Hou 0001, Houquan Zhou 0001, Zhenghua Li, Yu Zhang 0092, Min Zhang 0005, Zhefeng Wang 0001, Baoxing Huai, Nicholas Jing Yuan |
CoNLL | 4 |
| 2020 | Efficient Second-Order TreeCRF for Neural Dependency ParsingabstractIn the deep learning (DL) era, parsing models are extremely simplified with little hurt on performance, thanks to the remarkable capability of multi-layer BiLSTMs in context representation.As the most popular graphbased dependency parser due to its high efficiency and performance, the biaffine parser directly scores single dependencies under the arc-factorization assumption, and adopts a very simple local token-wise cross-entropy training loss.This paper for the first time presents a second-order TreeCRF extension to the biaffine parser.For a long time, the complexity and inefficiency of the inside-outside algorithm hinder the popularity of TreeCRF.To address this issue, we propose an effective way to batchify the inside and Viterbi algorithms for direct large matrix operation on GPUs, and to avoid the complex outside algorithm via efficient back-propagation.Experiments and analysis on 27 datasets from 13 languages clearly show that techniques developed before the DL era, such as structural learning (global TreeCRF loss) and high-order modeling are still useful, and can further boost parsing performance over the state-of-the-art biaffine parser, especially for partially annotated training data.We release our code at https: //github.com/yzhangcs/crfpar. Yu Zhang 0092, Zhenghua Li, Min Zhang 0005 |
ACL | 1 |
| 2020 | Fast and Accurate Neural CRF Constituency ParsingabstractEstimating probability distribution is one of the core issues in the NLP field. However, in both deep learning (DL) and pre-DL eras, unlike the vast applications of linear-chain CRF in sequence labeling tasks, very few works have applied tree-structure CRF to constituency parsing, mainly due to the complexity and inefficiency of the inside-outside algorithm. This work presents a fast and accurate neural CRF constituency parser. The key idea is to batchify the inside algorithm for loss computation by direct large tensor operations on GPU, and meanwhile avoid the outside algorithm for gradient computation via efficient back-propagation. We also propose a simple two-stage bracketing-then-labeling parsing approach to improve efficiency further. To improve the parsing performance, inspired by recent progress in dependency parsing, we introduce a new scoring architecture based on boundary representation and biaffine attention, and a beneficial dropout strategy. Experiments on PTB, CTB5.1, and CTB7 show that our two-stage CRF parser achieves new state-of-the-art performance on both settings of w/o and w/ BERT, and can parse over 1,000 sentences per second. We release our code at https://github.com/yzhangcs/crfpar. Yu Zhang 0092, Houquan Zhou 0001, Zhenghua Li |
IJCAI | 1 |
| 2020 | Is POS Tagging Necessary or Even Helpful for Neural Dependency Parsing?
Houquan Zhou 0001, Yu Zhang 0092, Zhenghua Li, Min Zhang 0005 |
NLPCC (1) | 2 |