Zhenghua Li

dblp:72/8937 · DBLP profile ↗
← Back
67ranked-venue papers
15as first author
27since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 64 · 14 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 FGNet: Leveraging Feature-Guided Attention to Refine SAM2 for 3D EM Neuron Segmentation
abstract
Accurate segmentation of neural structures in Electron Microscopy (EM) images is paramount for neuroscience. However, this task is challenged by intricate morphologies, low signal-to-noise ratios, and scarce annotations, limiting the accuracy and generalization of existing methods. To address these challenges, we seek to leverage the priors learned by visual foundation models on a vast amount of natural images to better tackle this task. Specifically, we propose a novel framework that can effectively transfer knowledge from Segment Anything 2 (SAM2), which is pre-trained on natural images, to the EM domain. We first use SAM2 to extract powerful, general-purpose features. To bridge the domain gap, we introduce a Feature-Guided Attention module that leverages semantic cues from SAM2 to guide a lightweight encoder, the Fine-Grained Encoder (FGE), in focusing on these challenging regions. Finally, a dual-affinity decoder generates both coarse and refined affinity maps. Experimental results demonstrate that our method achieves performance comparable to state-of-the-art (SOTA) approaches with the SAM2 weights frozen. Upon further fine-tuning on EM data, our method significantly outperforms existing SOTA methods. This study validates that transferring representations pre-trained on natural images, when combined with targeted domain-adaptive guidance, can effectively address the specific challenges in neuron segmentation.
Zhenghua Li, Hang Chen 0004, Kai Li 0047, Xiaolin Hu 0001
AAAI1
2026 Filter-based predefined-time optimal fault-tolerant consensus control for nonlinear multi-agent systems via reinforcement learning
Zhenghua Li, Guangjing Song
Neurocomputing1
2025 A Training-free LLM-based Approach to General Chinese Character Error Correction
abstract
Chinese spelling correction (CSC) is a crucial task that aims to correct character errors in Chinese text.While conventional CSC focuses on character substitution errors caused by mistyping, two other common types of character errors, missing and redundant characters, have received less attention.These errors are often excluded from CSC datasets during the annotation process or ignored during evaluation, even when they have been annotated.This issue limits the practicality of the CSC task.To address this issue, we introduce the task of General Chinese Character Error Correction (C2EC), which focuses on all three types of character errors.We construct a high-quality C2EC benchmark by combining and manually verifying data from CCTC and Lemon datasets.We extend the training-free prompt-free CSC method to C2EC by using Levenshtein distance for handling length changes and leveraging an additional prompt-based large language model (LLM) to improve performance.Experiments show that our method enables a 14B-parameter LLM to be on par with models nearly 50 times larger on both conventional CSC and C2EC tasks, without any fine-tuning.
Houquan Zhou 0001, Bo Zhang 0071, Zhenghua Li, Ming Yan 0008, Min Zhang 0005
ACL (1)3
2025 Dynamic Head Selection for Neural Lexicalized Constituency Parsing
abstract
Lexicalized parsing, which associates constituent nodes with lexical heads, has historically played a crucial role in constituency parsing by bridging constituency and dependency structures. Nevertheless, with the advent of neural networks, lexicalized structures have generally been neglected in favor of unlexicalized, span-based methods. In this paper, we revisit lexicalized parsing and propose a novel latent lexicalization framework that dynamically infers lexical heads during training without relying on predefined head-finding rules. Our method enables the model to learn lexical dependencies directly from data, offering greater adaptability across languages and datasets. Experiments on multiple treebanks demonstrate state-of-the-art or comparable performance. We also analyze the learned dependency structures, headword preferences, and linguistic biases.
Yang Hou 0001, Zhenghua Li
ACL (1)2
2025 Mixture of Small and Large Models for Chinese Spelling Check
abstract
In the era of large language models (LLMs), the Chinese Spelling Check (CSC) task has seen various LLM methods developed, yet their performance remains unsatisfactory.In contrast, fine-tuned BERT-based models, relying on high-quality in-domain data, show excellent performance but suffer from edit pattern overfitting.This paper proposes a novel dynamic mixture approach that effectively combines the probability distributions of small models and LLMs during the beam search decoding phase, achieving a balanced enhancement of precise corrections from small models and the fluency of LLMs.This approach also eliminates the need for fine-tuning LLMs, saving significant time and resources, and facilitating domain adaptation.Comprehensive experiments demonstrate that our mixture approach significantly boosts error correction capabilities, achieving state-of-the-art results across multiple datasets.Our code is available at https://github.com/zhqiao-nlp/MSLLM.
Ziheng Qiao, Houquan Zhou 0001, Zhenghua Li
ACL (1)3
2025 DISC: Plug-and-Play Decoding Intervention with Similarity of Characters for Chinese Spelling Check
abstract
Ziheng Qiao, Houquan Zhou, Yumeng Liu, Zhenghua Li, Min Zhang, Bo Zhang, Chen Li, Ji Zhang, Fei Huang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Ziheng Qiao, Houquan Zhou 0001, Zhenghua Li, Min Zhang 0005, Bo Zhang 0071, Chen Li 0001, Ji Zhang 0011, Fei Huang 0002
ACL (1)4
2025 Mining Word Boundaries from Speech-Text Parallel Data for Cross-domain Chinese Word Segmentation
abstract
Inspired by early research on exploring naturally annotated data for Chinese Word Segmentation (CWS), and also by recent research on integration of speech and text processing, this work for the first time proposes to explicitly mine word boundaries from parallel speech-text data. We employ the Montreal Forced Aligner (MFA) toolkit to perform character-level alignment on speech-text data, giving pauses as candidate word boundaries. Based on detailed analysis of collected pauses, we propose an effective probability-based strategy for filtering unreliable word boundaries. To more effectively utilize word boundaries as extra training data, we also propose a robust complete-then-train (CTT) strategy. We conduct cross-domain CWS experiments on two target domains, i.e., ZX and AISHELL2. We have annotated about 1K sentences as the evaluation data of AISHELL2. Experiments demonstrate the effectiveness of our proposed approach.
Zhenghua Li, Shilin Zhou 0002, Chen Gong 0004, Yang Hou 0001
COLING3
2025 Data Augmentation for Cross-domain Parsing via Lightweight LLM Generation and Tree Hybridization
abstract
Cross-domain constituency parsing remains a challenging task due to the lack of high-quality out-of-domain data. In this paper, we propose a data augmentation method via lightweight large language model (LLM) generation and tree hybridization. We utilize LLM to generate phrase structures (subtrees) for the target domain by incorporating grammar rules and lexical head information into the prompt. To better leverage LLM-generated target-domain subtrees, we hybridize them with existing source-domain subtrees to efficiently produce a large number of structurally diverse instances. Experimental results demonstrate that our method achieves significant improvements on five target domains with a lightweight LLM generation cost.
Yang Hou 0001, Chen Gong 0004, Zhenghua Li
COLING4
2025 ULDC: uncertainty-based learning for deep clustering
Luyao Chang, Xinzheng Niu, Zhenghua Li, Shenshen Li, Philippe Fournier-Viger
Appl. Intell.3
2025 Annotation error detection in painstakingly annotated data: Part-of-speech tagging as a case study
Zhenghua Li, Chen Gong 0004, Shilin Zhou 0002, Min Zhang 0005
Expert Syst. Appl.2
2024 CopyNE: Better Contextual ASR by Copying Named Entities
abstract
End-to-end automatic speech recognition (ASR) systems have made significant progress in general scenarios.However, it remains challenging to transcribe contextual named entities (NEs) in the contextual ASR scenario.Previous approaches have attempted to address this by utilizing the NE dictionary.These approaches treat entities as individual tokens and generate them token-by-token, which may result in incomplete transcriptions of entities.In this paper, we treat entities as indivisible wholes and introduce the idea of copying into ASR.We design a systematic mechanism called CopyNE, which can copy entities from the NE dictionary.By copying all tokens of an entity at once, we can reduce errors during entity transcription, ensuring the completeness of the entity.Experiments demonstrate that CopyNE consistently improves the accuracy of transcribing entities compared to previous approaches.Even when based on the strong Whisper, CopyNE still achieves notable improvements.
Shilin Zhou 0002, Zhenghua Li, Yu Hong 0001, Min Zhang 0005, Zhefeng Wang 0001, Baoxing Huai
ACL (1)2
2024 Improving Chinese Named Entity Recognition with Multi-grained Words and Part-of-Speech Tags via Joint Modeling
abstract
Nowadays, character-based sequence labeling becomes the mainstream Chinese named entity recognition (CNER) approach, instead of word-based methods, since the latter degrades performance due to propagation of word segmentation (WS) errors. To make use of WS information, previous studies usually learn CNER and WS simultaneously with multi-task learning (MTL) framework, or treat WS information as extra guide features for CNER model, in which the utilization of WS information is indirect and shallow. In light of the complementary information inside multi-grained words, and the close connection between named entities and part-of-speech (POS) tags, this work proposes a tree parsing approach for joint modeling CNER, multi-grained word segmentation (MWS) and POS tagging tasks simultaneously. Specifically, we first propose a unified tree representation for MWS, POS tagging, and CNER.Then, we automatically construct the MWS-POS-NER data based on the unified tree representation for model training. Finally, we present a two-stage joint tree parsing framework. Experimental results on OntoNotes4 and OntoNotes5 show that our proposed approach of jointly modeling CNER with MWS and POS tagging achieves better or comparable performance with latest methods.
Chenhui Dou, Chen Gong 0004, Zhenghua Li, Zhefeng Wang 0001, Baoxing Huai, Min Zhang 0005
LREC/COLING3
2024 High-order Joint Constituency and Dependency Parsing
abstract
This work revisits the topic of jointly parsing constituency and dependency trees, i.e., to produce compatible constituency and dependency trees simultaneously for input sentences, which is attractive considering that the two types of trees are complementary in representing syntax. The original work of Zhou and Zhao (2019) performs joint parsing only at the inference phase. They train two separate parsers under the multi-task learning framework (i.e., one shared encoder and two independent decoders). They design an ad-hoc dynamic programming-based decoding algorithm of O(n^5) time complexity for finding optimal compatible tree pairs. Compared to their work, we make progress in three aspects: (1) adopting a much more efficient decoding algorithm of O(n^4) time complexity, (2) exploring joint modeling at the training phase, instead of only at the inference phase, (3) proposing high-order scoring components to promote constituent-dependency interaction. We conduct experiments and analysis on seven languages, covering both rich-resource and low-resource scenarios. Results and analysis show that joint modeling leads to a modest overall performance boost over separate modeling, but substantially improves the complete matching ratio of whole trees, thanks to the explicit modeling of tree compatibility.
Yanggang Gu, Yang Hou 0001, Zhefeng Wang 0001, Xinyu Duan, Zhenghua Li
LREC/COLING5
2024 A Simple yet Effective Training-free Prompt-free Approach to Chinese Spelling Correction Based on Large Language Models
abstract
This work proposes a simple training-free prompt-free approach to leverage large language models (LLMs) for the Chinese spelling correction (CSC) task, which is totally different from all previous CSC approaches.The key idea is to use an LLM as a pure language model in a conventional manner.The LLM goes through the input sentence from the beginning, and at each inference step, produces a distribution over its vocabulary for deciding the next token, given a partial sentence.To ensure that the output sentence remains faithful to the input sentence, we design a minimal distortion model that utilizes pronunciation or shape similarities between the original and replaced characters.Furthermore, we propose two useful reward strategies to address practical challenges specific to the CSC task.Experiments on five public datasets demonstrate that our approach significantly improves LLM performance, enabling them to compete with state-of-the-art domain-general CSC models.
Houquan Zhou 0001, Zhenghua Li, Bo Zhang 0071, Chen Li 0001, Shaopeng Lai, Ji Zhang 0011, Fei Huang 0002, Min Zhang 0005
EMNLP2
2023 SeSQL: A High-Quality Large-Scale Session-Level Chinese Text-to-SQL Dataset
Saihao Huang, Zhenghua Li, Chenhui Dou, Fukang Yan, Xinyan Xiao, Hua Wu 0003, Min Zhang 0005
NLPCC (1)3
2022 Fast and Accurate End-to-End Span-based Semantic Role Labeling as Word-based Graph Parsing
abstract
This paper proposes to cast end-to-end span-based SRL as a word-based graph parsing task. The major challenge is how to represent spans at the word level. Borrowing ideas from research on Chinese word segmentation and named entity recognition, we propose and compare four different schemata of graph representation, i.e., BES, BE, BIES, and BII, among which we find that the BES schema performs the best. We further gain interesting insights through detailed analysis. Moreover, we propose a simple constrained Viterbi procedure to ensure the legality of the output graph according to the constraints of the SRL structure. We conduct experiments on two widely used benchmark datasets, i.e., CoNLL05 and CoNLL12. Results show that our word-based graph parsing approach achieves consistently better performance than previous results, under all settings of end-to-end and predicate-given, without and with pre-trained language models (PLMs). More importantly, our model can parse 669/252 sentences per second, without and with PLMs respectively.
Shilin Zhou 0002, Qingrong Xia, Zhenghua Li, Yu Zhang 0092, Yu Hong 0001, Min Zhang 0005
COLING3
2022 SynGEC: Syntax-Enhanced Grammatical Error Correction with a Tailored GEC-Oriented Parser
abstract
This work proposes a syntax-enhanced grammatical error correction (GEC) approach named SynGEC that effectively incorporates dependency syntactic information into the encoder part of GEC models. 1 The key challenge for this idea is that off-the-shelf parsers are unreliable when processing ungrammatical sentences.To confront this challenge, we propose to build a tailored GEC-oriented parser (GOPar) using parallel GEC training data as a pivot.First, we design an extended syntax representation scheme that allows us to represent both grammatical errors and syntax in a unified tree structure.Then, we obtain parse trees of the source incorrect sentences by projecting trees of the target correct sentences.Finally, we train GOPar with such projected trees.For GEC, we employ the graph convolution network to encode source-side syntactic information produced by GOPar, and fuse them with the outputs of the Transformer encoder.Experiments on mainstream English and Chinese GEC datasets show that our proposed SynGEC approach consistently and substantially outperforms strong baselines and achieves competitive performance.
Yue Zhang 0004, Bo Zhang 0071, Zhenghua Li, Zuyi Bao, Chen Li 0001, Min Zhang 0005
EMNLP3
2022 MuCGEC: a Multi-Reference Multi-Source Evaluation Dataset for Chinese Grammatical Error Correction
abstract
Yue Zhang, Zhenghua Li, Zuyi Bao, Jiacheng Li, Bo Zhang, Chen Li, Fei Huang, Min Zhang. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Yue Zhang 0004, Zhenghua Li, Zuyi Bao, Bo Zhang 0071, Chen Li 0001, Fei Huang 0002, Min Zhang 0005
NAACL-HLT2
2022 MuCPAD: A Multi-Domain Chinese Predicate-Argument Dataset
abstract
Yahui Liu, Haoping Yang, Chen Gong, Qingrong Xia, Zhenghua Li, Min Zhang. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Haoping Yang, Chen Gong 0004, Qingrong Xia, Zhenghua Li, Min Zhang 0005
NAACL-HLT5
2022 Faster and Better Grammar-Based Text-to-SQL Parsing via Clause-Level Parallel Decoding and Alignment Loss
Kun Wu 0009, Zhenghua Li, Xinyan Xiao
NLPCC (2)3
2022 Neural Coupled Sequence Labeling for Heterogeneous Annotation Conversion
abstract
Supervised statistical models rely on large-scale high-quality labeled data, which is important for model training but expensive to construct. Therefore, instead of constructing new dataset, researchers have attempted to make full use of various existing heterogeneous datasets to boost model performance, considering it is ubiquitous that the same task may have multiple annotated data following different and incompatible annotation guidelines. Representative methods include the guide-feature method which use the knowledge projected from the source-side to the target-side as extra features for target model guidance, and the multi-task learning (MTL) method which simultaneously train on multiple heterogeneous annotations with shared parameters to gain resource-share knowledge. Though effective, the guide-feature method fails to directly use the source-side data as training data, and the MTL method ignores the implicit mappings between heterogeneous datasets. Compared with the above methods, directly converting the heterogeneous datasets into homogeneous datasets for target model training is a more straightforward and effective way to fully exploit heterogeneous resources. In this work, we propose a neural coupled sequence labeling model for heterogeneous annotation conversion. First, for each token, we map a given one-side tag into a set of bundled tags by concatenating the tag with all the possible tags at the other side. Then, we build a neural coupled model over the bundled tag space. Finally, we convert heterogeneous annotations into homogeneous annotations by performing constraint decoding on the coupled model. We also propose a pruning strategy to address the oversize issue of the bundled tag space, which improves efficiency without hurting model performance.Experiments for part-of-speech (POS) tagging, word segmentation (WS), and WS&POS tagging tasks show that our proposed neural coupled model consistently outperforms several benchmark models for all the three tasks by large margin.
Chen Gong 0004, Zhenghua Li, Min Zhang 0005
IEEE ACM Trans. Audio Speech Lang. Process.2
2021 An In-depth Study on Internal Structure of Chinese Words
abstract
Chen Gong, Saihao Huang, Houquan Zhou, Zhenghua Li, Min Zhang, Zhefeng Wang, Baoxing Huai, Nicholas Jing Yuan. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Chen Gong 0004, Saihao Huang, Houquan Zhou 0001, Zhenghua Li, Min Zhang 0005, Zhefeng Wang 0001, Baoxing Huai, Nicholas Jing Yuan
ACL/IJCNLP (1)4
2021 A Coarse-to-Fine Labeling Framework for Joint Word Segmentation, POS Tagging, and Constituent Parsing
abstract
The most straightforward approach to joint word segmentation (WS), part-of-speech (POS) tagging, and constituent parsing (PAR) is converting a word-level tree into a char-level tree, which, however, leads to two severe challenges.First, a larger label set (e.g., ≥ 600) and longer inputs both increase computational cost.Second, it is difficult to rule out illegal trees containing conflicting production rules, which is important for reliable model evaluation.If a POS tag (like VV) is above a phrase tag (like VP) in the output tree, it becomes quite complex to decide word boundaries.To deal with both challenges, this work proposes a two-stage coarse-to-fine labeling framework for joint WS-POS-PAR.In the coarse labeling stage, the joint model outputs a bracketed tree, in which each node corresponds to one of four labels (i.e., phrase, subphrase, word, subword).The tree is guaranteed to be legal via constrained CKY decoding.In the fine labeling stage, the model expands each coarse label into a final label (such as VP, VP * , VV, VV * ).Experiments on Chinese Penn Treebank 5.1 and 7.0 show that our joint model consistently outperforms the pipeline approach on both settings of without and with BERT, and achieves new state-of-the-art performance.
Yang Hou 0001, Houquan Zhou 0001, Zhenghua Li, Yu Zhang 0092, Min Zhang 0005, Zhefeng Wang 0001, Baoxing Huai, Nicholas Jing Yuan
CoNLL3
2021 Data Augmentation with Hierarchical SQL-to-Question Generation for Cross-domain Text-to-SQL Parsing
abstract
Data augmentation has attracted a lot of research attention in the deep learning era for its ability in alleviating data sparseness.The lack of labeled data for unseen evaluation databases is exactly the major challenge for cross-domain text-to-SQL parsing.Previous works either require human intervention to guarantee the quality of generated data, or fail to handle complex SQL queries.This paper presents a simple yet effective data augmentation framework.First, given a database, we automatically produce a large number of SQL queries based on an abstract syntax tree grammar.For better distribution matching, we require that at least 80% of SQL patterns in the training data are covered by generated queries.Second, we propose a hierarchical SQL-to-question generation model to obtain high-quality natural language questions, which is the major contribution of this work.Finally, we design a simple sampling strategy that can greatly improve training efficiency given large amounts of generated data.Experiments on three cross-domain datasets, i.e., WikiSQL and Spider in English, and DuSQL in Chinese, show that our proposed data augmentation framework can consistently improve performance over strong baselines, and the hierarchical generation component is the key for the improvement.
Kun Wu 0009, Zhenghua Li, Xinyan Xiao, Hua Wu 0003, Min Zhang 0005, Haifeng Wang 0001
EMNLP (1)3
2021 A Unified Span-Based Approach for Opinion Mining with Syntactic Constituents
abstract
Qingrong Xia, Bo Zhang, Rui Wang, Zhenghua Li, Yue Zhang, Fei Huang, Luo Si, Min Zhang. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Qingrong Xia, Bo Zhang 0071, Rui Wang 0005, Zhenghua Li, Yue Zhang 0004, Fei Huang 0002, Luo Si, Min Zhang 0005
NAACL-HLT4
2021 Dependency-based syntax-aware word representations
Meishan Zhang, Zhenghua Li, Guohong Fu, Min Zhang 0005
Artif. Intell.2
2021 Efficiency or Innovation?: The Long-Run Payoff of Cloud Computing
abstract
Considering the mixed arguments and uncertainty about the payoff of cloud computing, this paper empirically studies the long-term cloud computing impact on the financial performance, specifically from the perspective of efficiency and innovation. Taking 253 pairs of listed companies in China as the research sample, propensity score matching and difference in differences techniques combined with OLS regression are conducted to analyze a rolling 5-year panel data. The analysis results show that cloud computing adoption leads to years of financial performance decline followed by an upturn. The downward trend is more pronounced when it is adopted with innovation. This paper contributes to the existing literatures by leveraging archival performance data to verify the long-term business value and revealing the value realization difference between efficiency- and innovation-oriented cloud computing adoptions. The findings remind the managers to see the two sides of cloud computing and make rational adoption decisions, especially cloud-based innovation, according to their actual situations.
Zhenghua Li, Huigang Liang, Nianxin Wang, Yajiong Xue
J. Glob. Inf. Manag.1
2020 Efficient Second-Order TreeCRF for Neural Dependency Parsing
abstract
In the deep learning (DL) era, parsing models are extremely simplified with little hurt on performance, thanks to the remarkable capability of multi-layer BiLSTMs in context representation.As the most popular graphbased dependency parser due to its high efficiency and performance, the biaffine parser directly scores single dependencies under the arc-factorization assumption, and adopts a very simple local token-wise cross-entropy training loss.This paper for the first time presents a second-order TreeCRF extension to the biaffine parser.For a long time, the complexity and inefficiency of the inside-outside algorithm hinder the popularity of TreeCRF.To address this issue, we propose an effective way to batchify the inside and Viterbi algorithms for direct large matrix operation on GPUs, and to avoid the complex outside algorithm via efficient back-propagation.Experiments and analysis on 27 datasets from 13 languages clearly show that techniques developed before the DL era, such as structural learning (global TreeCRF loss) and high-order modeling are still useful, and can further boost parsing performance over the state-of-the-art biaffine parser, especially for partially annotated training data.We release our code at https: //github.com/yzhangcs/crfpar.
Yu Zhang 0092, Zhenghua Li, Min Zhang 0005
ACL2
2020 Syntax-Aware Opinion Role Labeling with Dependency Graph Convolutional Networks
abstract
Opinion role labeling (ORL) is a fine-grained opinion analysis task and aims to answer "who expressed what kind of sentiment towards what?".Due to the scarcity of labeled data, ORL remains challenging for data-driven methods.In this work, we try to enhance neural ORL models with syntactic knowledge by comparing and integrating different representations.We also propose dependency graph convolutional networks (DEPGCN) to encode parser information at different processing levels.In order to compensate for parser inaccuracy and reduce error propagation, we introduce multi-task learning (MTL) to train the parser and the ORL model simultaneously.We verify our methods on the benchmark MPQA corpus.The experimental results show that syntactic information is highly valuable for ORL, and our final MTL model effectively boosts the F1 score by 9.29 over the syntaxagnostic baseline.In addition, we find that the contributions from syntactic knowledge do not fully overlap with contextualized word representations (BERT).Our best model achieves 4.34 higher F1 score than the current state-ofthe-art.
Bo Zhang 0071, Yue Zhang 0004, Rui Wang 0005, Zhenghua Li, Min Zhang 0005
ACL4
2020 Multi-grained Chinese Word Segmentation with Weakly Labeled Data
abstract
In contrast with the traditional single-grained word segmentation (SWS), where a sentence corresponds to a single word sequence, multi-grained Chinese word segmentation (MWS) aims to segment a sentence into multiple word sequences to preserve all words of different granularities.Due to the lack of manually annotated MWS data, previous work train and tune MWS models only on automatically generated pseudo MWS data.In this work, we further take advantage of the rich word boundary information in existing SWS data and naturally annotated data from dictionary example (DictEx) sentences, to advance the state-of-the-art MWS model based on the idea of weak supervision.Particularly, we propose to accommodate two types of weakly labeled data for MWS, i.e., SWS data and DictEx data by employing a simple yet competitive graph-based parser with local loss.Besides, we manually annotate a high-quality MWS dataset according to our newly compiled annotation guideline, consisting of over 9,000 sentences from two types of texts, i.e., canonical newswire (NEWS) and non-canonical web (BAIKE) data for better evaluation.Detailed evaluation shows that our proposed model with weakly labeled data significantly outperforms the state-of-the-art MWS model by 1.12 and 5.97 on NEWS and BAIKE data in F1.
Chen Gong 0004, Zhenghua Li, Bowei Zou, Min Zhang 0005
COLING2
2020 Semi-supervised Domain Adaptation for Dependency Parsing via Improved Contextualized Word Representations
abstract
In recent years, parsing performance is dramatically improved on in-domain texts thanks to the rapid progress of deep neural network models.The major challenge for current parsing research is to improve parsing performance on out-of-domain texts that are very different from the indomain training data when there is only a small-scale out-domain labeled data.To deal with this problem, we propose to improve the contextualized word representations via adversarial learning and fine-tuning BERT processes.Concretely, we apply adversarial learning to three representative semi-supervised domain adaption methods, i.e., direct concatenation (CON), feature augmentation (FA), and domain embedding (DE) with two useful strategies, i.e., fused targetdomain word representations and orthogonality constraints, thus enabling to model more pure yet effective domain-specific and domain-invariant representations.Simultaneously, we utilize a large-scale target-domain unlabeled data to fine-tune BERT with only the language model loss, thus obtaining reliable contextualized word representations that benefit for the cross-domain dependency parsing.Experiments on a benchmark dataset show that our proposed adversarial approaches achieve consistent improvements, and fine-tuning BERT further boosts the parsing accuracy by a large margin.Our single model achieves the same state-of-the-art performance as the top submitted system in the NLPCC-2019 shared task, which uses ensemble models and BERT.
Ying Li 0065, Zhenghua Li, Min Zhang 0005
COLING2
2020 Semantic Role Labeling with Heterogeneous Syntactic Knowledge
abstract
Recently, due to the interplay between syntax and semantics, incorporating syntactic knowledge into neural semantic role labeling (SRL) has achieved much attention.Most of the previous syntax-aware SRL works focus on explicitly modeling homogeneous syntactic knowledge over tree outputs.In this work, we propose to encode heterogeneous syntactic knowledge for SRL from both explicit and implicit representations.First, we introduce graph convolutional networks to explicitly encode multiple heterogeneous dependency parse trees.Second, we extract the implicit syntactic representations from syntactic parser trained with heterogeneous treebanks.Finally, we inject the two types of heterogeneous syntax-aware representations into the base SRL model as extra inputs.We conduct experiments on two widely-used benchmark datasets, i.e., Chinese Proposition Bank 1.0 and English CoNLL-2005 dataset.Experimental results show that incorporating heterogeneous syntactic knowledge brings significant improvements over strong baselines.We further conduct detailed analysis to gain insights on the usefulness of heterogeneous (vs.homogeneous) syntactic knowledge and the effectiveness of our proposed approaches for modeling such knowledge.
Qingrong Xia, Rui Wang 0005, Zhenghua Li, Yue Zhang 0004, Min Zhang 0005
COLING3
2020 DuSQL: A Large-Scale and Pragmatic Chinese Text-to-SQL Dataset
abstract
Due to the lack of labeled data, previous research on text-to-SQL parsing mainly focuses on English.Representative English datasets include ATIS, WikiSQL, Spider, etc.This paper presents DuSQL, a larges-scale and pragmatic Chinese dataset for the cross-domain text-to-SQL task, containing 200 databases, 813 tables, and 23,797 question/SQL pairs.Our new dataset has three major characteristics.First, by manually analyzing questions from several representative applications, we try to figure out the true distribution of SQL queries in real-life needs.Second, DuSQL contains a considerable proportion of SQL queries involving row or column calculations, motivated by our analysis on the SQL query distributions.Finally, we adopt an effective data construction framework via human-computer collaboration.The basic idea is automatically generating SQL queries based on the SQL grammar and constrained by the given database.This paper describes in detail the construction process and data statistics of DuSQL.Moreover, we present and compare performance of several open-source textto-SQL parsers with minor modification to accommodate Chinese, including a simple yet effective extension to IRNet for handling calculation SQL queries.
Kun Wu 0009, Ke Sun 0005, Zhenghua Li, Hua Wu 0003, Min Zhang 0005, Haifeng Wang 0001
EMNLP (1)5
2020 Fast and Accurate Neural CRF Constituency Parsing
abstract
Estimating probability distribution is one of the core issues in the NLP field. However, in both deep learning (DL) and pre-DL eras, unlike the vast applications of linear-chain CRF in sequence labeling tasks, very few works have applied tree-structure CRF to constituency parsing, mainly due to the complexity and inefficiency of the inside-outside algorithm. This work presents a fast and accurate neural CRF constituency parser. The key idea is to batchify the inside algorithm for loss computation by direct large tensor operations on GPU, and meanwhile avoid the outside algorithm for gradient computation via efficient back-propagation. We also propose a simple two-stage bracketing-then-labeling parsing approach to improve efficiency further. To improve the parsing performance, inspired by recent progress in dependency parsing, we introduce a new scoring architecture based on boundary representation and biaffine attention, and a beneficial dropout strategy. Experiments on PTB, CTB5.1, and CTB7 show that our two-stage CRF parser achieves new state-of-the-art performance on both settings of w/o and w/ BERT, and can parse over 1,000 sentences per second. We release our code at https://github.com/yzhangcs/crfpar.
Yu Zhang 0092, Houquan Zhou 0001, Zhenghua Li
IJCAI3
2020 Dependency Parsing with Noisy Multi-annotation Data
Yu Zhao 0043, Mingyue Zhou, Zhenghua Li, Min Zhang 0005
NLPCC (2)3
2020 Is POS Tagging Necessary or Even Helpful for Neural Dependency Parsing?
Houquan Zhou 0001, Yu Zhang 0092, Zhenghua Li, Min Zhang 0005
NLPCC (1)3
2020 Hierarchical LSTM with char-subword-word tree-structure representation for Chinese named entity recognition
Chen Gong 0004, Zhenghua Li, Qingrong Xia, Wenliang Chen, Min Zhang 0005
Sci. China Inf. Sci.2
2019 Syntax-Aware Neural Semantic Role Labeling
abstract
Semantic role labeling (SRL), also known as shallow semantic parsing, is an important yet challenging task in NLP. Motivated by the close correlation between syntactic and semantic structures, traditional discrete-feature-based SRL approaches make heavy use of syntactic features. In contrast, deep-neural-network-based approaches usually encode the input sentence as a word sequence without considering the syntactic structures. In this work, we investigate several previous approaches for encoding syntactic trees, and make a thorough study on whether extra syntax-aware representations are beneficial for neural SRL models. Experiments on the benchmark CoNLL-2005 dataset show that syntax-aware SRL approaches can effectively improve performance over a strong baseline with external word representations from ELMo. With the extra syntax-aware representations, our approaches achieve new state-of-the-art 85.6 F1 (single model) and 86.6 F1 (ensemble) on the test data, outperforming the corresponding strong baselines with ELMo by 0.8 and 1.0, respectively. Detailed error analysis are conducted to gain more insights on the investigated approaches.
Qingrong Xia, Zhenghua Li, Min Zhang 0005, Meishan Zhang, Guohong Fu, Rui Wang 0005, Luo Si
AAAI2
2019 Semi-supervised Domain Adaptation for Dependency Parsing
abstract
During the past decades, due to the lack of sufficient labeled data, most studies on crossdomain parsing focus on unsupervised domain adaptation, assuming there is no targetdomain training data.However, unsupervised approaches make limited progress so far due to the intrinsic difficulty of both domain adaptation and parsing.This paper tackles the semi-supervised domain adaptation problem for Chinese dependency parsing, based on two newly-annotated large-scale domain-specific datasets.1 We propose a simple domain embedding approach to merge the sourceand target-domain training data, which is shown to be more effective than both direct corpus concatenation and multi-task learning.In order to utilize unlabeled target-domain data, we employ the recent contextualized word representations and show that a simple fine-tuning procedure can further boost cross-domain parsing accuracy by large margins.
Zhenghua Li, Xue Peng, Min Zhang 0005, Rui Wang 0005, Luo Si
ACL (1)1
2019 A Syntax-aware Multi-task Learning Framework for Chinese Semantic Role Labeling
abstract
Qingrong Xia, Zhenghua Li, Min Zhang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Qingrong Xia, Zhenghua Li, Min Zhang 0005
EMNLP/IJCNLP (1)2
2019 Self-attentive Biaffine Dependency Parsing
abstract
The current state-of-the-art dependency parsing approaches employ BiLSTMs to encode input sentences.Motivated by the success of the transformer-based machine translation, this work for the first time applies the self-attention mechanism to dependency parsing as the replacement of the BiLSTM-based encoders, leading to competitive performance on both English and Chinese benchmark data. Based on the detailed error analysis, we then combine the power of both BiLSTM and self-attention via model ensembles, demonstrating their complementary capability of capturing contextual information. Finally, we explore the recently proposed contextualized word representations as extra input features, and further improve the parsing performance.
Ying Li 0065, Zhenghua Li, Min Zhang 0005, Rui Wang 0005, Sheng Li 0017, Luo Si
IJCAI2
2019 Overview of the NLPCC 2019 Shared Task: Cross-Domain Dependency Parsing
Xue Peng, Zhenghua Li, Min Zhang 0005, Rui Wang 0005, Yue Zhang 0004, Luo Si
NLPCC (2)2
2019 Conversion and Exploitation of Dependency Treebanks with Full-Tree LSTM
Bo Zhang 0071, Zhenghua Li, Min Zhang 0005
NLPCC (2)2
2019 Syntax-aware entity representations for neural relation extraction
Zhengqiu He, Wenliang Chen, Zhenghua Li, Wei Zhang 0027, Hao Shao, Min Zhang 0005
Artif. Intell.3
2018 SEE: Syntax-Aware Entity Embedding for Neural Relation Extraction
abstract
Distant supervised relation extraction is an efficient approach to scale relation extraction to very large corpora, and has been widely used to find novel relational facts from plain text. Recent studies on neural relation extraction have shown great progress on this task via modeling the sentences in low-dimensional spaces, but seldom considered syntax information to model the entities. In this paper, we propose to learn syntax-aware entity embedding for neural relation extraction. First, we encode the context of entities on a dependency tree as sentence-level entity embedding based on tree-GRU. Then, we utilize both intra-sentence and inter-sentence attentions to obtain sentence set-level entity embedding over all sentences containing the focus entity pair. Finally, we combine both sentence embedding and entity embedding for relation classification. We conduct experiments on a widely used real-world dataset and the experimental results show that our model can make full use of all informative instances and achieve state-of-the-art performance of relation extraction.
Zhengqiu He, Wenliang Chen, Zhenghua Li, Meishan Zhang, Wei Zhang 0027, Min Zhang 0005
AAAI3
2018 Supervised Treebank Conversion: Data and Approaches
abstract
Treebank conversion is a straightforward and effective way to exploit various heterogeneous treebanks for boosting parsing accuracy.However, previous work mainly focuses on unsupervised treebank conversion and makes little progress due to the lack of manually labeled data where each sentence has two syntactic trees complying with two different guidelines at the same time, referred as bi-tree aligned data.In this work, we for the first time propose the task of supervised treebank conversion.First, we manually construct a bi-tree aligned dataset containing over ten thousand sentences.Then, we propose two simple yet effective treebank conversion approaches (pattern embedding and treeLSTM) based on the state-of-the-art deep biaffine parser.Experimental results show that 1) the two approaches achieve comparable conversion accuracy, and 2) treebank conversion is superior to the widely used multi-task learning framework in multiple treebank exploitation and leads to significantly higher parsing accuracy.* The first two (student) authors make equal contributions to this work.Zhenghua is the correspondence author.
Xinzhou Jiang, Zhenghua Li, Bo Zhang 0071, Min Zhang 0005, Sheng Li 0017, Luo Si
ACL (1)2
2018 Distantly Supervised NER with Partial Annotation Learning and Reinforcement Learning
abstract
A bottleneck problem with Chinese named entity recognition (NER) in new domains is the lack of annotated data. One solution is to utilize the method of distant supervision, which has been widely used in relation extraction, to automatically populate annotated training data without humancost. The distant supervision assumption here is that if a string in text is included in a predefined dictionary of entities, the string might be an entity. However, this kind of auto-generated data suffers from two main problems: incomplete and noisy annotations, which affect the performance of NER models. In this paper, we propose a novel approach which can partially solve the above problems of distant supervision for NER. In our approach, to handle the incomplete problem, we apply partial annotation learning to reduce the effect of unknown labels of characters. As for noisy annotation, we design an instance selector based on reinforcement learning to distinguish positive sentences from auto-generated annotations. In experiments, we create two datasets for Chinese named entity recognition in two domains with the help of distant supervision. The experimental results show that the proposed approach obtains better performance than the comparison systems on both two datasets.
YaoSheng Yang, Wenliang Chen, Zhenghua Li, Zhengqiu He, Min Zhang 0005
COLING3
2018 M-CNER: A Corpus for Chinese Named Entity Recognition in Multi-Domains
YaoSheng Yang, Zhenghua Li, Wenliang Chen, Min Zhang 0005
LREC3
2017 Multi-Grained Chinese Word Segmentation
abstract
Traditionally, word segmentation (WS) adopts the single-granularity formalism, where a sentence corresponds to a single word sequence.However, Sproat et al. (1996) show that the inter-nativespeaker consistency ratio over Chinese word boundaries is only 76%, indicating single-grained WS (SWS) imposes unnecessary challenges on both manual annotation and statistical modeling.Moreover, WS results of different granularities can be complementary and beneficial for high-level applications.This work proposes and addresses multi-grained WS (MWS).First, we build a large-scale pseudo MWS dataset for model training and tuning by leveraging the annotation heterogeneity of three SWS datasets.Then we manually annotate 1,500 test sentences with true MWS annotations.Finally, we propose three benchmark approaches by casting MWS as constituent parsing and sequence labeling.Experiments and analysis lead to many interesting findings.
Chen Gong 0004, Zhenghua Li, Min Zhang 0005, Xinzhou Jiang
EMNLP2
2017 Dependency Parsing with Partial Annotations: An Empirical Comparison
abstract
This paper describes and compares two straightforward approaches for dependency parsing with partial annotations (PA). The first approach is based on a forest-based training objective for two CRF parsers, i.e., a biaffine neural network graph-based parser (Biaffine) and a traditional log-linear graph-based parser (LLGPar). The second approach is based on the idea of constrained decoding for three parsers, i.e., a traditional linear graph-based parser (LGPar), a globally normalized neural network transition-based parser (GN3Par) and a traditional linear transition-based parser (LTPar). For the test phase, constrained decoding is also used for completing partial trees. We conduct experiments on Penn Treebank under three different settings for simulating PA, i.e., random, most uncertain, and divergent outputs from the five parsers. The results show that LLGPar is most effective in directly learning from PA, and other parsers can achieve best performance when PAs are completed into full trees by LLGPar.
Yue Zhang 0004, Zhenghua Li, Jun Lang 0001, Qingrong Xia, Min Zhang 0005
IJCNLP(1)2
2017 Coupled POS Tagging on Heterogeneous Annotations
abstract
The limited scale and genre coverage of labeled data greatly hinders the effectiveness of supervised models, especially when analyzing spoken languages, such as texts transcribed from speech and informal text including tweets and product comments in Internet. In order to effectively utilize multiple labeled datasets with heterogeneous annotations for the same task, this paper proposes a coupled sequence labeling model that can directly learn and infer two heterogeneous annotations simultaneously, using Chinese part-of-speech (POS) tagging as our case study. The key idea is to bundle two sets of POS tags together (e.g., “[NN, n]n), and build a conditional random field (CRF) based tagging model in the enlarged space of bundled tags with the help of ambiguous labeling. To train our model on two nonoverlapping datasets that each has only one-side tags, we transform a one-side tag into a set of bundled tags by concatenating the tag with every possible tag at the missing side according to a predefined context-free tag-to-tag mapping function, thus producing ambiguous labeling as weak supervision. We design and investigate four different context-free tag-to-tag mapping functions, and find out that the coupled model achieves its best performance when each one-side tag is mapped to all tags at the other side (namely complete mapping), indicating that the model can effectively learn the loose mapping between the two heterogeneous annotations, without the need of manually designed mapping rules. Moreover, we propose a context-aware online pruning strategy that can more accurately capture mapping relationships between annotations based on contextual evidences and thus effectively solve the severe inefficiency problem with our coupled model under complete mapping, making it comparable with the baseline CRF model. Experiments on benchmark datasets show that our coupled model significantly outperforms the state-of-the-art baselines on both one-side POS tagging and annotation conversion tasks. The codes and newly annotated data are released for research usage.1
Zhenghua Li, Jiayuan Chao, Min Zhang 0005, Wenliang Chen, Meishan Zhang, Guohong Fu
IEEE ACM Trans. Audio Speech Lang. Process.1
2016 Active Learning for Dependency Parsing with Partial Annotation
abstract
Zhenghua Li, Min Zhang, Yue Zhang, Zhanyi Liu, Wenliang Chen, Hua Wu, Haifeng Wang. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2016.
Zhenghua Li, Min Zhang 0005, Yue Zhang 0004, Zhanyi Liu, Wenliang Chen, Hua Wu 0003, Haifeng Wang 0001
ACL (1)1
2016 Distributed Representations for Building Profiles of Users and Items from Text Reviews
abstract
In this paper, we propose an approach to learn distributed representations of users and items from text comments for recommendation systems. Traditional recommendation algorithms, e.g. collaborative filtering and matrix completion, are not designed to exploit the key information hidden in the text comments, while existing opinion mining methods do not provide direct support to recommendation systems with useful features on users and items. Our approach attempts to construct vectors to represent profiles of users and items under a unified framework to maximize word appearance likelihood. Then, the vector representations are used for a recommendation task in which we predict scores on unobserved user-item pairs without given texts. The recommendation-aware distributed representation approach is fully supported by effective and efficient learning algorithms over massive text archive. Our empirical evaluations on real datasets show that our system outperforms the state-of-the-art baseline systems.
Wenliang Chen, Zhenghua Li, Min Zhang 0005
COLING3
2016 Fast Coupled Sequence Labeling on Heterogeneous Annotations via Context-aware Pruning
abstract
The recently proposed coupled sequence labeling is shown to be able to effectively exploit multiple labeled data with heterogeneous annotations but suffer from severe inefficiency problem due to the large bundled tag space (Li et al., 2015).In their case study of part-ofspeech (POS) tagging, Li et al. (2015) manually design context-free tag-to-tag mapping rules with a lot of effort to reduce the tag space.This paper proposes a context-aware pruning approach that performs token-wise constraints on the tag space based on contextual evidences, making the coupled approach efficient enough to be applied to the more complex task of joint word segmentation (WS) and POS tagging for the first time.Experiments show that using the large-scale People Daily as auxiliary heterogeneous data, the coupled approach can improve F-score by 95.55 -94.88 = 0.67% on WS, and by 90.58 -89.49= 1.09% on joint WS&POS on Penn Chinese Treebank.All codes are released at http://hlt.suda.edu.cn/~zhli.
Zhenghua Li, Jiayuan Chao, Min Zhang 0005, Jiwen Yang
EMNLP1
2015 Coupled Sequence Labeling on Heterogeneous Annotations: POS Tagging as a Case Study
abstract
Zhenghua Li, Jiayuan Chao, Min Zhang, Wenliang Chen. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Zhenghua Li, Jiayuan Chao, Min Zhang 0005, Wenliang Chen
ACL (1)1
2015 Exploiting Heterogeneous Annotations for Weibo Word Segmentation and POS Tagging
abstract
This paper describes our system designed for the NLPCC 2015 shared task on Chinese word segmentation (WS) and POS tagging for Weibo Text. We treat WS and POS tagging as two separate tasks and use a cascaded approach. Our major focus is how to effectively exploit multiple heterogeneous data to boost performance of statistical models. This work considers three sets of heterogeneous data, i.e., Weibo ( \(\textit{WB}\) , 10K sentences), Penn Chinese Treebank 7.0 ( \(\textit{CTB7}\) , 50K), and People’s Daily ( \(\textit{PD}\) , 280K). For WS, we adopt the recently proposed coupled sequence labeling to combine \(\textit{WB}\) , \(\textit{CTB7}\) , and \(\textit{PD}\) , boosting F1 score from \(93.76\%\) (baseline model trained on only \(\textit{WB}\) ) to \(95.58\%\) ( \(+1.82\%\) ). For POS tagging, we adopt an ensemble approach combining coupled sequence labeling and the guide-feature based method, since the three datasets have three different annotation standards. First, we convert \(\textit{PD}\) into the annotation style of \(\textit{CTB7}\) based on coupled sequence labeling, denoted by \(\textit{PD}^{\textit{CTB}}\) . Then, we merge CTB 7 and \(\textit{PD}^{\textit{CTB}}\) to train a POS tagger, denoted by \(\textit{Tag}_{\textit{CTB7}+\textit{PD}^{\textit{CTB}}}\) , which is further used to produce guide features on \(\textit{WB}\) . Finally, the tagging F1 score is improved from 87.93% to 88.99% (+1.06%).
Jiayuan Chao, Zhenghua Li, Wenliang Chen, Min Zhang 0005
NLPCC2
2014 Ambiguity-aware Ensemble Training for Semi-supervised Dependency Parsing
abstract
This paper proposes a simple yet effective framework for semi-supervised dependency parsing at entire tree level, referred to as ambiguity-aware ensemble training.Instead of only using 1best parse trees in previous work, our core idea is to utilize parse forest (ambiguous labelings) to combine multiple 1-best parse trees generated from diverse parsers on unlabeled data.With a conditional random field based probabilistic dependency parser, our training objective is to maximize mixed likelihood of labeled data and auto-parsed unlabeled data with ambiguous labelings.This framework offers two promising advantages. 1) ambiguity encoded in parse forests compromises noise in 1-best parse trees.During training, the parser is aware of these ambiguous structures, and has the flexibility to distribute probability mass to its preferred parse trees as long as the likelihood improves.2) diverse syntactic structures produced by different parsers can be naturally compiled into forest, offering complementary strength to our single-view parser.Experimental results on benchmark data show that our method significantly outperforms the baseline supervised parser and other entire-tree based semi-supervised methods, such as self-training, co-training and tri-training.
Zhenghua Li, Min Zhang 0005, Wenliang Chen
ACL (1)1
2014 Soft Cross-lingual Syntax Projection for Dependency Parsing
Zhenghua Li, Min Zhang 0005, Wenliang Chen
COLING1
2014 Joint Optimization for Chinese POS Tagging and Dependency Parsing
abstract
Dependency parsing has gained more and more interest in natural language processing in recent years due to its simplicity and general applicability for diverse languages. Previous work demonstrates that part-of-speech (POS) is an indispensable feature in dependency parsing since pure lexical features suffer from serious data sparseness problem. However, due to little morphological changes, Chinese POS tagging has proven to be much more challenging than morphology-richer languages such as English (94% vs. 97% on POS tagging accuracy). This leads to severe error propagation for Chinese dependency parsing. Our experiments show that parsing accuracy drops by about 6% when replacing manual POS tags of the input sentence with automatic ones generated by a state-of-the-art statistical POS tagger. To address this issue, this paper proposes a solution by jointly optimizing POS tagging and dependency parsing in a unique model. We propose for our joint models several dynamic programming based decoding algorithms which can incorporate rich POS tagging and syntactic features. Then we present an effective pruning strategy to reduce the search space of candidate POS tags, leading to significant improvement of parsing speed. Experimental results on two Chinese data sets, i.e. Penn Chinese Treebank 5.1 and Penn Chinese Treebank 7, demonstrate that our joint models significantly improve both the state-of-the-art tagging and parsing accuracies. Detailed analysis shows that the joint method can help resolve syntax-sensitive POS ambiguities$\{{\ssr{NN}},{\ssr{VV}}\}$. In return, the POS tags become more reliable and helpful for parsing since the syntactic features are used in POS tagging. This is the fundamental reason for the performance improvement.
Zhenghua Li, Min Zhang 0005, Wanxiang Che, Ting Liu 0001, Wenliang Chen
IEEE ACM Trans. Audio Speech Lang. Process.1
2013 Learning Sentence Representation for Emotion Classification on Microblogs
Duyu Tang, Bing Qin 0001, Ting Liu 0001, Zhenghua Li
NLPCC4
2013 Fast data in the era of big data: Twitter's real-time related query suggestion architecture
abstract
We present the architecture behind Twitter's real-time related query suggestion and spelling correction service. Although these tasks have received much attention in the web search literature, the Twitter context introduces a real-time "twist": after significant breaking news events, we aim to provide relevant results within minutes. This paper provides a case study illustrating the challenges of real-time data processing in the era of "big data". We tell the story of how our system was built twice: our first implementation was built on a typical Hadoop-based analytics stack, but was later replaced because it did not meet the latency requirements necessary to generate meaningful real-time results. The second implementation, which is the system deployed in production today, is a custom in-memory processing engine specifically designed for the task. This experience taught us that the current typical usage of Hadoop as a "big data" platform, while great for experimentation, is not well suited to low-latency processing, and points the way to future work on data analytics platforms that can handle "big" as well as "fast" data.
Gilad Mishne, Jeff Dalton 0001, Zhenghua Li, Aneesh Sharma, Jimmy Lin
SIGMOD Conference3
2012 Exploiting Multiple Treebanks for Parsing with Quasi-synchronous Grammars
Zhenghua Li, Ting Liu 0001, Wanxiang Che
ACL (1)1
2012 A Separately Passive-Aggressive Training Algorithm for Joint POS Tagging and Dependency Parsing
Zhenghua Li, Min Zhang 0005, Wanxiang Che, Ting Liu 0001
COLING1
2012 Stacking Heterogeneous Joint Models of Chinese POS Tagging and Dependency Parsing
Meishan Zhang, Wanxiang Che, Ting Liu 0001, Zhenghua Li
COLING4
2011 Joint Models for Chinese POS Tagging and Dependency Parsing
Zhenghua Li, Min Zhang 0005, Wanxiang Che, Ting Liu 0001, Wenliang Chen, Haizhou Li 0001
EMNLP1
2011 Improving Chinese POS Tagging with Dependency Parsing
Zhenghua Li, Wanxiang Che, Ting Liu 0001
IJCNLP1
2008 A Cascaded Syntactic and Semantic Dependency Parsing System
Wanxiang Che, Zhenghua Li, Yuxuan Hu 0001, Bing Qin 0001, Ting Liu 0001, Sheng Li 0003
CoNLL2