EDBT 2026 Demo / reviewers in the wild / expert
Zhisong Zhang
dblp:174/7415
· DBLP profile ↗
31ranked-venue papers
11as first author
18since 2021 · last 2026
0009-0005-2108-1246ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 31 · 11 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | WebAggregator: Enhancing Compositional Reasoning Capabilities of Deep Research Agent Foundation ModelsabstractRui Wang, Ce Zhang, Jun-Yu Ma, Jianshu Zhang, Hongru Wang, Yi Chen, Boyang Xue, Tianqing Fang, Zhisong Zhang, Hongming Zhang, Haitao Mi, Dong Yu, Kam-Fai Wong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Rui Wang 0015, Ce Zhang 0009, Jun-Yu Ma, Hongru Wang 0003, Yi Chen 0007, Boyang Xue, Tianqing Fang, Zhisong Zhang, Hongming Zhang 0009, Haitao Mi, Dong Yu 0001, Kam-Fai Wong |
ACL (1) | 9 |
| 2025 | A Silver Bullet or a Compromise for Full Attention? A Comprehensive Study of Gist Token-based Context CompressionabstractChenlong Deng, Zhisong Zhang, Kelong Mao, Shuaiyi Li, Xinting Huang, Dong Yu, Zhicheng Dou. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Chenlong Deng, Zhisong Zhang, Kelong Mao, Shuaiyi Li, Xinting Huang, Dong Yu 0001, Zhicheng Dou |
ACL (1) | 2 |
| 2025 | Low-Bit Quantization Favors Undertrained LLMsabstractLow-bit quantization improves machine learning model efficiency but surprisingly favors undertrained large language models (LLMs).Larger models or those trained on fewer tokens exhibit less quantization-induced degradation (QiD), while smaller, well-trained models face significant performance losses.To gain deeper insights into this trend, we study over 1500+ quantized LLM checkpoints of various sizes and at different training levels (undertrained or fully trained) in a controlled setting, deriving scaling laws for understanding the relationship between QiD and factors: the number of training tokens, model size and bit width.With our derived scaling laws, we propose a novel perspective that we can use QiD to measure an LLM's training levels and determine the number of training tokens required for fully training LLMs of various sizes.Moreover, we use the scaling laws to predict the quantization performance of different-sized LLMs trained with 100 trillion tokens.Our projection shows that the low-bit quantization performance of future models, which are expected to be trained with over 100 trillion tokens, may NOT be desirable.This poses a potential challenge for low-bit quantization in the future and highlights the need for awareness of a model's training level when evaluating lowbit quantization research.To facilitate future research on this problem, we release all the 1500+ quantized checkpoints used in this work at https://huggingface.co/Xu-Ouyang. Xu Ouyang, Tao Ge 0001, Thomas Hartvigsen, Zhisong Zhang, Haitao Mi, Dong Yu 0001 |
ACL (1) | 4 |
| 2025 | LoGU: Long-form Generation with Uncertainty ExpressionsabstractRuihan Yang, Caiqi Zhang, Zhisong Zhang, Xinting Huang, Sen Yang, Nigel Collier, Dong Yu, Deqing Yang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Ruihan Yang, Caiqi Zhang, Zhisong Zhang, Xinting Huang, Nigel Collier, Dong Yu 0001, Deqing Yang |
ACL (1) | 3 |
| 2025 | Attention Entropy is a Key Factor: An Analysis of Parallel Context Encoding with Full-attention-based Pre-trained Language ModelsabstractLarge language models have shown remarkable performance across a wide range of language tasks, owing to their exceptional capabilities in context modeling. The most commonly used method of context modeling is full self-attention, as seen in standard decoder-only Transformers. Although powerful, this method can be inefficient for long sequences and may overlook inherent input structures. To address these problems, an alternative approach is parallel context encoding, which splits the context into sub-pieces and encodes them parallelly. Because parallel patterns are not encountered during training, naively applying parallel encoding leads to performance degradation. However, the underlying reasons and potential mitigations are unclear. In this work, we provide a detailed analysis of this issue and identify that unusually high attention entropy can be a key factor. Furthermore, we adopt two straightforward methods to reduce attention entropy by incorporating attention sinks and selective mechanisms. Experiments on various tasks reveal that these methods effectively lower irregular attention entropy and narrow performance gaps. We hope this study can illuminate ways to enhance context modeling mechanisms. Zhisong Zhang, Yan Wang 0060, Xinting Huang, Tianqing Fang, Hongming Zhang 0009, Chenlong Deng, Shuaiyi Li, Dong Yu 0001 |
ACL (1) | 1 |
| 2025 | WebEvolver: Enhancing Web Agent Self-Improvement with Co-evolving World ModelabstractAgent self-improvement, where agents autonomously train their underlying Large Language Model (LLM) on self-sampled trajectories, shows promising results but often stagnates in web environments due to limited exploration and under-utilization of pretrained web knowledge.To improve the performance of self-improvement, we propose a novel framework that introduces a co-evolving World Model LLM.This world model predicts the next observation based on the current observation and action within the web environment.The World Model serves dual roles: (1) as a virtual web server generating self-instructed training data to continuously refine the agent's policy, and (2) as an imagination engine during inference, enabling look-ahead simulation to guide action selection for the agent LLM.Experiments in real-world web environments (Mind2Web-Live, WebVoyager, and GAIAweb) show a 10% performance gain over existing self-evolving agents, demonstrating the efficacy and generalizability of our approach, without using any distillation from more powerful close-sourced models 1 . Tianqing Fang, Hongming Zhang 0009, Zhisong Zhang, Kaixin Ma, Wenhao Yu 0002, Haitao Mi, Dong Yu 0001 |
EMNLP | 3 |
| 2025 | Recall with Reasoning: Chain-of-Thought Distillation for Mamba's Long-Context Memory and ExtrapolationabstractMamba's theoretical infinite-context potential is limited in practice when sequences far exceed training lengths.This work explores unlocking Mamba's long-context memory ability by a simple-yet-effective method, Recall with Reasoning (RwR), by distilling chain-ofthought (CoT) summarization from a teacher model.Specifically, RwR prepends these summarization as CoT prompts during fine-tuning, teaching Mamba to actively recall and reason over long contexts.Experiments on LONG-MEMEVAL and HELMET show that RwR outperforms existing long-term memory methods on the Mamba model.Furthermore, under similar pre-training conditions, RwR improves the long-context performance of Mamba relative to comparable Transformer/hybrid baselines while preserving short-context capabilities, all without changing the architecture. Jun-Yu Ma, Tianqing Fang, Zhisong Zhang, Hongming Zhang 0009, Haitao Mi, Dong Yu 0001 |
EMNLP | 3 |
| 2025 | UNCLE: Benchmarking Uncertainty Expressions in Long-Form GenerationabstractLarge Language Models (LLMs) are prone to hallucination, particularly in long-form generations.A promising direction to mitigate hallucination is to teach LLMs to express uncertainty explicitly when they lack sufficient knowledge.However, existing work lacks direct and fair evaluation of LLMs' ability to express uncertainty effectively in long-form generation.To address this gap, we first introduce UNCLE, a benchmark designed to evaluate uncertainty expression in both long-and short-form question answering (QA).UNCLE covers five domains and includes more than 1,000 entities, each with paired short-and long-form QA items.Our dataset is the first to directly link short-and long-form QA through aligned questions and gold-standard answers.Along with UNCLE, we propose a suite of new metrics to assess the models' capabilities to selectively express uncertainty.We then demonstrate that current models fail to convey uncertainty appropriately in long-form generation.We further explore both prompt-based and training-based methods to improve models' performance, with the training-based methods yielding greater gains.Further analysis of alignment gaps between short-and long-form uncertainty expression highlights promising directions for future research using UNCLE. Ruihan Yang, Caiqi Zhang, Zhisong Zhang, Xinting Huang, Dong Yu 0001, Nigel Collier, Deqing Yang |
EMNLP | 3 |
| 2025 | UniGist: Towards General and Hardware-aligned Sequence-level Long Context CompressionabstractLarge language models are increasingly capable of handling long-context inputs, but the memory overhead of KV cache remains a major bottleneck for general-purpose deployment. While many compression strategies have been explored, sequence-level compression is particularly challenging due to its tendency to lose important details. We present UniGist, a gist token-based long context compression framework that removes the need for chunk-wise training, enabling the model to learn how to compress and utilize long-range context during training. To fully exploit the sparsity, we introduce a gist shift trick that transforms the attention layout into a right-aligned block structure and develop a block-table-free sparse attention kernel based on it. UniGist further supports one-pass training and flexible chunk sizes during inference, allowing efficient and adaptive context processing. Experiments across multiple long-context tasks show that UniGist significantly improves compression quality, with especially strong performance in recalling details and long-range dependency modeling. Chenlong Deng, Zhisong Zhang, Kelong Mao, Shuaiyi Li, Tianqing Fang, Hongming Zhang 0009, Haitao Mi, Dong Yu 0001, Zhicheng Dou |
NeurIPS | 2 |
| 2024 | A Thorough Examination of Decoding Methods in the Era of LLMsabstractDecoding methods play an indispensable role in converting language models from next-token predictors into practical task solvers.Prior research on decoding methods, primarily focusing on task-specific models, may not extend to the current era of general-purpose large language models (LLMs).Moreover, the recent influx of decoding strategies has further complicated this landscape.This paper provides a comprehensive and multifaceted analysis of various decoding methods within the context of LLMs, evaluating their performance, robustness to hyperparameter changes, and decoding speeds across a wide range of tasks, models, and deployment environments.Our findings reveal that decoding method performance is notably task-dependent and influenced by factors such as alignment, model size, and quantization.Intriguingly, sensitivity analysis exposes that certain methods achieve superior performance at the cost of extensive hyperparameter tuning, highlighting the trade-off between attaining optimal results and the practicality of implementation in varying contexts. Chufan Shi, Deng Cai 0002, Zhisong Zhang, Yujiu Yang 0001, Wai Lam |
EMNLP | 4 |
| 2024 | On the Worst Prompt Performance of Large Language ModelsabstractThe performance of large language models (LLMs) is acutely sensitive to the phrasing of prompts, which raises significant concerns about their reliability in real-world scenarios. Existing studies often divide prompts into task-level instructions and case-level inputs and primarily focus on evaluating and improving robustness against variations in tasks-level instructions. However, this setup fails to fully address the diversity of real-world user queries and assumes the existence of task-specific datasets. To address these limitations, we introduce RobustAlpacaEval, a new benchmark that consists of semantically equivalent case-level queries and emphasizes the importance of using the worst prompt performance to gauge the lower bound of model performance. Extensive experiments on RobustAlpacaEval with ChatGPT and six open-source LLMs from the Llama, Mistral, and Gemma families uncover substantial variability in model performance; for instance, a difference of 45.48% between the worst and best performance for the Llama-2-70B-chat model, with its worst performance dipping as low as 9.38%. We further illustrate the difficulty in identifying the worst prompt from both model-agnostic and model-dependent perspectives, emphasizing the absence of a shortcut to characterize the worst prompt. We also attempt to enhance the worst prompt performance using existing prompt engineering and prompt consistency methods, but find that their impact is limited. These findings underscore the need to create more resilient LLMs that can maintain high performance across diverse prompts. Bowen Cao, Deng Cai 0002, Zhisong Zhang, Yuexian Zou, Wai Lam |
NeurIPS | 3 |
| 2024 | Self-playing Adversarial Language Game Enhances LLM ReasoningabstractWe explore the potential of self-play training for large language models (LLMs) in a two-player adversarial language game called Adversarial Taboo. In this game, an attacker and a defender communicate around a target word only visible to the attacker. The attacker aims to induce the defender to speak the target word unconsciously, while the defender tries to infer the target word from the attacker's utterances. To win the game, both players must have sufficient knowledge about the target word and high-level reasoning ability to infer and express in this information-reserved conversation. Hence, we are curious about whether LLMs' reasoning ability can be further enhanced by Self-Playing this Adversarial language Game (SPAG). With this goal, we select several open-source LLMs and let each act as the attacker and play with a copy of itself as the defender on an extensive range of target words. Through reinforcement learning on the game outcomes, we observe that the LLMs' performances uniformly improve on a broad range of reasoning benchmarks. Furthermore, iteratively adopting this self-play process can continuously promote LLMs' reasoning abilities. The code is available at https://github.com/Linear95/SPAG. Pengyu Cheng, Tianhao Hu, Zhisong Zhang, Yong Dai 0001 |
NeurIPS | 4 |
| 2023 | Towards More Efficient Insertion Transformer with Fractional Positional EncodingabstractAuto-regressive neural sequence models have been shown to be effective across text generation tasks.However, their left-to-right decoding order prevents generation from being parallelized.Insertion Transformer (Stern et al., 2019) is an attractive alternative that allows outputting multiple tokens in a single generation step.Nevertheless, due to the incompatibility between absolute positional encoding and insertion-based generation schemes, it needs to refresh the encoding of every token in the generated partial hypothesis at each step, which could be costly.We design a novel reusable positional encoding scheme for Insertion Transformers called Fractional Positional Encoding (FPE), which allows reusing representations calculated in previous steps.Empirical studies on various text generation tasks demonstrate the effectiveness of FPE, which leads to floating-point operation reduction and latency improvements on batched decoding. Zhisong Zhang, Yizhe Zhang 0002, William B. Dolan |
EACL | 1 |
| 2022 | Transfer Learning from Semantic Role Labeling to Event Argument Extraction with Template-based Slot QueryingabstractIn this work, we investigate transfer learning from semantic role labeling (SRL) to event argument extraction (EAE), considering their similar argument structures.We view the extraction task as a role querying problem, unifying various methods into a single framework.There are key discrepancies on role labels and distant arguments between semantic role and event argument annotations.To mitigate these discrepancies, we specify natural language-like queries to tackle the label mismatch problem and devise argument augmentation to recover distant arguments.We show that SRL annotations can serve as a valuable resource for EAE, and a template-based slot querying strategy is especially effective for facilitating the transfer.In extensive evaluations on two English EAE benchmarks, our proposed model obtains impressive zero-shot results by leveraging SRL annotations, reaching nearly 80% of the fullysupervised scores.It further provides benefits in low-resource cases, where few EAE annotations are available.Moreover, we show that our approach generalizes to cross-domain and multilingual scenarios. Zhisong Zhang, Emma Strubell, Eduard H. Hovy |
EMNLP | 1 |
| 2022 | A Survey of Active Learning for Natural Language ProcessingabstractIn this work, we provide a literature review of active learning (AL) for its applications in natural language processing (NLP).In addition to a fine-grained categorization of query strategies, we also investigate several other important aspects of applying AL to NLP problems.These include AL for structured prediction tasks, annotation cost, model learning (especially with deep neural models), and starting and stopping AL.Finally, we conclude with a discussion of related topics and future directions. Zhisong Zhang, Emma Strubell, Eduard H. Hovy |
EMNLP | 1 |
| 2022 | Neural Character-Level Syntactic Parsing for ChineseabstractIn this work, we explore character-level neural syntactic parsing for Chinese with two typical syntactic formalisms: the constituent formalism and a dependency formalism based on a newly released character-level dependency treebank. Prior works in Chinese parsing have struggled with whether to de ne words when modeling character interactions. We choose to integrate full character-level syntactic dependency relationships using neural representations from character embeddings and richer linguistic syntactic information from human-annotated character-level Parts-Of-Speech and dependency labels. This has the potential to better understand the deeper structure of Chinese sentences and provides a better structural formalism for avoiding unnecessary structural ambiguities. Specifically, we first compare two different character-level syntax annotation styles: constituency and dependency. Then, we discuss two key problems for character-level parsing: (1) how to combine constituent and dependency syntactic structure in full character-level trees and (2) how to convert from character-level to word-level for both constituent and dependency trees. In addition, we also explore several other key parsing aspects, including di erent character-level dependency annotations and joint learning of Parts-Of-Speech and syntactic parsing. Finally, we evaluate our models on the Chinese Penn Treebank (CTB) and our published Shanghai Jiao Tong University Chinese Character Dependency Treebank (SCDT). The results show the e effectiveness of our model on both constituent and dependency parsing. We further provide empirical analysis and suggest several directions for future study. Zuchao Li, Junru Zhou, Hai Zhao 0001, Zhisong Zhang, Haonan Li 0002, Yuqi Ju |
J. Artif. Intell. Res. | 4 |
| 2021 | On the Benefit of Syntactic Supervision for Cross-lingual Transfer in Semantic Role LabelingabstractAlthough recent developments in neural architectures and pre-trained representations have greatly increased state-of-the-art model performance on fully-supervised semantic role labeling (SRL), the task remains challenging for languages where supervised SRL training data are not abundant.Cross-lingual learning can improve performance in this setting by transferring knowledge from high-resource languages to low-resource ones.Moreover, we hypothesize that annotations of syntactic dependencies can be leveraged to further facilitate cross-lingual transfer.In this work, we perform an empirical exploration of the helpfulness of syntactic supervision for crosslingual SRL within a simple multitask learning scheme.With comprehensive evaluations across ten languages (in addition to English) and three SRL benchmark datasets, including both dependency-and span-based SRL, we show the effectiveness of syntactic supervision in low-resource scenarios.Experiments Target Languages SRL Style Same Frames?Compatible Roles?Main SRL Setting EWT/UPB † ( §3.2) de,fr,it,es,pt,fi Dependency-based Yes Yes Zero-shot EWT/FiPB ( §3.3) fi Dependency-based No Yes Semi-supervised CoNLL-2009 ( §3.4) cs,zh,es,ca Dependency-based No No Semi-supervised OntoNotes ( §3.5) zh,ar Span-based No Yes Semi-supervised Zhisong Zhang, Emma Strubell, Eduard H. Hovy |
EMNLP (1) | 1 |
| 2021 | IA-CNN: A generalised interpretable convolutional neural network with attention mechanismabstractIn recent years, convolutional neural network (CNN) has been widely used in security, autonomous driving, and healthcare. Even though CNN has achieved a great performance, the results produced by CNN are difficult to explain and sometimes irresponsible. The black-box nature of CNN makes it lack trust. In this paper, we propose an attention based CNN structure, named IA -CNN, which highly improves the interpretability of the CNN models. Each feature map of the last conv-layer only has one response (one key point) of the target object, which is directly connected to the output. We also combine the attention mechanism to weakly supervise the last conv-layer. In this way, our model can clearly show that which features the model extracted are the keys to the output prediction. Meanwhile, our IA-CNN structure can be used in various classical models with higher performance in the fine-grained classification and comparative performance in the ordinary classification task. Note that our IA-CNN structure is an end-to-end model, the last conv-layer of which can extract key points from images automatically and is connected to the output prediction linearly. Zhisong Zhang, Yaran Chen, Haoran Li 0010 |
IJCNN | 1 |
| 2020 | A Two-Step Approach for Implicit Event Argument DetectionabstractIn this work, we explore the implicit event argument detection task, which studies event arguments beyond sentence boundaries.The addition of cross-sentence argument candidates imposes great challenges for modeling.To reduce the number of candidates, we adopt a two-step approach, decomposing the problem into two sub-problems: argument head-word detection and head-to-span expansion.Evaluated on the recent RAMS dataset (Ebner et al., 2020), our model achieves overall better performance than a strong sequence labeling baseline.We further provide detailed error analysis, presenting where the model mainly makes errors and indicating directions for future improvements.It remains a challenge to detect implicit arguments, calling for more future work of document-level modeling for this task. Zhisong Zhang, Xiang Kong, Zhengzhong Liu 0001, Xuezhe Ma, Eduard H. Hovy |
ACL | 1 |
| 2020 | Incorporating a Local Translation Mechanism into Non-autoregressive TranslationabstractIn this work, we introduce a novel local autoregressive translation (LAT) mechanism into non-autoregressive translation (NAT) models so as to capture local dependencies among target outputs.Specifically, for each target decoding position, instead of only one token, we predict a short sequence of tokens in an autoregressive way.We further design an efficient merging algorithm to align and merge the output pieces into one final output sequence.We integrate LAT into the conditional masked language model (CMLM; Ghazvininejad et al., 2019) and similarly adopt iterative decoding.Empirical results on five translation tasks show that compared with CMLM, our method achieves comparable or better performance with fewer decoding iterations, bringing a 2.5x speedup.Further analysis indicates that our method reduces repeated translations and performs better at longer sentences.The code for our model is available at https://github. com/shawnkx/NAT-with-Local-AT. Xiang Kong, Zhisong Zhang, Eduard H. Hovy |
EMNLP (1) | 2 |
| 2019 | Cross-Lingual Syntactic Transfer through Unsupervised Adaptation of Invertible ProjectionsabstractCross-lingual transfer is an effective way to build syntactic analysis tools in low-resource languages.However, transfer is difficult when transferring to typologically distant languages, especially when neither annotated target data nor parallel corpora are available.In this paper, we focus on methods for cross-lingual transfer to distant languages and propose to learn a generative model with a structured prior that utilizes labeled source data and unlabeled target data jointly.The parameters of source model and target model are softly shared through a regularized log likelihood objective.An invertible projection is employed to learn a new interlingual latent embedding space that compensates for imperfect crosslingual word embedding input.We evaluate our method on two syntactic tasks: part-ofspeech (POS) tagging and dependency parsing.On the Universal Dependency Treebanks, we use English as the only source corpus and transfer to a wide range of target languages.On the 10 languages in this dataset that are distant from English, our method yields an average of 5.2% absolute improvement on POS tagging and 8.3% absolute improvement on dependency parsing over a direct transfer method using state-of-the-art discriminative models. 1 3 Following Ahmad et al. (2019), we use the offline pre-trained alignment matrix present in https://github.com/Babylonpartners/ fastText_multilingual, which contains alignment matrices for 78 languages, which also allows comparison with their numbers in Section 4.3. Junxian He, Zhisong Zhang, Taylor Berg-Kirkpatrick, Graham Neubig |
ACL (1) | 2 |
| 2019 | Choosing Transfer Languages for Cross-Lingual LearningabstractYu-Hsiang Lin, Chian-Yu Chen, Jean Lee, Zirui Li, Yuyan Zhang, Mengzhou Xia, Shruti Rijhwani, Junxian He, Zhisong Zhang, Xuezhe Ma, Antonios Anastasopoulos, Patrick Littell, Graham Neubig. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019. Yu-Hsiang Lin, Chian-Yu Chen, Jean Lee, Mengzhou Xia, Shruti Rijhwani, Junxian He, Zhisong Zhang, Xuezhe Ma, Antonios Anastasopoulos, Patrick Littell, Graham Neubig |
ACL (1) | 9 |
| 2019 | An Empirical Investigation of Structured Output Modeling for Graph-based Neural Dependency ParsingabstractIn this paper, we investigate the aspect of structured output modeling for the state-ofthe-art graph-based neural dependency parser (Dozat and Manning, 2017).With evaluations on 14 treebanks, we empirically show that global output-structured models can generally obtain better performance, especially on the metric of sentence-level Complete Match.However, probably because neural models already learn good global views of the inputs, the improvement brought by structured output modeling is modest. Zhisong Zhang, Xuezhe Ma, Eduard H. Hovy |
ACL (1) | 1 |
| 2019 | Cross-Lingual Dependency Parsing with Unlabeled Auxiliary LanguagesabstractCross-lingual transfer learning has become an important weapon to battle the unavailability of annotated resources for low-resource languages.One of the fundamental techniques to transfer across languages is learning language-agnostic representations, in the form of word embeddings or contextual encodings.In this work, we propose to leverage unannotated sentences from auxiliary languages to help learning language-agnostic representations.Specifically, we explore adversarial training for learning contextual encoders that produce invariant representations across languages to facilitate cross-lingual transfer.We conduct experiments on cross-lingual dependency parsing where we train a dependency parser on a source language and transfer it to a wide range of target languages.Experiments on 28 target languages demonstrate that adversarial training significantly improves the overall transfer performances under several different settings.We conduct a careful analysis to evaluate the language-agnostic representations resulted from adversarial training. Wasi Uddin Ahmad, Zhisong Zhang, Xuezhe Ma, Kai-Wei Chang 0001, Nanyun Peng 0001 |
CoNLL | 2 |
| 2018 | Neural Character-level Dependency Parsing for ChineseabstractThis paper presents a truly full character-level neural dependency parser together with a newly released character-level dependency treebank for Chinese, which has suffered a lot from the dilemma of defining word or not to model character interactions. Integrating full character-level dependencies with character embedding and human annotated character-level part-of-speech and dependency labels for the first time, we show an extra performance enhancement from the evaluation on Chinese Penn Treebank and SJTU (Shanghai Jiao Tong University) Chinese Character Dependency Treebank and the potential of better understanding deeper structure of Chinese sentences. Haonan Li 0002, Zhisong Zhang, Yuqi Ju, Hai Zhao 0001 |
AAAI | 2 |
| 2018 | Exploring Recombination for Efficient Decoding of Neural Machine TranslationabstractIn Neural Machine Translation (NMT), the decoder can capture the features of the entire prediction history with neural connections and representations.This means that partial hypotheses with different prefixes will be regarded differently no matter how similar they are.However, this might be inefficient since some partial hypotheses can contain only local differences that will not influence future predictions.In this work, we introduce recombination in NMT decoding based on the concept of the "equivalence" of partial hypotheses.Heuristically, we use a simple n-gram suffix based equivalence function and adapt it into beam search decoding.Through experiments on large-scale Chinese-to-English and English-to-Germen translation tasks, we show that the proposed method can obtain similar translation quality with a smaller beam size, making NMT decoding more efficient. Zhisong Zhang, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita, Hai Zhao 0001 |
EMNLP | 1 |
| 2017 | Adversarial Connective-exploiting Networks for Implicit Discourse Relation ClassificationabstractImplicit discourse relation classification is of great challenge due to the lack of connectives as strong linguistic cues, which motivates the use of annotated implicit connectives to improve the recognition.We propose a feature imitation framework in which an implicit relation network is driven to learn from another neural network with access to connectives, and thus encouraged to extract similarly salient features for accurate classification.We develop an adversarial model to enable an adaptive imitation scheme through competition between the implicit network and a rival feature discriminator.Our method effectively transfers discriminability of connectives to the implicit features, and achieves state-of-the-art performance on the PDTB benchmark. Lianhui Qin, Zhisong Zhang, Hai Zhao 0001, Zhiting Hu, Eric P. Xing |
ACL (1) | 2 |
| 2016 | Probabilistic Graph-based Dependency Parsing with Convolutional Neural NetworkabstractThis paper presents neural probabilistic parsing models which explore up to thirdorder graph-based parsing with maximum likelihood training criteria.Two neural network extensions are exploited for performance improvement.Firstly, a convolutional layer that absorbs the influences of all words in a sentence is used so that sentence-level information can be effectively captured.Secondly, a linear layer is added to integrate different order neural models and trained with perceptron method.The proposed parsers are evaluated on English and Chinese Penn Treebanks and obtain competitive accuracies. Zhisong Zhang, Hai Zhao 0001, Lianhui Qin |
ACL (1) | 1 |
| 2016 | Implicit Discourse Relation Recognition with Context-aware Character-enhanced EmbeddingsabstractFor the task of implicit discourse relation recognition, traditional models utilizing manual features can suffer from data sparsity problem. Neural models provide a solution with distributed representations, which could encode the latent semantic information, and are suitable for recognizing semantic relations between argument pairs. However, conventional vector representations usually adopt embeddings at the word level and cannot well handle the rare word problem without carefully considering morphological information at character level. Moreover, embeddings are assigned to individual words independently, which lacks of the crucial contextual information. This paper proposes a neural model utilizing context-aware character-enhanced embeddings to alleviate the drawbacks of the current word level representation. Our experiments show that the enhanced embeddings work well and the proposed model obtains state-of-the-art results. Lianhui Qin, Zhisong Zhang, Hai Zhao 0001 |
COLING | 2 |
| 2016 | A Stacking Gated Neural Architecture for Implicit Discourse Relation Classification
Lianhui Qin, Zhisong Zhang, Hai Zhao 0001 |
EMNLP | 2 |
| 2015 | High-order Graph-based Neural Dependency Parsing
Zhisong Zhang, Hai Zhao 0001 |
PACLIC | 1 |