Guohong Fu

dblp:23/5204 · DBLP profile ↗
← Back
56ranked-venue papers
4as first author
29since 2021 · last 2026
0000-0001-6882-6181ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 42 · 3 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021
YearPublicationVenuePosition
2026 Interleaved Tool-Call Reasoning for Protein Function Understanding
abstract
Chuanliu Fan, Zicheng Ma, Huanran Meng, Aijia Zhang, Wenjie Du, Jun Zhang, Ziqiang Cao, Guohong Fu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Chuanliu Fan, Zicheng Ma, Huanran Meng, Jun Zhang 0069, Ziqiang Cao, Guohong Fu
ACL (1)8
2026 Locate and Explain: Joint Multimodal Emotion Cause Extraction and Summarization in Conversation
abstract
Multimodal emotion cause analysis in conversation aims to identify the causes of emotions by leveraging multimodal information.Existing studies mainly formulate this problem as either utterance-level emotion cause extraction, which provides clear cause localization but limited explanation, or multimodal emotion cause generation, which offers fine-grained explanations but lacks explicit traceability to source utterances.Moreover, existing datasets rely heavily on human judgment and lack well-defined structured theoretical guidance, leading to subjective and inconsistent annotations.To address these issues, we introduce joint Multimodal Emotion Cause Extraction and Summarization in conversation (MECES), a new task that simultaneously extracts emotion cause utterances and generates cause summaries, enabling both precise localization and interpretable explanations of emotion cause.We further construct a MECES dataset guided by the Activating events-Beliefs-Consequences theory from psychology.This dataset consists of 5,787 emotion utterances annotated with causes, comprising 12,231 emotion-cause pairs and 6,040 cause summaries.We also propose an effective endto-end joint learning approach for MECES task, establishing strong benchmark results for this newly introduced task and dataset.( U1,U2 , "Chuan Bai handed Guang Shi a gift, and Guang Shi didn't expect to receive one too.") Guang Shi:"I got a gift too!"
Jikun Wan, Chen Gong 0004, Guohong Fu
ACL (1)3
2026 Bridging modalities: a unified framework for textual and multimodal dialogue discourse parsing
Chen Gong 0004, Guohong Fu
Frontiers Comput. Sci.3
2026 Bidirectional GPT
Chuanliu Fan, Zicheng Ma, Jun Zhang 0071, Yiqin Gao, Ziqiang Cao, Guohong Fu
Inf. Process. Manag.8
2026 Enhancing multi-modal aspect-based sentiment classification via emotional semantic-aware cross-modal relation inference
Chen Gong 0004, Guohong Fu
Inf. Process. Manag.3
2026 Rethinking hard training sample generation for medical image segmentation
Zhibin Wan, Mingjie Sun, Cao Min, Guohong Fu
Pattern Recognit.7
2025 Multimodal Coreference Resolution for Chinese Social Media Dialogues: Dataset and Benchmark Approach
abstract
Multimodal coreference resolution (MCR) aims to identify mentions referring to the same entity across different modalities, such as text and visuals, and is essential for understanding multimodal content.In the era of rapidly growing multimodal content and social media, MCR is particularly crucial for interpreting user interactions and bridging text-visual references to improve communication and personalization.However, MCR research for real-world dialogues remains unexplored due to the lack of sufficient data resources.To address this gap, we introduce TikTalkCoref, the first Chinese multimodal coreference dataset for social media in real-world scenarios, derived from the popular Douyin short-video platform.This dataset pairs short videos with corresponding textual dialogues from user comments and includes manually annotated coreference clusters for both person mentions in the text and the coreferential person head regions in the corresponding video frames.We also present an effective benchmark approach for MCR, focusing on the celebrity domain, and conduct extensive experiments on our dataset, providing reliable benchmark results for this newly constructed dataset.We release the TikTalk-Coref dataset to facilitate future research on MCR for real-world social media dialogues at https://github.com/lxystaruni/TikTalkCoref.
Chen Gong 0004, Guohong Fu
ACL (1)3
2025 LlaMol: A Unified Molecule Designer via Preference Ranking and Numerical Enhancement
abstract
Goal-oriented de novo molecule design, namely generating molecules with specific property or substructure constraints from scratch, is a crucial yet challenging task in drug discovery. Existing research often relies on separate predictors for distinct properties and struggles with integrating substructure constraints due to the complexities involved in modeling structural information via multitask learning. This separation necessitates a dedicated prediction model for each constraint, limiting the flexibility and posing challenges for realworld applications. To address these limitations, we propose a unified framework for molecular design that incorporates multiple property and substructure constraints, leveraging LLMs to handle diverse constraint settings within a single model. We first integrate feedback learning derived from preference ranking to eliminate the need for separate property predictors. Then, we enhance the model's ability to follow numerical instructions by introducing a unified numerical encoding into the prompt. We conduct extensive experiments across single-property, substructureproperty, and multi-property constrained tasks. Experimental results demonstrate that LlaMol consistently outperforms state-of-the-art baselines across various constraint settings. Notably, in the multi-objective binding affinity maximization task, LlaMol achieves a significantly lower$\mathrm{K}_{\mathrm{D}}$value of 0.25 for the protein target ESR1, while maintaining the highest overall performance, surpassing previous methods by 4.76 %. These results underscore the effectiveness and versatility of LLM-based frameworks for molecule generation under complex constraints.
Chuanliu Fan, Zicheng Ma, Jun Zhang 0069, Ziqiang Cao, Yiqin Gao, Guohong Fu
BIBM8
2025 2S-DGM4: A Two-Stage Framework for Detecting and Grounding Multi-Modal Media Manipulation
abstract
Detecting and grounding multi-modal media manipulation (DGM4) aims to identify the authenticity of media content in the form of image-text pairs, and locate forgery contents including text tokens or image regions. Existing methods adopt joint optimization frameworks to train multiple dedicated heads for different sub-tasks. This paradigm is prone to yielding models with insufficient learning regarding manipulation traces. In this paper, we propose a two-stage learning framework, named 2S-DGM4, exploiting correlations between manipulation detection and grounding for debunking manipulation content. It introduces modality-specific manipulation types into the subsequent manipulation grounding, and naturally conducts the mutual agreement among various sub-tasks. Apart from the global context in cross-modal interaction, we develop a fine-grained refinement module to capture subtle manipulation traces. Experiments demonstrate that 2S-DGM4outperforms state-of-the-art methods on the public benchmark dataset DGM4. Further analyses verify the efficacy of 2S-DGM4on manipulation reasoning.
Junjie Wu 0005, Yumeng Fu, Chen Gong 0004, Guohong Fu
ICME5
2025 Speaker Intention Enhanced Dialogue Discourse Parsing
abstract
Dialogue Discourse Parsing (DDP) task focuses on capturing the structural relations between utterances in a dialogue, represented as a dependency tree. Previous studies have utilized speaker information to enhance the understanding of dialogue semantics and improve model performance on the DDP task. However, the role of speaker intention, which reflects the psychological states of speakers, has been largely overlooked. Speaker intention is crucial for understanding dialogue semantics and has proven beneficial for various conversational tasks. To address this gap, we propose the Speaker Intention Enhanced Dialogue Discourse Parsing Model (SIEDDP), which integrates speaker intention into the DDP process. We utilize the Large Language Model (LLM) such as LLaMA3 to extract speaker intention from the dialogue and compare the intention with the traditional method leveraging the COMmonsEnse Transformer (COMET) model. Experimental results show that incorporating speaker intention significantly improves DDP performance. Further analysis demonstrate that speaker intention assists in inferring the link and relation between current utterance and the context.
Suxian Zhao, Chen Gong 0004, Guohong Fu
IJCNN4
2025 CA-SAM2: SAM2-Based Context-Aware Network with Auto-prompting for Nuclei Instance Segmentation
Hanbin Huang, Liying Xu, Siwei Feng, Guohong Fu
MICCAI (9)6
2025 APSeg: Auto-prompt Model with Acquired and Injected Knowledge for Nuclear Instance Segmentation and Classification
Liying Xu, Hanbin Huang, Siwei Feng, Guohong Fu
PRCV (14)6
2025 OpenBA: an open-sourced 15B bilingual asymmetric Seq2Seq model pre-trained from scratch
Juntao Li 0005, Zecheng Tang, Yuyang Ding, Pinzheng Wang, Pei Guo, Wangjie You, Wenliang Chen, Guohong Fu, Qiaoming Zhu, Guodong Zhou 0001, Min Zhang 0005
Sci. China Inf. Sci.10
2024 PromptCD: Coupled and Decoupled Prompt Learning for Vision-Language Models
abstract
Large-scale pre-trained vision-language models (VLMs), like CLIP, have presented striking generalizability for adapting to image classification in a few shot setting. Most existing methods explore a set of learnable tokens, such as prompt learning, on data-efficient utilization for task adaptation. However, they focus on either the coupled-modality property by prompt projection or decoupled-modality characteristic by prompt consistency, which ignores effective interaction between prompts. To model the deep yet sufficient cross-modal interaction and enhance the generalization between both seen and unseen tasks, in this paper, we propose a novel coupled and decoupled prompt learning framework, dubbed PromptCD, for vision-language models. Specifically, we introduce a bi-directional coupled-modality mechanism to intensify the interaction between both vision and language branches. Additionally, we propose mixture consistency to further improve the generalization and discrimination of the models on unseen tasks. The integration of such a mechanism and consistency facilitates the proposed framework adaptation for various downstream tasks. We conduct extensive experiments on 11 image classification datasets under a range of evaluation protocols, including base-to-novel and domain generalization, and cross-dataset recognition. Experimental results demonstrate that our proposed PromptCD overall outperforms state-of-the-art methods.
Junjie Wu 0005, Mingjie Sun, Chen Gong 0004, Guohong Fu
ECAI5
2024 Gated Slot Attention for Efficient Linear-Time Sequence Modeling
abstract
Linear attention Transformers and their gated variants, celebrated for enabling parallel training and efficient recurrent inference, still fall short in recall-intensive tasks compared to traditional Transformers and demand significant resources for training from scratch. This paper introduces Gated Slot Attention (GSA), which enhances Attention with Bounded-memory-Control (ABC) by incorporating a gating mechanism inspired by Gated Linear Attention (GLA). Essentially, GSA comprises a two-layer GLA linked via $\operatorname{softmax}$, utilizing context-aware memory reading and adaptive forgetting to improve memory capacity while maintaining compact recurrent state size. This design greatly enhances both training and inference efficiency through GLA's hardware-efficient training algorithm and reduced state size. Additionally, retaining the $\operatorname{softmax}$ operation is particularly beneficial in ``finetuning pretrained Transformers to RNNs'' (T2R) settings, reducing the need for extensive training from scratch. Extensive experiments confirm GSA's superior performance in scenarios requiring in-context recall and in T2R settings.
Yu Zhang 0092, Rui-Jie Zhu 0003, Yue Zhang 0004, Leyang Cui, Yiqiao Wang 0005, Bolun Wang, Freda Shi, Bailin Wang, Wei Bi, Peng Zhou 0017, Guohong Fu
NeurIPS12
2023 Non-autoregressive Text Editing with Copy-aware Latent Alignments
abstract
Recent work has witnessed a paradigm shift from Seq2Seq to Seq2Edit in the field of text editing, with the aim of addressing the slow autoregressive inference problem posed by the former.Despite promising results, Seq2Edit approaches still face several challenges such as inflexibility in generation and difficulty in generalizing to other languages.In this work, we propose a novel non-autoregressive text editing method to circumvent the above issues, by modeling the edit process with latent CTC alignments.We make a crucial extension to CTC by introducing the copy operation into the edit space, thus enabling more efficient management of textual overlap in editing.We conduct extensive experiments on GEC and sentence fusion tasks, showing that our proposed method significantly outperforms existing Seq2Edit models and achieves similar or even better results than Seq2Seq with over 4ˆ speedup.Moreover, it demonstrates good generalizability on German and Russian.In-depth analyses reveal the strengths of our method in terms of the robustness under various scenarios and generating fluent and flexible outputs.
Yu Zhang 0092, Yue Zhang 0004, Leyang Cui, Guohong Fu
EMNLP4
2023 Multi-grained Aspect Fusion for Review Response Generation
Chen Gong 0004, Dexin Kong, Guohong Fu
ICANN (9)5
2023 Discourse-Aware Causal Emotion Entailment
Dexin Kong, Chen Gong 0004, Guohong Fu
ICONIP (9)6
2023 MCG-MNER: A Multi-Granularity Cross-Modality Generative Framework for Multimodal NER with Instruction
abstract
Multimodal named entity recognition (MNER) is an essential task of vision and language, which aims to locate named entities and classify them to the predefined categories using visual scenarios. However, existing MNER studies often suffer from bias issues with fine-grained visual cue fusion, which may produce noisy coarse-grained visual cues for MNER. To accurately capture text-image relations and better refine multimodal representations, we propose a novel instruction-based Multi-granularity Cross-modality Generative framework for MNER, namely MCG-MNER. Concretely, we introduce a multi-granularity relation propagation to infer visual clues relevant to text. Then, we propose a method to jnject multi-granularity visual information into cross-modality interaction and fusion to learn a unified representation. Finally, we integrate task-specific instructions and answers for MCG-MNER. Comprehensive experimental results on three benchmark datasets, such as Twitter2015, Twitter2017 and WikiDiverse, demonstrate the superiority of our proposed method over several state-of-the-art MNER methods. We will publicly release our codes for future studies.
Junjie Wu 0005, Chen Gong 0004, Ziqiang Cao, Guohong Fu
ACM Multimedia4
2023 RSpell: Retrieval-Augmented Framework for Domain Adaptive Chinese Spelling Check
Siqi Song, Qi Lv 0001, Lei Geng, Ziqiang Cao, Guohong Fu
NLPCC (1)5
2023 General and Domain-adaptive Chinese Spelling Check with Error-consistent Pretraining
abstract
The lack of label data is one of the significant bottlenecks for Chinese Spelling Check. Existing researches use the automatic generation method by exploiting unlabeled data to expand the supervised corpus. However, there is a big gap between the real input scenario and automatically generated corpus. Thus, we develop a competitive general speller ECSpell, which adopts the Error-consistent masking strategy to create data for pretraining. This error-consistency masking strategy is used to specify the error types of automatically generated sentences consistent with the real scene. The experimental result indicates that our model outperforms previous state-of-the-art models on the general benchmark. Moreover, spellers often work within a particular domain in real life. Due to many uncommon domain terms, experiments on our built domain-specific datasets show that general models perform terribly. Inspired by the common practice of input methods, we propose to add an alterable user dictionary to handle the zero-shot domain-adaption problem. Specifically, we attach a User Dictionary guided inference module (UD) to a general token classification-based speller. Our experiments demonstrate that ECSpell UD , namely, ECSpell combined with UD, surpasses all the other baselines broadly, even approaching the performance on the general benchmark. 1
Qi Lv 0001, Ziqiang Cao, Lei Geng, Chunhui Ai, Guohong Fu
ACM Trans. Asian Low Resour. Lang. Inf. Process.6
2022 RST Discourse Parsing with Second-Stage EDU-Level Pre-training
abstract
Pre-trained language models (PLMs) have shown great potentials in natural language processing (NLP) including rhetorical structure theory (RST) discourse parsing.Current PLMs are obtained by sentence-level pre-training, which is different from the basic processing unit, i.e. element discourse unit (EDU).To this end, we propose a second-stage EDU-level pretraining approach in this work, which presents two novel tasks to learn effective EDU representations continually based on well pre-trained language models.Concretely, the two tasks are (1) next EDU prediction (NEP) and ( 2) discourse marker prediction (DMP).We take a state-of-the-art transition-based neural parser as baseline, and adopt it with a light bi-gram EDU modification to effectively explore the EDU-level pre-trained EDU representation.Experimental results on a benckmark dataset show that our method is highly effective, leading a 2.1-point improvement in F1-score.All codes and pre-trained models will be released publicly to facilitate future studies. 1
Meishan Zhang, Guohong Fu, Min Zhang 0005
ACL (1)3
2022 Tracking Satisfaction States for Customer Satisfaction Prediction in E-commerce Service Chatbots
abstract
Due to the increasing use of service chatbots in E-commerce platforms in recent years, customer satisfaction prediction (CSP) is gaining more and more attention. CSP is dedicated to evaluating subjective customer satisfaction in conversational service and thus helps improve customer service experience. However, previous methods focus on modeling customer-chatbot interaction across different turns, which are hard to represent the important dynamic satisfaction states throughout the customer journey. In this work, we investigate the problem of satisfaction states tracking and its effects on CSP in E-commerce service chatbots. To this end, we propose a dialogue-level classification model named DialogueCSP to track satisfaction states for CSP. In particular, we explore a novel two-step interaction module to represent the dynamic satisfaction states at each turn. In order to capture dialogue-level satisfaction states for CSP, we further introduce dialogue-aware attentions to integrate historical informative cues into the interaction module. To evaluate the proposed approach, we also build a Chinese E-commerce dataset for CSP. Experiment results demonstrate that our model significantly outperforms multiple baselines, illustrating the benefits of satisfaction states tracking on CSP.
Liangqing Wu, Shuangyong Song, Xiaoguang Yu, Xiaodong He 0001, Guohong Fu
COLING6
2022 Speaker-Aware Discourse Parsing on Multi-Party Dialogues
abstract
Discourse parsing on multi-party dialogues is an important but difficult task in dialogue systems and conversational analysis. It is believed that speaker interactions are helpful for this task. However, most previous research ignores speaker interactions between different speakers. To this end, we present a speaker-aware model for this task. Concretely, we propose a speaker-context interaction joint encoding (SCIJE) approach, using the interaction features between different speakers. In addition, we propose a second-stage pre-training task, same speaker prediction (SSP), enhancing the conversational context representations by predicting whether two utterances are from the same speaker. Experiments on two standard benchmark datasets show that the proposed model achieves the best-reported performance in the literature. We will release the codes of this paper to facilitate future research.
Guohong Fu, Min Zhang 0005
COLING2
2022 Semantic Role Labeling as Dependency Parsing: Exploring Latent Tree Structures inside Arguments
abstract
Semantic role labeling (SRL) is a fundamental yet challenging task in the NLP community. Recent works of SRL mainly fall into two lines: 1) BIO-based; 2) span-based. Despite ubiquity, they share some intrinsic drawbacks of not considering internal argument structures, potentially hindering the model’s expressiveness. The key challenge is arguments are flat structures, and there are no determined subtree realizations for words inside arguments. To remedy this, in this paper, we propose to regard flat argument spans as latent subtrees, accordingly reducing SRL to a tree parsing task. In particular, we equip our formulation with a novel span-constrained TreeCRF to make tree structures span-aware and further extend it to the second-order case. We conduct extensive experiments on CoNLL05 and CoNLL12 benchmarks. Results reveal that our methods perform favorably better than all previous syntax-agnostic works, achieving new state-of-the-art under both end-to-end and w/ gold predicates settings.
Yu Zhang 0092, Qingrong Xia, Shilin Zhou 0002, Guohong Fu, Min Zhang 0005
COLING5
2022 A Dual-Pointer guided transition system for end-to-end structured sentiment analysis with global graph reasoning
Qiujing Xu, Bobo Li 0001, Fei Li 0021, Guohong Fu, Donghong Ji
Inf. Process. Manag.4
2021 Chinese Opinion Role Labeling with Corpus Translation: A Pivot Study
abstract
Opinion Role Labeling (ORL), aiming to identify the key roles of opinion, has received increasing interest.Unlike most of the previous works focusing on the English language, in this paper, we present the first work of Chinese ORL.We construct a Chinese dataset by manually translating and projecting annotations from a standard English MPQA dataset.Then, we investigate the effectiveness of cross-lingual transfer methods, including model transfer and corpus translation.We exploit multilingual BERT with Contextual Parameter Generator and Adapter methods to examine the potentials of unsupervised crosslingual learning and our experiments and analyses for both bilingual and multilingual transfers establish a foundation for the future research of this task 1 .
Ranran Zhen, Rui Wang 0005, Guohong Fu, Chengguo Lv, Meishan Zhang
EMNLP (1)3
2021 Integrating Rich Utterance Features for Emotion Recognition in Multi-party Conversations
Guohong Fu
ICONIP (4)3
2021 Dependency-based syntax-aware word representations
Meishan Zhang, Zhenghua Li, Guohong Fu, Min Zhang 0005
Artif. Intell.3
2020 Sentence Matching with Syntax- and Semantics-Aware BERT
abstract
Sentence matching aims to identify the special relationship between two sentences, and plays a key role in many natural language processing tasks.However, previous studies mainly focused on exploiting either syntactic or semantic information for sentence matching, and no studies consider integrating both of them.In this study, we propose integrating syntax and semantics into BERT with sentence matching.In particular, we use an implicit syntax and semantics integration method that is less sensitive to the output structure information.Thus the implicit integration can alleviate the error propagation problem.The experimental results show that our approach has achieved state-of-the-art or competitive performance on several sentence matching datasets, demonstrating the benefits of implicitly integrating syntactic and semantic features in sentence matching.
Chengguo Lv, Ranran Zhen, Guohong Fu
COLING5
2020 Unseen Target Stance Detection with Adversarial Domain Generalization
abstract
Although stance detection has made great progress in the past few years, it is still facing the problem of unseen targets. In this study, we investigate the domain difference between targets and thus incorporate attention-based conditional encoding with adversarial domain generalization to perform unseen target stance detection. Experimental results show that our approach achieves new state-of-the-art performance on the SemEval-2016 dataset, demonstrating the importance of domain difference between targets in unseen target stance detection.
Qiansheng Wang, Chengguo Lv, Xue Cao, Guohong Fu
IJCNN5
2019 Syntax-Aware Neural Semantic Role Labeling
abstract
Semantic role labeling (SRL), also known as shallow semantic parsing, is an important yet challenging task in NLP. Motivated by the close correlation between syntactic and semantic structures, traditional discrete-feature-based SRL approaches make heavy use of syntactic features. In contrast, deep-neural-network-based approaches usually encode the input sentence as a word sequence without considering the syntactic structures. In this work, we investigate several previous approaches for encoding syntactic trees, and make a thorough study on whether extra syntax-aware representations are beneficial for neural SRL models. Experiments on the benchmark CoNLL-2005 dataset show that syntax-aware SRL approaches can effectively improve performance over a strong baseline with external word representations from ELMo. With the extra syntax-aware representations, our approaches achieve new state-of-the-art 85.6 F1 (single model) and 86.6 F1 (ensemble) on the test data, outperforming the corresponding strong baselines with ELMo by 0.8 and 1.0, respectively. Detailed error analysis are conducted to gain more insights on the investigated approaches.
Qingrong Xia, Zhenghua Li, Min Zhang 0005, Meishan Zhang, Guohong Fu, Rui Wang 0005, Luo Si
AAAI5
2019 Cross-Lingual Dependency Parsing Using Code-Mixed TreeBank
abstract
Meishan Zhang, Yue Zhang, Guohong Fu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Meishan Zhang, Yue Zhang 0004, Guohong Fu
EMNLP/IJCNLP (1)3
2019 Domain Information Enhanced Dependency Parser
Ranran Zhen, Meishan Zhang, Guohong Fu
NLPCC (2)6
2019 End-to-end neural opinion extraction with a transition-based model
Meishan Zhang, Qiansheng Wang, Guohong Fu
Inf. Syst.3
2019 Effective Subword Segmentation for Text Comprehension
abstract
Representation learning is the foundation of machine reading comprehension and inference. In state-of-the-art models, character-level representations have been broadly adopted to alleviate the problem of effectively representing rare or complex words. However, character itself is not a natural minimal linguistic unit for representation or word embedding composing due to ignoring the linguistic coherence of consecutive characters inside word. This paper presents a general subword-augmented embedding framework for learning and composing computationally derived subword-level representations. We survey a series of unsupervised segmentation methods for subword acquisition and different subword-augmented strategies for text understanding, showing that subword-augmented embedding significantly improves our baselines in various types of text understanding tasks on both English and Chinese benchmarks.
Zhuosheng Zhang 0001, Hai Zhao 0001, Kangwei Ling, Jiangtong Li, Zuchao Li, Shexia He, Guohong Fu
IEEE ACM Trans. Audio Speech Lang. Process.7
2018 Transition-based Neural RST Parsing with Implicit Syntax Features
abstract
Syntax has been a useful source of information for statistical RST discourse parsing. Under the neural setting, a common approach integrates syntax by a recursive neural network (RNN), requiring discrete output trees produced by a supervised syntax parser. In this paper, we propose an implicit syntax feature extraction approach, using hidden-layer vectors extracted from a neural syntax parser. In addition, we propose a simple transition-based model as the baseline, further enhancing it with dynamic oracle. Experiments on the standard dataset show that our baseline model with dynamic oracle is highly competitive. When implicit syntax features are integrated, we are able to obtain further improvements, better than using explicit Tree-RNN.
Meishan Zhang, Guohong Fu
COLING3
2018 Transition-Based Neural Word Segmentation Using Word-Level Features
abstract
Character-based and word-based methods are two different solutions for Chinese word segmentation, the former exploiting sequence labeling models over characters and the latter using word-level features. Neural models have been exploited for character-based Chinese word segmentation, giving high accuracies by making use of external character embeddings, yet requiring less feature engineering. In this paper, we study a neural model for word-based Chinese word segmentation, by replacing the manually-designed discrete features with neural features in a transition-based word segmentation framework. Experimental results demonstrate that word features lead to comparable performance to the best systems in the literature, and a further combination of discrete and neural features obtains top accuracies on several benchmarks.
Meishan Zhang, Yue Zhang 0004, Guohong Fu
J. Artif. Intell. Res.3
2018 Recognizing irregular entities in biomedical text via deep neural networks
Fei Li 0021, Meishan Zhang, Bo Chen 0024, Guohong Fu, Donghong Ji
Pattern Recognit. Lett.5
2018 Joint POS Tagging and Dependence Parsing With Transition-Based Neural Networks
abstract
While part-of-speech (POS) tagging and dependency parsing are observed to be closely related, existing work on joint modeling with manually crafted feature templates suffers from the feature sparsity and incompleteness problems. In this paper, we propose an approach to joint POS tagging and dependency parsing using transition-based neural networks. Three neural network based classifiers are designed to resolve shift/reduce, tagging, and labeling conflicts. Experiments show that our approach significantly outperforms previous methods for joint POS tagging and dependency parsing across a variety of natural languages.
Liner Yang, Meishan Zhang, Yang Liu 0005, Maosong Sun 0001, Guohong Fu
IEEE ACM Trans. Audio Speech Lang. Process.6
2018 A Simple and Effective Neural Model for Joint Word Segmentation and POS Tagging
abstract
Joint models have shown stronger capabilities for Chinese word segmentation and POS tagging, and have received great interests in the community of Chinese natural language processing. In this paper, we follow this line of work, presenting a simple yet effective sequence-to-sequence neural model for the joint task, based on a well-defined transition system, by using long short term memory neural network structures. We conduct experiments on five different datasets. The results demonstrate that our proposed model is highly competitive. By using well-trained character-level embeddings, the proposed neural joint model is able to obtain the best-reported performances in the literature.
Meishan Zhang, Guohong Fu
IEEE ACM Trans. Audio Speech Lang. Process.3
2017 End-to-End Neural Relation Extraction with Global Optimization
abstract
Neural networks have shown promising results for relation extraction.State-ofthe-art models cast the task as an end-toend problem, solved incrementally using a local classifier.Yet previous work using statistical models have demonstrated that global optimization can achieve better performances compared to local classification.We build a globally optimized neural model for end-to-end relation extraction, proposing novel LSTM features in order to better learn context representations.In addition, we present a novel method to integrate syntactic information to facilitate global learning, yet requiring little background on syntactic grammars thus being easy to extend.Experimental results show that our proposed model is highly effective, achieving the best performances on two standard benchmarks.
Meishan Zhang, Yue Zhang 0004, Guohong Fu
EMNLP3
2017 Segmenting Chinese Microtext: Joint Informal-Word Detection and Segmentation with Neural Networks
abstract
State-of-the-art Chinese word segmentation systems typically exploit supervised modelstrained on a standard manually-annotated corpus,achieving performances over 95% on a similar standard testing corpus.However, the performances may drop significantly when the same models are applied onto Chinese microtext.One major challenge is the issue of informal words in the microtext.Previous studies show that informal word detection can be helpful for microtext processing.In this work, we investigate it under the neural setting, by proposing a joint segmentation model that integrates the detection of informal words simultaneously.In addition, we generate training corpus for the joint model by using existing corpus automatically.Experimental results show that the proposed model is highly effective for segmentation of Chinese microtext.
Meishan Zhang, Guohong Fu
IJCAI2
2017 A Neural Joint Model for Extracting Bacteria and Their Locations
Fei Li 0021, Meishan Zhang, Guohong Fu, Donghong Ji
PAKDD (2)3
2017 A neural joint model for entity and relation extraction from biomedical text
abstract
BACKGROUND: Extracting biomedical entities and their relations from text has important applications on biomedical research. Previous work primarily utilized feature-based pipeline models to process this task. Many efforts need to be made on feature engineering when feature-based models are employed. Moreover, pipeline models may suffer error propagation and are not able to utilize the interactions between subtasks. Therefore, we propose a neural joint model to extract biomedical entities as well as their relations simultaneously, and it can alleviate the problems above. RESULTS: Our model was evaluated on two tasks, i.e., the task of extracting adverse drug events between drug and disease entities, and the task of extracting resident relations between bacteria and location entities. Compared with the state-of-the-art systems in these tasks, our model improved the F1 scores of the first task by 5.1% in entity recognition and 8.0% in relation extraction, and that of the second task by 9.2% in relation extraction. CONCLUSIONS: The proposed model achieves competitive performances with less work on feature engineering. We demonstrate that the model based on neural networks is effective for biomedical entity and relation extraction. In addition, parameter sharing is an alternative method for neural models to jointly process this task. Our work can facilitate the research on biomedical text mining.
Fei Li 0021, Meishan Zhang, Guohong Fu, Donghong Ji
BMC Bioinform.3
2017 Coupled POS Tagging on Heterogeneous Annotations
abstract
The limited scale and genre coverage of labeled data greatly hinders the effectiveness of supervised models, especially when analyzing spoken languages, such as texts transcribed from speech and informal text including tweets and product comments in Internet. In order to effectively utilize multiple labeled datasets with heterogeneous annotations for the same task, this paper proposes a coupled sequence labeling model that can directly learn and infer two heterogeneous annotations simultaneously, using Chinese part-of-speech (POS) tagging as our case study. The key idea is to bundle two sets of POS tags together (e.g., “[NN, n]n), and build a conditional random field (CRF) based tagging model in the enlarged space of bundled tags with the help of ambiguous labeling. To train our model on two nonoverlapping datasets that each has only one-side tags, we transform a one-side tag into a set of bundled tags by concatenating the tag with every possible tag at the missing side according to a predefined context-free tag-to-tag mapping function, thus producing ambiguous labeling as weak supervision. We design and investigate four different context-free tag-to-tag mapping functions, and find out that the coupled model achieves its best performance when each one-side tag is mapped to all tags at the other side (namely complete mapping), indicating that the model can effectively learn the loose mapping between the two heterogeneous annotations, without the need of manually designed mapping rules. Moreover, we propose a context-aware online pruning strategy that can more accurately capture mapping relationships between annotations based on contextual evidences and thus effectively solve the severe inefficiency problem with our coupled model under complete mapping, making it comparable with the baseline CRF model. Experiments on benchmark datasets show that our coupled model significantly outperforms the state-of-the-art baselines on both one-side POS tagging and annotation conversion tasks. The codes and newly annotated data are released for research usage.1
Zhenghua Li, Jiayuan Chao, Min Zhang 0005, Wenliang Chen, Meishan Zhang, Guohong Fu
IEEE ACM Trans. Audio Speech Lang. Process.6
2016 Transition-Based Neural Word Segmentation
Meishan Zhang, Yue Zhang 0004, Guohong Fu
ACL (1)3
2016 Tweet Sarcasm Detection Using Deep Neural Network
abstract
Sarcasm detection has been modeled as a binary document classification task, with rich features being defined manually over input documents. Traditional models employ discrete manual features to address the task, with much research effect being devoted to the design of effective feature templates. We investigate the use of neural network for tweet sarcasm detection, and compare the effects of the continuous automatic features with discrete manual features. In particular, we use a bi-directional gated recurrent neural network to capture syntactic and semantic information over tweets locally, and a pooling neural network to extract contextual features automatically from history tweets. Results show that neural features give improved accuracies for sarcasm detection, with different error distributions compared with discrete manual features.
Meishan Zhang, Yue Zhang 0004, Guohong Fu
COLING3
2015 Polarity Classification of Short Product Reviews via Multiple Cluster-based SVM Classifiers
Jiaying Song, Guohong Fu
PACLIC3
2012 Learning Lexical Subjectivity Strength for Chinese Opinionated Sentence Identification
Guohong Fu
CICLing (1)2
2012 A CRF Sequence Labeling Approach to Chinese Punctuation Prediction
Guohong Fu
PACLIC3
2008 A Morpheme-based Part-of-Speech Tagger for Chinese
Guohong Fu, Jonathan J. Webster
IJCNLP1
2008 Chinese word segmentation as morpheme-based lexical chunking
Guohong Fu, Chunyu Kit, Jonathan J. Webster
Inf. Sci.1
2004 Chinese Unknown Word Identification Using Class-Based LM
Guohong Fu, Kang-Kwong Luke
IJCNLP1
2004 Grapheme-to-phoneme conversion for Chinese text-to-speech
abstract
This paper reports a study of grapheme-to-phoneme (G2P) conversion for Chinese text-to-speech (TTS) system. As Chinese is a syllabic language, syllable is commonly adopted as the phonetic unit in TTS, which is represented by pinyin, the standard Chinese romanization. A Chinese G2P conversion is to find correct pinyin for polyphonic graphemes in the input text. In this paper, a complete G2P framework is presented, which includes a two-stage statistical word segmentation module, a hidden Markov model (HMM) based part-of-speech (POS) tagging module and a word-to-pinyin conversion module. In the word-to-pinyin conversion, a word grapheme is augmented by its POS tag in an effort to resolve the pronunciation disambiguation in G2P. The G2P experiments show that the polyphone G2P accuracy is improved by 9.41% after introducing POS module and further improved by 1.39% while applying the proposed word-topinyin method.
Guohong Fu, Haizhou Li 0001
INTERSPEECH2
2003 An integrated approach for Chinese word segmentation
Guohong Fu, Kang-Kwong Luke
PACLIC1