Junwen Duan

dblp:153/9564 · DBLP profile ↗
← Back
37ranked-venue papers
13as first author
27since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 4 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 5 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 OW-CLIP: Data-Efficient Visual Supervision for Open-World Object Detection via Human-AI Collaboration
abstract
Open-world object detection (OWOD) extends traditional object detection to identifying both known and unknown object, necessitating continuous model adaptation as new annotations emerge. Current approaches face significant limitations: 1) data-hungry training due to reliance on a large number of crowdsourced annotations, 2) susceptibility to "partial feature overfitting," and 3) limited flexibility due to required model architecture modifications. To tackle these issues, we present OW-CLIP, a visual analytics system that provides curated data and enables data-efficient OWOD model incremental training. OW-CLIP implements plug-and-play multimodal prompt tuning tailored for OWOD settings and introduces a novel "Crop-Smoothing" technique to mitigate partial feature overfitting. To meet the data requirements for the training methodology, we propose dual-modal data refinement methods that leverage large language models and cross-modal similarity for data generation and filtering. Simultaneously, we develope a visualization interface that enables users to explore and deliver high-quality annotations-including class-specific visual feature phrases and fine-grained differentiated images. Quantitative evaluation demonstrates that OW-CLIP achieves competitive performance at 89% of state-of-the-art performance while requiring only 3.8% self-generated data, while outperforming SOTA approach when trained with equivalent data volumes. A case study shows the effectiveness of the developed method and the improved annotation quality of our visualization system.
Junwen Duan, Ziyao Kang, Shixia Liu, Jiazhi Xia
IEEE Trans. Vis. Comput. Graph.1
2025 ICA-RAG: Information Completeness Guided Adaptive Retrieval-Augmented Generation for Disease Diagnosis
abstract
Retrieval-Augmented Large Language Models, which integrate external knowledge, have shown remarkable performance in medical domains, including clinical diagnosis. However, existing RAG methods often struggle to tailor retrieval strategies to diagnostic difficulty and input sample informativeness. This limitation leads to excessive and often unnecessary retrieval, impairing computational efficiency and increasing the risk of introducing noise that can degrade diagnostic accuracy. To address this, we propose ICA-RAG (Information Completeness Guided Adaptive Retrieval-Augmented Generation), a novel framework for enhancing RAG reliability in disease diagnosis. ICA-RAG utilizes an adaptive control module to assess the necessity of retrieval based on the input's information completeness. By optimizing retrieval and incorporating knowledge filtering, ICARAG better aligns retrieval operations with clinical requirements. Experiments on three Chinese electronic medical record datasets demonstrate that ICA-RAG significantly outperforms baseline methods, highlighting its effectiveness in clinical diagnosis.
Mingyi Jia, Junwen Duan
BIBM4
2025 MultiPepDec: Decoupled Prompt Learning for Multi-Activity Therapeutic Peptides
abstract
Therapeutic peptides demonstrate significant potential in anti-infection, antitumor, and immunomodulation therapies owing to their high specificity and low toxicity. However, existing computational methods are predominantly limited to single-activity design, restricting their clinical applicability. Here, we present MultiPepDec, a novel decoupled prompt learning framework based on protein language model for concurrent generation of multifunctional peptides, which includes antimicrobial, anticancer, toxic, and metabolic activities. Our approach employs: i) Shared-prompts capturing universal therapeutic patterns via adversarial purification; ii) Private-prompts encoding activity-specific knowledge through contrastive learning, ensuring functional decoupling between four activities. Experimental results demonstrate that generated antimicrobial peptides achieve 80.38% predicted efficacy against E. coli, with comparable performance against most clinically relevant pathogens. This confirms robust broad-spectrum capabilities without requiring pathogen-specific training, while maintaining low computational costs. For other therapeutic activities, the designed sequences not only exhibit the intended biological functions but also show significantly improved diversity. This work establishes a new paradigm for efficient multi-activity peptide design, with potential extensions to other biomolecular engineering domains.
Xingdan Wang, Diya Zhang, Chengwei Ai, Shiqiang Ma, Qiaozhen Meng, Junwen Duan, Fei Guo 0001
BIBM6
2025 medIKAL: Integrating Knowledge Graphs as Assistants of LLMs for Enhanced Clinical Diagnosis on EMRs
abstract
Electronic Medical Records (EMRs), while integral to modern healthcare, present challenges for clinical reasoning and diagnosis due to their complexity and information redundancy. To address this, we proposed medIKAL (Integrating Knowledge Graphs as Assistants of LLMs), a framework that combines Large Language Models (LLMs) with knowledge graphs (KGs) to enhance diagnostic capabilities. medIKAL assigns weighted importance to entities in medical records based on their type, enabling precise localization of candidate diseases within KGs. It innovatively employs a residual network-like approach, allowing initial diagnosis by the LLM to be merged into KG search results. Through a path-based reranking algorithm and a fill-in-the-blank style prompt template, it further refined the diagnostic process. We validated medIKAL’s effectiveness through extensive experiments on a newly introduced open-sourced Chinese EMR dataset, demonstrating its potential to improve clinical diagnosis in real-world settings.
Mingyi Jia, Junwen Duan
COLING2
2025 RUIE: Retrieval-based Unified Information Extraction using Large Language Model
abstract
Unified information extraction (UIE) aims to extract diverse structured information from unstructured text. While large language models (LLMs) have shown promise for UIE, they require significant computational resources and often struggle to generalize to unseen tasks. We propose RUIE (Retrieval-based Unified Information Extraction), a framework that leverages in-context learning for efficient task generalization. RUIE introduces a novel demonstration selection mechanism combining LLM preferences with a keyword-enhanced reward model, and employs a bi-encoder retriever trained through contrastive learning and knowledge distillation. As the first trainable retrieval framework for UIE, RUIE serves as a universal plugin for various LLMs. Experimental results on eight held-out datasets demonstrate RUIE’s effectiveness, with average F1-score improvements of 19.22 and 3.22 compared to instruction-tuning methods and other retrievers, respectively.
Xincheng Liao, Junwen Duan, Yixi Huang, Jianxin Wang 0001
COLING2
2025 DDO: Dual-Decision Optimization for LLM-Based Medical Consultation via Multi-Agent Collaboration
abstract
Large Language Models (LLMs) demonstrate strong generalization and reasoning abilities, making them well-suited for complex decisionmaking tasks such as medical consultation (MC).However, existing LLM-based methods often fail to capture the dual nature of MC, which entails two distinct sub-tasks: symptom inquiry, a sequential decision-making process, and disease diagnosis, a classification problem.This mismatch often results in ineffective symptom inquiry and unreliable disease diagnosis.To address this, we propose DDO, a novel LLM-based framework that performs Dual-Decision Optimization by decoupling the two sub-tasks and optimizing them with distinct objectives through a collaborative multi-agent workflow.Experiments on three real-world MC datasets show that DDO consistently outperforms existing LLM-based approaches and achieves competitive performance with state-ofthe-art generation-based methods, demonstrating its effectiveness in the MC task.The code is available at https://github.com/zh-jia/DDO.
Mingyi Jia, Junwen Duan
EMNLP3
2025 GMD: A Multimodal Framework for AI-Generated Misinformation Detection
abstract
The sophistication of AI-generated misinformation has escalated markedly, with modern systems exploiting multimodal compositions—text, images, and videos—to bypass conventional detection methods reliant on Unimodal analysis. Contemporary detection frameworks must evolve to address these challenges by exploiting cross-modal synergies. This work introduces a bilingual multimodal detection architecture that combines deep feature extraction from visual and textual modalities. Specifically, ResNet-derived encoders process images, while language-specific transformers (Qwen for Chinese, BERT for English) generate contextual embeddings. Empirical validation on benchmark datasets (e.g., Weibo) demonstrates state-of-the-art performance, with 3.6% accuracy gains over existing methods. Ablation tests confirm the necessity of both synthetic augmentation and adaptive fusion, underscoring their roles in refining feature discriminability. These advancements highlight the framework’s capability to address evolving threats in AI-generated misinformation.
Junwen Duan, Wuyue Zhang, Zhende Liu
ICIP1
2025 SigChord: Sniffing Wide Non-sparse Multiband Signals for Terrestrial and Non-terrestrial Networks
abstract
While unencrypted information inspection in physical layer (e.g., open headers) can provide deep insights for optimizing wireless networks, the state-of-the-art (SOTA) methods heavily depend on full sampling rate (a.k.a Nyquist rate), and high-cost radios, due to terrestrial and non-terrestrial networks densely occupying multiple bands across large bandwidth (e.g., from 4G/5G at 0.4–7 GHz to LEO satellite at 4–40 GHz). To this end, we present SigChord, an efficient physical layer inspection system built on low-cost and sub-Nyquist sampling radios. We first design a deep and rule-based interleaving algorithm based on Transformer network to perform spectrum sensing and signal recovery under sub-Nyquist sampling rate, and second, cascade protocol identifier and decoder based on Transformer neural networks to help physical layer packets analysis. We implement SigChord using software-defined radio platforms, and extensively evaluate it on over-the-air terrestrial and non-terrestrial wireless signals. The experiments demonstrate that SigChord delivers over 99% accuracy in detecting and decoding, while still decreasing 34% sampling rate, compared with the SOTA approaches.
Jinbo Peng, Junwen Duan, Zheng Lin 0001, Haoxuan Yuan, Yue Gao 0001, Zhe Chen 0015
MobiSys2
2025 Efficient word segmentation for enhancing Chinese spelling check in pre-trained language model
Dafu Tang, Youran Shan, Junwen Duan
Knowl. Inf. Syst.5
2024 Faster Stochastic Variance Reduction Methods for Compositional MiniMax Optimization
abstract
This paper delves into the realm of stochastic optimization for compositional minimax optimization—a pivotal challenge across various machine learning domains, including deep AUC and reinforcement learning policy evaluation. Despite its significance, the problem of compositional minimax optimization is still under-explored. Adding to the complexity, current methods of compositional minimax optimization are plagued by sub-optimal complexities or heavy reliance on sizable batch sizes. To respond to these constraints, this paper introduces a novel method, called Nested STOchastic Recursive Momentum (NSTORM), which can achieve the optimal sample complexity and obtain the nearly accuracy solution, matching the existing minimax methods. We also demonstrate that NSTORM can achieve the same sample complexity under the Polyak-Lojasiewicz (PL)-condition—an insightful extension of its capabilities. Yet, NSTORM encounters an issue with its requirement for low learning rates, potentially constraining its real-world applicability in machine learning. To overcome this hurdle, we present ADAptive NSTORM (ADA-NSTORM) with adaptive learning rates. We demonstrate that ADA-NSTORM can achieve the same sample complexity but the experimental results show its more effectiveness. All the proposed complexities indicate that our proposed methods can match lower bounds to existing minimax optimizations, without requiring a large batch size in each iteration. Extensive experiments support the efficiency of our proposed methods.
Jin Liu 0012, Xiaokang Pan, Junwen Duan, Hongdong Li, Youqi Li
AAAI3
2024 MARE: Multi-Aspect Rationale Extractor on Unsupervised Rationale Extraction
abstract
Unsupervised rationale extraction aims to extract text snippets to support model predictions without explicit rationale annotation.Researchers have made many efforts to solve this task.Previous works often encode each aspect independently, which may limit their ability to capture meaningful internal correlations between aspects.While there has been significant work on mitigating spurious correlations, our approach focuses on leveraging the beneficial internal correlations to improve multi-aspect rationale extraction.In this paper, we propose a Multi-Aspect Rationale Extractor (MARE) to explain and predict multiple aspects simultaneously.Concretely, we propose a Multi-Aspect Multi-Head Attention (MAMHA) mechanism based on hard deletion to encode multiple text chunks simultaneously.Furthermore, multiple special tokens are prepended in front of the text with each corresponding to one certain aspect.Finally, multi-task training is deployed to reduce the training overhead.Experimental results on two unsupervised rationale extraction benchmarks show that MARE achieves state-ofthe-art performance.Ablation studies further demonstrate the effectiveness of our method.
Junwen Duan, Jianxin Wang 0001
EMNLP2
2024 Pre-trained Feature Fusion and Matching for Mild Cognitive Impairment Detection
abstract
Effective diagnosis of Mild Cognitive Impairment (MCI), a preclinical stage of cognitive decline, is significant for delaying disease progression.While most current spontaneous speechbased diagnostic methods focus on English speech, the Interspeech 2024 TAUKADIAL Challenge proposed an innovative research direction to develop a language-agnostic approach to diagnose MCI.This paper proposes an MCI diagnosis method by analyzing and combining linguistic and acoustic features using the bilingual Chinese-English speech dataset provided by the challenge.We employed a pre-trained multilingual model and expressivity encoder to extract language-agnostic speech features.To overcome the challenges of data scarcity and language diversity, we implemented data augmentation and alignment to enhance the model's generalization.Our approach achieved 77.5% accuracy, demonstrating its effectiveness and potential on cross-lingual data.
Junwen Duan, Fangyuan Wei, Hong-Dong Li, Jin Liu 0012
INTERSPEECH1
2024 HyperMatch: long-form text matching via hypergraph convolutional networks
Junwen Duan, Mingyi Jia, Jianbo Liao, Jianxin Wang 0001
Knowl. Inf. Syst.1
2024 Boundary-Aware Dual Biaffine Model for Sequential Sentence Classification in Biomedical Documents
abstract
Assigning appropriate rhetorical roles, such as "background," "intervention," and "outcome," to sentences in biomedical documents can streamline the process for physicians to locate evidence and resources for medical treatment and decision-making. While sequence labeling and span-based methods are frequently employed for this task, the former disregards a document's semantic structure, resulting in a lack of semantic coherence across continuous sentences. Span-based approaches, on the other hand, either necessitate the enumeration of all potential spans, which can be time-consuming, or may lead to the misclassification of sentences over extended spans. Consequently, an approach is required that models the semantic structure of documents explicitly and captures boundary information to achieve precise and effective sentence labeling in biomedical documents. To address these challenges, we propose a new approach, the boundary-aware dual biaffine model, which explicitly models the semantic structure of documents and incorporates boundary information via a dual biaffine layer. We introduce a dynamic programming algorithm to minimize missing labels and overlapping predictions, and achieve globally optimal decoding results. We evaluate our approach on three benchmark datasets, namely PubMed 20 k RCT, PubMed-PICO and NICTA-PIBOSO. The experimental results demonstrate that our approach outperforms strong baselines and achieves state-of-the-art performance on PubMed 20 k RCT and PubMed-PICO. Additionally, our method also achieves competitive results on NICTA-PIBOSO.
Junwen Duan, Huai Guo, Fei Guo 0001, Jianxin Wang 0001
IEEE ACM Trans. Comput. Biol. Bioinform.1
2024 Chinese EMR Named Entity Recognition Using Fused Label Relations Based on Machine Reading Comprehension Framework
abstract
Chinese electronic medical record (EMR) presents significant challenges for named entity recognition (NER) due to their specialized nature, unique language features, and diverse expressions. Traditionally, NER is treated as a sequence labeling task, where each token is assigned a label. Recent research has reframed NER within the machine reading comprehension (MRC) framework, extracting entities in a question-answer format, achieving state-of-the-art performance. However, these MRC-based methods have a significant limitation: they extract entities of various types independently, ignoring their interrelations. To address this, we introduce the Fusion Label Relations with MRC (FLR-MRC) model, which enhances the MRC model by implicitly capturing dependencies among entity types. FLR-MRC models interrelations between labels using graph attention networks, integrating these with textual data to identify entities. On the benchmark CMeEE and CCKS2017-CNER datasets, FLR-MRC achieves F1-scores of 0.6652 and 0.9101, respectively, outperforming existing clinical NER methods.
Junwen Duan, Shuyue Liu, Xincheng Liao, Feng Gong, Hailin Yue, Jianxin Wang 0001
IEEE ACM Trans. Comput. Biol. Bioinform.1
2024 Exploiting Conversation-Branch-Tweet HyperGraph Structure to Detect Misinformation on Social Media
abstract
The spread of misinformation on social media is a serious issue that can have negative consequences for public health and political stability. While detecting and identifying misinformation can be challenging, many attempts have been made to address this problem. However, traditional models that focus on pairwise relationships on misinformation propagation paths may not be effective in capturing the underlying connections among multiple tweets. To address this limitation, the proposed “Conversation-Branch-Tweet” hypergraph convolutional network (CBT-HGCN) uses a hypergraph to represent the internal structure and content of tweet data, with tweets and their replies viewed as nodes and hyperedges, respectively. The model first pre-processes the tweets of a conversation and then uses a pre-trained model as an encoder to extract node information. Finally, a hypergraph convolution network is used as an information fuser for classification. Experimental results on three benchmark datasets (Twitter15, Twitter16, and Pheme) show that the proposed model outperforms several strong baseline models and achieves state-of-the-art performance. This indicates that the CBT-HGCN approach is effective in detecting and identifying misinformation on social media by capturing the underlying connections among multiple tweets.
Fangfang Li 0004, Junwen Duan, Xingliang Mao, Heyuan Shi, Shichao Zhang 0001
ACM Trans. Knowl. Discov. Data3
2023 FBC: Fusing Bi-Encoder and Cross-Encoder for Long-Form Text Matching
abstract
Semantic text matching has a wide range of applications in natural language processing. Recently proposed models that have achieved excellent results on short text matching tasks are not well suited to long-form text matching problems due to input length limitations and increased noise. On the other hand, long-form texts contain a large amount of information at different granularities after encoding, which cannot be fully interacted and utilized by existing methods. To address above issues, we propose a novel long-form text-matching framework which fuses Bi-Encoder and Cross-Encoder (FBC). Specially, it first employs an entity-driven key sentence extraction method to obtain the crucial content of the text and filter out noise. Subsequently, it integrates Bi-Encoder and Cross-Encoder to better capture semantic features and matching signals. Extensive experiments on several publicly available datasets demonstrate the effectiveness of our approach, compared with strong baselines. Furthermore, our model exhibits greater stability and accuracy in determining the matching relationship between documents describing the same event, which outperforms previously established approaches. The code is released at https://github.com/CSU-NLP-Group/FBC.
Jianbo Liao, Mingyi Jia, Junwen Duan, Jianxin Wang 0001
ECAI3
2023 MHLAT: Multi-Hop Label-Wise Attention Model for Automatic ICD Coding
abstract
International Classification of Diseases (ICD) coding is the task of assigning ICD diagnosis codes to clinical notes. This can be challenging given the large quantity of labels (nearly 9,000) and lengthy texts (up to 8,000 tokens). However, unlike the single-pass reading process in previous works, humans tend to read the text and label definitions again to get more confident answers. Moreover, although pretrained language models have been used to address these problems, they suffer from huge memory usage. To address the above problems, we propose a simple but effective model called the Multi-Hop Label-wise ATtention (MHLAT), in which multi-hop label-wise attention is deployed to get more precise and informative representations. Extensive experiments on three benchmark MIMIC datasets indicate that our method achieves significantly better or competitive performance on all seven metrics, with much fewer parameters to optimize.
Junwen Duan
ICASSP1
2023 IK-DDI: a novel framework based on instance position embedding and key external text for DDI extraction
abstract
Determining drug-drug interactions (DDIs) is an important part of pharmacovigilance and has a vital impact on public health. Compared with drug trials, obtaining DDI information from scientific articles is a faster and lower cost but still a highly credible approach. However, current DDI text extraction methods consider the instances generated from articles to be independent and ignore the potential connections between different instances in the same article or sentence. Effective use of external text data could improve prediction accuracy, but existing methods cannot extract key information from external data accurately and reasonably, resulting in low utilization of external data. In this study, we propose a DDI extraction framework, instance position embedding and key external text for DDI (IK-DDI), which adopts instance position embedding and key external text to extract DDI information. The proposed framework integrates the article-level and sentence-level position information of the instances into the model to strengthen the connections between instances generated from the same article or sentence. Moreover, we introduce a comprehensive similarity-matching method that uses string and word sense similarity to improve the matching accuracy between the target drug and external text. Furthermore, the key sentence search method is used to obtain key information from external data. Therefore, IK-DDI can make full use of the connection between instances and the information contained in external text data to improve the efficiency of DDI extraction. Experimental results show that IK-DDI outperforms existing methods on both macro-averaged and micro-averaged metrics, which suggests our method provides complete framework that can be used to extract relationships between biomedical entities and process external text data.
Mingliang Dou, Jiaqi Ding, Genlang Chen, Junwen Duan, Fei Guo 0001, Jijun Tang
Briefings Bioinform.4
2023 LncLocFormer: a Transformer-based deep learning model for multi-label lncRNA subcellular localization prediction by using localization-specific attention mechanism
abstract
MOTIVATION: There is mounting evidence that the subcellular localization of lncRNAs can provide valuable insights into their biological functions. In the real world of transcriptomes, lncRNAs are usually localized in multiple subcellular localizations. Furthermore, lncRNAs have specific localization patterns for different subcellular localizations. Although several computational methods have been developed to predict the subcellular localization of lncRNAs, few of them are designed for lncRNAs that have multiple subcellular localizations, and none of them take motif specificity into consideration. RESULTS: In this study, we proposed a novel deep learning model, called LncLocFormer, which uses only lncRNA sequences to predict multi-label lncRNA subcellular localization. LncLocFormer utilizes eight Transformer blocks to model long-range dependencies within the lncRNA sequence and shares information across the lncRNA sequence. To exploit the relationship between different subcellular localizations and find distinct localization patterns for different subcellular localizations, LncLocFormer employs a localization-specific attention mechanism. The results demonstrate that LncLocFormer outperforms existing state-of-the-art predictors on the hold-out test set. Furthermore, we conducted a motif analysis and found LncLocFormer can capture known motifs. Ablation studies confirmed the contribution of the localization-specific attention mechanism in improving the prediction performance. AVAILABILITY AND IMPLEMENTATION: The LncLocFormer web server is available at http://csuligroup.com:9000/LncLocFormer. The source code can be obtained from https://github.com/CSUBioGroup/LncLocFormer.
Min Zeng 0004, Yifan Wu 0008, Rui Yin 0002, Chengqian Lu, Junwen Duan, Min Li 0007
Bioinform.6
2022 ASNet: An Adversarial Sparse Network for Multi-task Biomedical Named Entity Recognition
abstract
Biomedical named entity recognition (BioNER) is to extract entities, such as genes and proteins, from biomedical texts, where there is often a lack of high-quality training data. Recent work addresses this issue by multi-task learning with multiple datasets. However, these methods are usually over-parameterized, and some even suffer from negative transfer issues. To address above problems, we propose adversarial sparse sharing mechanism, which trains a sparsely shared encoder on multiple tasks with both task-agnostic and task-specific subnetworks. With adversarial training, we guide the task-agnostic subnetwork to learn shared task-invariant features and the task-specific subnetwork to learn task-dependent features. For a particular task, only the shared and its subnetwork are activated, which greatly reduces the number of parameters and avoids interference among tasks. Experimental results on 15 benchmark BioNER datasets show that our proposed method outperforms or is competitive with baseline methods with fewer parameters. Our code is released at: https://github.com/CSU-NLP-Group/ASNet
Junwen Duan, Huai Guo, Min Zeng 0004, Jianxin Wang 0001
BIBM1
2022 Fusing Label Relations for Chinese EMR Named Entity Recognition with Machine Reading Comprehension
Shuyue Liu, Junwen Duan, Feng Gong, Hailin Yue, Jianxin Wang 0001
ISBRA2
2022 Automatic ICD Coding Based on Multi-granularity Feature Fusion
Junwen Duan, Jianxin Wang 0001
ISBRA2
2022 Multi-task deep learning model based on hierarchical relations of address elements for semantic address matching
Fangfang Li 0004, Yiheng Lu, Xingliang Mao, Junwen Duan, Xiyao Liu 0001
Neural Comput. Appl.4
2021 EFCA: An Extended Formal Concept Analysis Method for Aspect Extraction in Healthcare Informatics
abstract
With the popularity of social media platforms, patients tend to share their experiences and opinions on them, and patient feedback is key to improving health services. Sentiment analysis techniques have been applied to automatically analyze the patients’ opinions to understand the quality of healthcare. Aspect extraction, which aims to identify the opinion targets in the text, is an important step towards understanding the patient’s opinion towards particular target or entity. However, due to the complex nature of medical domain data, existing approaches take much execution time. To address this, we presents a new approach for aspect extraction and refinement to smooth the sentiment analysis process. Furthermore, this work also introduces an intelligent weighting scheme for classifying the final aspects. For the experimental evaluation, a dataset from Yelp and RateMDs has been utilized. Experimental results show that the proposed model outperforms existing methods.
Zohair Ahmed, Junwen Duan, Fang-Xiang Wu, Jianxin Wang 0001
BIBM2
2021 nPTAS: A Novel Platform for Text Annotation and Service
abstract
Natural Language Processing (NLP) is a critical research area in artificial intelligence, which has spawned a variety of useful applications. However, creating an online NLP service from scratch is still challenging for non-expert users, which has to go through complex steps such as corpus annotation, model training and deployment. Existing tools mainly focus on corpus annotation, which rarely support collaborative annotation or online model training and deployment. To meet the requirement of rapid deployment of NLP services, we develop a novel full process platform nPTAS, which supports collaborative annotation, online model training and model deploying with RESTful interfaces. To validate the effectiveness of the platform, we build a clinical entities recognition service from Chinese Electronic Medical Records on it. The experimental results show that nPTAS improves the efficiency and quality of corpus annotation and greatly reduces the effort to build online NLP services. The platform is available at http://nptas.c2cloud.cn.
Junwen Duan, Min Li 0007
BIBM2
2021 Quantifying the effects of long-term news on stock markets on the basis of the multikernel Hawkes process
Jihao Shi, Junwen Duan, Bing Qin 0001, Ting Liu 0001
Sci. China Inf. Sci.3
2019 Event Representation Learning Enhanced with External Commonsense Knowledge
abstract
Xiao Ding, Kuo Liao, Ting Liu, Zhongyang Li, Junwen Duan. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Kuo Liao, Ting Liu 0001, Junwen Duan
EMNLP/IJCNLP (1)5
2019 TEND: A Target-Dependent Representation Learning Framework for News Document
abstract
Real-time news documents published on the Internet have global financial and political impacts. Pioneering statistical approaches investigate manually defined features to capture lexical, sentiment, and event information, which suffer from feature sparsity. As a remedy, recent work has considered learning dense vector representations for documents. Such representations are general, which can not model target-dependent scenarios, such as stance detection towards a specific claim. There has been work on target-specific word and sentence representations, but little was done on target-dependent document representation. Moreover, documents contain more potentially helpful information, but also noise compared to events and sentences. To address the above issues, we focus on models that are: 1. task-driven, which optimize the neural network representations for the end task; 2. target-specific, learning news representations by considering the influence of specific targets. In particular, we propose a novel document-level target-dependent learning framework TEND. The framework employs the information of the target and the news abstract as clues, obtaining relatively informative sentences from the entire document for our objectives. The framework assembles a document representation by integrating the news abstract representation and a weighted sum of sentence representations in the document. To the best of our knowledge, we are among the first to investigate target-dependent document representation. Existing text representation models can be easily integrated into our TEND framework, and it is general enough to be applied to different target-dependent document representation tasks. We empirically evaluate our framework on two target-dependent document-level tasks, including a cumulative abnormal return prediction task and a news stance detection task. Results show that our models give the best performances compared to state-of-the-art document embedding methods, yielding robust and consistent performances across datasets.
Junwen Duan, Yue Zhang 0004, Ting Liu 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2018 Learning Target-Specific Representations of Financial News Documents For Cumulative Abnormal Return Prediction
abstract
Texts from the Internet serve as important data sources for financial market modeling. Early statistical approaches rely on manually defined features to capture lexical, sentiment and event information, which suffers from feature sparsity. Recent work has considered learning dense representations for news titles and abstracts. Compared to news titles, full documents can contain more potentially helpful information, but also noise compared to events and sentences, which has been less investigated in previous work. To fill this gap, we propose a novel target-specific abstract-guided news document representation model. The model uses a target-sensitive representation of the news abstract to weigh sentences in the news content, so as to select and combine the most informative sentences for market modeling. Results show that document representations can give better performance for estimating cumulative abnormal returns of companies when compared to titles and abstracts. Our model is especially effective when it used to combine information from multiple document sources compared to the sentence-level baselines.
Junwen Duan, Yue Zhang 0004, Ching-Yun Chang, Ting Liu 0001
COLING1
2018 Learning Sentence Representations over Tree Structures for Target-Dependent Classification
abstract
Junwen Duan, Xiao Ding, Ting Liu. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Junwen Duan, Ting Liu 0001
NAACL-HLT1
2017 A Gaussian copula regression model for movie box-office revenues prediction
Junwen Duan, Ting Liu 0001
Sci. China Inf. Sci.1
2016 Knowledge-Driven Event Embedding for Stock Prediction
abstract
Representing structured events as vectors in continuous space offers a new way for defining dense features for natural language processing (NLP) applications. Prior work has proposed effective methods to learn event representations that can capture syntactic and semantic information over text corpus, demonstrating their effectiveness for downstream tasks such as event-driven stock prediction. On the other hand, events extracted from raw texts do not contain background knowledge on entities and relations that they are mentioned. To address this issue, this paper proposes to leverage extra information from knowledge graph, which provides ground truth such as attributes and properties of entities and encodes valuable relations between entities. Specifically, we propose a joint model to combine knowledge graph information into the objective function of an event embedding learning model. Experiments on event similarity and stock market prediction show that our model is more capable of obtaining better event embeddings and making more accurate prediction on stock market volatilities.
Yue Zhang 0004, Ting Liu 0001, Junwen Duan
COLING4
2015 Mining User Consumption Intention from Social Media Using Domain Adaptive Convolutional Neural Network
abstract
Social media platforms are often used by people to express their needs and desires. Such data offer great opportunities to identify users’ consumption intention from user-generated contents, so that better tailored products or services can be recommended. However, there have been few efforts on mining commercial intents from social media contents. In this paper, we investigate the use of social media data to identify consumption intentions for individuals. We develop a Consumption Intention Mining Model (CIMM) based on convolutional neural network (CNN), for identifying whether the user has a consumption intention. The task is domain-dependent, and learning CNN requires a large number of annotated instances, which can be available only in some domains. Hence, we investigate the possibility of transferring the CNN mid-level sentence representation learned from one domain to another by adding an adaptation layer. To demonstrate the effectiveness of CIMM, we conduct experiments on two domains. Our results show that CIMM offers a powerful paradigm for effectively identifying users’ consumption intention based on their social media data. Moreover, our results also confirm that the CNN learned in one domain can be effectively transferred to another domain. This suggests that a great potential for our model to significantly increase effectiveness of product recommendations and targeted advertising.
Ting Liu 0001, Junwen Duan, Jian-Yun Nie
AAAI3
2015 Deep Learning for Event-Driven Stock Prediction
Yue Zhang 0004, Ting Liu 0001, Junwen Duan
IJCAI4
2015 Mining Intention-Related Products on Online Q&A Community
Junwen Duan, Yiheng Chen, Ting Liu 0001
J. Comput. Sci. Technol.1
2014 Using Structured Events to Predict Stock Price Movement: An Empirical Investigation
abstract
It has been shown that news events influence the trends of stock price movements.However, previous work on news-driven stock market prediction rely on shallow features (such as bags-of-words, named entities and noun phrases), which do not capture structured entity-relation information, and hence cannot represent complete and exact events.Recent advances in Open Information Extraction (Open IE) techniques enable the extraction of structured events from web-scale data.We propose to adapt Open IE technology for event-based stock price movement prediction, extracting structured events from large-scale public news without manual efforts.Both linear and nonlinear models are employed to empirically investigate the hidden and complex relationships between events and the stock market.Largescale experiments show that the accuracy of S&P 500 index prediction is 60%, and that of individual stock prediction can be over 70%.Our event-based system outperforms bags-of-words-based baselines, and previously reported systems trained on S&P 500 stock historical data.
Yue Zhang 0004, Ting Liu 0001, Junwen Duan
EMNLP4