EDBT 2026 Demo / reviewers in the wild / expert
Jinlan Fu
dblp:218/7289
· DBLP profile ↗
36ranked-venue papers
11as first author
23since 2021 · last 2026
0000-0002-0370-1238ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 33 · 9 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video UnderstandingabstractRecent advancements in Multimodal Large Language Models (MLLMs) have demonstrated significant improvement in offline video understanding. However, extending these capabilities to streaming video inputs, remains challenging, as existing models struggle to simultaneously maintain stable understanding performance, real-time responses, and low GPU memory overhead. To address this challenge, we propose HERMES, a novel training-free architecture for real-time and accurate understanding of video streams. Based on a mechanistic attention investigation, we conceptualize KV cache as a hierarchical memory framework that encapsulates video information across multiple granularities. During inference, HERMES reuses a compact KV cache, enabling efficient streaming understanding under resource constraints. Notably, HERMES requires no auxiliary computations upon the arrival of user queries, thereby guaranteeing real-time responses for continuous video stream interactions. HERMES achieves 10\times faster TTFT compared to prior SOTA. Even when reducing video tokens by up to 68% compared with uniform sampling, HERMES achieves superior or comparable accuracy across all benchmarks, with up to 11.4% gains on streaming datasets. Shudong Yang, Jinlan Fu, See-Kiong Ng, Xipeng Qiu |
ACL (1) | 3 |
| 2026 | S HARING B EYOND D ECISION : Deep Collaboration between Large Language Models via Representation EnsembleabstractAbstract Large Language Models (LLMs) exhibit unique strengths arising from differences in model architecture, training data, and strategies. Ensemble learning has been explored to leverage these complementary strengths through decision-level sharing (i.e.,Decision Ensemble), which combines the predictions from multiple LLMs. However, such methods integrate only shallow decisions and overlook the exchange of deeper levels of information within the internal representations of LLMs, such as problem understanding, world knowledge, and latent reasoning patterns. In this work, we propose Representation Ensemble (RISE), a novel ensemble framework that enables cross-LLM representation sharing for richer information exchange. To address challenges of representation-level interaction caused by layer misalignment and latent-space incompatibility across LLMs, we introduce a representation alignment method based on relational similarity measures and an orthogonal latent-space transformation. Experimental results show that (1) RISE achieves performance competitive with existing decision ensemble methods, and (2) RISE is strongly complementary to decision ensemble, with their combination boosting collaboration gains by 14%–41%. Finally, we further compare ensemble of small LLMs to a single larger LLM and to model merging and composition approaches, and find that ensemble learning consistently generalizes well without additional training. Yichong Huang, Jinlan Fu, Xiachong Feng, Baohang Li, Zekai Ye, Libo Qin 0001, Hao Fei 0001, See-Kiong Ng, Bing Qin 0001 |
Trans. Assoc. Comput. Linguistics | 3 |
| 2025 | Investigating and Enhancing the Robustness of Large Multimodal Models Against Temporal InconsistencyabstractJiafeng Liang, Shixin Jiang, Xuan Dong, Ning Wang, Zheng Chu, Hui Su, Jinlan Fu, Ming Liu, See-Kiong Ng, Bing Qin. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Jiafeng Liang, Shixin Jiang, Ning Wang 0020, Hui Su, Jinlan Fu, Ming Liu 0004, See-Kiong Ng, Bing Qin 0001 |
ACL (1) | 7 |
| 2025 | World Modeling Makes a Better Planner: Dual Preference Optimization for Embodied Task PlanningabstractRecent advances in large vision-language models (LVLMs) have shown promise for embodied task planning, yet they struggle with fundamental challenges like dependency constraints and efficiency. Existing approaches either solely optimize action selection or directly leverage pre-trained models as world models during inference, overlooking the benefits of learning to model the world as a way to enhance planning capabilities. We propose Dual Preference Optimization (D^2PO), a new learning framework that jointly optimizes state prediction and action selection through preference learning, enabling LVLMs to understand environment dynamics for better planning. To automatically collect trajectories and stepwise preference data without human annotation, we introduce a tree search mechanism for extensive exploration via trial-and-error. Extensive experiments on VoTa-Bench demonstrate that our D^2PO-based method significantly outperforms existing methods and GPT-4o when applied to Qwen2-VL (7B), LLaVA-1.6 (7B), and LLaMA-3.2 (11B), achieving superior task success rates with more efficient execution paths. Siyin Wang, Zhaoye Fei, Qinyuan Cheng, Shiduo Zhang, Panpan Cai, Jinlan Fu, Xipeng Qiu |
ACL (1) | 6 |
| 2025 | Multi-Layer Visual Feature Fusion in Multimodal LLMs: Methods, Analysis, and Best PracticesabstractMultimodal Large Language Models (MLLMs) have made significant advancements in recent years, with visual features playing an increasingly critical role in enhancing model performance. However, the integration of multi-layer visual features in MLLMs remains underexplored, particularly with regard to optimal layer selection and fusion strategies. Existing methods often rely on arbitrary design choices, leading to suboptimal outcomes. In this paper, we systematically investigate two core aspects of multi-layer visual feature fusion: (1) selecting the most effective visual layers and (2) identifying the best fusion approach with the language model. Our experiments reveal that while combining visual features from multiple stages improves generalization, incorporating additional features from the same stage typically leads to diminished performance. Furthermore, we find that direct fusion of multi-layer visual features at the input stage consistently yields superior and more stable performance across various configurations. We make all our code publicly available: https://github.com/EIT-NLP/Layer_Select_Fuse_for_MLLM. Junyan Lin, Yingqi Fan, Hui Su, Jinlan Fu |
CVPR | 7 |
| 2025 | SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language ModelsabstractThe emergence of Vision Language Models (VLMs) has brought unprecedented advances in understanding multi-modal information. The combination of textual and visual semantics in VLMs is highly complex and diverse, making the safety alignment of these models challenging. Furthermore, due to the limited study on the safety alignment of VLMs, there is a lack of large-scale, high-quality datasets. To address these limitations, we propose a Safety Preference Alignment dataset for Vision Language Models named SPA-VL. In terms of breadth, SPA-VL covers 6 harmfulness domains, 13 categories, and 53 subcategories, and contains 100,788 samples of the quadruple (question, image, chosen response, rejected response). In terms of depth, the responses are collected from 12 open-source (e.g., QwenVL) and closed-source (e.g., Gemini) VLMs to ensure diversity. The construction of preference data is fully automated, and the experimental results indicate that models trained with alignment techniques on the SPA-VL dataset exhibit substantial improvements in harmlessness and helpfulness while maintaining core capabilities. SPA-VL, as a large-scale, high-quality, and diverse dataset, represents a significant milestone in ensuring that VLMs achieve both harmlessness and helpfulness. Yongting Zhang, Lu Chen 0001, Guodong Zheng, Yifeng Gao 0002, Jinlan Fu, Zhenfei Yin, Senjie Jin, Yu Qiao 0001, Xuanjing Huang 0001, Feng Zhao 0004, Tao Gui |
CVPR | 6 |
| 2025 | Multimodal Language Models See Better When They Look ShallowerabstractMultimodal large language models (MLLMs) typically extract visual features from the final layers of a pretrained Vision Transformer (ViT).This widespread deep-layer bias, however, is largely driven by empirical convention rather than principled analysis.While prior studies suggest that different ViT layers capture different types of information-shallower layers focusing on fine visual details and deeper layers aligning more closely with textual semantics, the impact of this variation on MLLM performance remains underexplored.We present the first comprehensive study of visual layer selection for MLLMs, analyzing representation similarity across ViT layers to establish shallow, middle, and deep layer groupings.Through extensive evaluation of MLLMs (1.4B-7B parameters) across 10 benchmarks encompassing 60+ tasks, we find that while deep layers excel in semantic-rich tasks like OCR, shallow and middle layers significantly outperform them on fine-grained visual tasks including counting, positioning, and object localization.Building on these insights, we propose a lightweight feature fusion method that strategically incorporates shallower layers, achieving consistent improvements over both single-layer and specialized fusion baselines.Our work offers the first principled study of visual layer selection in MLLMs, showing that MLLMs can often see better when they look shallower. Junyan Lin, Xinghao Chen 0009, Jianfeng Dong, Xin Jin 0014, Hui Su, Jinlan Fu, Xiaoyu Shen 0001 |
EMNLP | 8 |
| 2025 | VisiPruner: Decoding Discontinuous Cross-Modal Dynamics for Efficient Multimodal LLMsabstractMultimodal Large Language Models (MLLMs) have achieved strong performance across vision-language tasks, but suffer from significant computational overhead due to the quadratic growth of attention computations with the number of multimodal tokens.Though efforts have been made to prune tokens in MLLMs, they lack a fundamental understanding of how MLLMs process and fuse multimodal information.Through systematic analysis, we uncover a three-stage cross-modal interaction process: (1) Shallow layers recognize task intent, with visual tokens acting as passive attention sinks; (2) Cross-modal fusion occurs abruptly in middle layers, driven by a few critical visual tokens; (3) Deep layers discard vision tokens, focusing solely on linguistic refinement.Based on these findings, we propose VisiPruner, a training-free pruning framework that reduces up to 99% of visionrelated attention computations and 53.9% of FLOPs on LLaVA-v1.5 7B.It significantly outperforms existing token pruning methods and generalizes across diverse MLLMs.Beyond pruning, our insights further provide actionable guidelines for training efficient MLLMs by aligning model architecture with its intrinsic layer-wise processing dynamics. Yingqi Fan, Anhao Zhao, Jinlan Fu, Junlong Tong, Hui Su, Yijie Pan, Wei Zhang 0185, Xiaoyu Shen 0001 |
EMNLP | 3 |
| 2025 | CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMsabstractMultimodal Large Language Models (MLLMs) still struggle with hallucinations despite their impressive capabilities. Recent studies have attempted to mitigate this by applying Direct Preference Optimization (DPO) to multimodal scenarios using preference pairs from text-based responses. However, our analysis of representation distributions reveals that multimodal DPO struggles to align image and text representations and to distinguish between hallucinated and non-hallucinated descriptions. To address these challenges,
In this work, we propose a Cross-modal Hierarchical Direct Preference Optimization (CHiP) to address these limitations.
We introduce a visual preference optimization module within the DPO framework, enabling MLLMs to learn from both textual and visual preferences simultaneously. Furthermore, we propose a hierarchical textual preference optimization module that allows the model to capture preferences at multiple granular levels, including response, segment, and token levels. We evaluate CHiP through both quantitative and qualitative analyses, with results across multiple benchmarks demonstrating its effectiveness in reducing hallucinations. On the Object HalBench dataset, CHiP outperforms DPO in hallucination reduction, achieving improvements of 52.7% and 55.5% relative points based on the base model Muffin and LLaVA models, respectively. We make all our datasets and code publicly available. Jinlan Fu, Shenzhen Huangfu, Hao Fei 0001, Xiaoyu Shen 0001, Bryan Hooi, Xipeng Qiu, See-Kiong Ng |
ICLR | 1 |
| 2025 | FlipAttack: Jailbreak LLMs via FlippingabstractThis paper proposes a simple yet effective jailbreak attack named FlipAttack against black-box LLMs. First, from the autoregressive nature, we reveal that LLMs tend to understand the text from left to right and find that they struggle to comprehend the text when the perturbation is added to the left side. Motivated by these insights, we propose to disguise the harmful prompt by constructing a left-side perturbation merely based on the prompt itself, then generalize this idea to 4 flipping modes. Second, we verify the strong ability of LLMs to perform the text-flipping task and then develop 4 variants to guide LLMs to understand and execute harmful behaviors accurately. These designs keep FlipAttack universal, stealthy, and simple, allowing it to jailbreak black-box LLMs within only 1 query. Experiments on 8 LLMs demonstrate the superiority of FlipAttack. Remarkably, it achieves $\sim$78.97% attack success rate across 8 LLMs on average and $\sim$98% bypass rate against 5 guard models on average. Yue Liu 0008, Xiao-Xin He, Miao Xiong, Jinlan Fu, Shumin Deng, Yingwei Ma, Jiaheng Zhang, Bryan Hooi |
ICML | 4 |
| 2025 | VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models
Haojian Huang, Shengqiong Wu, Meng Luo 0010, Jinlan Fu, Xinya Du, Hanwang Zhang, Hao Fei 0001 |
ICML | 5 |
| 2025 | MCM-DPO: Multifaceted Cross-Modal Direct Preference Optimization for Alt-text Generation
Jinlan Fu, Shenzhen Huangfu, Hao Fei 0001, Yichong Huang, Xiaoyu Shen 0001, Xipeng Qiu, See-Kiong Ng |
ACM Multimedia | 1 |
| 2024 | Unveiling In-Context Learning: A Coordinate System to Understand Its Working MechanismabstractLarge language models (LLMs) exhibit remarkable in-context learning (ICL) capabilities.However, the underlying working mechanism of ICL remains poorly understood.Recent research presents two conflicting views on ICL: One emphasizes the impact of similar examples in the demonstrations, stressing the need for label correctness and more shots.The other attributes it to LLMs' inherent ability of task recognition, deeming label correctness and shot numbers of demonstrations as not crucial.In this work, we provide a Two-Dimensional Coordinate System that unifies both views into a systematic framework.The framework explains the behavior of ICL through two orthogonal variables: whether similar examples are presented in the demonstrations (perception) and whether LLMs can recognize the task (cognition).We propose the peak inverse rank metric to detect the task recognition ability of LLMs and study LLMs' reactions to different definitions of similarity.Based on these, we conduct extensive experiments to elucidate how ICL functions across each quadrant on multiple representative classification tasks.Finally, we extend our analyses to generation tasks, showing that our coordinate system can also be used to interpret ICL for generation tasks effectively. Anhao Zhao, Fanghua Ye 0001, Jinlan Fu, Xiaoyu Shen 0001 |
EMNLP | 3 |
| 2024 | GPTScore: Evaluate as You DesireabstractJinlan Fu, See-Kiong Ng, Zhengbao Jiang, Pengfei Liu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, Pengfei Liu 0003 |
NAACL-HLT | 1 |
| 2023 | CET2: Modelling Topic Transitions for Coherent and Engaging Knowledge-Grounded ConversationsabstractKnowledge-grounded dialogue systems aim to generate coherent and engaging responses based on the dialogue contexts and selected external knowledge. Previous knowledge selection methods tend to rely too heavily on the dialogue contexts or over-emphasize the new information in the selected knowledge, resulting in the selection of repetitious or incongruous knowledge and further generating repetitive or incoherent responses, as the generation of the response depends on the chosen knowledge. To address these shortcomings, we introduce a Coherent and Engaging Topic Transition (CET2) framework to model topic transitions for selecting knowledge that is coherent to the context of the conversations while providing adequate knowledge diversity for topic development. Our CET2 framework considers multiple factors for knowledge selection, including valid transition logic from dialogue contexts to the following topics and systematic comparisons between available knowledge candidates. Extensive experiments on two public benchmarks demonstrate the superiority and the better generalization ability of CET2 on knowledge selection. This is due to our well-designed transition features and comparative knowledge selection strategy, which are more transferable to conversations about unseen topics. Analysis of fine-grained knowledge selection accuracy also shows that CET2 can better balance topic entailment (contextual coherence) and development (knowledge diversity) in dialogue than existing approaches. Qixian Zhou, Jinlan Fu, See-Kiong Ng |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | CorefDiffs: Co-referential and Differential Knowledge Flow in Document Grounded ConversationsabstractKnowledge-grounded dialog systems need to incorporate smooth transitions among knowledge selected for generating responses, to ensure that dialog flows naturally. For document-grounded dialog systems, the inter- and intra-document knowledge relations can be used to model such conversational flows. We develop a novel Multi-Document Co-Referential Graph (Coref-MDG) to effectively capture the inter-document relationships based on commonsense and similarity and the intra-document co-referential structures of knowledge segments within the grounding documents. We propose CorefDiffs, a Co-referential and Differential flow management method, to linearize the static Coref-MDG into conversational sequence logic. CorefDiffs performs knowledge selection by accounting for contextual graph structures and the knowledge difference sequences. CorefDiffs significantly outperforms the state-of-the-art by 9.5%, 7.4% and 8.2% on three public benchmarks. This demonstrates that the effective modeling of co-reference and knowledge difference for dialog flows are critical for transitions in document-grounded conversation. Qixian Zhou, Jinlan Fu, Min-Yen Kan, See-Kiong Ng |
COLING | 3 |
| 2022 | Polyglot Prompt: Multilingual Multitask Prompt TrainingabstractThis paper aims for a potential architectural improvement for multilingual learning and asks: Can different tasks from different languages be modeled in a monolithic framework, i.e. without any task/language-specific module?The benefit of achieving this could open new doors for future multilingual research, including allowing systems trained on low resources to be further assisted by other languages as well as other tasks.We approach this goal by developing a learning framework named Polyglot Prompting to exploit prompting methods for learning a unified semantic space for different languages and tasks with multilingual prompt engineering.We performed a comprehensive evaluation of 6 tasks, namely topic classification, sentiment classification, named entity recognition, question answering, natural language inference, and summarization, covering 24 datasets and 49 languages.The experimental results demonstrated the efficacy of multilingual multitask prompt-based learning and led to inspiring observations.We also present an interpretable multilingual evaluation methodology and show how the proposed framework, multilingual multitask prompt training, works.We release all datasets prompted in the best setting and code. 1 Jinlan Fu, See-Kiong Ng, Pengfei Liu 0003 |
EMNLP | 1 |
| 2022 | Are All the Datasets in Benchmark Necessary? A Pilot Study of Dataset Evaluation for Text ClassificationabstractIn this paper, we ask the research question of whether all the datasets in the benchmark are necessary.We approach this by first characterizing the distinguishability of datasets when comparing different systems.Experiments on 9 datasets and 36 systems show that several existing benchmark datasets contribute little to discriminating top-scoring systems, while those less used datasets exhibit impressive discriminative power.We further, taking the text classification task as a case study, investigate the possibility of predicting dataset discrimination based on its properties (e.g., average sentence length).Our preliminary experiments promisingly show that given a sufficient number of training experimental records, a meaningful predictor can be learned to estimate dataset discrimination over unseen datasets.We released all datasets with features explored in this work on DataLab. Jinlan Fu, See-Kiong Ng, Pengfei Liu 0003 |
NAACL-HLT | 2 |
| 2021 | SpanNER: Named Entity Re-/Recognition as Span PredictionabstractJinlan Fu, Xuanjing Huang, Pengfei Liu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Jinlan Fu, Xuanjing Huang 0001, Pengfei Liu 0003 |
ACL/IJCNLP (1) | 1 |
| 2021 | Towards More Fine-grained and Reliable NLP Performance PredictionabstractPerformance prediction, the task of estimating a system's performance without performing experiments, allows us to reduce the experimental burden caused by the combinatorial explosion of different datasets, languages, tasks, and models.In this paper, we make two contributions to improving performance prediction for NLP tasks.First, we examine performance predictors not only for holistic measures of accuracy like F1 or BLEU, but also fine-grained performance measures such as accuracy over individual classes of examples.Second, we propose methods to understand the reliability of a performance prediction model from two angles: confidence intervals and calibration.We perform an analysis of four types of NLP tasks, and both demonstrate the feasibility of fine-grained performance prediction and the necessity to perform reliability analysis for performance prediction methods in the future.We make our code publicly available Zihuiwen Ye, Pengfei Liu 0003, Jinlan Fu, Graham Neubig |
EACL | 3 |
| 2021 | XTREME-R: Towards More Challenging and Nuanced Multilingual EvaluationabstractSebastian Ruder, Noah Constant, Jan Botha, Aditya Siddhant, Orhan Firat, Jinlan Fu, Pengfei Liu, Junjie Hu, Dan Garrette, Graham Neubig, Melvin Johnson. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Sebastian Ruder, Noah Constant, Jan A. Botha, Aditya Siddhant, Orhan Firat, Jinlan Fu, Pengfei Liu 0003, Junjie Hu 0001, Dan Garrette, Graham Neubig, Melvin Johnson |
EMNLP (1) | 6 |
| 2021 | A Partition Filter Network for Joint Entity and Relation ExtractionabstractIn joint entity and relation extraction, existing work either sequentially encode task-specific features, leading to an imbalance in inter-task feature interaction where features extracted later have no direct contact with those that come first. Or they encode entity features and relation features in a parallel manner, meaning that feature representation learning for each task is largely independent of each other except for input sharing. We propose a partition filter network to model two-way interaction between tasks properly, where feature encoding is decomposed into two steps: partition and filter. In our encoder, we leverage two gates: entity and relation gate, to segment neurons into two task partitions and one shared partition. The shared partition represents inter-task information valuable to both tasks and is evenly shared across two tasks to ensure proper two-way interaction. The task partitions represent intra-task information and are formed through concerted efforts of both gates, making sure that encoding of task-specific features is dependent upon each other. Experiment results on six public datasets show that our model performs significantly better than previous approaches. In addition, contrary to what previous work has claimed, our auxiliary experiments suggest that relation prediction is contributory to named entity prediction in a non-negligible way. The source code can be found at https://github.com/Coopercoppers/PFN. Zhiheng Yan, Jinlan Fu, Qi Zhang 0001, Zhongyu Wei |
EMNLP (1) | 3 |
| 2021 | Larger-Context Tagging: When and Why Does It Work?abstractJinlan Fu, Liangjing Feng, Qi Zhang, Xuanjing Huang, Pengfei Liu. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Jinlan Fu, Liangjing Feng, Qi Zhang 0001, Xuanjing Huang 0001, Pengfei Liu 0003 |
NAACL-HLT | 1 |
| 2020 | Rethinking Generalization of Neural Models: A Named Entity Recognition Case StudyabstractWhile neural network-based models have achieved impressive performance on a large body of NLP tasks, the generalization behavior of different models remains poorly understood: Does this excellent performance imply a perfect generalization model, or are there still some limitations? In this paper, we take the NER task as a testbed to analyze the generalization behavior of existing models from different perspectives and characterize the differences of their generalization abilities through the lens of our proposed measures, which guides us to better design models and training methods. Experiments with in-depth analyses diagnose the bottleneck of existing neural NER models in terms of breakdown performance analysis, annotation errors, dataset bias, and category relationships, which suggest directions for improvement. We have released the datasets: (ReCoNLL, PLONER) for the future research at our project page: http://pfliu.com/InterpretNER/. Jinlan Fu, Pengfei Liu 0003, Qi Zhang 0001 |
AAAI | 1 |
| 2020 | Interpretable Multi-dataset Evaluation for Named Entity RecognitionabstractWith the proliferation of models for natural language processing tasks, it is even harder to understand the differences between models and their relative merits.Simply looking at differences between holistic metrics such as accuracy, BLEU, or F1 does not tell us why or how particular methods perform differently and how diverse datasets influence the model design choices.In this paper, we present a general methodology for interpretable evaluation for the named entity recognition (NER) task.The proposed evaluation method enables us to interpret the differences in models and datasets, as well as the interplay between them, identifying the strengths and weaknesses of current systems.By making our analysis tool available, we make it easy for future researchers to run similar analyses and drive progress in this area: https: //github.com/neulab/InterpretEval. Jinlan Fu, Pengfei Liu 0003, Graham Neubig |
EMNLP (1) | 1 |
| 2020 | RethinkCWS: Is Chinese Word Segmentation a Solved Task?abstractThe performance of the Chinese Word Segmentation (CWS) systems has gradually reached a plateau with the rapid development of deep neural networks, especially the successful use of large pre-trained models.In this paper, we take stock of what we have achieved and rethink what's left in the CWS task.Methodologically, we propose a finegrained evaluation for existing CWS systems, which not only allows us to diagnose the strengths and weaknesses of existing models (under the in-dataset setting), but enables us to quantify the discrepancy between different criterion and alleviate the negative transfer problem when doing multi-criteria learning.Strategically, despite not aiming to propose a novel model in this paper, our comprehensive experiments on eight models and seven datasets, as well as thorough analysis, could search for some promising direction for future research.We make all codes publicly available and release an interface that can quickly evaluate and diagnose user's models: https://github. com/neulab/InterpretEval. Jinlan Fu, Pengfei Liu 0003, Qi Zhang 0001, Xuanjing Huang 0001 |
EMNLP (1) | 1 |
| 2020 | A Knowledge-Aware Sequence-to-Tree Network for Math Word Problem SolvingabstractWith the advancements in natural language processing tasks, math word problem solving has received increasing attention.Previous methods have achieved promising results but ignore background common-sense knowledge not directly provided by the problem.In addition, during generation, they focus on local features while neglecting global information.To incorporate external knowledge and global expression information, we propose a novel knowledge-aware sequence-to-tree (KA-S2T) network in which the entities in the problem sequences and their categories are modeled as an entity graph.Based on this entity graph, a graph attention network is used to capture knowledge-aware problem representations.Further, we use a tree-structured decoder with a state aggregation mechanism to capture the long-distance dependency and global expression information.Experimental results on the Math23K dataset revealed that the KA-S2T model can achieve better performance than previously reported best results. Qinzhuo Wu, Qi Zhang 0001, Jinlan Fu, Xuanjing Huang 0001 |
EMNLP (1) | 3 |
| 2020 | Recurrent Memory Reasoning Network for Expert Finding in Community Question AnsweringabstractExpert finding is a task designed to enable recommendation of the right person who can provide high-quality answers to a requester's question. Most previous works try to involve a content-based recommendation, which only superficially comprehends the relevance between a requester's question and the expertise of candidate experts by exploring the content or topic similarity between the requester's question and the candidate experts' historical answers. However, if a candidate expert has never answered a question similar to the requester's question, then existing methods have difficulty making a correct recommendation. Therefore, exploring the implicit relevance between a requester's question and a candidate expert's historical records by perception and reasoning should be taken into consideration. In this study, we propose a novel \textslrecurrent memory reasoning network (RMRN) to perform this task. This method focuses on different parts of a question, and accordingly retrieves information from the histories of the candidate expert.Since only a small percentage of historical records are relevant to any requester's question, we introduce a Gumbel-Softmax-based mechanism to select relevant historical records from candidate experts' answering histories. To evaluate the proposed method, we constructed two large-scale datasets drawn from Stack Overflow and Yahoo! Answer. Experimental results on the constructed datasets demonstrate that the proposed method could achieve better performance than existing state-of-the-art methods. Jinlan Fu, Qi Zhang 0001, Qinzhuo Wu, Renfeng Ma, Xuanjing Huang 0001, Yu-Gang Jiang 0001 |
WSDM | 1 |
| 2019 | Distantly Supervised Named Entity Recognition using Positive-Unlabeled LearningabstractIn this work, we explore the way to perform named entity recognition (NER) using only unlabeled data and named entity dictionaries.To this end, we formulate the task as a positive-unlabeled (PU) learning problem and accordingly propose a novel PU learning algorithm to perform the task.We prove that the proposed algorithm can unbiasedly and consistently estimate the task loss as if there is fully labeled data.A key feature of the proposed method is that it does not require the dictionaries to label every entity within a sentence, and it even does not require the dictionaries to label all of the words constituting an entity.This greatly reduces the requirement on the quality of the dictionaries and makes our method generalize well with quite simple dictionaries.Empirical studies on four public NER datasets demonstrate the effectiveness of our proposed method.We have published the source code at https:// github.com/v-mipeng/LexiconNER. Minlong Peng, Qi Zhang 0001, Jinlan Fu, Xuanjing Huang 0001 |
ACL (1) | 4 |
| 2019 | A Lexicon-Based Graph Neural Network for Chinese NERabstractTao Gui, Yicheng Zou, Qi Zhang, Minlong Peng, Jinlan Fu, Zhongyu Wei, Xuanjing Huang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Tao Gui, Yicheng Zou, Qi Zhang 0001, Minlong Peng, Jinlan Fu, Zhongyu Wei, Xuanjing Huang 0001 |
EMNLP/IJCNLP (1) | 5 |
| 2019 | Learning Task-Specific Representation for Novel Words in Sequence LabelingabstractWord representation is a key component in neural-network-based sequence labeling systems. However, representations of unseen or rare words trained on the end task are usually poor for appreciable performance. This is commonly referred to as the out-of-vocabulary (OOV) problem. In this work, we address the OOV problem in sequence labeling using only training data of the task. To this end, we propose a novel method to predict representations for OOV words from their surface-forms (e.g., character sequence) and contexts. The method is specifically designed to avoid the error propagation problem suffered by existing approaches in the same paradigm. To evaluate its effectiveness, we performed extensive empirical studies on four part-of-speech tagging (POS) tasks and four named entity recognition (NER) tasks. Experimental results show that the proposed method can achieve better or competitive performance on the OOV problem compared with existing state-of-the-art methods. Minlong Peng, Qi Zhang 0001, Tao Gui, Jinlan Fu, Xuanjing Huang 0001 |
IJCAI | 5 |
| 2019 | Model the Long-Term Post History for Hashtag Recommendation
Minlong Peng, Qiyuan Bian, Qi Zhang 0001, Tao Gui, Jinlan Fu, Lanjun Zeng, Xuanjing Huang 0001 |
NLPCC (1) | 5 |
| 2019 | Adaptive Multi-Attention Network Incorporating Answer Information for Duplicate Question DetectionabstractCommunity-based question answering (CQA), which provides a platform for people with diverse backgrounds to share information and knowledge, has become increasingly popular. With the accumulation of site data, methods to detect duplicate questions in CQA sites have attracted considerable attention. Existing methods typically use only questions to complete the task. However, the paired answers may also provide valuable information. In this paper, we propose an answer information- enhanced adaptive multi-attention network (AMAN) to perform this task. AMAN takes full advantage of the semantic information in the paired answers while alleviating the noise problem caused by adding the answers. To evaluate the proposed method, we use a CQADupStack set and the Quora question-pair dataset expanded with paired answers. Experimental results demonstrate that the proposed model can achieve state-of-the-art performance on the above two data sets. Di Liang, Fubao Zhang, Qi Zhang 0001, Jinlan Fu, Minlong Peng, Tao Gui, Xuanjing Huang 0001 |
SIGIR | 5 |
| 2019 | Implicit discourse relation detection using concatenated word embeddings and a gated relevance network
Jinlan Fu, Qi Zhang 0001, Jifan Chen, Minlong Peng, Tao Gui, Xipeng Qiu, Xuanjing Huang 0001 |
Sci. China Inf. Sci. | 1 |
| 2018 | Adaptive Co-attention Network for Named Entity Recognition in TweetsabstractIn this study, we investigate the problem of named entity recognition for tweets. Named entity recognition is an important task in natural language processing and has been carefully studied in recent decades. Previous named entity recognition methods usually only used the textual content when processing tweets. However, many tweets contain not only textual content, but also images. Such visual information is also valuable in the name entity recognition task. To make full use of textual and visual information, this paper proposes a novel method to process tweets that contain multimodal information. We extend a bi-directional long short term memory network with conditional random fields and an adaptive co-attention network to achieve this task. To evaluate the proposed methods, we constructed a large scale labeled dataset that contained multimodal tweets. Experimental results demonstrated that the proposed method could achieve a better performance than the previous methods in most cases. Qi Zhang 0001, Jinlan Fu, Xuanjing Huang 0001 |
AAAI | 2 |
| 2018 | Neural Networks Incorporating Dictionaries for Chinese Word SegmentationabstractIn recent years, deep neural networks have achieved significant success in Chinese word segmentation and many other natural language processing tasks. Most of these algorithms are end-to-end trainable systems and can effectively process and learn from large scale labeled datasets. However, these methods typically lack the capability of processing rare words and data whose domains are different from training data. Previous statistical methods have demonstrated that human knowledge can provide valuable information for handling rare cases and domain shifting problems. In this paper, we seek to address the problem of incorporating dictionaries into neural networks for the Chinese word segmentation task. Two different methods that extend the bi-directional long short-term memory neural network are proposed to perform the task. To evaluate the performance of the proposed methods, state-of-the-art supervised models based methods and domain adaptation approaches are compared with our methods on nine datasets from different domains. The experimental results demonstrate that the proposed methods can achieve better performance than other state-of-the-art neural network methods and domain adaptation approaches in most cases. Qi Zhang 0001, Jinlan Fu |
AAAI | 3 |