VLDB 2026 Research / reviewers in the wild / expert
Yunbo Cao
dblp:33/4066
· DBLP profile ↗
66ranked-venue papers
7as first author
30since 2021 · last 2026
0009-0005-2558-5206ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 47 · 3 first-author · 23 since 2021Databases, data management, data science and information retrieval · 23 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 1 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SHARP: Self-adaptive Harmful Category-aware Prompt Generation for Black-box JailbreakingabstractLarge Language Models (LLMs) have been widely applied in various domains such as education and healthcare, making safety assurance crucial.Jailbreak attacks, a method used in red-teaming, can help evaluate and improve the defensive strategies of LLMs.However, existing jailbreak methods often overlook the semantic differences across categories of harmful questions, leading to inconsistent success rates and reduced overall attack effectiveness.We propose the first category-aware jailbreak framework, SHARP, which incorporates the semantic category of harmful questions into prompt generation.Trained on a verified jailbreak dataset, SHARP enables the model to learn category-specific semantic features and adaptively generate prompts that bypass safety mechanisms.The method combines two-stage LoRA fine-tuning, and DPO-based reinforcement learning to optimize both attack success and category alignment.Experiments show that SHARP significantly improves attack success rates and achieves better cross-category robustness compared to the state-of-the-art (SOTA) baselines, providing an efficient and scalable tool for evaluating LLM safety. Yingjie Xue, Xingyou Xia, Yunbo Cao, Dengpan Ye, Guotong Geng |
ACL (1) | 4 |
| 2026 | DisCal: Distribution-Aware Calibration for Mathematical Reasoning Under Character-Level Noisy InputsabstractBo Zhang, Jiawei Zhang, Cong Gao, Bingxu Han, Minghao Hu, Jun Zhang, Yunbo Cao, Zhunchen Luo, Wen Yao, Guotong Geng, Zhong Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Bingxu Han, Yunbo Cao, Zhunchen Luo, Guotong Geng |
ACL (1) | 7 |
| 2026 | TRACE: Checklist-Driven Dynamic Multi-Agent Coordination for Autonomous Data Science Report Generation
Yinlong Xiao, Zhunchen Luo, Long Sheng, Shuai Lei, Yunbo Cao, Guotong Geng |
ICIC (24) | 8 |
| 2025 | AdaDARE-gamma: Balancing Stability and Plasticity in Multi-modal LLMs through Efficient AdaptationabstractAdapting Multi-modal Large Language Models (MLLMs) to target tasks often suffers from catastrophic forgetting, where acquiring new task-specific knowledge compromises performance on pre-trained tasks. In this paper, we introduce AdaDARE-γ, an efficient approach that alleviates catastrophic forgetting by controllably injecting new task-specific knowledge through adaptive parameter selection from fine-tuned models without requiring retraining procedures. This approach consists two key innovations: (1) an adaptive parameter selection mechanism that identifies and retains the most task-relevant parameters from fine-tuned models, and (2) a controlled task-specific information injection strategy that precisely balances the preservation of pre-trained knowledge with the acquisition of new capabilities. Theoretical analysis proves the optimality of our parameter selection strategy and establishes bounds for the task-specific information injection factor. Extensive experiments on InstructBLIP and LLaVA-1.5 across image captioning and visual question answering tasks demonstrate that AdaDARE-γ establishes new state-of-the-art results in balancing model performance. Specifically, it maintains 98.2% of pre-training effectiveness on original tasks while achieving 98.7% of standard fine-tuning performance on target tasks. Jintao Yang, Zhunchen Luo, Yunbo Cao, Wenpeng Hu |
CVPR | 4 |
| 2025 | D-PathVer: Dynamic Reasoning Pathways for Complex Claim Verification
Lingxiao Zheng, Zhunchen Luo, Wenpeng Hu, Yunbo Cao |
NLPCC (3) | 4 |
| 2024 | Large Language Models are not Fair EvaluatorsabstractPeiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Lingpeng Kong, Qi Liu, Tianyu Liu, Zhifang Sui. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Peiyi Wang, Lei Li 0039, Liang Chen 0024, Zefan Cai, Binghuai Lin, Yunbo Cao, Lingpeng Kong, Qi Liu 0049, Tianyu Liu 0001, Zhifang Sui |
ACL (1) | 7 |
| 2024 | SkillNet-X: A Multilingual Multitask Model with Sparsely Activated SkillsabstractTraditional multitask learning methods typically can only leverage shared knowledge within specific tasks or languages, resulting in a loss of either cross-language or cross-task knowledge. This paper proposes a general multilingual multitask model, named SkillNet-X, which enables a single model to tackle many different tasks from different languages. To this end, we define several language-specific skills and task-specific skills, each of which corresponds to a skill module. SkillNet-X sparsely activates parts of the skill modules which are relevant to eitherthe target task or the target language. Acting as knowledge transit hubs, skill modules are capable of absorbing task-related knowledge and language-related knowledge consecutively. We evaluate SkillNet-X on eleven natural language understanding datasets in four languages. Results show that SkillNet-X performs better than task-specific and two multitask learning baselines.To investigate the generalization of our model, we conduct experiments on two new tasks and find that SkillNet-X significantly outperforms baselines. Zhangyin Feng, Yong Dai 0001, Fan Zhang 0092, Duyu Tang, Shuangzhi Wu, Bing Qin 0001, Yunbo Cao, Shuming Shi 0001 |
ICASSP | 8 |
| 2024 | DialogVCS: Robust Natural Language Understanding in Dialogue System UpgradeabstractZefan Cai, Xin Zheng, Tianyu Liu, Haoran Meng, Jiaqi Han, Gang Yuan, Binghuai Lin, Baobao Chang, Yunbo Cao. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Zefan Cai, Tianyu Liu 0001, Haoran Meng, Binghuai Lin, Baobao Chang, Yunbo Cao |
NAACL-HLT | 9 |
| 2024 | Unifying Token- and Span-level Supervisions for Few-shot Sequence LabelingabstractFew-shot sequence labeling aims to identify novel classes based on only a few labeled samples. Existing methods solve the data scarcity problem mainly by designing token-level or span-level labeling models based on metric learning. However, these methods are only trained at a single granularity (i.e., either token-level or span-level) and have some weaknesses of the corresponding granularity. In this article, we first unify token- and span-level supervisions and propose a Consistent Dual Adaptive Prototypical (CDAP) network for few-shot sequence labeling. CDAP contains the token- and span-level networks, jointly trained at different granularities. To align the outputs of two networks, we further propose a consistent loss to enable them to learn from each other. During the inference phase, we propose a consistent greedy inference algorithm that first adjusts the predicted probability and then greedily selects non-overlapping spans with maximum probability. Extensive experiments show that our model achieves new state-of-the-art results on three benchmark datasets. All the code and data of this work will be released at https://github.com/zifengcheng/CDAP . Zifeng Cheng, Qingyu Zhou, Zhiwei Jiang 0001, Xuemin Zhao, Yunbo Cao, Qing Gu 0001 |
ACM Trans. Inf. Syst. | 5 |
| 2023 | Denoising Bottleneck with Mutual Information Maximization for Video Multimodal FusionabstractShaoxiang Wu, Damai Dai, Ziwei Qin, Tianyu Liu, Binghuai Lin, Yunbo Cao, Zhifang Sui. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Shaoxiang Wu, Damai Dai, Ziwei Qin, Tianyu Liu 0001, Binghuai Lin, Yunbo Cao, Zhifang Sui |
ACL (1) | 6 |
| 2023 | Soft Language Clustering for Multilingual Model Pre-trainingabstractJiali Zeng, Yufan Jiang, Yongjing Yin, Yi Jing, Fandong Meng, Binghuai Lin, Yunbo Cao, Jie Zhou. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Jiali Zeng, Yufan Jiang, Yongjing Yin, Yi Jing, Fandong Meng, Binghuai Lin, Yunbo Cao, Jie Zhou 0016 |
ACL (1) | 7 |
| 2023 | Read Then Respond: Multi-granularity Grounding Prediction for Knowledge-Grounded Dialogue Generation
Yiyang Du, Shi-Wei Zhang, Xianjie Wu, Yunbo Cao, Zhoujun Li 0001 |
ADMA (2) | 5 |
| 2023 | Contextual Similarity is More Valuable Than Character Similarity: An Empirical Study for Chinese Spell CheckingabstractChinese Spell Checking (CSC) task aims to detect and correct Chinese spelling errors. Recently, related researches focus on introducing character similarity from confusion set to enhance the CSC models, ignoring the context of characters that contain richer information. To make better use of contextual information, we propose a simple yet effective Curriculum Learning (CL) framework for the CSC task. With the help of our model-agnostic CL framework, existing CSC models will be trained from easy to difficult as humans learn Chinese characters and achieve further performance improvements. Extensive experiments and detailed analyses on widely used SIGHAN datasets show that our method outperforms previous state-of-the-art methods. More instructively, our study empirically suggests that contextual similarity is more valuable than character similarity for the CSC task. Qingyu Zhou, Shirong Ma, Yangning Li, Yunbo Cao, Hai-Tao Zheng 0002 |
ICASSP | 6 |
| 2023 | QURG: Question Rewriting Guided Context-Dependent Text-to-SQL Semantic Parsing
Linzheng Chai, Dongling Xiao, Jian Yang 0030, Liqun Yang, Qian-Wen Zhang, Yunbo Cao, Zhoujun Li 0001 |
PRICAI (2) | 7 |
| 2023 | Automatic Context Pattern Generation for Entity Set ExpansionabstractEntity Set Expansion (ESE) is a valuable task that aims to find entities of the target semantic class described by given seed entities. Various Natural Language Processing (NLP) and Information Retrieval (IR) downstream applications have benefited from ESE due to its ability to discover knowledge. Although existing corpus-based ESE methods have achieved great progress, they still rely on corpora with high-quality entity information annotated, because most of them need to obtain the context patterns through the position of the entity in a sentence. Therefore, the quality of the given corpora and their entity annotation has become the bottleneck that limits the performance of such methods. To overcome this dilemma and make the ESE models free from the dependence on entity annotation, our work aims to explore a new ESE paradigm, namely corpus-independent ESE. Specifically, we devise a context pattern generation module that utilizes autoregressive language models (e.g., GPT-2) to automatically generate high-quality context patterns for entities. In addition, we propose the GAPA, a novel ESE framework that leverages the aforementionedGenerAtedPAtterns to expand target entities. Extensive experiments and detailed analyses on three widely used datasets demonstrate the effectiveness of our method. All the codes of our experiments are available athttps://github.com/geekjuruo/GAPA. Shulin Huang, Xinwei Zhang 0009, Qingyu Zhou, Yangning Li, Ruiyang Liu, Yunbo Cao, Hai-Tao Zheng 0002, Ying Shen 0001 |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2022 | Pre-training and Fine-tuning Neural Topic Model: A Simple yet Effective Approach to Incorporating External KnowledgeabstractRecent years have witnessed growing interests in incorporating external knowledge such as pre-trained word embeddings (PWEs) or pretrained language models (PLMs) into neural topic modeling.However, we found that employing PWEs and PLMs for topic modeling only achieved limited performance improvements but with huge computational overhead.In this paper, we propose a novel strategy to incorporate external knowledge into neural topic modeling where the neural topic model is pretrained on a large corpus and then fine-tuned on the target dataset.Experiments have been conducted on three datasets and results show that the proposed approach significantly outperforms both current state-of-the-art neural topic models and some topic modeling approaches enhanced with PWEs or PLMs.Moreover, further study shows that the proposed approach greatly reduces the need for the huge size of training data. Linhai Zhang, Xuemeng Hu, Qian-Wen Zhang, Yunbo Cao |
ACL (1) | 6 |
| 2022 | Knowledge-Sensed Cognitive Diagnosis for Intelligent Education PlatformsabstractCognitive diagnosis is a fundamental issue of intelligent education platforms, whose goal is to reveal the mastery of students on knowledge concepts. Recently, certain efforts have been made to improve the diagnosis precision, by designing deep neural networks-based diagnostic functions or incorporating more rich context features to enhance the representation of students and exercises. However, how to interpretably infer the student's mastery over non-interactive knowledge concepts (i.e., knowledge concepts not related to his/her exercising records) still remains challenging, especially when not giving relations between knowledge concepts. To this end, we propose a Knowledge-Sensed Cognitive Diagnosis (KSCD) framework, aiming at learning intrinsic relations among knowledge concepts from student response logs and incorporating them for inferring students' mastery over all knowledge concepts in an end-to-end manner. Specifically, we firstly project students, exercises and knowledge concepts into embedding representation matrices, where the intrinsic relations among knowledge concepts are reflected in the knowledge embedding representation matrix. Then, the knowledge-sensed student knowledge mastery vector and exercise factor vectors are obtained by the multiply product of their embedding representations and the knowledge embedding representation matrix, which make the student's mastery of non-interactive knowledge concepts be interpretably inferred. Finally, we can utilize classical student-exercise interaction functions to predict student's exercising performance and jointly train the model. In additional, we also design a new function to better model the student-exercise interactions. Extensive experimental results on two real-world datasets clearly show the significant performance gain of our KSCD framework, especially in predicting students' mastery over non-interactive knowledge concepts, by comparing to state-of-the-art cognitive diagnosis models (CDMs). Haiping Ma, Manwei Li, Le Wu 0001, Haifeng Zhang 0003, Yunbo Cao, Xingyi Zhang 0001, Xuemin Zhao |
CIKM | 5 |
| 2022 | A Prerequisite Attention Model for Knowledge Proficiency Diagnosis of StudentsabstractWith the rapid development of intelligent education platforms, how to enhance the performance of diagnosing students' knowledge proficiency has become an important issue, e.g., by incorporating the prerequisite relation of knowledge concepts. Unfortunately, the differentiated influence from different predecessor concepts to successor concepts is still underexplored in existing approaches. To this end, we propose a Prerequisite Attention model for Knowledge Proficiency diagnosis of students (PAKP) to learn the attentive weights of precursor concepts on successor concepts and model it for inferring the knowledge proficiency. Specifically, given the student response records and knowledge prerequisite graph, we design an embedding layer to output the representations of students, exercises, and concepts. Influence coefficient among concepts is calculated via an efficient attention mechanism in a fusion layer. Finally, the performance of each student is predicted based on the mined student and exercise factors. Extensive experiments on real-data sets demonstrate that PAKP exhibits great efficiency and interpretability advantages without accuracy loss. Haiping Ma, Shangshang Yang, Qi Liu 0003, Haifeng Zhang 0003, Xingyi Zhang 0001, Yunbo Cao, Xuemin Zhao |
CIKM | 7 |
| 2022 | AiM: Taking Answers in Mind to Correct Chinese Cloze Tests in Educational ApplicationsabstractTo automatically correct handwritten assignments, the traditional approach is to use an OCR model to recognize characters and compare them to answers. The OCR model easily gets confused on recognizing handwritten Chinese characters, and the textual information of the answers is missing during the model inference. However, teachers always have these answers in mind to review and correct assignments. In this paper, we focus on the Chinese cloze tests correction and propose a multimodal approach(named AiM). The encoded representations of answers interact with the visual information of students’ handwriting. Instead of predicting ‘right’ or ‘wrong’, we perform the sequence labeling on the answer text to infer which answer character differs from the handwritten content in a fine-grained way. We take samples of OCR datasets as the positive samples for this task, and develop a negative sample augmentation method to scale up the training data. Experimental results show that AiM outperforms OCR-based methods by a large margin. Extensive studies demonstrate the effectiveness of our multimodal approach. Zhongli Li, Qingyu Zhou, Chao Li 0063, Mina Ma, Yunbo Cao, Hongzhi Liu 0001 |
COLING | 7 |
| 2022 | Learning Robust Representations for Continual Relation Extraction via Adversarial Class AugmentationabstractContinual relation extraction (CRE) aims to continually learn new relations from a classincremental data stream.CRE model usually suffers from catastrophic forgetting problem, i.e., the performance of old relations seriously degrades when the model learns new relations.Most previous work attributes catastrophic forgetting to the corruption of the learned representations as new relations come, with an implicit assumption that the CRE models have adequately learned the old relations.In this paper, through empirical studies we argue that this assumption may not hold, and an important reason for catastrophic forgetting is that the learned representations do not have good robustness against the appearance of analogous relations in the subsequent learning process.To address this issue, we encourage the model to learn more precise and robust representations through a simple yet effective adversarial class augmentation mechanism (ACA), which is easy to implement and model-agnostic.Experimental results show that ACA can consistently improve the performance of state-of-theart CRE models on two popular benchmarks. Peiyi Wang, Yifan Song 0002, Tianyu Liu 0001, Binghuai Lin, Yunbo Cao, Sujian Li, Zhifang Sui |
EMNLP | 5 |
| 2022 | HPT: Hierarchy-aware Prompt Tuning for Hierarchical Text ClassificationabstractHierarchical text classification (HTC) is a challenging subtask of multi-label classification due to its complex label hierarchy.Recently, the pretrained language models (PLM) have been widely adopted in HTC through a finetuning paradigm.However, in this paradigm, there exists a huge gap between the classification tasks with sophisticated label hierarchy and the masked language model (MLM) pretraining tasks of PLMs and thus the potential of PLMs cannot be fully tapped.To bridge the gap, in this paper, we propose HPT, a Hierarchy-aware Prompt Tuning method to handle HTC from a multi-label MLM perspective.Specifically, we construct a dynamic virtual template and label words that take the form of soft prompts to fuse the label hierarchy knowledge and introduce a zero-bounded multi-label cross-entropy loss to harmonize the objectives of HTC and MLM.Extensive experiments show HPT achieves state-of-the-art performances on 3 popular HTC datasets and is adept at handling the imbalance and low resource situations. Peiyi Wang, Tianyu Liu 0001, Binghuai Lin, Yunbo Cao, Zhifang Sui, Houfeng Wang |
EMNLP | 5 |
| 2022 | A Non-Hierarchical Attention Network with Modality Dropout for Textual Response Generation in Multimodal Dialogue SystemsabstractExisting text- and image-based multimodal dialogue systems use the traditional Hierarchical Recurrent Encoder-Decoder (HRED) framework, which has an utterance-level encoder to model utterance representation and a context-level encoder to model context representation. Although pioneer efforts have shown promising performances, they still suffer from the following challenges: (1) the interaction between textual features and visual features is not fine-grained enough. (2) the context representation can not provide a complete representation for the context. To address the issues mentioned above, we propose a non-hierarchical attention network with modality dropout, which abandons the HRED framework and utilizes attention modules to encode each utterance and model the context representation. To evaluate our proposed model, we conduct comprehensive experiments on a public multimodal dialogue dataset. Automatic and human evaluation demonstrate that our proposed model outperforms the existing methods and achieves state-of-the-art performance. Rongyi Sun, Borun Chen, Qingyu Zhou, Yunbo Cao, Hai-Tao Zheng 0002 |
ICASSP | 5 |
| 2022 | An Enhanced Span-based Decomposition Method for Few-Shot Sequence LabelingabstractPeiyi Wang, Runxin Xu, Tianyu Liu, Qingyu Zhou, Yunbo Cao, Baobao Chang, Zhifang Sui. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Peiyi Wang, Runxin Xu, Tianyu Liu 0001, Qingyu Zhou, Yunbo Cao, Baobao Chang, Zhifang Sui |
NAACL-HLT | 5 |
| 2021 | RadarMath: An Intelligent Tutoring System for Math EducationabstractWe propose and implement a novel intelligent tutoring system, called RadarMath, to support intelligent and personalized learning for math education. The system provides the services including automatic grading and personalized learning guidance. Specifically, two automatic grading models are designed to accomplish the tasks for scoring the text-answer and formula-answer questions respectively. An education-oriented knowledge graph with the individual learner’s knowledge state is used as the key tool for guiding the personalized learning process. The system demonstrates how the relevant AI techniques could be applied in today's intelligent tutoring systems. Yu Lu 0003, Yang Pian, Penghe Chen, Qinggang Meng, Yunbo Cao |
AAAI | 5 |
| 2021 | MMKE: A Multi-Model Knowledge Extraction System from Unstructured TextsabstractIn this work, we present a Multi-Model Knowledge Extraction (MMKE) System which consists of two unstructured text extraction models (RelationSO model and SubjectRO model) based on a multi-task learning framework. Instead of recognizing entity first and then predicting relationships between entity pairs in previous works, MMKE detects subject and corresponding relationships before extracting objects to cope with the diverse object-type problem, overlapping problem and non-predefined relation problem. Our system accepts unstructured text as input, from which it automatically extracts triplets knowledge (subject, relation, object). More importantly, we incorporate a number of user-friendly extraction functionalities, such as multi-format uploading, one-click extractions, knowledge editing and graphical displays. The demonstration video is available at this link: https://youtu.be/HtOPJrGhSxk. Qian-Wen Zhang, Shi-Wei Zhang, Meng-Liang Rao, Yunbo Cao |
AAAI | 7 |
| 2021 | Exploiting Unlabeled Data via Partial Label Assignment for Multi-Class Semi-Supervised LearningabstractIn semi-supervised learning, one key strategy in exploiting unlabeled data is trying to estimate its pseudo-label based on current predictive model, where the unlabeled data assigned with pseudo-label is further utilized to enlarge labeled data set for model update. Nonetheless, the supervision information conveyed by pseudo-label is prone to error especially when the performance of initial predictive model is mediocre due to limited amount of labeled data. In this paper, an intermediate unlabeled data exploitation strategy is investigated via partial label assignment, i.e. a set of candidate labels other than a single pseudo-label are assigned to the unlabeled data. We only assume that the ground-truth label of unlabeled data resides in the assigned candidate label set, which is less error-prone than trying to identify the single ground-truth label via pseudo-labeling. Specifically, a multi-class classifier is induced from the partial label examples with candidate labels to facilitate model induction with labeled examples. An iterative procedure is designed to enable labeling information communication between the classifiers induced from partial label examples and labeled examples, whose classification outputs are integrated to yield the final prediction. Comparative studies against state-of-the-art approaches clearly show the effectiveness of the proposed unlabeled data exploitation strategy for multi-class semi-supervised learning. Zhen-Ru Zhang, Qian-Wen Zhang, Yunbo Cao, Min-Ling Zhang |
AAAI | 3 |
| 2021 | A Unified Multi-Task Learning Framework for Joint Extraction of Entities and RelationsabstractJoint extraction of entities and relations focuses on detecting entity pairs and their relations simultaneously with a unified model. Based on the extraction order, previous works mainly solve this task through relation-last, relation-first and relation-middle manner. However, these methods still suffer from the template-dependency, non-entity detection and non-predefined relation prediction problem. To overcome these challenges, in this paper, we propose a unified multi-task learning framework to divide the task into three interacted sub-tasks. Specifically, we first introduce the type-attentional method for subject extraction to provide prior type information explicitly. Then, the subject-aware relation prediction is presented to select useful relations based on the combination of global and local semantics. Third, we propose a question generation based QA method for object extraction to obtain diverse queries automatically. Notably, our method detects subjects or objects without relying on NER models and thus it is capable of dealing with the non-entity scenario. Finally, three sub-tasks are integrated into a unified model through parameter sharing. Extensive experiments demonstrate that the proposed framework outperforms all the baseline methods on two benchmark datasets, and further achieve excellent performance for non-predefined relations. Tianyang Zhao 0003, Yunbo Cao, Zhoujun Li 0001 |
AAAI | 3 |
| 2021 | Dialogue Response Selection with Hierarchical Curriculum LearningabstractYixuan Su, Deng Cai, Qingyu Zhou, Zibo Lin, Simon Baker, Yunbo Cao, Shuming Shi, Nigel Collier, Yan Wang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yixuan Su, Deng Cai 0002, Qingyu Zhou, Zibo Lin, Simon Baker, Yunbo Cao, Shuming Shi 0001, Nigel Collier, Yan Wang 0060 |
ACL/IJCNLP (1) | 6 |
| 2021 | LANA: Towards Personalized Deep Knowledge Tracing Through Distinguishable Interactive Sequences
Yuhao Zhou 0004, Xihua Li 0002, Yunbo Cao, Xuemin Zhao, Jiancheng Lv 0001 |
EDM | 3 |
| 2021 | Correlation-Guided Representation for Multi-Label Text ClassificationabstractMulti-label text classification is an essential task in natural language processing. Existing multi-label classification models generally consider labels as categorical variables and ignore the exploitation of label semantics. In this paper, we view the task as a correlation-guided text representation problem: an attention-based two-step framework is proposed to integrate text information and label semantics by jointly learning words and labels in the same space. In this way, we aim to capture high-order label-label correlations as well as context-label correlations. Specifically, the proposed approach works by learning token-level representations of words and labels globally through a multi-layer Transformer and constructing an attention vector through word-label correlation matrix to generate the text representation. It ensures that relevant words receive higher weights than irrelevant words and thus directly optimizes the classification performance. Extensive experiments over benchmark multi-label datasets clearly validate the effectiveness of the proposed approach, and further analysis demonstrates that it is competitive in both predicting low-frequency labels and convergence speed. Qian-Wen Zhang, Ruifang Liu, Yunbo Cao, Min-Ling Zhang |
IJCAI | 5 |
| 2020 | Asking Effective and Diverse Questions: A Machine Reading Comprehension based Framework for Joint Entity-Relation ExtractionabstractRecent advances cast the entity-relation extraction to a multi-turn question answering (QA) task and provide an effective solution based on the machine reading comprehension (MRC) models. However, they use a single question to characterize the meaning of entities and relations, which is intuitively not enough because of the variety of context semantics. Meanwhile, existing models enumerate all relation types to generate questions, which is inefficient and easily leads to confusing questions. In this paper, we improve the existing MRC-based entity-relation extraction model through diverse question answering. First, a diversity question answering mechanism is introduced to detect entity spans and two answering selection strategies are designed to integrate different answers. Then, we propose to predict a subset of potential relations and filter out irrelevant ones to generate questions effectively. Finally, entity and relation extractions are integrated in an end-to-end way and optimized through joint learning. Experiment results show that the proposed method significantly outperforms baseline models, which improves the relation F1 to 62.1% (+1.9%) on ACE05 and 71.9% (+3.0%) on CoNLL04. Our implementation is available at https://github.com/TanyaZhao/MRC4ERE. Tianyang Zhao 0003, Yunbo Cao, Zhoujun Li 0001 |
IJCAI | 3 |
| 2020 | LARQ: Learning to Ask and Rewrite Questions for Community Question Answering
Huiyang Zhou, Haoyan Liu 0001, Yunbo Cao, Zhoujun Li 0001 |
NLPCC (2) | 4 |
| 2018 | Mention and Entity Description Co-Attention for Entity DisambiguationabstractFor the task of entity disambiguation, mention contexts and entity descriptions both contain various kinds of information content while only a subset of them are helpful for disambiguation. In this paper, we propose a type-aware co-attention model for entity disambiguation, which tries to identify the most discriminative words from mention contexts and most relevant sentences from corresponding entity descriptions simultaneously. To bridge the semantic gap between mention contexts and entity descriptions, we further incorporate entity type information to enhance the co-attention mechanism. Our evaluation shows that the proposed model outperforms the state-of-the-arts on three public datasets. Further analysis also confirms that both the co-attention mechanism and the type-aware mechanism are effective. Feng Nie, Yunbo Cao, Jinpeng Wang 0001, Chin-Yew Lin |
AAAI | 2 |
| 2018 | Overview of the NLPCC 2018 Shared Task: Spoken Language Understanding in Task-Oriented Dialog Systems
Xuemin Zhao, Yunbo Cao |
NLPCC (2) | 2 |
| 2014 | Collective Tweet Wikification based on Semi-supervised Graph RegularizationabstractWikification for tweets aims to automatically identify each concept mention in a tweet and link it to a concept referent in a knowledge base (e.g., Wikipedia).Due to the shortness of a tweet, a collective inference model incorporating global evidence from multiple mentions and concepts is more appropriate than a noncollecitve approach which links each mention at a time.In addition, it is challenging to generate sufficient high quality labeled data for supervised models with low cost.To tackle these challenges, we propose a novel semi-supervised graph regularization model to incorporate both local and global evidence from multiple tweets through three fine-grained relations.In order to identify semanticallyrelated mentions for collective inference, we detect meta path-based semantic relations through social networks.Compared to the state-of-the-art supervised model trained from 100% labeled data, our proposed approach achieves comparable performance with 31% labeled data and obtains 5% absolute F1 gain with 50% labeled data.Stay up Hawk Fans.We are going through a slump now, but we have to stay positive.Go Hawks!Congrats to UCONN and Kemba Walker.5 wins in 5 days, very impressive... Just getting to the Arena, we play the Bucks tonight. Hongzhao Huang, Yunbo Cao, Xiaojiang Huang, Heng Ji 0001, Chin-Yew Lin |
ACL (1) | 2 |
| 2013 | Learning a Replacement Model for Query Segmentation with Consistency in Search Logs
Wei Zhang 0038, Yunbo Cao, Chin-Yew Lin, Jian Su 0002, Chew Lim Tan |
IJCNLP | 2 |
| 2012 | A Lazy Learning Model for Entity Linking using Query-Specific Information
Wei Zhang 0038, Jian Su 0002, Chew Lim Tan, Yunbo Cao, Chin-Yew Lin |
COLING | 4 |
| 2011 | Learning to Suggest Questions in Online ForumsabstractOnline forums contain interactive and semantically related discussions on various questions. Extracted question-answer archive is invaluable knowledge, which can be used to improve Question Answering services. In this paper, we address the problem of Question Suggestion, which targets at suggesting questions that are semantically related to a queried question. Existing bag-of-words approaches suffer from the shortcoming that they could not bridge the lexical chasm between semantically related questions. Therefore, we present a new framework to suggest questions, and propose the Topicenhanced Translation-based Language Model (TopicTRLM) which fuses both the lexical and latent semantic knowledge. Extensive experiments have been conducted with a large real world data set. Experimental results indicate our approach is very effective and outperforms other popular methods in several metrics. Tom Chao Zhou, Chin-Yew Lin, Irwin King, Michael R. Lyu, Young-In Song, Yunbo Cao |
AAAI | 6 |
| 2011 | Leveraging Unlabeled Data to Scale Blocking for Record Linkage
Yunbo Cao, Jiamin Zhu, Pei Yue, Chin-Yew Lin, Yong Yu 0001 |
IJCAI | 1 |
| 2011 | A structural support vector method for extracting contexts and answers of questions from online forums
Yunbo Cao, Wen-Yun Yang, Chin-Yew Lin, Yong Yu 0001 |
Inf. Process. Manag. | 1 |
| 2011 | Re-ranking question search results by clustering questionsabstractIn this article, we address the problem of question clustering and study its use for re-ranking question search results. In question clustering we have to organize question search results into certain meaningful and condensed groups. Specifically, we propose to use a data structure consisting of question topic and question focus for modeling questions, and then cluster questions on the basis of the data structure. Experimental results show that our approach to question clustering improves the performance of question search significantly better than the approach not utilizing the topic–focus structure. Yunbo Cao, Huizhong Duan, Chin-Yew Lin, Yong Yu 0001 |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2010 | Automatic extraction of web data records containing user-generated contentabstractIn this paper, we are concerned with the problem of automatically extracting web data records that contain user-generated content (UGC). In previous work, web data records are usually assumed to be well-formed with a limited amount of UGC, and thus can be extracted by testing repetitive structure similarity. However, when a web data record includes a large portion of free-format UGC, the similarity test between records may fail, which in turn results in lower performance. In our work, we find that certain domain constraints (e.g., post-date) can be used to design better similarity measures capable of circumventing the influence of UGC. In addition, we also use anchor points provided by the domain constraints to improve the extraction process, which ends in an algorithm called MiBAT (Mining data records Based on Anchor Trees). We conduct extensive experiments on a dataset consisting of forum thread pages which are collected from 307 sites that cover 219 different forum software packages. Our approach achieves a precision of 98.9% and a recall of 97.3% with respect to post record extraction. On page level, it perfectly handles 91.7% of pages without extracting any wrong posts or missing any golden posts. We also apply our approach to comment extraction and achieve good results as well. Xinying Song, Jing Liu 0022, Yunbo Cao, Chin-Yew Lin, Hsiao-Wuen Hon |
CIKM | 3 |
| 2009 | Learning to recommend questions based on user ratingsabstractAt community question answering services, users are usually encouraged to rate questions by votes. The questions with the most votes are then recommended and ranked on the top when users browse questions by category. As users are not obligated to rate questions, usually only a small proportion of questions eventually gets rating. Thus, in this paper, we are concerned with learning to recommend questions from user ratings of a limited size. To overcome the data sparsity, we propose to utilize questions without users rating as well. Further, as there exist certain noises within user ratings (the preference of some users expressed in their ratings diverges from that of the majority of users), we design a new algorithm called 'majority-based perceptron algorithm' which can avoid the influence of noisy instances by emphasizing its learning over data instances from the majority users. Experimental results from a large collection of real questions confirm the effectiveness of our proposals. Ke Sun 0007, Yunbo Cao, Xinying Song, Young-In Song, Xiaolong Wang 0001, Chin-Yew Lin |
CIKM | 2 |
| 2009 | A Structural Support Vector Method for Extracting Contexts and Answers of Questions from Online Forums
Wen-Yun Yang, Yunbo Cao, Chin-Yew Lin |
EMNLP | 2 |
| 2008 | Question Utility: A Novel Static Ranking of Question Search
Young-In Song, Chin-Yew Lin, Yunbo Cao, Hae-Chang Rim |
AAAI | 3 |
| 2008 | A Probabilistic Model for Fine-Grained Expert Search
Shenghua Bao, Huizhong Duan, Qi Zhou 0001, Miao Xiong, Yunbo Cao, Yong Yu 0001 |
ACL | 5 |
| 2008 | Searching Questions by Identifying Question Topic and Question Focus
Huizhong Duan, Yunbo Cao, Chin-Yew Lin, Yong Yu 0001 |
ACL | 2 |
| 2008 | Understanding and Summarizing Answers in Community-Based Question Answering Services
Yuanjie Liu, Yunbo Cao, Chin-Yew Lin, Dingyi Han, Yong Yu 0001 |
COLING | 3 |
| 2008 | Recommending questions using the mdl-based tree cut modelabstractThe paper is concerned with the problem of question recommendation. Specifically, given a question as query, we are to retrieve and rank other questions according to their likelihood of being good recommendations of the queried question. A good recommendation provides alternative aspects around users' interest. We tackle the problem of question recommendation in two steps: first represent questions as graphs of topic terms, and then rank recommendations on the basis of the graphs. We formalize both steps as the tree-cutting problems and then employ the MDL (Minimum Description Length) for selecting the best cuts. Experiments have been conducted with the real questions posted at Yahoo! Answers. The questions are about two domains, 'travel' and 'computers & internet'. Experimental results indicate that the use of the MDL-based tree cut model can significantly outperform the baseline methods of word-based VSM or phrase-based VSM. The results also show that the use of the MDL-based tree cut model is essential to our approach. Yunbo Cao, Huizhong Duan, Chin-Yew Lin, Yong Yu 0001, Hsiao-Wuen Hon |
WWW | 1 |
| 2008 | Competitor Mining with the WebabstractThis paper is concerned with the problem of mining competitors from the Web automatically. Nowadays the fierce competition in the market necessitates every company not only to know which companies are its primary competitors, but also in which fields the company's rivals compete with itself and what its competitors' strength is in a specific competitive domain. The task of competitor mining that we address in the paper includes mining all the information such as competitors, competing fields and competitors' strength. A novel algorithm called CoMiner is proposed, which tries to conduct a Web-scale mining in a domain-independent manner. The CoMiner algorithm consists of three parts: 1) given an input entity, extracting a set of comparative candidates and then ranking them according to comparability; 2) extracting the fields in which the given entity and its competitors play against each other; 3) identifying and summarizing the competitive evidence that details the competitors' strength. As for evaluation, a prototype system implementing the CoMiner algorithm is presented. An evaluation data set consisting of 70 entities is constructed. 728 competitors and 3,640 competitive fields with 6,381 competitive evidences are discovered with the prototype. The experimental results show that the proposed algorithm is highly effective. Shenghua Bao, Rui Li 0049, Yong Yu 0001, Yunbo Cao |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2007 | Using social annotations to improve language model for information retrievalabstractThis poster is concerned with the problem of exploring the use of social annotations for improving language models for information retrieval (denoted as LMIR). Two properties of social annotations, namely keyword property and structure property are studied for this aim. The keyword property improves LMIR by concatenating all the annotations of a document to generate a summary of the document. The structure property can boost LMIR further when similarity among annotations and similarity among documents are taken into consideration simultaneously. The two properties of social annotations are leveraged for the use of language modeling with a mixture model named as "Language Annotation Model" (denoted as LAM). Evaluations using del.icio.us data show that LAM outperforms the traditional LMIR approaches significantly. Shengliang Xu, Shenghua Bao, Yunbo Cao, Yong Yu 0001 |
CIKM | 3 |
| 2007 | Searching Documents Based on Relevance and Type
Jun Xu 0001, Yunbo Cao, Hang Li 0001, Nick Craswell, Yalou Huang |
ECIR | 2 |
| 2007 | Low-Quality Product Review Detection in Opinion Summarization
Yunbo Cao, Chin-Yew Lin, Yalou Huang |
EMNLP-CoNLL | 2 |
| 2007 | Using Social Annotations to Smooth the Language Model for IR
Shengliang Xu, Shenghua Bao, Yong Yu 0001, Yunbo Cao |
PAKDD | 4 |
| 2007 | Web page title extraction and its application
Yewei Xue, Yunhua Hu, Guomao Xin, Ruihua Song, Shuming Shi 0001, Yunbo Cao, Chin-Yew Lin, Hang Li 0001 |
Inf. Process. Manag. | 6 |
| 2006 | Cost-Sensitive Learning of SVM for Ranking
Jun Xu 0001, Yunbo Cao, Hang Li 0001, Yalou Huang |
ECML | 2 |
| 2006 | Mining Latent Associations of Objects Using a Typed Mixture Model--A Case Study on Expert/Expertise MiningabstractThis paper studies the problem of discovering latent associations among objects in text documents. Specifically, given two sets of objects and various types of co-occurrence data concerning the objects existing in texts, we aim to discover the hidden or latent associative relationships between the two sets of objects. Existing methods are not directly applicable as they are unable to consider all this information. For example, the probabilistic mixture model called Separable Mixture Model (SMM) proposed by Hofmann can use only one type of co-occurrences to mine latent associations. This paper proposes a more general probabilistic mixture model called the Typed Separable Mixture Model (TSMM), which is able to use all types of co-occurrences within a single framework. Experimental results based on the expert/expertise mining task show that TSMM outperforms SMM significantly. Shenghua Bao, Yunbo Cao, Bing Liu 0001, Yong Yu 0001, Hang Li 0001 |
ICDM | 2 |
| 2006 | CoMiner: An Effective Algorithm for Mining Competitors from the WebabstractThis paper attempts to accomplish a novel task of mining competitive information with respect to an entity (such as a company, product, person) from the web. An algorithm called "CoMiner" is proposed, which first extracts a set of comparative candidates of the input entity and then ranks them according to the comparability, and finally extracts the competitive fields. The experimental results show that the proposed algorithm drafts a complete picture of competitive relation of a given entity effectively. Rui Li 0049, Shenghua Bao, Yong Yu 0001, Yunbo Cao |
ICDM | 5 |
| 2006 | Adapting ranking SVM to document retrievalabstractThe paper is concerned with applying learning to rank to document retrieval. Ranking SVM is a typical method of learning to rank. We point out that there are two factors one must consider when applying Ranking SVM, in general a "learning to rank" method, to document retrieval. First, correctly ranking documents on the top of the result list is crucial for an Information Retrieval system. One must conduct training in a way that such ranked results are accurate. Second, the number of relevant documents can vary from query to query. One must avoid training a model biased toward queries with a large number of relevant documents. Previously, when existing methods that include Ranking SVM were applied to document retrieval, none of the two factors was taken into consideration. We show it is possible to make modifications in conventional Ranking SVM, so it can be better used for document retrieval. Specifically, we modify the "Hinge Loss" function in Ranking SVM to deal with the problems described above. We employ two methods to conduct optimization on the loss function: gradient descent and quadratic programming. Experimental results show that our method, referred to as Ranking SVM for IR, can outperform the conventional Ranking SVM and other existing methods for document retrieval on two datasets. Yunbo Cao, Jun Xu 0001, Tie-Yan Liu, Hang Li 0001, Yalou Huang, Hsiao-Wuen Hon |
SIGIR | 1 |
| 2006 | Automatic extraction of titles from general documents using machine learning
Yunhua Hu, Hang Li 0001, Yunbo Cao, Li Teng 0002, Dmitriy Meyerzon |
Inf. Process. Manag. | 3 |
| 2006 | A Supervised Learning Approach to Search of Definitions
Jun Xu 0001, Yunbo Cao, Hang Li 0001, Yalou Huang |
J. Comput. Sci. Technol. | 2 |
| 2005 | A new approach to intranet search based on information extractionabstractThis paper is concerned with 'intranet search'. By intranet search, we mean searching for information on an intranet within an organization. We have found that search needs on an intranet can be categorized into types, through an analysis of survey results and an analysis of search log data. The types include searching for definitions, persons, experts, and homepages. Traditional information retrieval only focuses on search of relevant documents, but not on search of special types of information. We propose a new approach to intranet search in which we search for information in each of the special types, in addition to the traditional relevance search. Information extraction technologies can play key roles in such kind of 'search by type' approach, because we must first extract from the documents the necessary information in each type. We have developed an intranet search system called 'Information Desk'. In the system, we try to address the most important types of search first - finding term definitions, homepages of groups or topics, employees' personal information and experts on topics. For each type of search, we use information extraction technologies to extract, fuse, and summarize information in advance. The system is in operation on the intranet of Microsoft and receives accesses from about 500 employees per month. Feedbacks from users and system logs show that users consider the approach useful and the system can really help people to find information. This paper describes the architecture, features, component technologies, and evaluation results of the system. Hang Li 0001, Yunbo Cao, Jun Xu 0001, Yunhua Hu, Shenjie Li, Dmitriy Meyerzon |
CIKM | 2 |
| 2005 | Email data cleaningabstractAddressed in this paper is the issue of ‘email data cleaning ’ for text mining. Many text mining applications need take emails as input. Email data is usually noisy and thus it is necessary to clean it before mining. Several products offer email cleaning features, however, the types of noises that can be eliminated are restricted. Despite the importance of the problem, email cleaning has received little attention in the research community. A thorough and systematic investigation on the issue is thus needed. In this paper, email cleaning is formalized as a problem of non-text filtering and text normalization. In this way, email cleaning becomes independent from any specific text mining processing. A cascaded approach is proposed, which cleans up an email in four passes including non-text filtering, paragraph normalization, sentence normalization, and word normalization. As far as we know, non-text filtering and paragraph normalization have not been investigated previously. Methods for performing the tasks on the basis of Support Vector Machines (SVM) have also been proposed in this paper. Features in the models have been defined. Experimental results indicate that the proposed SVM based methods can significantly outperform the baseline methods for email cleaning. The proposed method has been applied to term extraction, a typical text mining processing. Experimental results show that the accuracy of term extraction can be significantly improved by using the data cleaning method. Jie Tang 0001, Hang Li 0001, Yunbo Cao, ZhaoHui Tang |
KDD | 3 |
| 2005 | Title extraction from bodies of HTML documents and its application to web page retrievalabstractThis paper is concerned with automatic extraction of titles from the bodies of HTML documents. Titles of HTML documents should be correctly defined in the title fields; however, in reality HTML titles are often bogus. It is desirable to conduct automatic extraction of titles from the bodies of HTML documents. This is an issue which does not seem to have been investigated previously. In this paper, we take a supervised machine learning approach to address the problem. We propose a specification on HTML titles. We utilize format information such as font size, position, and font weight as features in title extraction. Our method significantly outperforms the baseline method of using the lines in largest font size as title (20.9%-32.6% improvement in F1 score). As application, we consider web page retrieval. We use the TREC Web Track data for evaluation. We propose a new method for HTML documents retrieval using extracted titles. Experimental results indicate that the use of both extracted titles and title fields is almost always better than the use of title fields alone; the use of extracted titles is particularly helpful in the task of named page finding (23.1% -29.0% improvements). Yunhua Hu, Guomao Xin, Ruihua Song, Shuming Shi 0001, Yunbo Cao, Hang Li 0001 |
SIGIR | 6 |
| 2003 | Uncertainty Reduction in Collaborative Bootstrapping: Measure and AlgorithmabstractThis paper proposes the use of uncertainty reduction in machine learning methods such as co-training and bilingual boot-strapping, which are referred to, in a general term, as 'collaborative bootstrapping'. The paper indicates that uncertainty reduction is an important factor for enhancing the performance of collaborative bootstrapping. It proposes a new measure for representing the degree of uncertainty correlation of the two classifiers in collaborative bootstrapping and uses the measure in analysis of collaborative bootstrapping. Furthermore, it proposes a new algorithm of collaborative bootstrapping on the basis of uncertainty reduction. Experimental results have verified the correctness of the analysis and have demonstrated the significance of the new algorithm. Yunbo Cao, Hang Li 0001, Li Lian |
ACL | 1 |
| 2002 | Base Noun Phrase Translation Using Web Data and the EM Algorithm
Yunbo Cao, Hang Li 0001 |
COLING | 1 |