Chin-Yew Lin

dblp:64/6843 · DBLP profile ↗
← Back
122ranked-venue papers
11as first author
20since 2021 · last 2025
0000-0002-0798-6365ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 97 · 10 first-author · 16 since 2021Databases, data management, data science and information retrieval · 29 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 4Human-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2025 SeCom: On Memory Construction and Retrieval for Personalized Conversational Agents
abstract
To deliver coherent and personalized experiences in long-term conversations, existing approaches typically perform retrieval augmented response generation by constructing memory banks from conversation history at either the turn-level, session-level, or through summarization techniques. In this paper, we explore the impact of different memory granularities and present two key findings: (1) Both turn-level and session-level memory units are suboptimal, affecting not only the quality of final responses, but also the accuracy of the retrieval process. (2) The redundancy in natural language introduces noise, hindering precise retrieval. We demonstrate that *LLMLingua-2*, originally designed for prompt compression to accelerate LLM inference, can serve as an effective denoising method to enhance memory retrieval accuracy. Building on these insights, we propose **SeCom**, a method that constructs a memory bank with topical segments by introducing a conversation **Se**gmentation model, while performing memory retrieval based on **Com**pressed memory units. Experimental results show that **SeCom** outperforms turn-level, session-level, and several summarization-based methods on long-term conversation benchmarks such as *LOCOMO* and *Long-MT-Bench+*. Additionally, the proposed conversation segmentation method demonstrates superior performance on dialogue segmentation datasets such as *DialSeg711*, *TIAGE*, and *SuperDialSeg*.
Zhuoshi Pan, Qianhui Wu, Huiqiang Jiang, Xufang Luo, Hao Cheng 0002, Dongsheng Li 0002, Yuqing Yang 0001, Chin-Yew Lin, H. Vicky Zhao, Lili Qiu, Jianfeng Gao 0001
ICLR8
2024 Spot the Error: Non-autoregressive Graphic Layout Generation with Wireframe Locator
abstract
Layout generation is a critical step in graphic design to achieve meaningful compositions of elements. Most previous works view it as a sequence generation problem by concatenating element attribute tokens (i.e., category, size, position). So far the autoregressive approach (AR) has achieved promising results, but is still limited in global context modeling and suffers from error propagation since it can only attend to the previously generated tokens. Recent non-autoregressive attempts (NAR) have shown competitive results, which provides a wider context range and the flexibility to refine with iterative decoding. However, current works only use simple heuristics to recognize erroneous tokens for refinement which is inaccurate. This paper first conducts an in-depth analysis to better understand the difference between the AR and NAR framework. Furthermore, based on our observation that pixel space is more sensitive in capturing spatial patterns of graphic layouts (e.g., overlap, alignment), we propose a learning-based locator to detect erroneous tokens which takes the wireframe image rendered from the generated layout sequence as input. We show that it serves as a complementary modality to the element sequence in object space and contributes greatly to the overall performance. Experiments on two public datasets show that our approach outperforms both AR and NAR baselines. Extensive studies further prove the effectiveness of different modules with interesting findings. Our code will be available at https://github.com/ffffatgoose/SpotError.
Jieru Lin, Danqing Huang, Tiejun Zhao, Dechen Zhan, Chin-Yew Lin
AAAI5
2024 LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression
abstract
Huiqiang Jiang, Qianhui Wu, Xufang Luo, Dongsheng Li, Chin-Yew Lin, Yuqing Yang, Lili Qiu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Huiqiang Jiang, Qianhui Wu, Xufang Luo, Dongsheng Li 0002, Chin-Yew Lin, Yuqing Yang 0001, Lili Qiu
ACL (1)5
2024 Desigen: A Pipeline for Controllable Design Template Generation
abstract
Templates serve as a good starting point to implement a design (e.g., banner, slide) but it takes great effort from designers to manually create. In this paper, we present Desigen, an automatic template creation pipeline which generates background images as well as harmonious layout elements over the background. Different from natural images, a background image should preserve enough non-salient space for the overlaying layout elements. To equip existing advanced diffusion-based models with stronger spatial control, we propose two simple but effective techniques to constrain the saliency distribution and reduce the attention weight in desired regions during the background generation process. Then conditioned on the background, we synthesize the layout with a Transformer-based autoregressive generator. To achieve a more harmonious composition, we propose an iterative inference strategy to adjust the synthesized background and layout in multiple rounds. We constructed a design dataset with more than 40k advertisement banners to verify our approach. Extensive experiments demonstrate that the proposed pipeline generates high-quality templates comparable to human designers. More than a single-page design, we further show an application of presentation generation that outputs a set of theme-consistent slides. The data and code are available at https://whaohan.github.io/desigen.
Haohan Weng, Danqing Huang, Chin-Yew Lin, Tong Zhang 0015, C. L. Philip Chen
CVPR5
2024 MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
abstract
The computational challenges of Large Language Model (LLM) inference remain a significant barrier to their widespread deployment, especially as prompt lengths continue to increase. Due to the quadratic complexity of the attention computation, it takes 30 minutes for an 8B LLM to process a prompt of 1M tokens (i.e., the pre-filling stage) on a single A100 GPU. Existing methods for speeding up prefilling often fail to maintain acceptable accuracy or efficiency when applied to long-context LLMs. To address this gap, we introduce MInference (Milliontokens Inference), a sparse calculation method designed to accelerate pre-filling of long-sequence processing. Specifically, we identify three unique patterns in long-context attention matrices-the A-shape, Vertical-Slash, and Block-Sparse-that can be leveraged for efficient sparse computation on GPUs. We determine the optimal pattern for each attention head offline and dynamically build sparse indices based on the assigned pattern during inference. With the pattern and sparse indices, we perform efficient sparse attention calculations via our optimized GPU kernels to significantly reduce the latency in the pre-filling stage of longcontext LLMs. Our proposed technique can be directly applied to existing LLMs without any modifications to the pre-training setup or additional fine-tuning. By evaluating on a wide range of downstream tasks, including InfiniteBench, RULER, PG-19, and Needle In A Haystack, and models including LLaMA-3-1M, GLM-4-1M, Yi-200K, Phi-3-128K, and Qwen2-128K, we demonstrate that MInference effectively reduces inference latency by up to 10x for pre-filling on an A100, while maintaining accuracy. Our code is available at https://aka.ms/MInference.
Huiqiang Jiang, Chengruidong Zhang, Qianhui Wu, Xufang Luo, Surin Ahn, Zhenhua Han, Amir H. Abdi, Dongsheng Li 0002, Chin-Yew Lin, Yuqing Yang 0001, Lili Qiu
NeurIPS10
2024 Decomposed Meta-Learning for Few-Shot Sequence Labeling
abstract
Few-shot sequence labeling is a general problem formulation for many natural language understanding tasks in data-scarcity scenarios, which require models to generalize to new types via only a few labeled examples. Recent advances mostly adopt metric-based meta-learning and thus face the challenges of modeling the miscellaneousOtherprototype and the inability to generalize to classes with large domain gaps. To overcome these challenges, we propose a decomposed meta-learning framework for few-shot sequence labeling that breaks down the task into few-shot mention detection and few-shot type classification, and sequentially tackles them via meta-learning. Specifically, we employ model-agnostic meta-learning (MAML) to prompt the mention detection model to learn boundary knowledge shared across types. With the detected mention spans, we further leverage the MAML-enhanced span-level prototypical network for few-shot type classification. In this way, the decomposition framework bypasses the requirement of modeling the miscellaneousOtherprototype. Meanwhile, the adoption of the MAML algorithm enables us to explore the knowledge contained in support examples more efficiently, so that our model can quickly adapt to new types using only a few labeled examples. Under our framework, we explore a basic implementation that uses two separate models for the two subtasks. We further propose a joint model to reduce model size and inference time, making our framework more applicable for scenarios with limited resources. Extensive experiments on nine benchmark datasets, including named entity recognition, slot tagging, event detection, and part-of-speech tagging, show that the proposed approach achieves start-of-the-art performance across various few-shot sequence labeling tasks.
Qianhui Wu, Huiqiang Jiang, Jieru Lin, Börje Karlsson 0001, Tiejun Zhao, Chin-Yew Lin
IEEE ACM Trans. Audio Speech Lang. Process.7
2023 Layout Generation as Intermediate Action Sequence Prediction
abstract
Layout generation plays a crucial role in graphic design intelligence. One important characteristic of the graphic layouts is that they usually follow certain design principles. For example, the principle of repetition emphasizes the reuse of similar visual elements throughout the design. To generate a layout, previous works mainly attempt at predicting the absolute value of bounding box for each element, where such target representation has hidden the information of higher-order design operations like repetition (e.g. copy the size of the previously generated element). In this paper, we introduce a novel action schema to encode these operations for better modeling the generation process. Instead of predicting the bounding box values, our approach autoregressively outputs the intermediate action sequence, which can then be deterministically converted to the final layout. We achieve state-of-the-art performances on three datasets. Both automatic and human evaluations show that our approach generates high-quality and diverse layouts. Furthermore, we revisit the commonly used evaluation metric FID adapted in this task, and observe that previous works use different settings to train the feature extractor for obtaining real/generated data distribution, which leads to inconsistent conclusions. We conduct an in-depth analysis on this metric and settle for a more robust and reliable evaluation setting. Code is available at this website.
Huiting Yang, Danqing Huang, Chin-Yew Lin, Shengfeng He
AAAI3
2023 CoLaDa: A Collaborative Label Denoising Framework for Cross-lingual Named Entity Recognition
abstract
Tingting Ma, Qianhui Wu, Huiqiang Jiang, Börje Karlsson, Tiejun Zhao, Chin-Yew Lin. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Qianhui Wu, Huiqiang Jiang, Börje Karlsson 0001, Tiejun Zhao, Chin-Yew Lin
ACL (1)6
2023 Multi-Level Knowledge Distillation for Out-of-Distribution Detection in Text
abstract
Self-supervised representation learning has proved to be a valuable component for outof-distribution (OoD) detection with only the texts of in-distribution (ID) examples.These approaches either train a language model from scratch or fine-tune a pre-trained language model using ID examples, and then take the perplexity output by the language model as OoD scores.In this paper, we analyze the complementary characteristics of both OoD detection methods and propose a multi-level knowledge distillation approach that integrates their strengths while mitigating their limitations.Specifically, we use a fine-tuned model as the teacher to teach a randomly initialized student model on the ID examples.Besides the prediction layer distillation, we present a similarity-based intermediate layer distillation method to thoroughly explore the representation space of the teacher model.In this way, the learned student can better represent the ID data manifold while gaining a stronger ability to map OoD examples outside the ID data manifold with the regularization inherited from pre-training.Besides, the student model sees only ID examples during parameter learning, further promoting more distinguishable features for OoD detection.We conduct extensive experiments over multiple benchmark datasets, i.e., CLINC150, SST, ROSTD, 20 NewsGroups, and AG News; showing that the proposed method yields new state-of-the-art performance 1 .We also explore its application as an AIGC detector to distinguish between answers generated by ChatGPT and human experts.It is observed that our model exceeds human evaluators in the pair-expert task on the Human ChatGPT Comparison Corpus.
Qianhui Wu, Huiqiang Jiang, Haonan Yin, Börje Karlsson 0001, Chin-Yew Lin
ACL (1)5
2023 LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models
abstract
Large language models (LLMs) have been applied in various applications due to their astonishing capabilities.With advancements in technologies such as chain-of-thought (CoT) prompting and in-context learning (ICL), the prompts fed to LLMs are becoming increasingly lengthy, even exceeding tens of thousands of tokens.To accelerate model inference and reduce cost, this paper presents LLMLingua, a coarse-to-fine prompt compression method that involves a budget controller to maintain semantic integrity under high compression ratios, a token-level iterative compression algorithm to better model the interdependence between compressed contents, and an instruction tuning based method for distribution alignment between language models.We conduct experiments and analysis over four datasets from different scenarios, i.e., GSM8K, BBH, ShareGPT, and Arxiv-March23; showing that the proposed approach yields state-of-the-art performance and allows for up to 20x compression with little performance loss. 1
Huiqiang Jiang, Qianhui Wu, Chin-Yew Lin, Yuqing Yang 0001, Lili Qiu
EMNLP3
2023 Relation-enhanced DETR for Component Detection in Graphic Design Reverse Engineering
abstract
It is a common practice for designers to create digital prototypes from a mock-up/screenshot. Reverse engineering graphic design by detecting its components (e.g., text, icon, button) helps expedite this process. This paper first conducts a statistical analysis to emphasize the importance of relations in graphic layouts, which further motivates us to incorporate relation modeling into component detection. Built on the current state-of-the-art DETR (DEtection TRansformer), we introduce a learnable relation matrix to model class correlations. Specifically, the matrix will be added in the DETR decoder to update the query-to-query self-attention. Experiment results on three public datasets show that our approach achieves better performance than several strong baselines. We further visualize the learnt relation matrix and observe some reasonable patterns. Moreover, we show an application of component detection where we leverage the detection outputs as augmented training data for layout generation, which achieves promising results.
Xixuan Hao, Danqing Huang, Jieru Lin, Chin-Yew Lin
IJCAI4
2023 Learn and Sample Together: Collaborative Generation for Graphic Design Layout
abstract
In the process of graphic layout generation, user specifications including element attributes and their relationships are commonly used to constrain the layouts (e.g.,"put the image above the button''). It is natural to encode spatial constraints between elements using a graph. This paper presents a two-stage generation framework: a spatial graph generator and a subsequent layout decoder which is conditioned on the previous output graph. Training the two highly dependent networks separately as in previous work, we observe that the graph generator generates out-of-distribution graphs with a high frequency, which are unseen to the layout decoder during training and thus leads to huge performance drop in inference. To coordinate the two networks more effectively, we propose a novel collaborative generation strategy to perform round-way knowledge transfer between the networks in both training and inference. Experiment results on three public datasets show that our model greatly benefits from the collaborative generation and has achieved the state-of-the-art performance. Furthermore, we conduct an in-depth analysis to better understand the effectiveness of graph condition modeling.
Haohan Weng, Danqing Huang, Tong Zhang 0015, Chin-Yew Lin
IJCAI4
2022 TIARA: Multi-grained Retrieval for Robust Question Answering over Large Knowledge Base
abstract
Pre-trained language models (PLMs) have shown their effectiveness in multiple scenarios.However, KBQA remains challenging, especially regarding coverage and generalization settings.This is due to two main factors: i) understanding the semantics of both questions and relevant knowledge from the KB; ii) generating executable logical forms with both semantic and syntactic correctness.In this paper, we present a new KBQA model, TIARA, which addresses those issues by applying multi-grained retrieval to help the PLM focus on the most relevant KB contexts, viz., entities, exemplary logical forms, and schema items.Moreover, constrained decoding is used to control the output space and reduce generation errors.Experiments over important benchmarks demonstrate the effectiveness of our approach.TIARA outperforms previous SOTA, including those using PLMs or oracle entity annotations, by at least 4.1 and 1.1 F1 points on GrailQA and WebQuestionsSP, respectively.Specifically on GrailQA, TIARA outperforms previous models in all categories, with an improvement of 4.7 F1 points in zero-shot generalization. 1
Yiheng Shu, Zhiwei Yu 0001, Yuhan Li 0001, Börje Karlsson 0001, Yuzhong Qu, Chin-Yew Lin
EMNLP7
2022 Exploring Heterogeneous Feature Representation for Document Layout Understanding
abstract
There are increasing interests in document layout representation learning and understanding. Transformer, with its great power, has become the mainstream model architecture and achieved promising results in this area. As elements in a document layout consist of multi-modal and multi-dimensional features such as position, size, and its text content, prior works represent each element by summing all feature embeddings into one unified vector in the input layer, which is then fed into the self-attention for element-wise interaction. However, this simple summation would potentially raise mixed correlations among heterogeneous features and bring noise to the representation learning. In this paper, we propose a novel two-step disentangled attention mechanism to allow more flexible feature interactions in the self-attention. Furthermore, inspired by the principles of document design (e.g., contrast, proximity), we propose an unsupervised learning objective to constrain the layout representations. We verify our approach on two layout understanding tasks, namely element role labeling and image captioning. Experiment results show that our approach achieves state-of-the-art performances. Moreover, we conduct extensive studies and observe better interpretability using our approach.
Guosheng Feng, Danqing Huang, Chin-Yew Lin, Damjan Dakic, Milos Milunovic, Tamara Stankovic, Igor Ilic
ICTAI3
2022 On the Effectiveness of Sentence Encoding for Intent Detection Meta-Learning
abstract
Tingting Ma, Qianhui Wu, Zhiwei Yu, Tiejun Zhao, Chin-Yew Lin. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Qianhui Wu, Zhiwei Yu 0001, Tiejun Zhao, Chin-Yew Lin
NAACL-HLT5
2022 A Mixed-Initiative Approach to Reusing Infographic Charts
abstract
Infographic bar charts have been widely adopted for communicating numerical information because of their attractiveness and memorability. However, these infographics are often created manually with general tools, such as PowerPoint and Adobe Illustrator, and merely composed of primitive visual elements, such as text blocks and shapes. With the absence of chart models, updating or reusing these infographics requires tedious and error-prone manual edits. In this paper, we propose a mixed-initiative approach to mitigate this pain point. On one hand, machines are adopted to perform precise and trivial operations, such as mapping numerical values to shape attributes and aligning shapes. On the other hand, we rely on humans to perform subjective and creative tasks, such as changing embellishments or approving the edits made by machines. We encapsulate our technique in a PowerPoint add-in prototype and demonstrate the effectiveness by applying our technique on a diverse set of infographic bar chart examples.
Weiwei Cui 0001, Jinpeng Wang 0001, Yun Wang 0012, Chin-Yew Lin, Dongmei Zhang 0001
IEEE Trans. Vis. Comput. Graph.5
2022 LinkingPark: An automatic semantic table interpretation system
Shuang Chen 0003, Alperen Karaoglu, Carina Negreanu, Jin-Ge Yao, Jack Williams 0001, Feng Jiang 0001, Andrew D. Gordon 0001, Chin-Yew Lin
J. Web Semant.9
2021 Towards Topic-Aware Slide Generation For Academic Papers With Unsupervised Mutual Learning
abstract
Slides are commonly used to present information and tell stories. In academic and research communities, slides are typically used to summarize findings in accepted papers for presentation in meetings and conferences. These slides for academic papers usually contain common and essential topics such as major contributions, model design, experiment details and future work. In this paper, we aim to automatically generate slides for academic papers. We first conducted an in-depth analysis of how humans create slides. We then mined frequently used slide topics. Given a topic, our approach extracts relevant sentences in the paper to provide the draft slides. Due to the lack of labeling data, we integrate prior knowledge of ground truth sentences into a log-linear model to create an initial pseudo-target distribution. Two sentence extractors are learned collaboratively and bootstrap the performance of each other. Evaluation results on a labeled test set show that our model can extract more relevant sentences than baseline methods. Human evaluation also shows slides generated by our model can serve as a good basis for preparing the final presentations.
Danqing Huang, Chin-Yew Lin
AAAI4
2021 CanvasEmb: Learning Layout Representation with Large-scale Pre-training for Graphic Design
abstract
Layout representation, which models visual elements and their inter-relations in a canvas, plays a crucial role in graphic design intelligence. With a large variety of layout designs and the unique characteristic of layouts that visual elements are defined as a list of categorical (e.g., type) and numerical (e.g., position and size) properties, it is challenging to learn general and compact representations with limited data. Inspired by the recent success of self-supervised pre-training techniques in various natural language processing tasks, in this paper, we propose CanvasEmb (Canvas Embedding), which pre-trains deep representations from unlabeled graphic designs by jointly conditioning on all the context elements in a canvas, with a multi-dimensional feature encoder and a multi-task learning objective. The pre-trained CanvasEmb model can be fine-tuned with just one additional output layer and with a small size of training data to create models for a wide range of downstream tasks. We verify our approach with presentation slides data. We construct a large-scale dataset with more than one million slides and propose two layout understanding tasks with human-labeled sets, namely element role labeling and image captioning. Evaluation results on these two tasks show that our model with fine-tuning achieves state-of-the-art performance. Furthermore, we conduct a deep analysis aiming to understand the modeling mechanism of CanvasEmb and demonstrate its great potential with two extended applications: layout auto completion and layout retrieval.
Yuxi Xie, Danqing Huang, Jinpeng Wang 0001, Chin-Yew Lin
ACM Multimedia4
2021 ChartOCR: Data Extraction from Charts Images via a Deep Hybrid Framework
abstract
Chart images are commonly used for data visualization. Automatically reading the chart values is a key step for chart content understanding. Charts have a lot of variations in style (e.g. bar chart, line chart, pie chart and etc.), which makes pure rule-based data extraction methods difficult to handle. However, it is also improper to directly apply end- to-end deep learning solutions since these methods usually deal with specific types of charts. In this paper, we propose an unified method ChartOCR to extract data from various types of charts. We show that by combing deep framework and rule-based methods, we can achieve a satisfying generalization ability and obtain accurate and semantic-rich intermediate results. Our method extracts the key points that define the chart components. By adjusting the prior rules, the framework can be applied to different chart types. Experiments show that our method achieves state-of-the- art performance with fast processing speed on two public datasets. Besides, we also introduce and evaluate on a large dataset ExcelChart400K for training deep models on chart images. The code and the dataset are publicly available at https://github.com/soap117/DeepRule.
Junyu Luo 0001, Jinpeng Wang 0001, Chin-Yew Lin
WACV4
2020 Improving Entity Linking by Modeling Latent Entity Type Information
abstract
Existing state of the art neural entity linking models employ attention-based bag-of-words context model and pre-trained entity embeddings bootstrapped from word embeddings to assess topic level context compatibility. However, the latent entity type information in the immediate context of the mention is neglected, which causes the models often link mentions to incorrect entities with incorrect type. To tackle this problem, we propose to inject latent entity type information into the entity embeddings based on pre-trained BERT. In addition, we integrate a BERT-based entity similarity score into the local context model of a state-of-the-art model to better capture latent entity type information. Our model significantly outperforms the state-of-the-art entity linking models on standard benchmark (AIDA-CoNLL). Detailed experiment analysis demonstrates that our model corrects most of the type errors produced by the direct baseline.
Shuang Chen 0003, Jinpeng Wang 0001, Feng Jiang 0001, Chin-Yew Lin
AAAI4
2020 Enhanced Meta-Learning for Cross-Lingual Named Entity Recognition with Minimal Resources
abstract
For languages with no annotated resources, transferring knowledge from rich-resource languages is an effective solution for named entity recognition (NER). While all existing methods directly transfer from source-learned model to a target language, in this paper, we propose to fine-tune the learned model with a few similar examples given a test case, which could benefit the prediction by leveraging the structural and semantic information conveyed in such similar examples. To this end, we present a meta-learning algorithm to find a good model parameter initialization that could fast adapt to the given test case and propose to construct multiple pseudo-NER tasks for meta-training by computing sentence similarities. To further improve the model's generalization ability across different languages, we introduce a masking scheme and augment the loss function with an additional maximum term during meta-training. We conduct extensive experiments on cross-lingual named entity recognition with minimal resources over five target languages. The results show that our approach significantly outperforms existing state-of-the-art methods across the board.
Qianhui Wu, Zijia Lin, Hui Chen 0013, Börje Karlsson 0001, Biqing Huang, Chin-Yew Lin
AAAI7
2020 Learning Semantic Correspondences from Noisy Data-text Pairs by Local-to-Global Alignments
abstract
Learning semantic correspondences between structured input data (e.g., slot-value pairs) and associated texts is a core problem for many downstream NLP applications, e.g., data-to-text generation.Large-scale datasets recently proposed for generation contain loosely corresponding data text pairs, where part of spans in text cannot be aligned to its incomplete paired input.To learn semantic correspondences from such datasets, we propose a two-stage local-to-global alignment (L2GA) framework.First, a local model based on multi-instance learning is applied to build alignments for texts spans that can be directly grounded to its paired structured input.Then, a novel global model built upon a memory-guided conditional random field (CRF) layer aims to infer missing alignments for text spans not supported by paired incomplete inputs, where the memory is designed to leverage alignment clues provided by the local model to strengthen the global model.In this way, the local model and global model can work jointly to learn semantic correspondences in the same framework.Experimental results show that our proposed method can be generalized to both restaurant and computer domains and improve the alignment accuracy.
Feng Nie, Jinpeng Wang 0001, Chin-Yew Lin
COLING3
2020 Visual Style Extraction from Chart Images for Chart Restyling
abstract
Creating a good looking chart for better visualization is time consuming. There are plenty of well-designed charts on the Web, which are ideal references for imitation of chart style. However, stored as bitmap images, reference charts have hinder machine interpretation of style settings and thus difficult to be directly applied. In this paper, we extract visual properties from reference chart images as style templates to restyle charts. We first construct a large-scale dataset of 187,059 chart images from real world data, labeled with predefined visual property values. Then we introduce an end-to-end learning network to extract the properties based on two image-encoding approaches. Furthermore, in order to capture spatial relationships of chart objects, which are crucial in solving the task, we propose a novel positional encoding method to integrate clues of relative positions between objects. Experimental results show that our model significantly outperforms baseline models. By adding positional features, our model achieves better performance. Finally, we present the application for chart restyling based on our model.
Danqing Huang, Jinpeng Wang 0001, Chin-Yew Lin
ICPR4
2020 Hybrid Cascade Point Search Network for High Precision Bar Chart Component Detection
abstract
Bar charts are commonly used for data visualization. One common form of chart distribution is in its image form. To enable machine comprehension of chart images, precise detection of chart components in chart images is a critical step. Existing image object detection methods do not perform well in chart component detection which requires high boundary detection precision. And traditional rule-based approaches lack enough generalization ability. In order to address this problem, we design a novel two-stage component detection framework for bar charts that combines point-based and region-based ideas, by simulating the process that human creating bounding boxes for objects. The experiment on our labeled ChartDet dataset shows our method greatly improves the performance of chart object detection. We further extend our method to a general object detection task and get comparable performance.
Junyu Luo 0001, Jinpeng Wang 0001, Chin-Yew Lin
ICPR3
2019 Towards Improving Neural Named Entity Recognition with Gazetteers
abstract
Most of the recently proposed neural models for named entity recognition have been purely data-driven, with a strong emphasis on getting rid of the efforts for collecting external resources or designing hand-crafted features.This could increase the chance of overfitting since the models cannot access any supervision signal beyond the small amount of annotated data, limiting their power to generalize beyond the annotated entities.In this work, we show that properly utilizing external gazetteers could benefit segmental neural NER models.We add a simple module on the recently proposed hybrid semi-Markov CRF architecture and observe some promising results.
Tianyu Liu 0004, Jin-Ge Yao, Chin-Yew Lin
ACL (1)3
2019 A Simple Recipe towards Reducing Hallucination in Neural Surface Realisation
abstract
Recent neural language generation systems often hallucinate contents (i.e., producing irrelevant or contradicted facts), especially when trained on loosely corresponding pairs of the input structure and text.To mitigate this issue, we propose to integrate a language understanding module for data refinement with selftraining iterations to effectively induce strong equivalence between the input data and the paired text.Experiments on the E2E challenge dataset show that our proposed framework can reduce more than 50% relative unaligned noise from the original data-text pairs.A vanilla sequence-to-sequence neural NLG model trained on the refined data has improved on content correctness compared with the current state-of-the-art ensemble generator.
Feng Nie, Jin-Ge Yao, Jinpeng Wang 0001, Chin-Yew Lin
ACL (1)5
2019 Enhancing Neural Data-To-Text Generation Models with External Background Knowledge
abstract
Shuang Chen, Jinpeng Wang, Xiaocheng Feng, Feng Jiang, Bing Qin, Chin-Yew Lin. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Shuang Chen 0003, Jinpeng Wang 0001, Feng Jiang 0001, Bing Qin 0001, Chin-Yew Lin
EMNLP/IJCNLP (1)6
2019 An Encoder with non-Sequential Dependency for Neural Data-to-Text Generation
abstract
Data-to-text generation aims to generate descriptions given a structured input data (i.e., a table with multiple records).Existing neural methods for encoding input data can be divided into two categories: a) pooling based encoders which ignore dependencies between input records or b) recurrent encoders which model only sequential dependencies between input records.In our investigation, although the recurrent encoder generally outperforms the pooling based encoder by learning the sequential dependencies, it is sensitive to the order of the input records (i.e., performance decreases when injecting the random shuffling noise over input data).To overcome this problem, we propose to adopt the self-attention mechanism to learn dependencies between arbitrary input records.Experimental results show the proposed method achieves comparable results and remains stable under random shuffling over input data.
Feng Nie, Jinpeng Wang 0001, Chin-Yew Lin
INLG4
2019 Beyond Word for Word: Fact Guided Training for Neural Data-to-Document Generation
Feng Nie, Hailin Chen, Jinpeng Wang 0001, Chin-Yew Lin
NLPCC (1)5
2018 Mention and Entity Description Co-Attention for Entity Disambiguation
abstract
For the task of entity disambiguation, mention contexts and entity descriptions both contain various kinds of information content while only a subset of them are helpful for disambiguation. In this paper, we propose a type-aware co-attention model for entity disambiguation, which tries to identify the most discriminative words from mention contexts and most relevant sentences from corresponding entity descriptions simultaneously. To bridge the semantic gap between mention contexts and entity descriptions, we further incorporate entity type information to enhance the co-attention mechanism. Our evaluation shows that the proposed model outperforms the state-of-the-arts on three public datasets. Further analysis also confirms that both the co-attention mechanism and the type-aware mechanism are effective.
Feng Nie, Yunbo Cao, Jinpeng Wang 0001, Chin-Yew Lin
AAAI4
2018 Using Intermediate Representations to Solve Math Word Problems
abstract
To solve math word problems, previous statistical approaches attempt at learning a direct mapping from a problem description to its corresponding equation system.However, such mappings do not include the information of a few higher-order operations that cannot be explicitly represented in equations but are required to solve the problem.The gap between natural language and equations makes it difficult for a learned model to generalize from limited data.In this work we present an intermediate meaning representation scheme that tries to reduce this gap.We use a sequence-to-sequence model with a novel attention regularization term to generate the intermediate forms, then execute them to obtain the final answers.Since the intermediate forms are latent, we propose an iterative labeling framework for learning by leveraging supervision signals from both equations and answers.Our experiments show using intermediate forms outperforms directly predicting equations.
Danqing Huang, Jin-Ge Yao, Chin-Yew Lin, Qingyu Zhou, Jian Yin 0001
ACL (1)3
2018 Open-Schema Event Profiling for Massive News Corpora
abstract
With the rapid growth of online information services, a sheer volume of news data becomes available. To help people quickly digest the explosive information, we define a new problem - schema-based news event profiling - profiling events reported in open-domain news corpora, with a set of slots and slot-value pairs for each event, where the set of slots forms the schema of an event type. Such profiling not only provides readers with concise views of events, but also facilitates various applications such as information retrieval, knowledge graph construction and question answering. It is however a quite challenging task. The first challenge is to find out events and event types because they are both initially unknown. The second difficulty is the lack of pre-defined event-type schemas. Lastly, even with the schemas extracted, to generate event profiles from them is still essential yet demanding.
Quan Yuan 0001, Xiang Ren 0001, Wenqi He, Chao Zhang 0014, Xinhe Geng, Lifu Huang, Heng Ji 0001, Chin-Yew Lin, Jiawei Han 0001
CIKM8
2018 Neural Math Word Problem Solver with Reinforcement Learning
abstract
Sequence-to-sequence model has been applied to solve math word problems. The model takes math problem descriptions as input and generates equations as output. The advantage of sequence-to-sequence model requires no feature engineering and can generate equations that do not exist in training data. However, our experimental analysis reveals that this model suffers from two shortcomings: (1) generate spurious numbers; (2) generate numbers at wrong positions. In this paper, we propose incorporating copy and alignment mechanism to the sequence-to-sequence model (namely CASS) to address these shortcomings. To train our model, we apply reinforcement learning to directly optimize the solution accuracy. It overcomes the “train-test discrepancy” issue of maximum likelihood estimation, which uses the surrogate objective of maximizing equation likelihood during training while the evaluation metric is solution accuracy (non-differentiable) at test time. Furthermore, to explore the effectiveness of our neural model, we use our model output as a feature and incorporate it into the feature-based model. Experimental results show that (1) The copy and alignment mechanism is effective to address the two issues; (2) Reinforcement learning leads to better performance than maximum likelihood on this task; (3) Our neural model is complementary to the feature-based model and their combination significantly outperforms the state-of-the-art results.
Danqing Huang, Jing Liu 0022, Chin-Yew Lin, Jian Yin 0001
COLING3
2018 Aggregated Semantic Matching for Short Text Entity Linking
abstract
The task of entity linking aims to identify concepts mentioned in a text fragments and link them to a reference knowledge base.Entity linking in long text has been well studied in previous work.However, short text entity linking is more challenging since the texts are noisy and less coherent.To better utilize the local information provided in short texts, we propose a novel neural network framework, Aggregated Semantic Matching (ASM), in which two different aspects of semantic information between the local context and the candidate entity are captured via representationbased and interaction-based neural semantic matching models, and then two matching signals work jointly for disambiguation with a rank aggregation mechanism.Our evaluation shows that the proposed model outperforms the state-of-the-arts on public tweet datasets.
Feng Nie, Shuyan Zhou, Jing Liu 0022, Jinpeng Wang 0001, Chin-Yew Lin
CoNLL5
2018 Operation-guided Neural Networks for High Fidelity Data-To-Text Generation
abstract
Recent neural models for data-to-text generation are mostly based on data-driven end-toend training over encoder-decoder networks.Even though the generated texts are mostly fluent and informative, they often generate descriptions that are not consistent with the input structured data.This is a critical issue especially in domains that require inference or calculations over raw data.In this paper, we attempt to improve the fidelity of neural data-to-text generation by utilizing pre-executed symbolic operations.We propose a framework called Operationguided Attention-based sequence-to-sequence network (OpAtt), with a specifically designed gating mechanism as well as a quantization module for operation results to utilize information from pre-executed operations.Experiments on two sports datasets show our proposed method clearly improves the fidelity of the generated texts to the input structured data.
Feng Nie, Jinpeng Wang 0001, Jin-Ge Yao, Chin-Yew Lin
EMNLP5
2018 Learning Latent Semantic Annotations for Grounding Natural Language to Structured Data
abstract
Previous work on grounded language learning did not fully capture the semantics underlying the correspondences between structured world state representations and texts, especially those between numerical values and lexical terms.In this paper, we attempt at learning explicit latent semantic annotations from paired structured tables and texts, establishing correspondences between various types of values and texts.We model the joint probability of data fields, texts, phrasal spans, and latent annotations with an adapted semi-hidden Markov model, and impose a soft statistical constraint to further improve the performance.As a by-product, we leverage the induced annotations to extract templates for language generation.Experimental results suggest the feasibility of the setting in this study, as well as the effectiveness of our proposed framework.1
Guanghui Qin, Jin-Ge Yao, Jinpeng Wang 0001, Chin-Yew Lin
EMNLP5
2018 Design and Implementation of a Smart Headband for Epileptic Seizure Detection and Its Verification Using Clinical Database
abstract
Epilepsy is a common neural disorder disease; about 0.5% of the global population has epilepsy [1]. Most patients take antiepileptic drugs to reduce their seizures. Among them, nearly one-third of the patients are drug-resistant epilepsy. The alternative treatment is the resection surgery of removing the epileptogenic zone. However, all above patients will still suffer seizures occasionally, which will influence the patients' quality-of-life, and further introduce danger and inconvenience to patients and people around. This paper presents the design and development of the smart headband for epileptic seizure detection. The headband consists of a fabric headband with flexible print circuit (FPC) inside and fabric electrodes on it. The whole system includes the following circuits: an analog front-end circuitry, an epileptic seizure detection tag (ESDT), and a Bluetooth Low Power (BLE) chip. The result of designed circuits yields a compact and low-power design of smart headband for epileptic seizure detection which is suitable for wearable usage. The epileptic seizure detection algorithm is validated by Children's Hospital Boston-MIT (CHB-MIT) EEG database [2].
Shih-Kai Lin, Isti Qomah, Yu-Shan Lin, Herming Chiueh, Chin-Yew Lin
ISCAS5
2018 Revisiting Distant Supervision for Relation Extraction
Tingsong Jiang, Jing Liu 0022, Chin-Yew Lin, Zhifang Sui
LREC3
2018 INFAR: insight extraction from app reviews
abstract
App reviews play an essential role for users to convey their feedback about using the app. The critical information contained in app reviews can assist app developers for maintaining and updating mobile apps. However, the noisy nature and large-quantity of daily generated app reviews make it difficult to understand essential information carried in app reviews. Several prior studies have proposed methods that can automatically classify or cluster user reviews into a few app topics (e.g., security). These methods usually act on a static collection of user reviews. However, due to the dynamic nature of user feedback (i.e., reviews keep coming as new users register or new app versions being released) and multiple analysis dimensions (e.g., review quantity and user rating), developers still need to spend substantial effort in extracting contrastive information that can only be teased out by comparing data from multiple time periods or analysis dimensions. This is needed to answer questions such as: what kind of issues users are experiencing most? is there an unexpected rise in a particular kind of issue? etc. To address this need, in this paper, we introduce INFAR, a tool that automatically extracts INsights From App Reviews across time periods and analysis dimensions, and presents them in natural language supported by an interactive chart. The insights INFAR extracts include several perspectives: (1) salient topics (i.e., issue topics with significantly lower ratings), (2) abnormal topics (i.e., issue topics that experience a rapid rise in volume during a time period), (3) correlations between two topics, and (4) causal factors to rating or review quantity changes. To evaluate our tool, we conduct an empirical evaluation by involving six popular apps and 12 industrial practitioners, and 92% (11/12) of them approve the practical usefulness of the insights summarized by INFAR.
Cuiyun Gao 0001, Jichuan Zeng, David Lo 0001, Chin-Yew Lin, Michael R. Lyu, Irwin King
ESEC/SIGSOFT FSE4
2017 Trust, but Verify! Better Entity Linking through Automatic Verification
abstract
We introduce automatic verification as a post-processing step for entity linking (EL).The proposed method trusts EL system results collectively, by assuming entity mentions are mostly linked correctly, in order to create a semantic profile of the given text using geospatial and temporal information, as well as fine-grained entity types.This profile is then used to automatically verify each linked mention individually, i.e., to predict whether it has been linked correctly or not.Verification allows leveraging a rich set of global and pairwise features that would be prohibitively expensive for EL systems employing global inference.Evaluation shows consistent improvements across datasets and systems.In particular, when applied to state-of-theart systems, our method yields an absolute improvement in linking performance of up to 1.7 F 1 on AIDA/CoNLL'03 and up to 2.4 F 1 on the English TAC KBP 2015 TEDL dataset.* The majority of this work was done during an internship at Microsoft Research Asia. 1 We use entity to refer to both real-word entities and to their corresponding entries in the KB.
Benjamin Heinzerling, Michael Strube 0001, Chin-Yew Lin
EACL (1)3
2017 Learning Fine-Grained Expressions to Solve Math Word Problems
abstract
This paper presents a novel templatebased method to solve math word problems.This method learns the mappings between math concept phrases in math word problems and their math expressions from training data.For each equation template, we automatically construct a rich template sketch by aggregating information from various problems with the same template.Our approach is implemented in a two-stage system.It first retrieves a few relevant equation system templates and aligns numbers in math word problems to those templates for candidate equation generation.It then does a fine-grained inference to obtain the final answer.Experiment results show that our method achieves an accuracy of 28.4% on the linear Dolphin18K benchmark, which is 10% (54% relative) higher than previous stateof-the-art systems while achieving an accuracy increase of 12% (59% relative) on the TS6 benchmark subset.
Danqing Huang, Shuming Shi 0001, Chin-Yew Lin, Jian Yin 0001
EMNLP3
2016 How well do Computers Solve Math Word Problems? Large-Scale Dataset Construction and Evaluation
abstract
Recently a few systems for automatically solving math word problems have reported promising results. However, the datasets used for evaluation have limitations in both scale and diversity. In this paper, we build a large-scale dataset which is more than 9 times the size of previous ones, and contains many more problem types. Problems in the dataset are semi-automatically obtained from community question-answering (CQA) web pages. A ranking SVM model is trained to automatically extract problem answers from the answer text provided by CQA users, which significantly reduces human annotation cost. Experiments conducted on the new dataset lead to interesting and surprising results.
Danqing Huang, Shuming Shi 0001, Chin-Yew Lin, Jian Yin 0001, Wei-Ying Ma
ACL (1)3
2016 News Citation Recommendation with Implicit and Explicit Semantics
abstract
In this work, we focus on the problem of news citation recommendation. The task aims to recommend news citations for both authors and readers to create and search news references. Due to the sparsity issue of news citations and the engineering difficulty in obtaining information on authors, we focus on content similarity-based methods instead of collaborative filtering-based approaches. In this paper, we explore word embedding (i.e., implicit semantics) and grounded entities (i.e., explicit semantics) to address the variety and ambiguity issues of language. We formulate the problem as a reranking task and integrate different similarity measures under the learning to rank framework. We evaluate our approach on a real-world dataset. The experimental results show the efficacy of our method.
Jing Liu 0022, Chin-Yew Lin
ACL (1)3
2016 RBPB: Regularization-Based Pattern Balancing Method for Event Extraction
abstract
Event extraction is a particularly challenging information extraction task, which intends to identify and classify event triggers and arguments from raw text.In recent works, when determining event types (trigger classification), most of the works are either pattern-only or feature-only.However, although patterns cannot cover all representations of an event, it is still a very important feature.In addition, when identifying and classifying arguments, previous works consider each candidate argument separately while ignoring the relationship between arguments.This paper proposes a Regularization-Based Pattern Balancing Method (RBPB).Inspired by the progress in representation learning, we use trigger embedding, sentence-level embedding and pattern features together as our features for trigger classification so that the effect of patterns and other useful features can be balanced.In addition, RBPB uses a regularization method to take advantage of the relationship between arguments.Experiments show that we achieve results better than current state-of-art equivalents.
Lei Sha, Jing Liu 0022, Chin-Yew Lin, Sujian Li, Baobao Chang, Zhifang Sui
ACL (1)3
2016 Knowledge Base Completion via Coupled Path Ranking
abstract
Knowledge bases (KBs) are often greatly incomplete, necessitating a demand for KB completion. The path ranking algorithm (PRA) is one of the most promising approaches to this task. Previous work on PRA usually follows a single-task learning paradigm, building a prediction model for each relation independently with its own training data. It ignores meaningful associations among certain relations, and might not get enough training data for less frequent relations. This paper proposes a novel multi-task learning framework for PRA, referred to as coupled PRA (CPRA). It first devises an agglomerative clustering strategy to automatically discover relations that are highly correlated to each other, and then employs a multi-task learning strategy to effectively couple the prediction of such relations. As such, CPRA takes into account relation association and enables implicit data sharing among them. We empirically evaluate CPRA on benchmark data created from Freebase. Experimental results show that CPRA can effectively identify coherent clusters in which relations are highly correlated. By further coupling such relations, CPRA significantly outperforms PRA, in terms of both predictive accuracy and model interpretability.
Quan Wang 0002, Jing Liu 0022, Yuanfei Luo, Bin Wang 0004, Chin-Yew Lin
ACL (1)5
2016 Cold Start Cumulative Citation Recommendation for Knowledge Base Acceleration
Jingang Wang, Jingtian Jiang, Lejian Liao, Chin-Yew Lin
ECIR6
2016 App relationship calculation: An iterative process
abstract
Today, plenty of apps are released to help users make the best use of their mobile phones. Facing the large amount of apps, app retrieval and app recommendation are extensively adopted to help users obtain their favorite apps. To acquire the high-quality retrieval or recommending results, it needs to obtain the accurate app relationship calculating results in advance. Unfortunately, recent methods are conducted mostly depending on user's log or app's contexts, which can only detect whether two apps are downloaded, installed meanwhile or provide similar functions or not. In fact, apps contain many deep relationships other than similarity, e.g., one app needs another app to cooperate to fulfill its work. Obviously, app's reviews contain user's viewpoint. They are useful to help dig deep relationship between apps. Therefore, to calculate relationship between apps via reviews, we propose an iterative process by combining review similarity calculation and app relationship calculation together.
Ming Liu 0004, Chong Wu 0001, Xiang-Nan Zhao, Chin-Yew Lin, Xiaolong Wang 0001
ICDE4
2016 Who Will Reply to/Retweet This Tweet?: The Dynamics of Intimacy from Online Social Interactions
abstract
Friendships are dynamic. Previous studies have converged to suggest that social interactions, in both online and offline social networks, are diagnostic reflections of friendship relations (also called social ties). However, most existing approaches consider a social tie as either a binary relation, or a fixed value (named tie strength). In this paper, we investigate the dynamics of dyadic friend relationships through online social interactions, in terms of a variety of aspects, such as reciprocity, temporality, and contextuality. In turn, we propose a model to predict repliers and retweeters given a particular tweet posted at a certain time in a microblog-based social network. More specifically, we have devised a learning-to-rank approach to train a ranker that considers elaborate user-level and tweet-level features (like sentiment, self-disclosure, and responsiveness) to address these dynamics. In the prediction phase, a tweet posted by a user is deemed a query and the predicted repliers/retweeters are retrieved using the learned ranker. We have collected a large dataset containing 73.3 million dyadic relationships with their interactions (replies and retweets). Extensive experimental results based on this dataset show that by incorporating the dynamics of friendship relations, our approach significantly outperforms state-of-the-art models in terms of multiple evaluation metrics, such as MAP, NDCG and Topmost Accuracy. In particular, the advantage of our model is even more promising in predicting the exact sequence of repliers/retweeters considering their orders. Furthermore, the proposed approach provides emerging implications for many high-value applications in online social networks.
Nicholas Jing Yuan, Xing Xie 0001, Chin-Yew Lin, Yong Rui
WSDM5
2016 A computational approach to measuring the correlation between expertise and social media influence for celebrities on microblogs
Wayne Xin Zhao, Jing Liu 0022, Yulan He 0001, Chin-Yew Lin, Ji-Rong Wen
World Wide Web4
2015 Context-aware Entity Morph Decoding
abstract
Boliang Zhang, Hongzhao Huang, Xiaoman Pan, Sujian Li, Chin-Yew Lin, Heng Ji, Kevin Knight, Zhen Wen, Yizhou Sun, Jiawei Han, Bulent Yener. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Boliang Zhang, Hongzhao Huang, Xiaoman Pan, Sujian Li, Chin-Yew Lin, Heng Ji 0001, Kevin Knight, Yizhou Sun, Jiawei Han 0001, Bülent Yener
ACL (1)5
2015 Improving Ranking Consistency for Web Search by Leveraging a Knowledge Base and Search Logs
abstract
In this paper, we propose a new idea called ranking consistency in web search. Relevance ranking is one of the biggest problems in creating an effective web search system. Given some queries with similar search intents, conventional approaches typically only optimize ranking models by each query separately. Hence, there are inconsistent rankings in modern search engines. It is expected that the search results of different queries with similar search intents should preserve ranking consistency. The aim of this paper is to learn consistent rankings in search results for improving the relevance ranking in web search. We then propose a re-ranking model aiming to simultaneously improve relevance ranking and ranking consistency by leveraging knowledge bases and search logs. To the best of our knowledge, our work offers the first solution to improving relevance rankings with ranking consistency. Extensive experiments have been conducted using the Freebase knowledge base and the large-scale query-log of a commercial search engine. The experimental results show that our approach significantly improves relevance ranking and ranking consistency. Two user surveys on Amazon Mechanical Turk also show that users are sensitive and prefer the consistent ranking results generated by our model.
Jyun-Yu Jiang, Jing Liu 0022, Chin-Yew Lin, Pu-Jen Cheng
CIKM3
2015 Joint Entity Recognition and Disambiguation
abstract
Extracting named entities in text and linking extracted names to a given knowledge base are fundamental tasks in applications for text understanding.Existing systems typically run a named entity recognition (NER) model to extract entity names first, then run an entity linking model to link extracted names to a knowledge base.NER and linking models are usually trained separately, and the mutual dependency between the two tasks is ignored.We propose JERL, Joint Entity Recognition and Linking, to jointly model NER and linking tasks and capture the mutual dependency between them.It allows the information from each task to improve the performance of the other.To the best of our knowledge, JERL is the first model to jointly optimize NER and linking tasks together completely.In experiments on the CoNLL'03/AIDA data set, JERL outperforms state-of-art NER and linking systems, and we find improvements of 0.4% absolute F 1 for NER on CoNLL'03, and 0.36% absolute precision@1 for linking on AIDA.
Xiaojiang Huang, Chin-Yew Lin, Zaiqing Nie
EMNLP3
2015 Automatically Solving Number Word Problems by Semantic Parsing and Reasoning
abstract
This paper presents a semantic parsing and reasoning approach to automatically solving math word problems.A new meaning representation language is designed to bridge natural language text and math expressions.A CFG parser is implemented based on 9,600 semi-automatically created grammar rules.We conduct experiments on a test set of over 1,500 number word problems (i.e., verbally expressed number problems) and yield 95.4% precision and 60.2% recall.
Shuming Shi 0001, Yuehui Wang, Chin-Yew Lin, Xiaojiang Liu, Yong Rui
EMNLP3
2015 LDTM: A Latent Document Type Model for Cumulative Citation Recommendation
abstract
This paper studies Cumulative Citation Recommendation (CCR) -given an entity in Knowledge Bases, how to effectively detect its potential citations from volume text streams.Most previous approaches treated all kinds of features indifferently to build a global relevance model, in which the prior knowledge embedded in documents cannot be exploited adequately.To address this problem, we propose a latent document type discriminative model by introducing a latent layer to capture the correlations between documents and their underlying types.The model can better adjust to different types of documents and yield flexible performance when dealing with a broad range of document types.An extensive set of experiments has been conducted on TREC-KBA-2013 dataset, and the results demonstrate that this model can yield a significant performance gain in recommendation quality as compared to the state-of-the-art.
Jingang Wang, Lejian Liao, Luo Si, Chin-Yew Lin
EMNLP6
2015 Re-Ranking Voting-Based Answers by Discarding User Behavior Biases
Xiaochi Wei, Heyan Huang, Chin-Yew Lin, Xin Xin 0001, Xianling Mao, Shangguang Wang
IJCAI3
2015 Cross-Domain Collaborative Filtering with Review Text
Xin Xin 0001, Zhirun Liu, Chin-Yew Lin, Heyan Huang, Xiaochi Wei, Ping Guo 0002
IJCAI3
2015 Why Read if You Can Scan? Trigger Scoping Strategy for Biographical Fact Extraction
abstract
The rapid growth of information sources brings a unique challenge to biographical information extraction: how to find specific facts without having to read all the words. An effective solution is to follow the human scanning strategy which keeps a specific keyword in mind and searches within a specific scope. In this paper, we mimic a scanning process to extract biographical facts. We use event and relation triggers as
Dian Yu 0001, Heng Ji 0001, Sujian Li, Chin-Yew Lin
HLT-NAACL4
2015 An Entity Class-Dependent Discriminative Mixture Model for Cumulative Citation Recommendation
abstract
This paper studies Cumulative Citation Recommendation (CCR) for Knowledge Base Acceleration (KBA). The CCR task aims to detect potential citations of a set of target entities with priorities from a volume of temporally-ordered stream corpus. Previous approaches for CCR that build an individual relevance model for each entity fail to handle unseen entities without annotation. A baseline solution is to build a global entity-unspecific model for all entities regardless of the relationship information among entities, which cannot guarantee to achieve satisfactory result for each entity. In this paper, we propose a novel entity class-dependent discriminative mixture model by introducing a latent entity class layer to model the correlations between entities and latent entity classes. The model can better adjust to different types of entities and achieve better performance when dealing with a broad range of entities. An extensive set of experiments has been conducted on TREC-KBA-2013 dataset, and the experimental results demonstrate that the proposed model can achieve the state-of-the-art performance.
Jingang Wang, Qifan Wang 0001, Luo Si, Lejian Liao, Chin-Yew Lin
SIGIR7
2015 Resorting Relevance Evidences to Cumulative Citation Recommendation for Knowledge Base Acceleration
Jingang Wang, Lejian Liao, Lerong Ma, Chin-Yew Lin, Yong Rui
WAIM5
2015 When Factorization Meets Heterogeneous Latent Topics: An Interpretable Cross-Site Recommendation Framework
Xin Xin 0001, Chin-Yew Lin, Xiaochi Wei, Heyan Huang
J. Comput. Sci. Technol.2
2015 APP Relationship Calculation: An Iterative Process
abstract
Today, plenty of apps are released to enable users to make the best use of their cell phones. Facing the large amount of apps, app retrieval and app recommendation become important, since users can easily use them to acquire their desired apps. To obtain high-quality retrieval and recommending results, it needs to obtain the precise app relationship calculating results. Unfortunately, the recent methods are conducted mostly relying on user's log or app's description, which can only detect whether two apps are downloaded, installed meanwhile or provide similar functions or not. In fact, apps contain many general relationships other than similarity, such as one app needs another app as its tool. These relationships cannot be dug via user's log or app's description. Reviews contain user's viewpoint and judgment to apps, thus they can be used to calculate relationship between apps. To use reviews, this paper proposes an iterative process by combining review similarity and app relationship together. Experimental results demonstrate that via this iterative process, relationship between apps can be calculated exactly. Furthermore, this process is improved in two aspects. One is to obtain excellent results even with weak initialization. The other is to apply matrix product to reduce running time.
Ming Liu 0004, Chong Wu 0001, Xiang-Nan Zhao, Chin-Yew Lin, Xiaolong Wang 0001
IEEE Trans. Knowl. Data Eng.4
2014 Collective Tweet Wikification based on Semi-supervised Graph Regularization
abstract
Wikification for tweets aims to automatically identify each concept mention in a tweet and link it to a concept referent in a knowledge base (e.g., Wikipedia).Due to the shortness of a tweet, a collective inference model incorporating global evidence from multiple mentions and concepts is more appropriate than a noncollecitve approach which links each mention at a time.In addition, it is challenging to generate sufficient high quality labeled data for supervised models with low cost.To tackle these challenges, we propose a novel semi-supervised graph regularization model to incorporate both local and global evidence from multiple tweets through three fine-grained relations.In order to identify semanticallyrelated mentions for collective inference, we detect meta path-based semantic relations through social networks.Compared to the state-of-the-art supervised model trained from 100% labeled data, our proposed approach achieves comparable performance with 31% labeled data and obtains 5% absolute F1 gain with 50% labeled data.Stay up Hawk Fans.We are going through a slump now, but we have to stay positive.Go Hawks!Congrats to UCONN and Kemba Walker.5 wins in 5 days, very impressive... Just getting to the Arena, we play the Bucks tonight.
Hongzhao Huang, Yunbo Cao, Xiaojiang Huang, Heng Ji 0001, Chin-Yew Lin
ACL (1)5
2014 A computational approach to measuring the correlation between expertise and social media influence for celebrities on microblogs
abstract
Existing approaches of social influence analysis usually focus on how to develop effective algorithms to quantize users' influence scores. They rarely consider a person's expertise levels which are arguably important to influence measures. In this paper, we propose a computational approach to measuring the correlation between expertise and social media influence, and we take a new perspective to understand social media influence by incorporating expertise into influence analysis. We carefully constructed a large dataset of 13,684 Chinese celebrities from Sina Weibo (literally “Sina microblogging”). We found that there is a strong correlation between expertise levels and social media influence scores. In addition, different expertise levels showed influence variation patterns: high-expertise celebrities have stronger influence on the “audience” in their expertise domains.
Wayne Xin Zhao, Jing Liu 0022, Yulan He 0001, Chin-Yew Lin, Ji-Rong Wen
ASONAM4
2014 Self-disclosure topic model for classifying and analyzing Twitter conversations
abstract
Self-disclosure, the act of revealing one-self to others, is an important social be-havior that strengthens interpersonal rela-tionships and increases social support. Al-though there are many social science stud-ies of self-disclosure, they are based on manual coding of small datasets and ques-tionnaires. We conduct a computational analysis of self-disclosure with a large dataset of naturally-occurring conversa-tions, a semi-supervised machine learning algorithm, and a computational analysis of the effects of self-disclosure on subse-quent conversations. We use a longitu-dinal dataset of 17 million tweets, all of which occurred in conversations that con-sist of five or more tweets directly reply-ing to the previous tweet, and from dyads with twenty of more conversations each. We develop self-disclosure topic model (SDTM), a variant of latent Dirichlet al-location (LDA) for automatically classi-fying the level of self-disclosure for each tweet. We take the results of SDTM and analyze the effects of self-disclosure on subsequent conversations. Our model sig-nificantly outperforms several comparable methods on classifying the level of self-disclosure, and the analysis of the longitu-dinal data using SDTM uncovers signifi-cant and positive correlation between self-disclosure and conversation frequency and length. 1
JinYeong Bak, Chin-Yew Lin, Alice Oh
EMNLP2
2014 Unsupervised Template Mining for Semantic Category Understanding
abstract
We propose an unsupervised approach to constructing templates from a large collection of semantic category names, and use the templates as the semantic representation of categories.The main challenge is that many terms have multiple meanings, resulting in a lot of wrong templates.Statistical data and semantic knowledge are extracted from a web corpus to improve template generation.A nonlinear scoring function is proposed and demonstrated to be effective.Experiments show that our approach achieves significantly better results than baseline methods.As an immediate application, we apply the extracted templates to the cleaning of a category collection and see promising results (precision improved from 81% to 89%).
Lei Shi 0015, Shuming Shi 0001, Chin-Yew Lin, Yidong Shen, Yong Rui
EMNLP3
2014 COM: a generative model for group recommendation
abstract
With the rapid development of online social networks, a growing number of people are willing to share their group activities, e.g. having dinners with colleagues, and watching movies with spouses. This motivates the studies on group recommendation, which aims to recommend items for a group of users. Group recommendation is a challenging problem because different group members have different preferences, and how to make a trade-off among their preferences for recommendation is still an open problem. In this paper, we propose a probabilistic model named COM (COnsensus Model) to model the generative process of group activities, and make group recommendations. Intuitively, users in a group may have different influences, and those who are expert in topics relevant to the group are usually more influential. In addition, users in a group may behave differently as group members from as individuals. COM is designed based on these intuitions, and is able to incorporate both users' selection history and personal considerations of content factors. When making recommendations, COM estimates the preference of a group to an item by aggregating the preferences of the group members with different weights. We conduct extensive experiments on four datasets, and the results show that the proposed model is effective in making group recommendations, and outperforms baseline methods significantly.
Quan Yuan 0001, Gao Cong, Chin-Yew Lin
KDD3
2013 A Hierarchical Entity-Based Approach to Structuralize User Generated Content in Social Media: A Case of Yahoo! Answers
abstract
Social media like forums and microblogs have accumulated a huge amount of user generated content (UGC) containing human knowledge.Currently, most of UGC is listed as a whole or in pre-defined categories.This "list-based" approach is simple, but hinders users from browsing and learning knowledge of certain topics effectively.To address this problem, we propose a hierarchical entity-based approach for structuralizing UGC in social media.By using a large-scale entity repository, we design a three-step framework to organize UGC in a novel hierarchical structure called "cluster entity tree (CET)".With Yahoo!Answers as a test case, we conduct experiments and the results show the effectiveness of our framework in constructing CET.We further evaluate the performance of CET on UGC organization in both user and system aspects.From a user aspect, our user study demonstrates that, with CET-based structure, users perform significantly better in knowledge learning than using traditional list-based approach.From a system aspect, CET substantially boosts the performance of two information retrieval models (i.e., vector space model and query likelihood language model).
Baichuan Li, Jing Liu 0022, Chin-Yew Lin, Irwin King, Michael R. Lyu
EMNLP3
2013 Question Difficulty Estimation in Community Question Answering Services
abstract
In this paper, we address the problem of estimating question difficulty in community question answering services.We propose a competition-based model for estimating question difficulty by leveraging pairwise comparisons between questions and users.Our experimental results show that our model significantly outperforms a PageRank-based approach.Most importantly, our analysis shows that the text of question descriptions reflects the question difficulty.This implies the possibility of predicting question difficulty from the text of question descriptions.
Jing Liu 0022, Quan Wang 0002, Chin-Yew Lin, Hsiao-Wuen Hon
EMNLP3
2013 Learning a Replacement Model for Query Segmentation with Consistency in Search Logs
Wei Zhang 0038, Yunbo Cao, Chin-Yew Lin, Jian Su 0002, Chew Lim Tan
IJCNLP3
2013 What's in a name?: an unsupervised approach to link users across communities
abstract
In this paper, we consider the problem of linking users across multiple online communities. Specifically, we focus on the alias-disambiguation step of this user linking task, which is meant to differentiate users with the same usernames. We start quantitatively analyzing the importance of the alias-disambiguation step by conducting a survey on 153 volunteers and an experimental analysis on a large dataset of About.me (75,472 users). The analysis shows that the alias-disambiguation solution can address a major part of the user linking problem in terms of the coverage of true pairwise decisions (46.8%). To the best of our knowledge, this is the first study on human behaviors with regards to the usages of online usernames. We then cast the alias-disambiguation step as a pairwise classification problem and propose a novel unsupervised approach. The key idea of our approach is to automatically label training instances based on two observations: (a) rare usernames are likely owned by a single natural person, e.g. pennystar88 as a positive instance; (b) common usernames are likely owned by different natural persons, e.g. tank as a negative instance. We propose using the n-gram probabilities of usernames to estimate the rareness or commonness of usernames. Moreover, these two observations are verified by using the dataset of Yahoo! Answers. The empirical evaluations on 53 forums verify: (a) the effectiveness of the classifiers with the automatically generated training data and (b) that the rareness and commonness of usernames can help user linking. We also analyze the cases where the classifiers fail.
Jing Liu 0022, Fan Zhang 0092, Xinying Song, Young-In Song, Chin-Yew Lin, Hsiao-Wuen Hon
WSDM5
2013 FoCUS: Learning to Crawl Web Forums
abstract
In this paper, we present Forum Crawler Under Supervision (FoCUS), a supervised web-scale forum crawler. The goal of FoCUS is to crawl relevant forum content from the web with minimal overhead. Forum threads contain information content that is the target of forum crawlers. Although forums have different layouts or styles and are powered by different forum software packages, they always have similar implicit navigation paths connected by specific URL types to lead users from entry pages to thread pages. Based on this observation, we reduce the web forum crawling problem to a URL-type recognition problem. And we show how to learn accurate and effective regular expression patterns of implicit navigation paths from automatically created training sets using aggregated results from weak page type classifiers. Robust page type classifiers can be trained from as few as five annotated forums and applied to a large set of unseen forums. Our test results show that FoCUS achieved over 98 percent effectiveness and 97 percent coverage on a large set of test forums powered by over 150 different forum software packages. In addition, the results of applying FoCUS on more than 100 community Question and Answer sites and Blog sites demonstrated that the concept of implicit navigation path could apply to other social media sites.
Jingtian Jiang, Xinying Song, Nenghai Yu, Chin-Yew Lin
IEEE Trans. Knowl. Data Eng.4
2013 Comparable Entity Mining from Comparative Questions
abstract
Comparing one thing with another is a typical part of human decision making process. However, it is not always easy to know what to compare and what are the alternatives. In this paper, we present a novel way to automatically mine comparable entities from comparative questions that users posted online to address this difficulty. To ensure high precision and high recall, we develop a weakly supervised bootstrapping approach for comparative question identification and comparable entity extraction by leveraging a large collection of online question archive. The experimental results show our method achieves F1-measure of 82.5 percent in comparative question identification and 83.3 percent in comparable entity extraction. Both significantly outperform an existing state-of-the-art method. Additionally, our ranking results show highly relevance to user's comparison intents in web.
Shasha Li 0001, Chin-Yew Lin, Young-In Song, Zhoujun Li 0001
IEEE Trans. Knowl. Data Eng.2
2012 An unsupervised method for author extraction from web pages containing user-generated content
abstract
In this paper, we address the problem of author extraction (AE) from user generated content (UGC) pages. Most existing solutions for web information extraction, including AE, adopt supervised approaches, which require expensive manual annotation. We propose a novel unsupervised approach for automatically collecting and labeling training data based on two key observations of author names: (1) people tend to use a single name across sites if their preferred names are available; (2) people tend to create unique usernames to easily distinguish themselves from others, e.g. travelbug61. Our AE solution only requires features extracted from a single UGC page instead of relying on clues from multiple UGC pages. We conducted extensive experiments. (1) The evaluation of automatically labeled author field data shows 95.0% precision. (2) Our method achieves an F1 score of 96.1%, which significantly outperforms a state-of-the-art supervised approach with single page features (F1 score: 68.4%) and has a comparable performance to its multiple page solution (F1 score: 95.4%). (3) We also examine the robustness of our approach on various UGC pages from forums and review sites, and achieve promising results as well.
Jing Liu 0022, Xinying Song, Jingtian Jiang, Chin-Yew Lin
CIKM4
2012 A Lazy Learning Model for Entity Linking using Query-Specific Information
Wei Zhang 0038, Jian Su 0002, Chew Lim Tan, Yunbo Cao, Chin-Yew Lin
COLING5
2012 Ensemble Semantics for Large-scale Unsupervised Relation Extraction
Bonan Min, Shuming Shi 0001, Ralph Grishman, Chin-Yew Lin
EMNLP-CoNLL4
2012 Category hierarchy maintenance: a data-driven approach
abstract
Category hierarchies often evolve at a much slower pace than the documents reside in. With newly available documents kept adding into a hierarchy, new topics emerge and documents within the same category become less topically cohesive. In this paper, we propose a novel automatic approach to modifying a given category hierarchy by redistributing its documents into more topically cohesive categories. The modification is achieved with three operations (namely, sprout, merge, and assign) with reference to an auxiliary hierarchy for additional semantic information; the auxiliary hierarchy covers a similar set of topics as the hierarchy to be modified. Our user study shows that the modified category hierarchy is semantically meaningful. As an extrinsic evaluation, we conduct experiments on document classification using real data from Yahoo! Answers and AnswerBag hierarchies, and compare the classification accuracies obtained on the original and the modified hierarchies. Our experiments show that the proposed method achieves much larger classification accuracy improvement compared with several baseline methods for hierarchy modification.
Quan Yuan 0001, Gao Cong, Aixin Sun, Chin-Yew Lin, Nadia Magnenat-Thalmann
SIGIR4
2012 Towards Large-Scale Unsupervised Relation Extraction from the Web
abstract
The Web brings an open-ended set of semantic relations. Discovering the significant types is very challenging. Unsupervised algorithms have been developed to extract relations from a corpus without knowing the relation types in advance, but most rely on tagging arguments of predefined types. One recently reported system is able to jointly extract relations and their argument semantic classes, taking a set of relation instances extracted by an open IE (Information Extraction) algorithm as input. However, it cannot handle polysemy of relation phrases and fails to group many similar (“synonymous”) relation instances because of the sparseness of features. In this paper, the authors present a novel unsupervised algorithm that provides a more general treatment of the polysemy and synonymy problems. The algorithm incorporates various knowledge sources which they will show to be very effective for unsupervised relation extraction. Moreover, it explicitly disambiguates polysemous relation phrases and groups synonymous ones. While maintaining approximately the same precision, the algorithm achieves significant improvement on recall compared to the previous method. It is also very efficient. Experiments on a real-world dataset show that it can handle 14.7 million relation instances and extract a very large set of relations from the Web.
Bonan Min, Shuming Shi 0001, Ralph Grishman, Chin-Yew Lin
Int. J. Semantic Web Inf. Syst.4
2011 Learning to Suggest Questions in Online Forums
abstract
Online forums contain interactive and semantically related discussions on various questions. Extracted question-answer archive is invaluable knowledge, which can be used to improve Question Answering services. In this paper, we address the problem of Question Suggestion, which targets at suggesting questions that are semantically related to a queried question. Existing bag-of-words approaches suffer from the shortcoming that they could not bridge the lexical chasm between semantically related questions. Therefore, we present a new framework to suggest questions, and propose the Topicenhanced Translation-based Language Model (TopicTRLM) which fuses both the lexical and latent semantic knowledge. Extensive experiments have been conducted with a large real world data set. Experimental results indicate our approach is very effective and outperforms other popular methods in several metrics.
Tom Chao Zhou, Chin-Yew Lin, Irwin King, Michael R. Lyu, Young-In Song, Yunbo Cao
AAAI2
2011 Nonlinear Evidence Fusion and Propagation for Hyponymy Relation Mining
Fan Zhang 0092, Shuming Shi 0001, Jing Liu 0022, Shu-Qi Sun, Chin-Yew Lin
ACL5
2011 Participation Maximization Based on Social Influence in Online Discussion Forums
Wei Chen 0013, Zhenming Liu, Yajun Wang 0001, Xiaorui Sun, Ming Zhang 0004, Chin-Yew Lin
ICWSM7
2011 Leveraging Unlabeled Data to Scale Blocking for Record Linkage
Yunbo Cao, Jiamin Zhu, Pei Yue, Chin-Yew Lin, Yong Yu 0001
IJCAI5
2011 Unsupervised Modeling of Dialog Acts in Asynchronous Conversations
Shafiq R. Joty, Giuseppe Carenini, Chin-Yew Lin
IJCAI3
2011 Competition-based user expertise score estimation
abstract
In this paper, we consider the problem of estimating the relative expertise score of users in community question and answering services (CQA). Previous approaches typically only utilize the explicit question answering relationship between askers and an-swerers and apply link analysis to address this problem. The im-plicit pairwise comparison between two users that is implied in the best answer selection is ignored. Given a question and answering thread, it's likely that the expertise score of the best answerer is higher than the asker's and all other non-best answerers'. The goal of this paper is to explore such pairwise comparisons inferred from best answer selections to estimate the relative expertise scores of users. Formally, we treat each pairwise comparison between two users as a two-player competition with one winner and one loser. Two competition models are proposed to estimate user expertise from pairwise comparisons. Using the NTCIR-8 CQA task data with 3 million questions and introducing answer quality prediction based evaluation metrics, the experimental results show that the pairwise comparison based competition model significantly outperforms link analysis based approaches (PageRank and HITS) and pointwise approaches (number of best answers and best answer ratio) for estimating the expertise of active users. Furthermore, it's shown that pairwise comparison based competi-tion models have better discriminative power than other methods. It's also found that answer quality (best answer) is an important factor to estimate user expertise.
Jing Liu 0022, Young-In Song, Chin-Yew Lin
SIGIR3
2011 Using graded-relevance metrics for evaluating community QA answer selection
abstract
Community Question Answering (CQA) sites such as Yahoo! Answers have emerged as rich knowledge resources for information seekers. However, answers posted to CQA sites can be irrelevant, incomplete, redundant, incorrect, biased, ill-formed or even abusive. Hence, automatic selection of "good" answers for a given posted question is a practical research problem that will help us manage the quality of accumulated knowledge. One way to evaluate answer selection systems for CQA would be to use the Best Answers (BAs) that are readily available from the CQA sites. However, BAs may be biased, and even if they are not, there may be other good answers besides BAs. To remedy these two problems, we propose system evaluation methods that involve multiple answer assessors and graded-relevance information retrieval metrics. Our main findings from experiments using the NTCIR-8 CQA task data are that, using our evaluation methods, (a) we can detect many substantial differences between systems that would have been overlooked by BA-based evaluation; and (b) we can better identify hard questions (i.e. those that are handled poorly by many systems and therefore require focussed investigation) compared to BAbased evaluation. We therefore argue that our approach is useful for building effective CQA answer selection systems despite the cost of manual answer assessments.
Tetsuya Sakai, Daisuke Ishikawa, Noriko Kando, Yohei Seki, Kazuko Kuriyama, Chin-Yew Lin
WSDM6
2011 A structural support vector method for extracting contexts and answers of questions from online forums
Yunbo Cao, Wen-Yun Yang, Chin-Yew Lin, Yong Yu 0001
Inf. Process. Manag.3
2011 Re-ranking question search results by clustering questions
abstract
In this article, we address the problem of question clustering and study its use for re-ranking question search results. In question clustering we have to organize question search results into certain meaningful and condensed groups. Specifically, we propose to use a data structure consisting of question topic and question focus for modeling questions, and then cluster questions on the basis of the data structure. Experimental results show that our approach to question clustering improves the performance of question search significantly better than the approach not utilizing the topic–focus structure.
Yunbo Cao, Huizhong Duan, Chin-Yew Lin, Yong Yu 0001
J. Assoc. Inf. Sci. Technol.3
2010 Comparable Entity Mining from Comparative Questions
Shasha Li 0001, Chin-Yew Lin, Young-In Song, Zhoujun Li 0001
ACL2
2010 Automatic extraction of web data records containing user-generated content
abstract
In this paper, we are concerned with the problem of automatically extracting web data records that contain user-generated content (UGC). In previous work, web data records are usually assumed to be well-formed with a limited amount of UGC, and thus can be extracted by testing repetitive structure similarity. However, when a web data record includes a large portion of free-format UGC, the similarity test between records may fail, which in turn results in lower performance. In our work, we find that certain domain constraints (e.g., post-date) can be used to design better similarity measures capable of circumventing the influence of UGC. In addition, we also use anchor points provided by the domain constraints to improve the extraction process, which ends in an algorithm called MiBAT (Mining data records Based on Anchor Trees). We conduct extensive experiments on a dataset consisting of forum thread pages which are collected from 307 sites that cover 219 different forum software packages. Our approach achieves a precision of 98.9% and a recall of 97.3% with respect to post record extraction. On page level, it perfectly handles 91.7% of pages without extracting any wrong posts or missing any golden posts. We also apply our approach to comment extraction and achieve good results as well.
Xinying Song, Jing Liu 0022, Yunbo Cao, Chin-Yew Lin, Hsiao-Wuen Hon
CIKM4
2009 Learning to recommend questions based on user ratings
abstract
At community question answering services, users are usually encouraged to rate questions by votes. The questions with the most votes are then recommended and ranked on the top when users browse questions by category. As users are not obligated to rate questions, usually only a small proportion of questions eventually gets rating. Thus, in this paper, we are concerned with learning to recommend questions from user ratings of a limited size. To overcome the data sparsity, we propose to utilize questions without users rating as well. Further, as there exist certain noises within user ratings (the preference of some users expressed in their ratings diverges from that of the majority of users), we design a new algorithm called 'majority-based perceptron algorithm' which can avoid the influence of noisy instances by emphasizing its learning over data instances from the majority users. Experimental results from a large collection of real questions confirm the effectiveness of our proposals.
Ke Sun 0007, Yunbo Cao, Xinying Song, Young-In Song, Xiaolong Wang 0001, Chin-Yew Lin
CIKM6
2009 Semi-supervised Speech Act Recognition in Emails and Forums
Minwoo Jeong, Chin-Yew Lin, Gary Geunbae Lee
EMNLP2
2009 A Structural Support Vector Method for Extracting Contexts and Answers of Questions from Online Forums
Wen-Yun Yang, Yunbo Cao, Chin-Yew Lin
EMNLP3
2008 Question Utility: A Novel Static Ranking of Question Search
Young-In Song, Chin-Yew Lin, Yunbo Cao, Hae-Chang Rim
AAAI2
2008 Using Conditional Random Fields to Extract Contexts and Answers of Questions from Online Forums
Shilin Ding, Gao Cong, Chin-Yew Lin, Xiaoyan Zhu 0001
ACL3
2008 Searching Questions by Identifying Question Topic and Question Focus
Huizhong Duan, Yunbo Cao, Chin-Yew Lin, Yong Yu 0001
ACL3
2008 Understanding and Summarizing Answers in Community-Based Question Answering Services
Yuanjie Liu, Yunbo Cao, Chin-Yew Lin, Dingyi Han, Yong Yu 0001
COLING4
2008 Better Binarization for the CKY Parsing
Xinying Song, Shilin Ding, Chin-Yew Lin
EMNLP3
2008 Finding question-answer pairs from online forums
abstract
Online forums contain a huge amount of valuable user generated content. In this paper we address the problem of extracting question-answer pairs from forums. Question-answer pairs extracted from forums can be used to help Question Answering services (e.g. Yahoo! Answers) among other applications. We propose a sequential patterns based classification method to detect questions in a forum thread, and a graph based propagation method to detect answers for questions in the same thread. Experimental results show that our techniques are very promising.
Gao Cong, Chin-Yew Lin, Young-In Song, Yueheng Sun
SIGIR3
2008 Recommending questions using the mdl-based tree cut model
abstract
The paper is concerned with the problem of question recommendation. Specifically, given a question as query, we are to retrieve and rank other questions according to their likelihood of being good recommendations of the queried question. A good recommendation provides alternative aspects around users' interest. We tackle the problem of question recommendation in two steps: first represent questions as graphs of topic terms, and then rank recommendations on the basis of the graphs. We formalize both steps as the tree-cutting problems and then employ the MDL (Minimum Description Length) for selecting the best cuts. Experiments have been conducted with the real questions posted at Yahoo! Answers. The questions are about two domains, 'travel' and 'computers & internet'. Experimental results indicate that the use of the MDL-based tree cut model can significantly outperform the baseline methods of word-based VSM or phrase-based VSM. The results also show that the use of the MDL-based tree cut model is essential to our approach.
Yunbo Cao, Huizhong Duan, Chin-Yew Lin, Yong Yu 0001, Hsiao-Wuen Hon
WWW3
2007 Mining Sequential Patterns and Tree Patterns to Detect Erroneous Sentences
Guihua Sun, Gao Cong, Chin-Yew Lin, Ming Zhou 0001
AAAI4
2007 Detecting Erroneous Sentences using Automatically Mined Sequential Patterns
Guihua Sun, Gao Cong, Ming Zhou 0001, Zhongyang Xiong, John Lee 0001, Chin-Yew Lin
ACL7
2007 Topic Analysis for Psychiatric Document Retrieval
Liang-Chih Yu, Chung-Hsien Wu 0001, Chin-Yew Lin, Eduard H. Hovy, Chia-Ling Lin
ACL3
2007 Low-Quality Product Review Detection in Opinion Summarization
Yunbo Cao, Chin-Yew Lin, Yalou Huang
EMNLP-CoNLL3
2007 Web page title extraction and its application
Yewei Xue, Yunhua Hu, Guomao Xin, Ruihua Song, Shuming Shi 0001, Yunbo Cao, Chin-Yew Lin, Hang Li 0001
Inf. Process. Manag.7
2006 Re-evaluating Machine Translation Results with Paraphrase Support
Chin-Yew Lin, Eduard H. Hovy
EMNLP2
2006 Automated Summarization Evaluation with Basic Elements
Eduard H. Hovy, Chin-Yew Lin, Jun-ichi Fukumoto
LREC2
2006 Summarizing Answers for Complicated Questions
Chin-Yew Lin, Eduard H. Hovy
LREC2
2006 An Information-Theoretic Approach to Automatic Evaluation of Summaries
Chin-Yew Lin, Guihong Cao, Jianfeng Gao 0001, Jian-Yun Nie
HLT-NAACL1
2006 ParaEval: Using Paraphrases to Evaluate Summaries Automatically
Chin-Yew Lin, Dragos Stefan Munteanu, Eduard H. Hovy
HLT-NAACL2
2004 Automatic Evaluation of Machine Translation Quality Using Longest Common Subsequence and Skip-Bigram Statistics
abstract
In this paper we describe two new objective automatic evaluation methods for machine translation. The first method is based on longest common subsequence between a candidate translation and a set of reference translations. Longest common subsequence takes into account sentence level structure similarity naturally and identifies longest co-occurring in-sequence n-grams automatically. The second method relaxes strict n-gram matching to skip-bigram matching. Skip-bigram is any pair of words in their sentence order. Skip-bigram cooccurrence statistics measure the overlap of skip-bigrams between a candidate translation and a set of reference translations. The empirical results show that both methods correlate with human judgments very well in both adequacy and fluency.
Chin-Yew Lin, Franz Josef Och
ACL1
2004 ORANGE: a Method for Evaluating Automatic Evaluation Metrics for Machine Translation
Chin-Yew Lin, Franz Josef Och
COLING1
2004 Introduction to the special issue on statistical language modeling
abstract
introduction Share on Introduction to the special issue on statistical language modeling Authors: Jianfeng Gao Microsoft Research Asia, Beijing, China Microsoft Research Asia, Beijing, ChinaView Profile , Chin-Yew Lin Information sciences institute, university of southern california, CA Information sciences institute, university of southern california, CAView Profile Authors Info & Claims ACM Transactions on Asian Language Information ProcessingVolume 3Issue 2June 2004 pp 87–93https://doi.org/10.1145/1034780.1034781Published:01 June 2004Publication History 4citation859DownloadsMetricsTotal Citations4Total Downloads859Last 12 Months3Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my Alerts New Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Jianfeng Gao 0001, Chin-Yew Lin
ACM Trans. Asian Lang. Inf. Process.2
2003 Automatic Evaluation of Summaries Using N-gram Co-occurrence Statistics
Chin-Yew Lin, Eduard H. Hovy
HLT-NAACL1
2003 Cross-lingual C*ST*RD: English access to Hindi information
abstract
We present C*ST*RD, a cross-language information delivery system that supports cross-language information retrieval, information space visualization and navigation, machine translation, and text summarization of single documents and clusters of documents. C*ST*RD was assembled and trained within 1 month, in the context of DARPA's Surprise Language Exercise, that selected as source a heretofore unstudied language, Hindi. Given the brief time, we could not create deep Hindi capabilities for all the modules, but instead experimented with combining shallow Hindi capabilities, or even English-only modules, into one integrated system. Various possible configurations, with different tradeoffs in processing speed and ease of use, enable the rapid deployment of C*ST*RD to new languages under various conditions.
Anton Leuski, Chin-Yew Lin, Ulrich Germann, Franz Josef Och, Eduard H. Hovy
ACM Trans. Asian Lang. Inf. Process.2
2002 From Single to Multi-document Summarization
abstract
NeATS is a multi-document summarization system that attempts to extract relevant or interesting portions from a set of documents about some topic and present them in coherent order. NeATS is among the best performers in the large scale summarization evaluation DUC 2001.
Chin-Yew Lin, Eduard H. Hovy
ACL1
2002 Using Knowledge to Facilitate Factoid Answer Pinpointing
Eduard H. Hovy, Ulf Hermjakob, Chin-Yew Lin, Deepak Ravichandran
COLING3
2002 The Effectiveness of Dictionary and Web-Based Answer Reranking
Chin-Yew Lin
COLING1
2000 The Automated Acquisition of Topic Signatures for Text Summarization
Chin-Yew Lin, Eduard H. Hovy
COLING1
1999 Training a Selection Function for Extraction
abstract
In this paper we compare performance of several heuristics in generating informative generic/query-oriented extracts for newspaper articles in order to learn how topic prominence affects the performance of each heuristic. We study how different query types can affect the performance of each heuristic and discuss the possibility of using machine learning algorithms to automatically learn good combination functions to combine several heuristics. We also briefly describe the design, implementation, and performance of a multilingual text summarization system SUMMARIST.
Chin-Yew Lin
CIKM1
1999 Machine translation for information access across the language barrier: the MuST system
abstract
In this paper we describe the design and implementation of MuST, a multilingual information retrieval, summarization, and translation system. MuST integrates machine translation and other text processing services to enable users to perform cross-language information retrieval using available search services such as commercial Internet search engines. To handle non-standard languages, a new Internet indexing agent can be deployed, specialized local search services can be built, and shallow MT can be added to provide useful functionality. A case study of augmenting MuST with Indonesian is included. MuST adopts ubiquitous web browsers as its primary user interface, and provides tightly integrated automated shallow translation and user biased summarization to help users quickly judge the relevance of documents.
Chin-Yew Lin
MTSummit1
1999 Information Access Across the Language Barrier: The MuST System (demonstration abstract)
abstract
No abstract available.
Chin-Yew Lin
SIGIR1
1995 Knowledge-Based Automatic Topic Identification
abstract
As the first step in an automated text summarization algorithm, this work presents a new method for automatically identifying the central ideas in a text based on a knowledge-based concept counting paradigm. To represent and generalize concepts, we use the hierarchical concept taxonomy WordNet. By setting appropriate cutoff values for such parameters as concept generality and child-to-parent frequency ratio, we control the amount and level of generality of concepts extracted from the text.
Chin-Yew Lin
ACL1