EDBT 2026 Demo / reviewers in the wild / expert
Hou Pong Chan
dblp:178/3691
· DBLP profile ↗
29ranked-venue papers
6as first author
24since 2021 · last 2026
0000-0001-9207-4178ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 5 first-author · 22 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VisAidMath: Benchmarking Visual-Aided Mathematical ReasoningabstractJingkun Ma, Runzhe Zhan, Yang Li, Di Sun, Hou Pong Chan, Lidia S. Chao, Derek F. Wong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jingkun Ma, Runzhe Zhan, Hou Pong Chan, Lidia S. Chao, Derek F. Wong |
ACL (1) | 5 |
| 2026 | Understanding the Behaviors of Environment-aware Information RetrievalabstractRuifeng Yuan, Chaohao Yuan, David Dai, Yu Rong, Hong Cheng, Hou Pong Chan, Chenghao Xiao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Ruifeng Yuan, Chaohao Yuan, David Dai, Yu Rong 0001, Hong Cheng 0001, Hou Pong Chan, Chenghao Xiao |
ACL (1) | 6 |
| 2025 | FineReason: Evaluating and Improving LLMs' Deliberate Reasoning through Reflective Puzzle SolvingabstractGuizhen Chen, Weiwen Xu, Hao Zhang, Hou Pong Chan, Chaoqun Liu, Lidong Bing, Deli Zhao, Anh Tuan Luu, Yu Rong. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Guizhen Chen, Weiwen Xu, Hao Zhang 0048, Hou Pong Chan, Chaoqun Liu, Lidong Bing, Deli Zhao, Anh Tuan Luu, Yu Rong 0001 |
ACL (1) | 4 |
| 2025 | Analyzing LLMs' Knowledge Boundary Cognition Across Languages Through the Lens of Internal RepresentationsabstractWhile understanding the knowledge boundaries of LLMs is crucial to prevent hallucination, research on the knowledge boundaries of LLMs has predominantly focused on English. In this work, we present the first study to analyze how LLMs recognize knowledge boundaries across different languages by probing their internal representations when processing known and unknown questions in multiple languages. Our empirical studies reveal three key findings: 1) LLMs' perceptions of knowledge boundaries are encoded in the middle to middle-upper layers across different languages. 2) Language differences in knowledge boundary perception follow a linear structure, which motivates our proposal of a training-free alignment method that effectively transfers knowledge boundary perception ability across languages, thereby helping reduce hallucination risk in low-resource languages; 3) Fine-tuning on bilingual question pair translation further enhances LLMs' recognition of knowledge boundaries across languages. Given the absence of standard testbeds for cross-lingual knowledge boundary analysis, we construct a multilingual evaluation suite comprising three representative types of knowledge boundary data. Our code and datasets are publicly available at https://github.com/DAMO-NLP-SG/ LLM-Multilingual-Knowledge-Boundaries. Chenghao Xiao, Hou Pong Chan, Hao Zhang 0048, Mahani Aljunied, Lidong Bing, Noura Al Moubayed, Yu Rong 0001 |
ACL (1) | 2 |
| 2025 | Debate-to-Write: A Persona-Driven Multi-Agent Framework for Diverse Argument GenerationabstractWriting arguments is a challenging task for both humans and machines. It entails incorporating high-level beliefs from various perspectives on the topic, along with deliberate reasoning and planning to construct a coherent narrative. Current language models often generate outputs autoregressively, lacking explicit integration of these underlying controls, resulting in limited output diversity and coherence. In this work, we propose a persona-based multi-agent framework for argument writing. Inspired by the human debate, we first assign each agent a persona representing its high-level beliefs from a unique perspective, and then design an agent interaction process so that the agents can collaboratively debate and discuss the idea to form an overall plan for argument writing. Such debate process enables fluid and nonlinear development of ideas. We evaluate our framework on argumentative essay writing. The results show that our framework generates more diverse and persuasive arguments by both automatic and human evaluations. Hou Pong Chan, Jing Li 0049, Yu Yin 0001 |
COLING | 2 |
| 2025 | ManiTweet: A New Benchmark for Identifying Manipulation of News on Social MediaabstractConsiderable advancements have been made to tackle the misrepresentation of information derived from reference articles in the domains of fact-checking and faithful summarization. However, an unaddressed aspect remains - the identification of social media posts that manipulate information within associated news articles. This task presents a significant challenge, primarily due to the prevalence of personal opinions in such posts. We present a novel task, identifying manipulation of news on social media, which aims to detect manipulation in social media posts and identify manipulated or inserted information. To study this task, we have proposed a data collection schema and curated a dataset called ManiTweet, consisting of 3.6K pairs of tweets and corresponding articles. Our analysis demonstrates that this task is highly challenging, with large language models (LLMs) yielding unsatisfactory performance. Additionally, we have developed a simple yet effective basic model that outperforms LLMs significantly on the ManiTweet dataset. Finally, we have conducted an exploratory analysis of human-written tweets, unveiling intriguing connections between manipulation and the domain and factuality of news articles, as well as revealing that manipulated sentences are more likely to encapsulate the main story or consequences of a news outlet. Kung-Hsiang Huang, Hou Pong Chan, Kathy McKeown, Heng Ji 0001 |
COLING | 2 |
| 2025 | Persona-DB: Efficient Large Language Model Personalization for Response Prediction with Collaborative Data RefinementabstractThe increasing demand for personalized interactions with large language models (LLMs) calls for methodologies capable of accurately and efficiently identifying user opinions and preferences. Retrieval augmentation emerges as an effective strategy, as it can accommodate a vast number of users without the costs from fine-tuning. Existing research, however, has largely focused on enhancing the retrieval stage and devoted limited exploration toward optimizing the representation of the database, a crucial aspect for tasks such as personalization. In this work, we examine the problem from a novel angle, focusing on how data can be better represented for more data-efficient retrieval in the context of LLM customization. To tackle this challenge, we introduce Persona-DB, a simple yet effective framework consisting of a hierarchical construction process to improve generalization across task contexts and collaborative refinement to effectively bridge knowledge gaps among users. In the evaluation of response prediction, Persona-DB demonstrates superior context efficiency in maintaining accuracy with a significantly reduced retrieval size, a critical advantage in scenarios with extensive histories or limited context windows. Our experiments also indicate a marked improvement of over 10% under cold-start scenarios, when users have extremely sparse data. Furthermore, our analysis reveals the increasing importance of collaborative knowledge as the retrieval capacity expands. Chenkai Sun, Ke Yang 0003, Revanth Gangi Reddy, Yi R. Fung 0001, Hou Pong Chan, Kevin Small, ChengXiang Zhai, Heng Ji 0001 |
COLING | 5 |
| 2025 | M-LongDoc: A Benchmark For Multimodal Super-Long Document Understanding And A Retrieval-Aware Tuning FrameworkabstractYew Ken Chia, Liying Cheng, Hou Pong Chan, Maojia Song, Chaoqun Liu, Mahani Aljunied, Soujanya Poria, Lidong Bing. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yew Ken Chia, Liying Cheng, Hou Pong Chan, Maojia Song, Chaoqun Liu, Mahani Aljunied, Soujanya Poria, Lidong Bing |
EMNLP | 3 |
| 2025 | Praxis-VLM: Vision-Grounded Decision Making via Text-Driven Reinforcement LearningabstractVision Language Models exhibit impressive performance for various tasks, yet they often lack the sophisticated situational reasoning required for complex decision-making. This paper shows that VLMs can achieve surprisingly strong decision-making performance when visual scenes are replaced by textual descriptions, suggesting foundational reasoning can be effectively learned from language. Motivated by this insight, we propose Praxis-VLM, a reasoning VLM for vision-grounded decision-making. Praxis-VLM employs the GRPO algorithm on textual scenarios to instill robust reasoning capabilities, where models learn to evaluate actions and their consequences. These reasoning skills, acquired purely from text, successfully transfer to multimodal inference with visual inputs, significantly reducing reliance on scarce paired image-text training data. Experiments across diverse decision-making benchmarks demonstrate that Praxis-VLM substantially outperforms standard supervised fine-tuning, exhibiting superior performance and generalizability. Further analysis confirms that our models engage in explicit and effective reasoning, underpinning their enhanced performance and adaptability. Jing Li 0049, Zhongzhu Pu, Hou Pong Chan, Yu Yin 0001 |
NeurIPS | 4 |
| 2025 | Scaling Language-centric Omnimodal Representation LearningabstractRecent multimodal embedding approaches leveraging multimodal large language models (MLLMs) fine-tuned with contrastive learning (CL) have shown promising results, yet the underlying reasons behind their superiority remain underexplored. This work argues that a crucial advantage of MLLM-based approaches stems from implicit cross-modal alignment achieved during generative pretraining, where the language decoder learns to exploit multimodal signals within a shared representation space for generating unimodal outputs. Through analysis of anisotropy and kernel similarity structure, we empirically confirm that latent alignment emerges within MLLM representations, allowing CL to serve as a lightweight refinement stage. Leveraging this insight, we propose a Language-Centric Omnimodal Embedding framework, termed LCO-Embed. Extensive experiments across diverse backbones and benchmarks demonstrate its effectiveness, achieving state-of-the-art performance across modalities. Furthermore, we identify a Generation-Representation Scaling Law (GRSL), showing that the representational capabilities gained through contrastive refinement scale positively with the MLLM's generative capabilities. This suggests that improving generative abilities evolves as an effective paradigm for enhancing representation quality. We provide a theoretical explanation of GRSL, which formally links the MLLM's generative quality to the upper bound on its representation performance, and validate it on a challenging, low-resource visual-document retrieval task, showing that continual generative pretraining before CL can further enhance the potential of a model's embedding capabilities. Codes, models, and resources are available at https://github.com/LCO-Embedding/LCO-Embedding. Chenghao Xiao, Hou Pong Chan, Hao Zhang 0048, Weiwen Xu, Mahani Aljunied, Yu Rong 0001 |
NeurIPS | 2 |
| 2025 | From Pixels to Insights: A Survey on Automatic Chart Understanding in the Era of Large Foundation ModelsabstractData visualization in the form of charts plays a pivotal role in data analysis, offering critical insights and aiding in informed decision-making. Automatic chart understanding has witnessed significant advancements with the rise of large foundation models in recent years. Foundation models, such as large language models, have revolutionized various natural language processing tasks and are increasingly being applied to chart understanding tasks. This survey paper provides a comprehensive overview of the recent developments, challenges, and future directions in chart understanding within the context of these foundation models. We review fundamental building blocks crucial for studying chart understanding tasks. Additionally, we explore various tasks and their evaluation metrics and sources of both charts and textual inputs. Various modeling strategies are then examined, encompassing both classification-based and generation-based approaches, along with tool augmentation techniques that enhance chart understanding performance. Furthermore, we discuss the state-of-the-art performance of each task and discuss how we can improve the performance. Challenges and future directions are addressed, highlighting the importance of several topics, such as domain-specific charts, lack of efforts in developing evaluation metrics, and agent-oriented settings. This survey paper aims to provide valuable insights and directions for future research in chart understanding leveraging large foundation models. Kung-Hsiang Huang, Hou Pong Chan, May Fung, Haoyi Qiu, Shafiq R. Joty, Shih-Fu Chang, Heng Ji 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2024 | AMERICANO: Argument Generation with Discourse-driven Decomposition and Agent InteractionabstractArgument generation is a challenging task in natural language processing, which requires rigorous reasoning and proper content organization.Inspired by recent chain-of-thought prompting that breaks down a complex task into intermediate steps, we propose AMERI-CANO, a novel framework with agent interaction for argument generation.Our approach decomposes the generation process into sequential actions grounded on argumentation theory, which first executes actions sequentially to generate argumentative discourse components, and then produces a final argument conditioned on the components.To further mimic the human writing process and improve the left-to-right generation paradigm of current autoregressive language models, we introduce an argument refinement module that automatically evaluates and refines argument drafts based on feedback received.We evaluate our framework on the task of counterargument generation using a subset of Reddit/CMV dataset.The results show that our method outperforms both end-to-end and chain-of-thought prompting methods and can generate more coherent and persuasive arguments with diverse and rich contents. Hou Pong Chan, Yu Yin 0001 |
INLG | 2 |
| 2023 | SumREN: Summarizing Reported Speech about Events in NewsabstractA primary objective of news articles is to establish the factual record for an event, frequently achieved by conveying both the details of the specified event (i.e., the 5 Ws; Who, What, Where, When and Why regarding the event) and how people reacted to it (i.e., reported statements). However, existing work on news summarization almost exclusively focuses on the event details. In this work, we propose the novel task of summarizing the reactions of different speakers, as expressed by their reported statements, to a given event. To this end, we create a new multi-document summarization benchmark, SumREN, comprising 745 summaries of reported statements from various public figures obtained from 633 news articles discussing 132 events. We propose an automatic silver-training data generation approach for our task, which helps smaller models like BART achieve GPT-3 level performance on this task. Finally, we introduce a pipeline-based framework for summarizing reported speech, which we empirically show to generate summaries that are more abstractive and factual than baseline query-focused summarization approaches. Revanth Gangi Reddy, Heba Elfardy, Hou Pong Chan, Kevin Small, Heng Ji 0001 |
AAAI | 3 |
| 2023 | Zero-shot Faithful Factual Error CorrectionabstractFaithfully correcting factual errors is critical for maintaining the integrity of textual knowledge bases and preventing hallucinations in generative models.Drawing on humans' ability to identify and correct factual errors, we present a zero-shot framework that formulates questions about input claims, looks for correct answers in the given evidence, and assesses the faithfulness of each correction based on its consistency with the evidence.Our zero-shot framework outperforms fully-supervised approaches, as demonstrated by experiments on the FEVER and SCIFACT datasets, where our outputs are shown to be more faithful.More importantly, the decomposability nature of our framework inherently provides interpretability.Additionally, to reveal the most suitable metrics for evaluating factual error corrections, we analyze the correlation between commonly used metrics with human judgments in terms of three different dimensions regarding intelligibility and faithfulness.1 Kung-Hsiang Huang, Hou Pong Chan, Heng Ji 0001 |
ACL (1) | 2 |
| 2023 | Can LMs Generalize to Future Data? An Empirical Analysis on Text SummarizationabstractRecent pre-trained language models (PLMs) achieve promising results in existing abstractive summarization datasets.However, existing summarization benchmarks overlap in time with the standard pre-training corpora and finetuning datasets.Hence, the strong performance of PLMs may rely on the parametric knowledge that is memorized during pre-training and fine-tuning.Moreover, the knowledge memorized by PLMs may quickly become outdated, which affects the generalization performance of PLMs on future data.In this work, we propose TEMPOSUM, a novel benchmark that contains data samples from 2010 to 2022, to understand the temporal generalization ability of abstractive summarization models.Through extensive human evaluation, we show that parametric knowledge stored in summarization models significantly affects the faithfulness of the generated summaries on future data.Moreover, existing faithfulness enhancement methods cannot reliably improve the faithfulness of summarization models on future data.Finally, we discuss several recommendations to the research community on how to evaluate and improve the temporal generalization capability of text summarization models. 1 Chi Seng Cheang, Hou Pong Chan, Derek F. Wong, Xuebo Liu 0002, Zhaocong Li, Shudong Liu 0004, Lidia S. Chao |
EMNLP | 2 |
| 2023 | Decoding the Silent Majority: Inducing Belief Augmented Social Graph with Large Language Model for Response ForecastingabstractAutomatic response forecasting for news media plays a crucial role in enabling content producers to efficiently predict the impact of news releases and prevent unexpected negative outcomes such as social conflict and moral injury. To effectively forecast responses, it is essential to develop measures that leverage the social dynamics and contextual information surrounding individuals, especially in cases where explicit profiles or historical actions of the users are limited (referred to as lurkers). As shown in a previous study, 97% of all tweets are produced by only the most active 25% of users. However, existing approaches have limited exploration of how to best process and utilize these important features. To address this gap, we propose a novel framework, named SOCIALSENSE, that leverages a large language model to induce a belief-centered graph on top of an existent social network, along with graph-based propagation to capture social dynamics. We hypothesize that the induced graph that bridges the gap between distant users who share similar beliefs allows the model to effectively capture the response patterns. Our method surpasses existing state-of-the-art in experimental evaluations for both zero-shot and supervised settings, demonstrating its effectiveness in response forecasting. Moreover, the analysis reveals the framework's capability to effectively handle unseen user and lurker scenarios, further highlighting its robustness and practical applicability. Chenkai Sun, Jinning Li 0001, Yi R. Fung 0001, Hou Pong Chan, Tarek F. Abdelzaher, ChengXiang Zhai, Heng Ji 0001 |
EMNLP | 4 |
| 2023 | PDSum: Prototype-driven Continuous Summarization of Evolving Multi-document Sets StreamabstractSummarizing text-rich documents has been long studied in the literature, but most of the existing efforts have been made to summarize a static and predefined multi-document set. With the rapid development of online platforms for generating and distributing text-rich documents, there arises an urgent need for continuously summarizing dynamically evolving multi-document sets where the composition of documents and sets is changing over time. This is especially challenging as the summarization should be not only effective in incorporating relevant, novel, and distinctive information from each concurrent multi-document set, but also efficient in serving online applications. In this work, we propose a new summarization problem, Evolving Multi-Document sets stream Summarization (EMDS), and introduce a novel unsupervised algorithm PDSum with the idea of prototype-driven continuous summarization. PDSum builds a lightweight prototype of each multi-document set and exploits it to adapt to new documents while preserving accumulated knowledge from previous documents. To update new summaries, the most representative sentences for each multi-document set are extracted by measuring their similarities to the prototypes. A thorough evaluation with real multi-document sets streams demonstrates that PDSum outperforms state-of-the-art unsupervised multi-document summarization algorithms in EMDS in terms of relevance, novelty, and distinctiveness and is also robust to various evaluation settings. Susik Yoon, Hou Pong Chan, Jiawei Han 0001 |
WWW | 2 |
| 2023 | Controllable Dialogue Generation With Disentangled Multi-Grained Style Specification and Attribute Consistency RewardabstractControllable text generation is an appealing but challenging task, which allows users to specify particular attributes of the generated outputs. In this paper, we propose a controllable dialogue generation model to steer response generation under multi-attribute constraints. Specifically, we define and categorize the commonly-used control attributes into global and local ones, which possess different granularities of effects on response generation. Then, we significantly extend the conventional seq2seq framework by introducing a novel two-stage decoder, which first uses amulti-grainedstyle specification layerto impose the stylistic constraints and determine word-level control states of responses based on the attributes, and then employs aresponse generation layerto generate final responses maintaining both semantic relevancy to the contexts and fidelity to the attributes. Furthermore, we train our model with an attribute consistency reward to promote response control with explicit supervision signals. Extensive experiments and in-depth analyses on two datasets indicate that our model can significantly outperform competitive baselines in terms of response quality, content diversity and controllability. Hou Pong Chan, Xinyan Xiao, Jinsong Su, Hua Wu 0003 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | PLANET: Dynamic Content Planning in Autoregressive Transformers for Long-form Text GenerationabstractDespite recent progress of pre-trained language models on generating fluent text, existing methods still suffer from incoherence problems in long-form text generation tasks that require proper content control and planning to form a coherent high-level logical flow.In this work, we propose PLANET, a novel generation framework leveraging autoregressive self-attention mechanism to conduct content planning and surface realization dynamically.To guide the generation of output sentences, our framework enriches the Transformer decoder with latent representations to maintain sentence-level semantic plans grounded by bag-of-words.Moreover, we introduce a new coherence-based contrastive learning objective to further improve the coherence of output.Extensive experiments are conducted on two challenging longform text generation tasks including counterargument generation and opinion article generation.Both automatic and human evaluations show that our method significantly outperforms strong baselines and generates more coherent texts with richer contents. Hou Pong Chan, Xinyan Xiao, Hua Wu 0003, Lifu Huang |
ACL (1) | 2 |
| 2022 | MOCHA: A Multi-Task Training Approach for Coherent Text Generation from Cognitive PerspectiveabstractTeaching neural models to generate narrative coherent texts is a critical problem.Recent pretrained language models have achieved promising results, but there is still a gap between human written texts and machine-generated outputs.In this work, we propose a novel multitask training strategy for coherent text generation grounded on the cognitive theory of writing, which empowers the model to learn essential subskills needed for writing including planning and reviewing besides end-to-end generation.We extensively evaluate our model on three open-ended generation tasks including story generation, news article writing and argument generation.Experiments show that our model achieves better results on both few-shot and fully-supervised settings than strong baselines, and human evaluations confirm that our model can generate more coherent outputs. Hou Pong Chan, Lifu Huang |
EMNLP | 2 |
| 2022 | Grounding Commands for Autonomous Vehicles via Layer Fusion with Region-specific Dynamic Layer AttentionabstractGrounding a command to the visual environment is an essential ingredient for interactions between autonomous vehicles and humans. In this work, we study the problem of language grounding for autonomous vehicles, which aims to localize a region in a visual scene according to a natural language command from a passenger. Prior work only employs the top layer representations of a vision-and-language pretrained model to predict the region referred to by the command. However, such a method omits the useful features encoded in other layers, and thus results in inadequate understanding of the input scene and command. To tackle this limitation, we present the first layer fusion approach for this task. Since different visual regions may require distinct types of features to disambiguate them from each other, we further propose the region-specific dynamic (RSD) layer attention to adaptively fuse the multimodal information across layers for each region. Extensive experiments on the Talk2Car benchmark demonstrate that our approach helps predict more accurate regions and outperforms state-of-the-art methods. Hou Pong Chan, Mingxi Guo, Cheng-Zhong Xu 0001 |
IROS | 1 |
| 2021 | A condense-then-select strategy for text summarization
Hou Pong Chan, Irwin King |
Knowl. Based Syst. | 1 |
| 2021 | Dialogue summarization with supporting utterance flow modelling and fact regularization
Wang Chen 0001, Piji Li, Hou Pong Chan, Irwin King |
Knowl. Based Syst. | 3 |
| 2021 | Controllable Summarization with Constrained Markov Decision ProcessabstractAbstract We study controllable text summarization, which allows users to gain control on a particular attribute (e.g., length limit) of the generated summaries. In this work, we propose a novel training framework based on Constrained Markov Decision Process (CMDP), which conveniently includes a reward function along with a set of constraints, to facilitate better summarization control. The reward function encourages the generation to resemble the human-written reference, while the constraints are used to explicitly prevent the generated summaries from violating user-imposed requirements. Our framework can be applied to control important attributes of summarization, including length, covered entities, and abstractiveness, as we devise specific constraints for each of these aspects. Extensive experiments on popular benchmarks show that our CMDP framework helps generate informative summaries while complying with a given attribute’s requirement.1 Hou Pong Chan, Lu Wang 0008, Irwin King |
Trans. Assoc. Comput. Linguistics | 1 |
| 2020 | Exclusive Hierarchical Decoding for Deep Keyphrase GenerationabstractKeyphrase generation (KG) aims to summarize the main ideas of a document into a set of keyphrases.A new setting is recently introduced into this problem, in which, given a document, the model needs to predict a set of keyphrases and simultaneously determine the appropriate number of keyphrases to produce.Previous work in this setting employs a sequential decoding process to generate keyphrases.However, such a decoding method ignores the intrinsic hierarchical compositionality existing in the keyphrase set of a document.Moreover, previous work tends to generate duplicated keyphrases, which wastes time and computing resources.To overcome these limitations, we propose an exclusive hierarchical decoding framework that includes a hierarchical decoding process and either a soft or a hard exclusion mechanism.The hierarchical decoding process is to explicitly model the hierarchical compositionality of a keyphrase set.Both the soft and the hard exclusion mechanisms keep track of previouslypredicted keyphrases within a window size to enhance the diversity of the generated keyphrases.Extensive experiments on multiple KG benchmark datasets demonstrate the effectiveness of our method to generate less duplicated and more accurate keyphrases 1 . Wang Chen 0001, Hou Pong Chan, Piji Li, Irwin King |
ACL | 2 |
| 2020 | A Unified Dual-view Model for Review Summarization and Sentiment Classification with Inconsistency LossabstractAcquiring accurate summarization and sentiment from user reviews is an essential component of modern e-commerce platforms. Review summarization aims at generating a concise summary that describes the key opinions and sentiment of a review, while sentiment classification aims to predict a sentiment label indicating the sentiment attitude of a review. To effectively leverage the shared sentiment information in both review summarization and sentiment classification tasks, we propose a novel dual-view model that jointly improves the performance of these two tasks. In our model, an encoder first learns a context representation for the review, then a summary decoder generates a review summary word by word. After that, a source-view sentiment classifier uses the encoded context representation to predict a sentiment label for the review, while a summary-view sentiment classifier uses the decoder hidden states to predict a sentiment label for the generated summary. During training, we introduce an inconsistency loss to penalize the disagreement between these two classifiers. It helps the decoder to generate a summary to have a consistent sentiment tendency with the review and also helps the two sentiment classifiers learn from each other. Experiment results on four real-world datasets from different domains demonstrate the effectiveness of our model. Hou Pong Chan, Wang Chen 0001, Irwin King |
SIGIR | 1 |
| 2019 | Neural Keyphrase Generation via Reinforcement Learning with Adaptive RewardsabstractGenerating keyphrases that summarize the main points of a document is a fundamental task in natural language processing.Although existing generative models are capable of predicting multiple keyphrases for an input document as well as determining the number of keyphrases to generate, they still suffer from the problem of generating too few keyphrases.To address this problem, we propose a reinforcement learning (RL) approach for keyphrase generation, with an adaptive reward function that encourages a model to generate both sufficient and accurate keyphrases.Furthermore, we introduce a new evaluation method that incorporates name variations of the ground-truth keyphrases using the Wikipedia knowledge base.Thus, our evaluation method can more robustly evaluate the quality of predicted keyphrases.Extensive experiments on five real-world datasets of different scales demonstrate that our RL approach consistently and significantly improves the performance of the state-of-the-art generative models with both conventional and new evaluation methods. Document: DCE MRI data analysis for cancer area classification.The paper aims at improving the support of medical researchers in the context of in-vivo cancer imaging… The proposed approach is based on a three-step procedure: i) robust feature extraction from raw time-intensity curves, ii) voxel segmentation, and iii) voxel classification based on a learning-by-example approach… Finally, in the third step, a support vector machine (SVM) is trained to classify voxels according to the labels obtained by the clustering phase… Keyphrase labels: svm; Hou Pong Chan, Wang Chen 0001, Lu Wang 0008, Irwin King |
ACL (1) | 1 |
| 2019 | Topic-Aware Neural Keyphrase Generation for Social Media LanguageabstractA huge volume of user-generated content is daily produced on social media.To facilitate automatic language understanding, we study keyphrase prediction, distilling salient information from massive posts.While most existing methods extract words from source posts to form keyphrases, we propose a sequence-to-sequence (seq2seq) based neural keyphrase generation framework, enabling absent keyphrases to be created.Moreover, our model, being topic-aware, allows joint modeling of corpus-level latent topic representations, which helps alleviate the data sparsity that widely exhibited in social media language.Experiments on three datasets collected from English and Chinese social media platforms show that our model significantly outperforms both extraction and generation models that do not exploit latent topics. 1 Further discussions show that our model learns meaningful topics, which interprets its superiority in social media keyphrase generation. Yue Wang 0034, Jing Li 0049, Hou Pong Chan, Irwin King, Michael R. Lyu, Shuming Shi 0001 |
ACL (1) | 3 |
| 2018 | Thread Popularity Prediction and Tracking with a Permutation-invariant ModelabstractThe task of thread popularity prediction and tracking aims to recommend a few popular comments to subscribed users when a batch of new comments arrive in a discussion thread.This task has been formulated as a reinforcement learning problem, in which the reward of the agent is the sum of positive responses received by the recommended comments.In this work, we propose a novel approach to tackle this problem.First, we propose a deep neural network architecture to model the expected cumulative reward (Q-value) of a recommendation (action).Unlike the state-ofthe-art approach, which treats an action as a sequence, our model uses an attention mechanism to integrate information from a set of comments.Thus, the prediction of Q-value is invariant to the permutation of the comments, which leads to a more consistent agent behavior.Second, we employ a greedy procedure to approximate the action that maximizes the predicted Q-value from a combinatorial action space.Different from the state-of-the-art approach, this procedure does not require an additional pre-trained model to generate candidate actions.Experiments on five real-world datasets show that our approach outperforms the state-of-the-art. Hou Pong Chan, Irwin King |
EMNLP | 1 |