EDBT 2026 Demo / reviewers in the wild / expert
Lifu Huang
dblp:127/0072
· DBLP profile ↗
76ranked-venue papers
8as first author
50since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 63 · 7 first-author · 43 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 since 2021Systems, architecture and hardware · 2Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Endowing Vision-Language Models with System 2 Thinking for Fine-grained Visual RecognitionabstractVision-Language Models (VLMs) excel at extracting salient visual features from query images, thus exhibiting promising visual recognition performance. However, VLMs would encounter significant degradation in fine-grained scenarios due to their deficiency in distinguishing nuanced differences among candidate categories. As a remedy, we draw inspiration from the ``System 1 & System 2" cognitive theory of humans, paving the way to achieve fine-grained recognition for VLMs. To be specific, we observe that VLMs naturally align with System 1, quickly identifying candidate categories but leaving easily-confused ones unresolved. Based on the observation, we propose System-2 enhanCed visuAl recogNition (SCAN), a novel plug-and-play approach that makes VLMs aware of nuanced differences. In brief, SCAN first specifies and abstracts the discriminative attributes for the confused candidate categories and query images by resorting to off-the-shelf large foundation models, respectively. After that, SCAN adaptively integrates the salient visual features from System 1 with the nuanced differences derived from System 2, resolving confusion in candidates with estimated uncertainty. Extensive experiments on eight widely used fine-grained recognition benchmarks against 10 state-of-the-art baselines verify the effectiveness and superiority of SCAN. Yutong Yang, Lifu Huang, Yijie Lin 0001, Xi Peng 0001, Mouxing Yang |
AAAI | 2 |
| 2026 | From Vulnerable to Resilient: Examining Parent and Teen Perceptions on How to Respond to Unwanted Cybergrooming AdvancesabstractCybergrooming is a form of online abuse that threatens teens’ mental health and physical safety. Yet, most prior work has focused on detecting perpetrators’ behaviors, leaving a limited understanding of how teens might respond to such unwanted advances. To address this gap, we conducted an online survey with 74 participants—51 parents and 23 teens—who responded to simulated cybergrooming scenarios in two ways: responses that they think would make teens more vulnerable or resilient to unwanted sexual advances. Through a mixed-methods analysis, we identified four types of vulnerable responses (encouraging escalation, accepting an advance, displaying vulnerability, and negating risk concern) and four types of protective strategies (setting boundaries, directly declining, signaling risk awareness, and leveraging avoidance techniques). As the cybergrooming risk escalated, both vulnerable responses and protective strategies showed a corresponding progression. This study contributes a teen-centered understanding of cybergrooming, a labeled dataset, and a stage-based taxonomy of perceived protective strategies, while offering implications for educational programs and sociotechnical interventions. Xinyi Zhang 0007, Mamtaj Akter, Heajun An, Minqian Liu, Qi Zhang 0104, Lifu Huang, Jin-Hee Cho, Pamela J. Wisniewski, Sang Won Lee 0002 |
CHI | 6 |
| 2026 | StagePilot: Stage-Level Planning for Long-Horizon Dialogue Simulation in CybergroomingabstractCybergrooming is an evolving threat to youth, requiring proactive educational interventions. We address this by modeling dialogue progression as a structured planning problem over stage-wise interactions. We propose StagePilot, a dialogue framework that separates stage-level planning from response generation, in which the model selects the next stage under constrained transitions and generates responses conditioned on it, enabling coherent and realistic progression. Reinforcement learning is used to learn stage-level policies from offline data, optimizing for both emotional alignment and goal-consistent progression. Our empirical experiments show that StagePilot generates more structured, coherent dialogue trajectories and reduces conversational stagnation compared to baselines; notably, the IQL+AWAC variant reaches the final stage more often while maintaining over 70% positive or neutral responses, yielding a 43% relative improvement. Heajun An, Qi Zhang 0104, Minqian Liu, Xinyi Zhang 0007, Sang Won Lee 0002, Lifu Huang, Pamela J. Wisniewski, Jin-Hee Cho |
SIGDIAL | 6 |
| 2025 | LLM Braces: Straightening Out LLM Predictions with Relevant Sub-UpdatesabstractRecent findings reveal that much of the knowledge in a Transformer-based Large Language Model (LLM) is encoded in its feed-forward (FFN) layers, where each FFN layer can be interpreted as the summation of sub-updates, each corresponding to a weighted column vector from the FFN's value parameter matrix that often encodes human-interpretable concepts.In light of this, we hypothesize that model performance and behaviors can be further enhanced and controlled by modulating the contributions of these sub-updates based on their relevance to the input or target output style, and propose LLMBRACES, a novel and efficient method that computes relevance scores associated with value vectors in FFN layers and leverages these scores to dynamically adjust the contribution of sub-updates.By optimizing sub-update contributions, LLMBRACES refines the prediction process, leading to more accurate and reliable outputs, much like a 'brace' providing support and stability.Moreover, LLMBRACES can be extended to support conditional control over generation characteristics, such as sentiment, thereby offering fine-grained steering of LLM outputs.Extensive experiments on various LLMs-including Qwen2.5-1.5B,Llama2-7B, and Llama3-8B-demonstrate that LLM-BRACES outperforms baseline approaches in both fine-tuning and zero-shot settings while requiring significantly fewer tunable parameters, up to 75% fewer compared to LoRA.Furthermore, LLMBRACES excels in sentimentcontrolled generation and toxicity reduction, highlighting its potential for flexible, controlled text generation across applications. Lifu Huang |
ACL (1) | 2 |
| 2025 | Inference Compute-Optimal Video Vision Language ModelsabstractPeiqi Wang, ShengYun Peng, Xuewen Zhang, Hanchao Yu, Yibo Yang, Lifu Huang, Fujun Liu, Qifan Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Shengyun Peng, Xuewen Zhang, Hanchao Yu, Lifu Huang, Fujun Liu, Qifan Wang 0001 |
ACL (1) | 6 |
| 2025 | Error-driven Data-efficient Large Multimodal Model TuningabstractLarge Multimodal Models (LMMs) have demonstrated impressive performance across numerous academic benchmarks.However, fine-tuning still remains essential to achieve satisfactory performance on downstream tasks, while the task-specific tuning samples are usually not readily available or expensive and time-consuming to obtain.To address this, we propose an error-driven data-efficient tuning framework that aims to efficiently adapt generic LMMs to newly emerging tasks without requiring extensive task-specific training samples.In our approach, a generic LMM, acting as a student model, is first evaluated on a small validation set of the target task, and then a more powerful model, acting as a teacher model, identifies the erroneous steps within the student model's reasoning steps and analyzes its capability gaps from fully addressing the target task.Based on these gaps, targeted training samples are further retrieved from existing taskagnostic datasets to tune the student model and tailor it to the target task.We perform extensive experiments across three different training data scales and seven tasks, demonstrating that our training paradigm significantly and efficiently improves LMM's performance on downstream tasks, achieving an average performance boost of 7.01% 1 . Barry Menglong Yao, Qifan Wang 0001, Lifu Huang |
ACL (1) | 3 |
| 2025 | Sycophancy Mitigation Through Reinforcement Learning with Uncertainty-Aware Adaptive Reasoning TrajectoriesabstractMohammad Beigi, Ying Shen, Parshin Shojaee, Qifan Wang, Zichao Wang, Chandan K. Reddy, Ming Jin, Lifu Huang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Mohammad Beigi, Ying Shen 0006, Parshin Shojaee, Qifan Wang 0001, Zichao Wang 0001, Chandan K. Reddy, Ming Jin 0002, Lifu Huang |
EMNLP | 8 |
| 2025 | R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generationabstractbased on instance-specific, reasoning-oriented evaluation questions that assess three critical dimensions: text-image alignment, reasoning accuracy, and image quality.Extensive experiments with 17 representative T2I models, including a strong pipeline-based framework that decouples reasoning and generation using the state-of-the-art language and image generation models, demonstrate consistently limited reasoning performance, highlighting the need for more robust, reasoning-aware architectures in the next generation of T2I systems. Kaijie Chen, Zihao Lin 0003, Zhiyang Xu, Ying Shen 0006, Yuguang Yao, Joy Rimchala, Jiaxin Zhang 0005, Lifu Huang |
EMNLP | 8 |
| 2025 | MEPT: Mixture of Expert Prompt Tuning as a Manifold MapperabstractRunjia Zeng, Guangyan Sun, Qifan Wang, Tong Geng, Sohail Dianat, Xiaotian Han, Raghuveer Rao, Xueling Zhang, Cheng Han, Lifu Huang, Dongfang Liu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Runjia Zeng, Guangyan Sun, Qifan Wang 0001, Tong Geng, Sohail A. Dianat, Raghuveer M. Rao, Xueling Zhang, Cheng Han 0001, Lifu Huang, Dongfang Liu |
EMNLP | 10 |
| 2025 | Re-Imagining Multimodal Instruction Tuning: A Representation ViewabstractMultimodal instruction tuning has proven to be an effective strategy for achieving zero-shot generalization by fine-tuning pre-trained Large Multimodal Models (LMMs) with instruction-following data. However, as the scale of LMMs continues to grow, fully fine-tuning these models has become highly parameter-intensive. Although Parameter-Efficient Fine-Tuning (PEFT) methods have been introduced to reduce the number of tunable parameters, a significant performance gap remains compared to full fine-tuning. Furthermore, existing PEFT approaches are often highly parameterized, making them difficult to interpret and control. In light of this, we introduce Multimodal Representation Tuning (MRT), a novel approach that focuses on directly editing semantically rich multimodal representations to achieve strong performance and provide intuitive control over LMMs. Empirical results show that our method surpasses current state-of-the-art baselines with significant performance gains (e.g., 1580.40 MME score) while requiring substantially fewer tunable parameters (e.g., 0.03% parameters). Additionally, we conduct experiments on editing instrumental tokens within multimodal representations, demonstrating that direct manipulation of these representations enables simple yet effective control over network behavior. Yiyang Liu 0003, James Liang, Ruixiang Tang, Yugyung Lee, Majid Rabbani, Sohail A. Dianat, Raghuveer M. Rao, Lifu Huang, Dongfang Liu, Qifan Wang 0001, Cheng Han 0001 |
ICLR | 8 |
| 2025 | SPARTUN3D: Situated Spatial Understanding of 3D World in Large Language ModelabstractIntegrating the 3D world into large language models (3D-based LLMs) has been a promising research direction for 3D scene understanding. However, current 3D-based LLMs fall short in situated understanding due to two key limitations: 1) existing 3D datasets are constructed from a global perspective of the 3D scenes and lack situated context.
2) the architectures of the current 3D-based LLMs lack an explicit mechanism for aligning situated spatial information between 3D representations and natural language, limiting their performance in tasks requiring precise spatial reasoning.
In this work, we address these issues by introducing a scalable situated 3D dataset, named Spartun3D, that incorporates various situated spatial information.
In addition, we propose a situated spatial alignment module to enhance the learning between 3D visual representations and their corresponding textual descriptions. Our experimental results demonstrate that both our dataset and alignment module enhance situated spatial understanding ability. Yue Zhang 0004, Zhiyang Xu, Ying Shen 0001, Parisa Kordjamshidi, Lifu Huang |
ICLR | 5 |
| 2025 | Modality-Specialized Synergizers for Interleaved Vision-Language GeneralistsabstractRecent advancements in Vision-Language Models (VLMs) have led to the emergence of Vision-Language Generalists (VLGs) capable of understanding and generating both text and images. However, seamlessly generating an arbitrary sequence of text and images remains a challenging task for the current VLGs. One primary limitation lies in applying a unified architecture and the same set of parameters to simultaneously model discrete text tokens and continuous image features. Recent works attempt to tackle this fundamental problem by introducing modality-aware expert models. However, they employ identical architectures to process both text and images, disregarding the intrinsic inductive biases in these two modalities. In this work, we introduce Modality-Specialized Synergizers (MoSS), a novel design that efficiently optimizes existing unified architectures of VLGs with modality-specialized adaptation layers, i.e., a Convolutional LoRA for modeling the local priors of image patches and a Linear LoRA for processing sequential text. This design enables more effective modeling of modality-specific features while maintaining the strong cross-modal integration gained from pretraining. In addition, to improve the instruction-following capability on interleaved text-and-image generation, we introduce LeafInstruct, the first open-sourced interleaved instruction tuning dataset comprising 184,982 high-quality instances on more than 10 diverse domains. Extensive experiments show that VLGs integrated with MoSS achieve state-of-the-art performance, significantly surpassing baseline VLGs in complex interleaved generation tasks. Furthermore, our method exhibits strong generalizability on different VLGs. Zhiyang Xu, Minqian Liu, Ying Shen 0006, Joy Rimchala, Jiaxin Zhang 0005, Qifan Wang 0001, Lifu Huang |
ICLR | 8 |
| 2025 | AAAR-1.0: Assessing AI's Potential to Assist ResearchabstractNumerous studies have assessed the proficiency of AI systems, particularly large language models (LLMs), in facilitating everyday tasks such as email writing, question answering, and creative content generation. However, researchers face unique challenges and opportunities in leveraging LLMs for their own work, such as brainstorming research ideas, designing experiments, and writing or reviewing papers. In this study, we introduce AAAR-1.0, a benchmark dataset designed to evaluate LLM performance in three fundamental, expertise-intensive research tasks: (i) EquationInference, assessing the correctness of equations based on the contextual information in paper submissions; (ii) ExperimentDesign, designing experiments to validate research ideas and solutions; and (iii) PaperWeakness, identifying weaknesses in paper submissions. AAAR-1.0 differs from prior benchmarks in two key ways: first, it is explicitly research-oriented, with tasks requiring deep domain expertise; second, it is researcher-oriented, mirroring the primary activities that researchers engage in on a daily basis. An evaluation of both open-source and proprietary LLMs reveals their potential as well as limitations in conducting sophisticated research tasks. We will release the AAAR-1.0 and keep iterating it to new versions. Renze Lou, Hanzi Xu, Jiangshu Du, Ryo Kamoi, Xiaoxin Lu, Yuxuan Sun 0002, Yusen Zhang 0001, Jihyun Janice Ahn, Hongchao Fang, Zhuoyang Zou, Kai Zhang 0033, Congying Xia, Lifu Huang, Wenpeng Yin 0001 |
ICML | 17 |
| 2025 | LLMs Can Reason Faster Only If We Let ThemabstractLarge language models (LLMs) are making inroads into classical AI problems such as automated planning, yet key shortcomings continue to hamper their integration. Chain-of-Thought (CoT) struggles in complex multi-step reasoning, and Tree-of-Thoughts requires multiple queries that increase computational overhead. Recently, Algorithm-of-Thoughts (AoT) have shown promise using in-context examples, at the cost of significantly longer solutions compared to CoT. Aimed at bridging the solution length gap between CoT and AoT, this paper introduces AoT-O3, which combines supervised finetuning on AoT-style plans with a reinforcement learning (RL) framework designed to reduce solution length. The RL component uses a reward model that favors concise, valid solutions while maintaining planning accuracy. Empirical evaluations indicate that AoT-O3 shortens solution length by up to 80\% compared to baseline AoT while maintaining or surpassing prior performance. These findings suggest a promising pathway for more efficient, scalable LLM-based planning. Bilgehan Sel, Lifu Huang, Naren Ramakrishnan, Ruoxi Jia 0001, Ming Jin 0002 |
ICML | 2 |
| 2025 | MetamatBench: Integrating Heterogeneous Data, Computational Tools, and Visual Interface for Metamaterial DiscoveryabstractMetamaterials, engineered materials with architected structures across multiple length scales, offer unprecedented and tunable mechanical properties that surpass those of conventional materials. However, leveraging advanced machine learning (ML) for metamaterial discovery is hindered by three fundamental challenges: (C1) Data Heterogeneity Challenge arises from heterogeneous data sources, heterogeneous composition scales, and heterogeneous structure categories; (C2) Model Complexity Challenge stems from the intricate geometric constraints of ML models, which complicate their adaptation to metamaterial structures; and (C3) Human-AI Collaboration Challenge comes from the ''dual black-box'' nature of sophisticated ML models and the need for intuitive user interfaces. To tackle these challenges, we introduce a unified framework, named MetamatBench, that operates on three levels. (1) At the data level, we integrate and standardize 5 heterogeneous, multi-modal metamaterial datasets. (2) The ML level provides a comprehensive toolkit that adapts 17 state-of-the-art ML methods for metamaterial discovery. It also includes a comprehensive evaluation suite with 12 novel performance metrics plus a finite element-based assessment to ensure accurate and reliable model validation. (3) The user level features a visual-interactive interface that bridges the gap between complex ML techniques and non-ML researchers, advancing property prediction and inverse design of metamaterials for research and applications. MetamatBench offers a unified platform that enables machine learning researchers and practitioners to develop and evaluate new methodologies in metamaterial discovery. For accessibility and reproducibility, we open-source our benchmark and the codebase at https://github.com/cjpcool/Metamaterial-Benchmark. Jianpeng Chen, Wangzhi Zhan, Haohui Wang, Zian Jia, Jingru Gan, Jingyuan Qi, Lifu Huang, Muhao Chen 0001, Wei Wang 0010, Dawei Zhou 0003 |
KDD (2) | 9 |
| 2025 | UniHGKR: Unified Instruction-aware Heterogeneous Knowledge RetrieversabstractDehai Min, Zhiyang Xu, Guilin Qi, Lifu Huang, Chenyu You. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Dehai Min, Zhiyang Xu, Guilin Qi, Lifu Huang, Chenyu You |
NAACL (Long Papers) | 4 |
| 2025 | All You Need is One: Capsule Prompt Tuning with a Single VectorabstractPrompt-based learning has emerged as a parameter-efficient finetuning (PEFT) approach to facilitate Large Language Model (LLM) adaptation to downstream tasks by conditioning generation with task-aware guidance. Despite its successes, current prompt-based learning methods heavily rely on laborious grid searching for optimal prompt length and typically require considerable number of prompts, introducing additional computational burden. Worse yet, our pioneer findings indicate that the task-aware prompt design is inherently limited by its absence of instance-aware information, leading to a subtle attention interplay with the input sequence. In contrast, simply incorporating instance-aware information as a part of the guidance can enhance the prompt-tuned model performance without additional fine-tuning. Moreover, we find an interesting phenomenon, namely "attention anchor", that incorporating instance-aware tokens at the earliest position of the sequence can successfully preserve strong attention to critical structural information and exhibit more active attention interaction with all input tokens. In light of our observation, we introduce Capsule Prompt-Tuning (CaPT), an efficient and effective solution that leverages off-the-shelf, informative instance semantics into prompt-based learning. Our approach innovatively integrates both instance-aware and task-aware information in a nearly parameter-free manner (i.e., one single capsule prompt).
Empirical results demonstrate that our method can exhibit superior performance across various language tasks (e.g., 84.03\% average accuracy on T5-Large), serving as an "attention anchor," while enjoying high parameter efficiency (e.g., 0.003\% of model parameters on Llama3.2-1B). Yiyang Liu 0003, James Liang, Heng Fan 0001, Yiming Cui 0002, Lifu Huang, Dongfang Liu, Qifan Wang 0001, Cheng Han 0001 |
NeurIPS | 7 |
| 2025 | AR-RAG: Autoregressive Retrieval Augmentation for Image GenerationabstractWe introduce Autoregressive Retrieval Augmentation (AR-RAG), a novel paradigm that enhances image generation by autoregressively incorporating k-nearest neighbor retrievals at the patch level.
Unlike prior methods that perform a single, static retrieval before generation and condition the entire generation on fixed reference images, AR-RAG performs context-aware retrievals at each generation step, using prior-generated patches as queries to retrieve and incorporate the most relevant patch-level visual references,
enabling the model to respond to evolving generation needs while avoiding limitations (e.g., over-copying, stylistic bias, etc.) prevalent in existing methods. To realize AR-RAG, we propose two parallel frameworks: (1) Distribution-Augmentation in Decoding (DAiD), a training-free plug-and-use decoding strategy that directly merges the distribution of model-predicted patches with the distribution of retrieved patches, and (2) Feature-Augmentation in Decoding (FAiD), a parameter-efficient fine-tuning method that progressively smooths the features of retrieved patches via multi-scale convolution operations and leverages them to augment the image generation process. We validate the effectiveness of AR-RAG on widely adopted benchmarks, including Midjourney-30K, GenEval and DPG-Bench, demonstrating significant performance gains over state-of-the-art image generation models. Jingyuan Qi, Zhiyang Xu, Qifan Wang 0001, Lifu Huang |
NeurIPS | 4 |
| 2025 | Probabilistic Token Alignment for Large Language Model FusionabstractTraining large language models (LLMs) from scratch can yield models with unique functionalities and strengths, but it is costly and often leads to redundant capabilities. A more cost-effective alternative is to fuse existing pre-trained LLMs with different architectures into a more powerful model. However, a key challenge in existing model fusion is their dependence on manually predefined vocabulary alignment, which may not generalize well across diverse contexts, leading to performance degradation in several evaluation. To solve this, we draw inspiration from distribution learning and propose the probabilistic token alignment method as a general and soft mapping for alignment, named as PTA-LLM. Our approach innovatively reformulates token alignment into a classic mathematical problem: optimal transport, seamlessly leveraging distribution-aware learning to facilitate more coherent model fusion. Apart from its inherent generality, PTA-LLM exhibits interpretability from a distributional perspective, offering insights into the essence of the token alignment. Empirical results demonstrate that probabilistic token alignment enhances the target model's performance across multiple capabilities. Runjia Zeng, James Liang, Cheng Han 0001, Zhiwen Cao, Xiaojun Quan, Victor Y. Chen, Lifu Huang, Tong Geng, Qifan Wang 0001, Dongfang Liu |
NeurIPS | 8 |
| 2025 | Advancing Chart Question Answering with Robust Chart Component RecognitionabstractChart comprehension presents significant challenges for machine learning models due to the diverse and intricate shapes of charts. Existing multimodal methods often over-look these visual features or fail to integrate them effectively for Chart Question Answering. To address this, we introduce CHARTFORMER, a unified framework that enhances chart component recognition by accurately identifying and classifying components such as bars, lines, pies, titles, legends, and axes. Additionally, we propose a novel Question-guided Deformable Co-Attention (QDCAt) mechanism, which fuses chart features encoded by Chart-former with the given question, leveraging the question's guidance to ground the correct answer. Extensive experiments demonstrate a 3.2% improvement in mAP over the baselines for chart component recognition. For ChartQA and OpenCQA tasks, our approach achieves improvements of 15.4% in accuracy and 0.8 in BLEU score, respectively, underscoring the robustness of our solution for detailed visual data interpretation across various applications.11The source code and dataset are publicly available at https://github.com/VT-NLP/chartQA Hanwen Zheng, Christopher Thomas 0004, Lifu Huang |
WACV | 4 |
| 2025 | Trustworthy Visual-Textual RetrievalabstractVisual-textual retrieval, as a link between computer vision and natural language processing, aims at jointly learning visual-semantic relevance to bridge the heterogeneity gap across visual and textual spaces. Existing methods conduct retrieval only relying on the ranking of pairwise similarities, but they cannot self-evaluate the uncertainty of retrieved results, resulting in unreliable retrieval and hindering interpretability. To address this problem, we propose a novel Trust-Consistent Learning framework (TCL) to endow visual-textual retrieval with uncertainty evaluation for trustworthy retrieval. More specifically, TCL first models the matching evidence according to cross-modal similarity to estimate the uncertainty for cross-modal uncertainty-aware learning. Second, a simple yet effective consistency module is presented to enforce the subjective opinions of bidirectional learning to be consistent for high reliability and accuracy. Finally, extensive experiments are conducted to demonstrate the superiority and generalizability of TCL on six widely-used benchmark datasets, i.e., Flickr30K, MS-COCO, MSVD, MSR-VTT, ActivityNet, and DiDeMo. Furthermore, some qualitative experiments are carried out to provide comprehensive and insightful analyses for trustworthy visual-textual retrieval, verifying the reliability and interoperability of TCL. The code is available in https://github.com/QinYang79/TCL. Lifu Huang, Dezhong Peng, Bohan Jiang, Joey Tianyi Zhou, Xi Peng 0001, Peng Hu 0002 |
IEEE Trans. Image Process. | 2 |
| 2024 | MULTISCRIPT: Multimodal Script Learning for Supporting Open Domain Everyday TasksabstractAutomatically generating scripts (i.e. sequences of key steps described in text) from video demonstrations and reasoning about the subsequent steps are crucial to the modern AI virtual assistants to guide humans to complete everyday tasks, especially unfamiliar ones. However, current methods for generative script learning rely heavily on well-structured preceding steps described in text and/or images or are limited to a certain domain, resulting in a disparity with real-world user scenarios. To address these limitations, we present a new benchmark challenge – MULTISCRIPT, with two new tasks on task-oriented multimodal script learning: (1) multimodal script generation, and (2) subsequent step prediction. For both tasks, the input consists of a target task name and a video illustrating what has been done to complete the target task, and the expected output is (1) a sequence of structured step descriptions in text based on the demonstration video, and (2) a single text description for the subsequent step, respectively. Built from WikiHow, MULTISCRIPT covers multimodal scripts in videos and text descriptions for over 6,655 human everyday tasks across 19 diverse domains. To establish baseline performance on MULTISCRIPT, we propose two knowledge-guided multimodal generative frameworks that incorporate the task-related knowledge prompted from large language models such as Vicuna. Experimental results show that our proposed approaches significantly improve over the competitive baselines. Jingyuan Qi, Minqian Liu, Ying Shen 0006, Zhiyang Xu, Lifu Huang |
AAAI | 5 |
| 2024 | Navigating the Dual Facets: A Comprehensive Evaluation of Sequential Memory Editing in Large Language ModelsabstractZihao Lin, Mohammad Beigi, Hongxuan Li, Yufan Zhou, Yuxiang Zhang, Qifan Wang, Wenpeng Yin, Lifu Huang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Zihao Lin 0003, Mohammad Beigi, Yufan Zhou 0001, Qifan Wang 0001, Wenpeng Yin 0001, Lifu Huang |
ACL (1) | 8 |
| 2024 | Multimodal Instruction Tuning with Conditional Mixture of LoRAabstractMultimodal Large Language Models (MLLMs) have demonstrated remarkable proficiency in diverse tasks across different domains, with an increasing focus on improving their zeroshot generalization capabilities for unseen multimodal tasks.Multimodal instruction tuning has emerged as a successful strategy for achieving zero-shot generalization by fine-tuning pretrained models on diverse multimodal tasks through instructions.As MLLMs grow in complexity and size, the need for parameterefficient fine-tuning methods like Low-Rank Adaption (LoRA), which fine-tunes with a minimal set of parameters, becomes essential.However, applying LoRA in multimodal instruction tuning presents the challenge of task interference, which leads to performance degradation, especially when dealing with a broad array of multimodal tasks.To address this, this paper introduces a novel approach that integrates multimodal instruction tuning with Conditional Mixture-of-LoRA (MixLoRA).It innovates upon LoRA by dynamically constructing low-rank adaptation matrices tailored to the unique demands of each input instance, aiming to mitigate task interference.Experimental results on various multimodal evaluation datasets indicate that MixLoRA not only outperforms the conventional LoRA with the same or even higher ranks, demonstrating its efficacy and adaptability in diverse multimodal tasks 1 . Ying Shen 0001, Zhiyang Xu, Qifan Wang 0001, Wenpeng Yin 0001, Lifu Huang |
ACL (1) | 6 |
| 2024 | Ameli: Enhancing Multimodal Entity Linking with Fine-Grained AttributesabstractBarry Yao, Sijia Wang, Yu Chen, Qifan Wang, Minqian Liu, Zhiyang Xu, Licheng Yu, Lifu Huang. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Barry Menglong Yao, Yu Chen 0022, Qifan Wang 0001, Minqian Liu, Zhiyang Xu, Licheng Yu, Lifu Huang |
EACL (1) | 8 |
| 2024 | AMD: Automatic Multi-step Distillation of Large-Scale Vision Models
Cheng Han 0001, Qifan Wang 0001, Sohail A. Dianat, Majid Rabbani, Raghuveer M. Rao, Yi Fang 0008, Qiang Guan, Lifu Huang, Dongfang Liu |
ECCV (65) | 8 |
| 2024 | DA-Code: Agent Data Science Code Generation Benchmark for Large Language ModelsabstractYiming Huang, Jianwen Luo, Yan Yu, Yitong Zhang, Fangyu Lei, Yifan Wei, Shizhu He, Lifu Huang, Xiao Liu, Jun Zhao, Kang Liu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Fangyu Lei, Yifan Wei 0001, Shizhu He, Lifu Huang, Jun Zhao 0001, Kang Liu 0001 |
EMNLP | 8 |
| 2024 | Holistic Evaluation for Interleaved Text-and-Image GenerationabstractInterleaved text-and-image generation has been an intriguing research direction, where the models are required to generate both images and text pieces in an arbitrary order.Despite the emerging advancements in interleaved generation, the progress in its evaluation still significantly lags behind.Existing evaluation benchmarks do not support arbitrarily interleaved images and text for both inputs and outputs, and they only cover a limited number of domains and use cases.Also, current works predominantly use similarity-based metrics which fall short in assessing the quality in open-ended scenarios.To this end, we introduce INTER-LEAVEDBENCH, the first benchmark carefully curated for the evaluation of interleaved textand-image generation.INTERLEAVEDBENCH features a rich array of tasks to cover diverse real-world use cases.In addition, we present INTERLEAVEDEVAL, a strong reference-free metric powered by GPT-4o to deliver accurate and explainable evaluation.We carefully define five essential evaluation aspects for IN-TERLEAVEDEVAL, including text quality, perceptual quality, image coherence, text-image coherence, and helpfulness, to ensure a comprehensive and fine-grained assessment.Through extensive experiments and rigorous human evaluation, we show that our benchmark and metric can effectively evaluate the existing models with a strong correlation with human judgments surpassing previous reference-based metrics.We also provide substantial findings and insights to foster future research in interleaved generation and its evaluation. 1 Minqian Liu, Zhiyang Xu, Zihao Lin 0003, Trevor Ashby, Joy Rimchala, Jiaxin Zhang 0005, Lifu Huang |
EMNLP | 7 |
| 2024 | M²PT: Multimodal Prompt Tuning for Zero-shot Instruction LearningabstractTaowen Wang, Yiyang Liu, James Chenhao Liang, Junhan Zhao, Yiming Cui, Yuning Mao, Shaoliang Nie, Jiahao Liu, Fuli Feng, Zenglin Xu, Cheng Han, Lifu Huang, Qifan Wang, Dongfang Liu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Taowen Wang, Yiyang Liu 0003, James Liang, Junhan Zhao, Yiming Cui 0002, Yuning Mao, Shaoliang Nie, Fuli Feng, Zenglin Xu, Cheng Han 0001, Lifu Huang, Qifan Wang 0001, Dongfang Liu |
EMNLP | 12 |
| 2024 | Facing the Elephant in the Room: Visual Prompt Tuning or Full finetuning?abstractAs the scale of vision models continues to grow, the emergence of Visual Prompt Tuning (VPT) as a parameter-efficient transfer learning technique has gained attention due to its superior performance compared to traditional full-finetuning. However, the conditions favoring VPT (the "when") and the underlying rationale (the "why") remain unclear. In this paper, we conduct a comprehensive analysis across 19 distinct datasets and tasks. To understand the "when" aspect, we identify the scenarios where VPT proves favorable by two dimensions: task objectives and data distributions. We find that VPT is preferrable when there is 1) a substantial disparity between the original and the downstream task objectives ($e.g.$, transitioning from classification to counting), or 2) a notable similarity in data distributions between the two tasks ($e.g.$, both involve natural images). In exploring the "why" dimension, our results indicate VPT's success cannot be attributed solely to overfitting and optimization considerations. The unique way VPT preserves original features and adds parameters appears to be a pivotal factor. Our study provides insights into VPT's mechanisms, and offers guidance for its optimal utilization. Cheng Han 0001, Qifan Wang 0001, Yiming Cui 0002, Wenguan Wang, Lifu Huang, Siyuan Qi, Dongfang Liu |
ICLR | 5 |
| 2024 | Position: TrustLLM: Trustworthiness in Large Language ModelsabstractLarge language models (LLMs) have gained considerable attention for their excellent natural language processing capabilities. Nonetheless, these LLMs present many challenges, particularly in the realm of trustworthiness. This paper introduces TrustLLM, a comprehensive study of trustworthiness in LLMs, including principles for different dimensions of trustworthiness, established benchmark, evaluation, and analysis of trustworthiness for mainstream LLMs, and discussion of open challenges and future directions. Specifically, we first propose a set of principles for trustworthy LLMs that span eight different dimensions. Based on these principles, we further establish a benchmark across six dimensions including truthfulness, safety, fairness, robustness, privacy, and machine ethics. We then present a study evaluating 16 mainstream LLMs in TrustLLM, consisting of over 30 datasets. Our findings firstly show that in general trustworthiness and capability (i.e., functional effectiveness) are positively related. Secondly, our observations reveal that proprietary LLMs generally outperform most open-source counterparts in terms of trustworthiness, raising concerns about the potential risks of widely accessible open-source LLMs. However, a few open-source LLMs come very close to proprietary ones, suggesting that open-source models can achieve high levels of trustworthiness without additional mechanisms like moderator, offering valuable insights for developers in this field. Thirdly, it is important to note that some LLMs may be overly calibrated towards exhibiting trustworthiness, to the extent that they compromise their utility by mistakenly treating benign prompts as harmful and consequently not responding. Besides these observations, we’ve uncovered key insights into the multifaceted trustworthiness in LLMs. We emphasize the importance of ensuring transparency not only in the models themselves but also in the technologies that underpin trustworthiness. We advocate that the establishment of an AI alliance between industry, academia, the open-source community to foster collaboration is imperative to advance the trustworthiness of LLMs. Yue Huang 0001, Lichao Sun 0001, Haoran Wang 0005, Siyuan Wu 0001, Qihui Zhang, Chujie Gao, Wenhan Lyu, Yixuan Zhang 0001, Xiner Li, Hanchi Sun, Zhengliang Liu, Yixin Liu 0002, Yijue Wang, Bertie Vidgen, Bhavya Kailkhura, Caiming Xiong, Chaowei Xiao, Chunyuan Li, Eric P. Xing, Furong Huang, Heng Ji 0001, Hongyi Wang 0001, Huan Zhang 0001, Huaxiu Yao, Manolis Kellis, Marinka Zitnik, Meng Jiang 0001, Mohit Bansal, James Zou 0001, Jian Pei 0001, Jianfeng Gao 0001, Jiawei Han 0001, Jieyu Zhao 0001, Jiliang Tang, Jindong Wang 0001, Joaquin Vanschoren, John C. Mitchell, Kai Shu, Kaidi Xu, Kai-Wei Chang 0001, Lifang He 0001, Lifu Huang, Michael Backes 0001, Neil Zhenqiang Gong, Philip S. Yu, Quanquan Gu, Ran Xu 0001, Rex Ying, Shuiwang Ji, Suman Jana, Tianlong Chen 0001, Tianming Liu 0001, Tianyi Zhou 0001, William Yang Wang, Xiang Li 0001, Xiangliang Zhang 0001, Xiao Wang 0012, Xing Xie 0001, Xuyu Wang, Yan Liu 0002, Yanfang Ye 0001, Yinzhi Cao, Yong Chen 0016, Yue Zhao 0016 |
ICML | 47 |
| 2024 | Towards Effective Long Conversation Generation with Dynamic Topic Tracking and RecommendationabstractDuring conversations, the human flow of thoughts may result in topic shifts and evolution.In open-domain dialogue systems, it is crucial to track the topics discussed and recommend relevant topics to be included in responses to have effective conversations.Furthermore, topic evolution is needed to prevent stagnation as conversation length increases.Existing open-domain dialogue systems do not pay sufficient attention to topic evolution and shifting, resulting in performance degradation due to ineffective responses as conversation length increases.To address the shortcomings of existing approaches, we propose EVOLV-CONV.EVOLVCONV conducts real-time conversation topic and user preference tracking and utilizes the tracking information to evolve and shift topics depending on conversation status.We conduct extensive experiments to validate the topic evolving and shifting capabilities of EVOLVCONV as conversation length increases.Un-referenced evaluation metric UniEval compare EVOLVCONV with the baselines.Experimental results show that EVOLV-CONV maintains a smooth conversation flow without abruptly shifting topics; the probability of topic shifting ranges between 5%-8% throughout the conversation.EVOLVCONV recommends 4.77% more novel topics than the baselines, and the topic evolution follows balanced topic groupings.Furthermore, we conduct user surveys to test the practical viability of EVOLVCONV.User survey results reveal that responses generated by EVOLVCONV are preferred 47.8% of the time compared to the baselines and comes second to real human responses. Trevor Ashby, Adithya Kulkarni, Jingyuan Qi, Minqian Liu, Eunah Cho, Vaibhav Kumar, Lifu Huang |
INLG | 7 |
| 2024 | X-Eval: Generalizable Multi-aspect Text Evaluation via Augmented Instruction Tuning with Auxiliary Evaluation AspectsabstractMinqian Liu, Ying Shen, Zhiyang Xu, Yixin Cao, Eunah Cho, Vaibhav Kumar, Reza Ghanadan, Lifu Huang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Minqian Liu, Ying Shen 0006, Zhiyang Xu, Yixin Cao 0002, Eunah Cho, Vaibhav Kumar, Reza Ghanadan, Lifu Huang |
NAACL-HLT | 8 |
| 2024 | RE²: Region-Aware Relation Extraction from Visually Rich DocumentsabstractPritika Ramu, Sijia Wang, Lalla Mouatadid, Joy Rimchala, Lifu Huang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Pritika Ramu, Lalla Mouatadid, Joy Rimchala, Lifu Huang |
NAACL-HLT | 5 |
| 2024 | Visual Fourier Prompt TuningabstractWith the scale of vision Transformer-based models continuing to grow, finetuning these large-scale pretrained models for new tasks has become increasingly parameter-intensive. Visual prompt tuning is introduced as a parameter-efficient finetuning (PEFT) method to this trend. Despite its successes, a notable research challenge persists within almost all PEFT approaches: significant performance degradation is observed when there is a substantial disparity between the datasets applied in pretraining and finetuning phases. To address this challenge, we draw inspiration from human visual cognition, and propose the Visual Fourier Prompt Tuning (VFPT) method as a general and effective solution for adapting large-scale transformer-based models. Our approach innovatively incorporates the Fast Fourier Transform into prompt embeddings and harmoniously considers both spatial and frequency domain information. Apart from its inherent simplicity and intuitiveness, VFPT exhibits superior performance across all datasets, offering a general solution to dataset challenges, irrespective of data disparities. Empirical results demonstrate that our approach outperforms current state-of-the-art baselines on two benchmarks, with low parameter usage (e.g., 0.57% of model parameters on VTAB-1k) and notable performance enhancements (e.g., 73.20% of mean accuracy on VTAB-1k). Our code is avaliable at https://github.com/runtsang/VFPT. Runjia Zeng, Cheng Han 0001, Qifan Wang 0001, Chunshu Wu, Tong Geng, Lifu Huang, Ying Nian Wu, Dongfang Liu |
NeurIPS | 6 |
| 2024 | Trust it or not: Confidence-guided automatic radiology report generation
Yixin Wang 0003, Zihao Lin 0003, Zhe Xu 0012, Jie Luo 0003, Jiang Tian, Zhongchao Shi, Lifu Huang, Yang Zhang 0002, Jianping Fan 0007, Zhiqiang He 0002 |
Neurocomputing | 8 |
| 2023 | MultiInstruct: Improving Multi-Modal Zero-Shot Learning via Instruction TuningabstractInstruction tuning, a new learning paradigm that fine-tunes pre-trained language models on tasks specified through instructions, has shown promising zero-shot performance on various natural language processing tasks.However, it has yet to be explored for vision and multimodal tasks.In this work, we introduce MUL-TIINSTRUCT, the first multimodal instruction tuning benchmark dataset that consists of 62 diverse multimodal tasks in a unified seq-toseq format covering 10 broad categories.The tasks are derived from 21 existing open-source datasets and each task is equipped with 5 expertwritten instructions.We take OFA (Wang et al., 2022a) as the base pre-trained model for multimodal instruction tuning, and to further improve its zero-shot performance, we explore multiple transfer learning strategies to leverage the large-scale NATURAL INSTRUCTIONS dataset (Mishra et al., 2022).Experimental results demonstrate strong zero-shot performance on various unseen multimodal tasks and the benefit of transfer learning from a text-only instruction dataset.We also design a new evaluation metric -Sensitivity, to evaluate how sensitive the model is to the variety of instructions.Our results indicate that fine-tuning the model on a diverse set of tasks and instructions leads to a reduced sensitivity to variations in instructions for each task 1 . Zhiyang Xu, Ying Shen 0006, Lifu Huang |
ACL (1) | 3 |
| 2023 | The Art of SOCRATIC QUESTIONING: Recursive Thinking with Large Language ModelsabstractChain-of-Thought (CoT) prompting enables large language models to solve complex reasoning problems by generating intermediate steps.However, confined by its inherent singlepass and sequential generation process, CoT heavily relies on the initial decisions, causing errors in early steps to accumulate and impact the final answers.In contrast, humans adopt recursive thinking when tackling complex reasoning problems, i.e., iteratively breaking the original problem into approachable subproblems and aggregating their answers to resolve the original one.Inspired by the human cognitive process, we propose SOCRATIC QUESTIONING, a divide-and-conquer style algorithm that mimics the recursive thinking process.Specifically, SOCRATIC QUESTIONING leverages large language models to raise and answer sub-questions until collecting enough information to tackle the original question.Unlike CoT, SOCRATIC QUESTIONING explicitly navigates the thinking space, stimulates effective recursive thinking, and is more robust towards errors in the thinking process.Extensive experiments on several complex reasoning tasks, including MMLU, MATH, LogiQA, and visual question-answering demonstrate significant performance improvements over the stateof-the-art prompting methods, such as CoT, and Tree-of-Thought.The qualitative analysis clearly shows that the intermediate reasoning steps elicited by SOCRATIC QUESTIONING are similar to humans' recursively thinking process of complex reasoning problems 12 . Jingyuan Qi, Zhiyang Xu, Ying Shen 0006, Minqian Liu, Qifan Wang 0001, Lifu Huang |
EMNLP | 7 |
| 2023 | APrompt: Attention Prompt Tuning for Efficient Adaptation of Pre-trained Language ModelsabstractQifan Wang, Yuning Mao, Jingang Wang, Hanchao Yu, Shaoliang Nie, Sinong Wang, Fuli Feng, Lifu Huang, Xiaojun Quan, Zenglin Xu, Dongfang Liu. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Qifan Wang 0001, Yuning Mao, Jingang Wang, Hanchao Yu, Shaoliang Nie, Sinong Wang, Fuli Feng, Lifu Huang, Xiaojun Quan, Zenglin Xu, Dongfang Liu |
EMNLP | 8 |
| 2023 | Authentic Dialogue Generation to Improve Youth's Awareness of Cybergrooming for Online SafetyabstractThis paper deals with a cybergrooming and sexual misconduct topic in artificial intelligence-based educational programs. Although cybergrooming has been recognized as a cybercrime, there is a lack of programs to protect youth from cybergrooming. We present a generative chatbot framework, SERI (Stop cybERgroomIng), that can generate fluent authentic conversations in the context of cybergrooming between a perpetrator chatbot and a potential victim chatbot. Furthermore, we propose deep-reinforcement-learning-based dialogue generation with a stage-related reward to lead the conversation to the expected stage. We also minimize potential ethical issues introduced by the perverted languages when deploying the chatbots for cybersecurity education programs. We evaluated the conversations of SERI with open-source referenced, unreferenced metrics and human evaluation. We developed SERI as a platform for deploying perpetrator chatbot to interact with youth users to observe their responses and collect reactions when they are asked for private or sensitive information by the perpetrator. Zhen Guo 0002, Lifu Huang, Jin-Hee Cho |
ICTAI | 3 |
| 2023 | Understand the Dynamic World: An End-to-End Knowledge Informed Framework for Open Domain Entity State TrackingabstractOpen domain entity state tracking aims to predict reasonable state changes of entities (i.e., [attribute] of [entity] was [before_state] and [after_state] afterwards) given the action descriptions. It's important to many reasoning tasks to support human everyday activities. However, it's challenging as the model needs to predict an arbitrary number of entity state changes caused by the action while most of the entities are implicitly relevant to the actions and their attributes as well as states are from open vocabularies. To tackle these challenges, we propose a novel end-to-end Knowledge Informed framework for open domain Entity State Tracking, namely KIEST, which explicitly retrieves the relevant entities and attributes from external knowledge graph (i.e., ConceptNet) and incorporates them to autoregressively generate all the entity state changes with a novel dynamic knowledge grained encoder-decoder framework. To enforce the logical coherence among the predicted entities, attributes, and states, we design a new constraint decoding strategy and employ a coherence reward to improve the decoding process. Experimental results show that our proposed KIEST framework significantly outperforms the strong baselines on the public benchmark dataset - OpenPI Lifu Huang |
SIGIR | 2 |
| 2023 | End-to-End Multimodal Fact-Checking and Explanation Generation: A Challenging Dataset and ModelsabstractWe propose end-to-end multimodal fact-checking and explanation generation, where the input is a claim and a large collection of web sources, including articles, images, videos, and tweets, and the goal is to assess the truthfulness of the claim by retrieving relevant evidence and predicting a truthfulness label (e.g., support, refute or not enough information), and to generate a statement to summarize and explain the reasoning and ruling process. To support this research, we construct MOCHEG, a large-scale dataset consisting of 15,601 claims where each claim is annotated with a truthfulness label and a ruling statement, and 33,880 textual paragraphs and 12,112 images in total as evidence. To establish baseline performances on MOCHEG, we experiment with several state-of-the-art neural architectures on the three pipelined subtasks: multimodal evidence retrieval, claim verification, and explanation generation, and demonstrate that the performance of the state-of-the-art end-to-end multimodal fact-checking does not provide satisfactory outcomes. To the best of our knowledge, we are the first to build the benchmark dataset and solutions for end-to-end multimodal fact-checking and explanation generation. The dataset, source code and model checkpoints are available at https://github.com/VT-NLP/Mocheg. Barry Menglong Yao, Aditya Shah, Lichao Sun 0001, Jin-Hee Cho, Lifu Huang |
SIGIR | 5 |
| 2022 | MuMuQA: Multimedia Multi-Hop News Question Answering via Cross-Media Knowledge Extraction and GroundingabstractRecently, there has been an increasing interest in building question answering (QA) models that reason across multiple modalities, such as text and images. However, QA using images is often limited to just picking the answer from a pre-defined set of options. In addition, images in the real world, especially in news, have objects that are co-referential to the text, with complementary information from both modalities. In this paper, we present a new QA evaluation benchmark with 1,384 questions over news articles that require cross-media grounding of objects in images onto text. Specifically, the task involves multi-hop questions that require reasoning over image-caption pairs to identify the grounded visual object being referred to and then predicting a span from the news body text to answer the question. In addition, we introduce a novel multimedia data augmentation framework, based on cross-media knowledge extraction and synthetic question-answer generation, to automatically augment data that can provide weak supervision for this task. We evaluate both pipeline-based and end-to-end pretraining-based multimedia QA models on our benchmark, and show that they achieve promising performance, while considerably lagging behind human performance hence leaving large room for future work on this challenging new task. Revanth Gangi Reddy, Xilin Rui, Manling Li, Xudong Lin 0003, Haoyang Wen, Jaemin Cho 0001, Lifu Huang, Mohit Bansal, Avirup Sil, Shih-Fu Chang, Alexander G. Schwing, Heng Ji 0001 |
AAAI | 7 |
| 2022 | PLANET: Dynamic Content Planning in Autoregressive Transformers for Long-form Text GenerationabstractDespite recent progress of pre-trained language models on generating fluent text, existing methods still suffer from incoherence problems in long-form text generation tasks that require proper content control and planning to form a coherent high-level logical flow.In this work, we propose PLANET, a novel generation framework leveraging autoregressive self-attention mechanism to conduct content planning and surface realization dynamically.To guide the generation of output sentences, our framework enriches the Transformer decoder with latent representations to maintain sentence-level semantic plans grounded by bag-of-words.Moreover, we introduce a new coherence-based contrastive learning objective to further improve the coherence of output.Extensive experiments are conducted on two challenging longform text generation tasks including counterargument generation and opinion article generation.Both automatic and human evaluations show that our method significantly outperforms strong baselines and generates more coherent texts with richer contents. Hou Pong Chan, Xinyan Xiao, Hua Wu 0003, Lifu Huang |
ACL (1) | 6 |
| 2022 | Incremental Prompting: Episodic Memory Prompt for Lifelong Event DetectionabstractLifelong event detection aims to incrementally update a model with new event types and data while retaining the capability on previously learned old types. One critical challenge is that the model would catastrophically forget old types when continually trained on new data. In this paper, we introduce Episodic Memory Prompts (EMP) to explicitly retain the learned task-specific knowledge. Our method adopts continuous prompt for each task and they are optimized to instruct the model prediction and learn event-specific representation. The EMPs learned in previous tasks are carried along with the model in subsequent tasks, and can serve as a memory module that keeps the old knowledge and transferring to new tasks. Experiment results demonstrate the effectiveness of our method. Furthermore, we also conduct a comprehensive analysis of the new and old event types in lifelong learning. Minqian Liu, Shiyu Chang, Lifu Huang |
COLING | 3 |
| 2022 | MOCHA: A Multi-Task Training Approach for Coherent Text Generation from Cognitive PerspectiveabstractTeaching neural models to generate narrative coherent texts is a critical problem.Recent pretrained language models have achieved promising results, but there is still a gap between human written texts and machine-generated outputs.In this work, we propose a novel multitask training strategy for coherent text generation grounded on the cognitive theory of writing, which empowers the model to learn essential subskills needed for writing including planning and reviewing besides end-to-end generation.We extensively evaluate our model on three open-ended generation tasks including story generation, news article writing and argument generation.Experiments show that our model achieves better results on both few-shot and fully-supervised settings than strong baselines, and human evaluations confirm that our model can generate more coherent outputs. Hou Pong Chan, Lifu Huang |
EMNLP | 3 |
| 2022 | What Makes the Story Forward?: Inferring Commonsense Explanations as Prompts for Future Event GenerationabstractPrediction over event sequences is critical for many real-world applications in Information Retrieval and Natural Language Processing. Future Event Generation (FEG) is a challenging task in event sequence prediction because it requires not only fluent text generation but also commonsense reasoning to maintain the logical coherence of the entire event story. In this paper, we propose a novel explainable FEG framework, Coep. It highlights and integrates two types of event knowledge, sequential knowledge of direct event-event relations and inferential knowledge that reflects the intermediate character psychology between events, such as intents, causes, reactions, which intrinsically pushes the story forward. To alleviate the knowledge forgetting issue, we design two modules, IM and GM, for each type of knowledge, which are combined via prompt tuning. First, IM focuses on understanding inferential knowledge to generate commonsense explanations and provide a soft prompt vector for GM. We also design a contrastive discriminator for better generalization ability. Second, GM generates future events by modeling direct sequential knowledge with the guidance of IM. Automatic and human evaluation demonstrate that our approach can generate more coherent, specific, and logical future events. Li Lin 0011, Yixin Cao 0002, Lifu Huang, Shuang Li 0015, Xuming Hu, Lijie Wen 0001, Jianmin Wang 0001 |
SIGIR | 3 |
| 2021 | How Knowledge Graph and Attention Help? A Qualitative Analysis into Bag-level Relation ExtractionabstractZikun Hu, Yixin Cao, Lifu Huang, Tat-Seng Chua. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Zikun Hu, Yixin Cao 0002, Lifu Huang, Tat-Seng Chua |
ACL/IJCNLP (1) | 3 |
| 2021 | The Future is not One-dimensional: Complex Event Schema Induction by Graph Modeling for Event PredictionabstractEvent schemas encode knowledge of stereotypical structures of events and their connections.As events unfold, schemas are crucial to act as a scaffolding.Previous work on event schema induction focuses either on atomic events or linear temporal event sequences, ignoring the interplay between events via arguments and argument relations.We introduce a new concept of Temporal Complex Event Schema: a graph-based schema representation that encompasses events, arguments, temporal connections and argument relations.In addition, we propose a Temporal Event Graph Model that predicts event instances following the temporal complex event schema.To build and evaluate such schemas, we release a new schema learning corpus containing 6,399 documents accompanied with event graphs, and we have manually constructed gold-standard schemas.Intrinsic evaluations by schema matching and instance graph perplexity, prove the superior quality of our probabilistic graph schema library compared to linear representations.Extrinsic evaluation on schema-guided future event prediction further demonstrates the predictive power of our event graph model, significantly outperforming human schemas and baselines by more than 23.8% on HITS@1. 1 Manling Li, Zhenhailong Wang, Lifu Huang, Kyunghyun Cho, Heng Ji 0001, Jiawei Han 0001, Clare R. Voss |
EMNLP (1) | 4 |
| 2021 | Expertise-Aware Truth Analysis and Task Allocation in Mobile CrowdsourcingabstractIn mobile crowdsourcing, the accuracy of the collected data is usually hard to ensure. Researchers have proposed techniques to identify truth from noisy data by inferring and utilizing the reliability of mobile users, and allocate tasks to users with higher reliability. However, they neglect the fact that a user may only have expertise on some problems (in some domains), but not others, and hence causing two problems: low estimation accuracy in truth analysis and ineffective task allocation. To address these problems, we propose Expertise-aware Truth Analysis and Task Allocation (ETA2), which can effectively infer user expertise, and then estimate truth and allocate tasks based on the inferred expertise. ETA2relies on a novel semantic analysis method to identify the expertise, and an expertise-aware truth analysis method to find the truth. For expertise-aware task allocation in ETA2, we formalize and solve two problems based on the optimization objectives: max-qualitytask allocation which maximizes the probability fortasks to be allocated to users with high expertise and min-costtask allocation which minimizes the cost of task allocation while ensuring high-quality data are collected. Experimental results based on two real-world datasets and one synthetic dataset demonstrate that ETA2significantly outperforms existing solutions. Xiaomei Zhang 0001, Lifu Huang, Heng Ji 0001, Guohong Cao |
IEEE Trans. Mob. Comput. | 3 |
| 2020 | Semi-supervised New Event Type Induction and Event DetectionabstractMost previous event extraction studies assume a set of target event types and corresponding event annotations are given, which could be very expensive.In this paper, we work on a new task of semi-supervised event type induction, aiming to automatically discover a set of unseen types from a given corpus by leveraging annotations available for a few seen types.We design a Semi-Supervised Vector Quantized Variational Autoencoder framework to automatically learn a discrete latent type representation for each seen and unseen type and optimize them using seen type event annotations.A variational autoencoder is further introduced to enforce the reconstruction of each event mention conditioned on its latent type distribution.Experiments show that our approach can not only achieve state-of-the-art performance on supervised event detection but also discover high-quality new event types. 1 Lifu Huang, Heng Ji 0001 |
EMNLP (1) | 1 |
| 2020 | ReviewRobot: Explainable Paper Review Generation based on Knowledge SynthesisabstractTo assist human review process, we build a novel ReviewRobot to automatically assign a review score and write comments for multiple categories such as novelty and meaningful comparison.A good review needs to be knowledgeable, namely that the comments should be constructive and informative to help improve the paper; and explainable by providing detailed evidence.ReviewRobot achieves these goals via three steps: (1) We perform domainspecific Information Extraction to construct a knowledge graph (KG) from the target paper under review, a related work KG from the papers cited by the target paper, and a background KG from a large collection of previous papers in the domain.(2) By comparing these three KGs, we predict a review score and detailed structured knowledge as evidence for each review category.(3) We carefully select and generalize human review sentences into templates, and apply these templates to transform the review scores and evidence into natural language comments.Experimental results show that our review score predictor reaches 71.4%-100% accuracy.Human assessment by domain experts shows that 41.7%-70.5% of the comments generated by ReviewRobot are valid and constructive, and better than humanwritten ones for 20% of the time.Thus, Re-viewRobot can serve as an assistant for paper reviewers, program chairs and authors. 1 Qingyun Wang 0005, Qi Zeng 0001, Lifu Huang, Kevin Knight, Heng Ji 0001, Nazneen Fatema Rajani |
INLG | 3 |
| 2019 | PaperRobot: Incremental Draft Generation of Scientific IdeasabstractWe present a PaperRobot who performs as an automatic research assistant by (1) conducting deep understanding of a large collection of human-written papers in a target domain and constructing comprehensive background knowledge graphs (KGs); (2) creating new ideas by predicting links from the background KGs, by combining graph attention and contextual text attention; (3) incrementally writing some key elements of a new paper based on memory-attention networks: from the input title along with predicted related entities to generate a paper abstract, from the abstract to generate conclusion and future work, and finally from future work to generate a title for a follow-on paper.Turing Tests, where a biomedical domain expert is asked to compare a system output and a human-authored string, show PaperRobot generated abstracts, conclusion and future work sections, and new titles are chosen over human-written ones up to 30%, 24% and 12% of the time, respectively. 1 keeps almost the same across years.In 2012, US scientists estimated that they read, on average, only 264 papers per year (1 out of 5000 available papers), which is, statistically, not different from what they reported in an identical survey last conducted in 2005.PaperRobot automatically reads existing papers to build background knowledge graphs (KGs), in which nodes are entities/concepts and edges are the relations between these entities (Section 2.2). Qingyun Wang 0005, Lifu Huang, Zhiying Jiang, Kevin Knight, Heng Ji 0001, Mohit Bansal, Yi Luan |
ACL (1) | 2 |
| 2019 | Cosmos QA: Machine Reading Comprehension with Contextual Commonsense ReasoningabstractLifu Huang, Ronan Le Bras, Chandra Bhagavatula, Yejin Choi. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Lifu Huang, Ronan Le Bras 0001, Chandra Bhagavatula, Yejin Choi 0001 |
EMNLP/IJCNLP (1) | 1 |
| 2018 | Zero-Shot Transfer Learning for Event ExtractionabstractMost previous supervised event extraction methods have relied on features derived from manual annotations, and thus cannot be applied to new event types without extra annotation effort.We take a fresh look at event extraction and model it as a generic grounding problem: mapping each event mention to a specific type in a target event ontology.We design a transferable architecture of structural and compositional neural networks to jointly represent and map event mentions and types into a shared semantic space.Based on this new framework, we can select, for each event mention, the event type which is semantically closest in this space as its type.By leveraging manual annotations available for a small set of existing event types, our framework can be applied to new unseen event types without additional manual annotations.When tested on 23 unseen event types, this zeroshot framework, without manual annotations, achieves performance comparable to a supervised model trained from 3,000 sentences annotated with 500 event mentions.1 Lifu Huang, Heng Ji 0001, Kyunghyun Cho, Ido Dagan, Sebastian Riedel 0001, Clare R. Voss |
ACL (1) | 1 |
| 2018 | Open-Schema Event Profiling for Massive News CorporaabstractWith the rapid growth of online information services, a sheer volume of news data becomes available. To help people quickly digest the explosive information, we define a new problem - schema-based news event profiling - profiling events reported in open-domain news corpora, with a set of slots and slot-value pairs for each event, where the set of slots forms the schema of an event type. Such profiling not only provides readers with concise views of events, but also facilitates various applications such as information retrieval, knowledge graph construction and question answering. It is however a quite challenging task. The first challenge is to find out events and event types because they are both initially unknown. The second difficulty is the lack of pre-defined event-type schemas. Lastly, even with the schemas extracted, to generate event profiles from them is still essential yet demanding. Quan Yuan 0001, Xiang Ren 0001, Wenqi He, Chao Zhang 0014, Xinhe Geng, Lifu Huang, Heng Ji 0001, Chin-Yew Lin, Jiawei Han 0001 |
CIKM | 6 |
| 2018 | Global Attention for Name TaggingabstractMany name tagging approaches use local contextual information with much success, but fail when the local context is ambiguous or limited.We present a new framework to improve name tagging by utilizing local, documentlevel, and corpus-level contextual information.We retrieve document-level context from other sentences within the same document and corpus-level context from sentences in other topically related documents.We propose a model that learns to incorporate documentlevel and corpus-level contextual information alongside local contextual information via global attentions, which dynamically weight their respective contextual information, and gating mechanisms, which determine the influence of this information.Extensive experiments on benchmark datasets show the effectiveness of our approach, which achieves state-of-the-art results for Dutch, German, and Spanish on the CoNLL-2002 and CoNLL-2003 datasets.1 . Boliang Zhang, Spencer Whitehead, Lifu Huang, Heng Ji 0001 |
CoNLL | 3 |
| 2018 | Multi-lingual Common Semantic Space Construction via Cluster-Consistent Word EmbeddingabstractWe construct a multilingual common semantic space based on distributional semantics, where words from multiple languages are projected into a shared space via which all available resources and knowledge can be shared across multiple languages.Beyond word alignment, we introduce multiple cluster-level alignments and enforce the word clusters to be consistently distributed across multiple languages.We exploit three signals for clustering: (1) neighbor words in the monolingual word embedding space; (2) character-level information; and (3) linguistic properties (e.g., apposition, locative suffix) derived from linguistic structure knowledge bases available for thousands of languages.We introduce a new cluster-consistent correlational neural network to construct the common semantic space by aligning words as well as clusters.Intrinsic evaluation on monolingual and multilingual QVEC tasks shows our approach achieves significantly higher correlation with linguistic features which are extracted from manually crafted lexical resources than state-of-the-art multi-lingual embedding learning methods do.Using low-resource language name tagging as a case study for extrinsic evaluation, our approach achieves up to 14.6% absolute F-score gain over the state of the art on cross-lingual direct transfer.Our approach is also shown to be robust even when the size of bilingual dictionary is small.1 Lifu Huang, Kyunghyun Cho, Boliang Zhang, Heng Ji 0001, Kevin Knight |
EMNLP | 1 |
| 2018 | Entity-aware Image Caption GenerationabstractCurrent image captioning approaches generate descriptions which lack specific information, such as named entities that are involved in the images.In this paper we propose a new task which aims to generate informative image captions, given images and hashtags as input.We propose a simple but effective approach to tackle this problem.We first train a convolutional neural networks -long short term memory networks (CNN-LSTM) model to generate a template caption based on the input image.Then we use a knowledge graph based collective inference algorithm to fill in the template with specific named entities retrieved via the hashtags.Experiments on a new benchmark dataset collected from Flickr show that our model generates news-style image descriptions with much richer information.Our model outperforms unimodal baselines significantly with various evaluation metrics. 1 Di Lu 0003, Spencer Whitehead, Lifu Huang, Heng Ji 0001, Shih-Fu Chang |
EMNLP | 3 |
| 2018 | Genre Separation Network with Adversarial Training for Cross-genre Relation ExtractionabstractRelation Extraction suffers from dramatical performance decrease when training a model on one genre and directly applying it to a new genre, due to the distinct feature distributions.Previous studies address this problem by discovering a shared space across genres using manually crafted features, which requires great human effort.To effectively automate this process, we design a genre-separation network, which applies two encoders, one genreindependent and one genre-shared, to explicitly extract genre-specific and genre-agnostic features.Then we train a relation classifier using the genre-agnostic features on the source genre and directly apply to the target genre.Experiment results on three distinct genres of the ACE dataset show that our approach achieves up to 6.1% absolute F1-score gain compared to previous methods.By incorporating a set of external linguistic features, our approach outperforms the state-of-the-art by 1.7% absolute F1 gain.We make all programs of our model publicly available for research purpose 1 . Ge Shi 0002, Chong Feng 0001, Lifu Huang, Boliang Zhang, Heng Ji 0001, Lejian Liao, Heyan Huang |
EMNLP | 3 |
| 2018 | Describing a Knowledge BaseabstractWe aim to automatically generate natural language descriptions about an input structured knowledge base (KB).We build our generation framework based on a pointer network which can copy facts from the input KB, and add two attention mechanisms: (i) slot-aware attention to capture the association between a slot type and its corresponding slot value; and (ii) a new table position self-attention to capture the inter-dependencies among related slots.For evaluation, besides standard metrics including BLEU, METEOR, and ROUGE, we propose a KB reconstruction based metric by extracting a KB from the generation output and comparing it with the input KB.We also create a new data set which includes 106,216 pairs of structured KBs and their corresponding natural language descriptions for two distinct entity types.Experiments show that our approach significantly outperforms stateof-the-art methods.The reconstructed KB achieves 68.8% -72.6% F-score. 1 Qingyun Wang 0005, Xiaoman Pan, Lifu Huang, Boliang Zhang, Zhiying Jiang, Heng Ji 0001, Kevin Knight |
INLG | 3 |
| 2018 | Tracking State Changes in Procedural Text: a Challenge Dataset and Models for Process Paragraph ComprehensionabstractBhavana Dalvi, Lifu Huang, Niket Tandon, Wen-tau Yih, Peter Clark. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Bhavana Dalvi, Lifu Huang, Niket Tandon, Scott Yih, Peter Clark |
NAACL-HLT | 2 |
| 2017 | Bridge Text and Knowledge by Learning Multi-Prototype Entity Mention EmbeddingabstractIntegrating text and knowledge into a unified semantic space has attracted significant research interests recently.However, the ambiguity in the common space remains a challenge, namely that the same mention phrase usually refers to various entities.In this paper, to deal with the ambiguity of entity mentions, we propose a novel Multi-Prototype Mention Embedding model, which learns multiple sense embeddings for each mention by jointly modeling words from textual contexts and entities derived from a knowledge base.In addition, we further design an efficient language model based approach to disambiguate each mention to a specific sense.In experiments, both qualitative and quantitative analysis demonstrate the high quality of the word, entity and multi-prototype mention embeddings.Using entity linking as a study case, we apply our disambiguation method as well as the multi-prototype mention embeddings on the benchmark dataset, and achieve state-of-the-art performance. Yixin Cao 0002, Lifu Huang, Heng Ji 0001, Xu Chen 0017, Juan-Zi Li |
ACL (1) | 2 |
| 2017 | Improving Slot Filling Performance with Attentive Neural Networks on Dependency StructuresabstractSlot Filling (SF) aims to extract the values of certain types of attributes (or slots, such as person:cities of residence) for a given entity from a large collection of source documents.In this paper we propose an effective DNN architecture for SF with the following new strategies: (1).Take a regularized dependency graph instead of a raw sentence as input to DNN, to compress the wide contexts between query and candidate filler; (2).Incorporate two attention mechanisms: local attention learned from query and candidate filler, and global attention learned from external knowledge bases, to guide the model to better select indicative contexts to determine slot type.Experiments show that this framework outperforms state-of-the-art on both relation extraction (16% absolute F-score gain) and slot filling validation for each individual system (up to 8.5% absolute Fscore gain). Lifu Huang, Avirup Sil, Heng Ji 0001, Radu Florian |
EMNLP | 1 |
| 2017 | Expertise-Aware Truth Analysis and Task Allocation in Mobile CrowdsourcingabstractMobile crowdsourcing has received considerable attention as it enables people to collect and share large volume of data through their mobile devices. Since the accuracy of the collected data is usually hard to ensure, researchers have proposed techniques to identify truth from noisy data by inferring and utilizing the reliability of users, and allocate tasks to users with higher reliability. However, they neglect the fact that a user may only have expertise on some problems (in some domains), but not others. Neglecting this expertise diversity may cause two problems: low estimation accuracy in truth analysis and ineffective task allocation. To address these problems, we propose an Expertise-aware Truth Analysis and Task Allocation (ETA2) approach, which can effectively infer user expertise and then allocate tasks and estimate truth based on the inferred expertise. ETA2relies on a novel semantic analysis method to identify the expertise domains of the tasks and user expertise, an expertise-aware truth analysis solution to estimate truth and learn user expertise, and an expertise-aware task allocation method to maximize the probability that tasks are allocated to users with the right expertise while ensuring the work load does not exceed the processing capability at each user. Experimental results based on two real-world datasets demonstrate that ETA2significantly outperforms existing solutions. Xiaomei Zhang 0001, Lifu Huang, Heng Ji 0001, Guohong Cao |
ICDCS | 3 |
| 2017 | Open Relation Extraction and GroundingabstractPrevious open Relation Extraction (open RE) approaches mainly rely on linguistic patterns and constraints to extract important relational triples from large-scale corpora. However, they lack of abilities to cover diverse relation expressions or measure the relative importance of candidate triples within a sentence. It is also challenging to name the relation type of a relational triple merely based on context words, which could limit the usefulness of open RE in downstream applications. We propose a novel importance-based open RE approach by exploiting the global structure of a dependency tree to extract salient triples. We design an unsupervised relation type naming method by grounding relational triples to a large-scale Knowledge Base (KB) schema, leveraging KB triples and weighted context words associated with relational triples. Experiments on the English Slot Filling 2013 dataset demonstrate that our approach achieves 8.1% higher F-score over state-of-the-art open RE methods. Dian Yu 0001, Lifu Huang, Heng Ji 0001 |
IJCNLP(1) | 2 |
| 2017 | Improving Event Extraction via Multimodal IntegrationabstractIn this paper, we focus on improving Event Extraction (EE) by incorporating visual knowledge with words and phrases from text documents. We first discover visual patterns from large-scale text-image pairs in a weakly-supervised manner and then propose a multimodal event extraction algorithm where the event extractor is jointly trained with textual features and visual patterns. Extensive experimental results on benchmark data sets demonstrate that the proposed multimodal EE method can achieve significantly better performance on event extraction: absolute 7.1% F-score gain on event trigger labeling and 8.5% F-score gain on event argument labeling. Tongtao Zhang, Spencer Whitehead, Hanwang Zhang, Hongzhi Li 0001, Joseph G. Ellis, Lifu Huang, Wei Liu 0005, Heng Ji 0001, Shih-Fu Chang |
ACM Multimedia | 6 |
| 2016 | Liberal Event Extraction and Event Schema Induction
Lifu Huang, Taylor Cassidy, Heng Ji 0001, Clare R. Voss, Jiawei Han 0001, Avirup Sil |
ACL (1) | 1 |
| 2016 | AFET: Automatic Fine-Grained Entity Typing by Hierarchical Partial-Label EmbeddingabstractDistant supervision has been widely used in current systems of fine-grained entity typing to automatically assign categories (entity types) to entity mentions.However, the types so obtained from knowledge bases are often incorrect for the entity mention's local context.This paper proposes a novel embedding method to separately model "clean" and "noisy" mentions, and incorporates the given type hierarchy to induce loss functions.We formulate a joint optimization problem to learn embeddings for mentions and typepaths, and develop an iterative algorithm to solve the problem.Experiments on three public datasets demonstrate the effectiveness and robustness of the proposed method, with an average 15% improvement in accuracy over the next best compared method 1 . * Equal contribution.1 Codes and datasets used in this paper can be downloaded at https://github.com/shanzhenren/AFET. Xiang Ren 0001, Wenqi He, Meng Qu, Lifu Huang, Heng Ji 0001, Jiawei Han 0001 |
EMNLP | 4 |
| 2015 | Memory-Aware NoC Application Mapping Based on Adaptive Genetic Algorithm
Yizhuo Wang 0001, Zhibiao Zhang, Lifu Huang, Weixing Ji |
ICA3PP (1) | 3 |
| 2014 | Generating Supplementary Travel Guides from Social Media
Liu Yang 0005, Jing Jiang 0001, Lifu Huang, Minghui Qiu, Lizi Liao |
COLING | 3 |
| 2014 | Discovering Informative Contents of Web Pages
Qifeng Fan, Chunwei Yan, Lifu Huang, Lian'en Huang |
WAIM | 3 |
| 2014 | Comments-Oriented Summarization in Blogsphere Using a Two-Stage Sentence Similarity Measure
Lifu Huang, Qifeng Fan, Lian'en Huang |
WAIM | 2 |
| 2013 | Optimized Event Storyline Generation based on Mixture-Event-Aspect ModelabstractRecently, much research focuses on event storyline generation, which aims to produce a concise, global and temporal event summary from a collection of articles.Generally, each event contains multiple sub-events and the storyline should be composed by the component summaries of all the sub-events.However, different sub-events have different part-whole relationship with the major event, which is important to correspond to users' interests but seldom considered in previous work.To distinguish different types of sub-events, we propose a mixture-event-aspect model which models different sub-events into local and global aspects.Combining these local/global aspects with summarization requirements together, we utilize an optimization method to generate the component summaries along the timeline.We develop experimental systems on 6 distinctively different datasets.Evaluation and comparison results indicate the effectiveness of our proposed method. Lifu Huang, Lian'en Huang |
EMNLP | 1 |
| 2013 | Comments-Oriented Document Summarization Based on Multi-aspect Co-feedback Ranking
Lifu Huang, Lian'en Huang |
WAIM | 1 |
| 2012 | RelationListwise for Query-Focused Multi-Document Summarization
Wenpeng Yin 0001, Lifu Huang, Yulong Pei, Lian'en Huang |
COLING | 2 |