VLDB 2026 Research / reviewers in the wild / expert
Hongcheng Guo
dblp:84/8542
· DBLP profile ↗
29ranked-venue papers
7as first author
29since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 5 first-author · 22 since 2021Databases, data management, data science and information retrieval · 8 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Outcome-Grounded Advantage Reshaping for Fine-Grained Credit Assignment in Mathematical ReasoningabstractZiheng Li, Liu Kang, Feng Xiao, Luxi Xing, Qingyi Si, Zhuoran Li, Weikang Gong, Deqing Yang, Yanghua Xiao, Hongcheng Guo. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Liu Kang, Luxi Xing, Qingyi Si, Weikang Gong, Deqing Yang, Yanghua Xiao, Hongcheng Guo |
ACL (1) | 10 |
| 2026 | GatheringSense: AI-Generated Imagery and Embodied Experiences for Understanding Literati GatheringsabstractChinese literati gatherings (Wenren Yaji), as a situated form of Chinese traditional culture, remain underexplored in depth. Although generative AI supports powerful multimodal generation, current cultural applications largely emphasize aesthetic reproduction and struggle to convey the deeper meanings of cultural rituals and social frameworks. Based on embodied cognition, we propose an AI-driven dual-path framework for cultural understanding, which we instantiate through GatheringSense, a literati-gathering experience. We conduct a mixed-methods study (N = 48) to compare how AI-generated multimodal content and embodied participation complement each other in supporting the understanding of literati gatherings and fostering cultural resonance. Our results show that AI-generated content effectively improves the readability of cultural symbols and initial emotional attraction, yet limitations in physical coherence and micro-level credibility may affect users’ satisfaction. In contrast, embodied experience significantly deepens participants’ understanding of ritual rules and social roles, and increases their psychological closeness and presence. Based on these findings, we offer empirical evidence and five transferable design implications for generative experience in cultural heritage. Bingyuan Wang, Hongcheng Guo, Zeyu Wang 0003 |
CHI | 3 |
| 2026 | Building and Benchmarking Large Language Models for Machine Translation in Social Network Services
Hongcheng Guo, Fei Zhao 0012, Shaosheng Cao, Xinze Lyu, Zijie Meng, Yao Hu 0002, Zhoujun Li 0001, Zuozhu Liu |
ICDE | 1 |
| 2026 | Multi-branch semantic alignment for few-shot image classification
Heng Wu 0004, Laishui Lv, Changchun Zhang, Hongcheng Guo, Shanzhou Niu, Gaohang Yu |
Inf. Sci. | 5 |
| 2025 | Mitigating Hallucinations in Large Vision-Language Models by Adaptively Constraining Information FlowabstractLarge vision-language models show tremendous potential in understanding visual information through human languages. However, they are prone to suffer from object hallucination, i.e., the generated image descriptions contain objects that do not exist in the image. In this paper, we reveal that object hallucination can be attributed to overconfidence in irrelevant visual features when soft visual tokens map to the LLM's word embedding space. Specifically, by figuring out the semantic similarity between visual tokens and LLM's word embedding, we observe that the smoothness of similarity distribution strongly correlates with the emergence of object hallucinations. To mitigate hallucinations, we propose using the Variational Information Bottleneck (VIB) to alleviate overconfidence by introducing stochastic noise, facilitating the constraining of irrelevant information. Furthermore, we propose an entropy-based noise-controlling strategy to enable the injected noise to be adaptively constrained regarding the smoothness of the similarity distribution. We adapt the proposed AdaVIB across distinct model architectures. Experimental results demonstrate that the proposed AdaVIB mitigates object hallucinations by effectively alleviating the overconfidence in irrelevant visual features, with consistent improvements on two object hallucination benchmarks. Jiaqi Bai 0001, Hongcheng Guo, Zhongyuan Peng, Jian Yang 0030, Zhoujun Li 0001, Mohan Li, Zhihong Tian |
AAAI | 2 |
| 2025 | XCOT: Cross-lingual Instruction Tuning for Cross-lingual Chain-of-Thought ReasoningabstractChain-of-thought (CoT) has emerged as a powerful technique to elicit reasoning in large language models and improve a variety of downstream tasks. CoT mainly demonstrates excellent performance in English, but its usage in low-resource languages is constrained due to poor language generalization. To bridge the gap among different languages, we propose a cross-lingual instruction fine-tuning framework (xCoT) to transfer knowledge from high-resource languages to low-resource languages. Specifically, the multilingual instruction training data (xCoT-Instruct) is created to encourage the semantic alignment of multiple languages. We introduce cross-lingual in-context few-shot learning (xICL) to accelerate multilingual agreement in instruction tuning, where some fragments of source languages in examples are randomly substituted by their counterpart translations of target languages. During multilingual instruction tuning, we adopt the randomly online CoT strategy to enhance the multilingual reasoning ability of the large language model by first translating the query to another language and then answering in English. To further facilitate the language transfer, we leverage the high-resource CoT to supervise the training of low-resource languages with cross-lingual distillation. Experimental results demonstrate the superior performance of xCoT in reducing the gap among different languages, highlighting its potential to reduce the cross-lingual gap. Linzheng Chai, Jian Yang 0030, Tao Sun 0016, Hongcheng Guo, Xinnian Liang, Jiaqi Bai 0001, Tongliang Li, Qiyao Peng 0001, Zhoujun Li 0001 |
AAAI | 4 |
| 2025 | DinoCompanion: An Attachment-Theory Informed Multimodal Robot for Emotionally Responsive Child-AI InteractionabstractEmotional development of children fundamentally relies on secure attachment relationships, yet current AI companions lack the theoretical foundation to provide developmentally appropriate emotional support. We introduce DinoCompanion, the first attachment-theory-grounded multimodal robot for emotionally responsive child-AI interaction. We address three critical challenges in child-AI systems: the absence of developmentally-informed AI architectures, the need to balance engagement with safety, and the lack of standardized evaluation frameworks for attachment-based capabilities. Our contributions include: (i) a multimodal dataset of 128 caregiver-child dyads containing 125,382 annotated clips with paired preference-risk labels, (ii) CARPO (Child-Aware Risk-calibrated Preference Optimization), a novel training objective that maximizes engagement while applying epistemic-uncertainty-weighted risk penalties, and (iii) AttachSecure-Bench, a comprehensive evaluation benchmark covering ten attachment-centric competencies with strong expert consensus. AttachSecure-Bench achieves state-of-the-art performance (57.15%), outperforming GPT-4o and Gemini-2.5-Pro, with exceptional secure base behaviors and superior attachment risk detection. Ablations validate the critical importance of multimodal fusion, uncertainty-aware risk modeling, and hierarchical memory for coherent, emotionally attuned interactions. Boyang Wang 0006, Yuhao Song, Jinyuan Cao, Hongcheng Guo, Zhoujun Li 0001 |
CIKM | 5 |
| 2025 | Pet-Bench: Benchmarking the Abilities of Large Language Models as E-Pets in Social Network ServicesabstractAs interest in using Large Language Models for interactive and emotionally rich experiences grows, virtual pet companionship emerges as a novel yet underexplored application. Existing approaches focus on basic pet role-playing interactions without systematically benchmarking LLMs for comprehensive companionship. In this paper, we introduce PET-BENCH, a dedicated benchmark that evaluates LLMs across both self-interaction and human-interaction dimensions. Unlike prior work, PET-BENCH emphasizes self-evolution and developmental behaviors alongside interactive engagement, offering a more realistic reflection of pet companionship. It features diverse tasks such as intelligent scheduling, memory-based dialogues, and psychological conversations, with over 7,500 interaction instances designed to simulate pet behaviors. Evaluation of 28 LLMs reveals significant performance variations linked to model size and inherent capabilities, underscoring the need for specialized optimization in this domain. PET-BENCH serves as a foundational resource for benchmarking pet-related LLM abilities and advancing emotionally immersive human-pet interactions. Hongcheng Guo, Zheyong Xie, Shaosheng Cao, Boyang Wang 0006, Weiting Liu 0001, Zheyu Ye, Zhoujun Li 0001, Zuozhu Liu, Wei Lu 0011 |
CIKM | 1 |
| 2025 | ADC: Enhancing Function Calling Via Adversarial Datasets and Code Line-Level FeedbackabstractLarge Language Models (LLMs) have made significant strides in Natural Language Processing and coding, yet they struggle with robustness and accuracy in complex function calls. To tackle these challenges, this paper introduces ADC, an innovative approach that enhances LLMs’ ability to follow function formats and match complex parameters. ADC utilizes a high-quality code fine-tuning dataset with line-level execution feedback, providing granular process supervision that fosters strong logical reasoning and adherence to function formats. It also employs an adversarial dataset generation process to improve parameter matching. The staged training methodology capitalizes on both enriched code datasets and refined adversarial datasets, leading to marked improvements in function calling capabilities on the Berkeley Function-Calling Leaderboard (BFCL) Benchmark. The innovation of ADC lies in its strategic combination of process supervision, adversarial refinement, and incremental learning, setting a new standard for LLM proficiency in complex function calling. Wei Zhang 0384, Qianghuai Jia, Feijun Jiang, Hongcheng Guo, Zhoujun Li 0001, Mengping Zhou |
ICASSP | 6 |
| 2025 | McEval: Massively Multilingual Code EvaluationabstractCode large language models (LLMs) have shown remarkable advances in code understanding, completion, and generation tasks. Programming benchmarks, comprised of a selection of code challenges and corresponding test cases, serve as a standard to evaluate the capability of different LLMs in such tasks. However, most existing benchmarks primarily focus on Python and are still restricted to a limited number of languages, where other languages are translated from the Python samples degrading the data diversity. To further facilitate the research of code LLMs, we propose a massively multilingual code benchmark covering 40 programming languages (McEval) with 16K test samples, which substantially pushes the limits of code LLMs in multilingual scenarios. The benchmark contains challenging code completion, understanding, and generation evaluation tasks with finely curated massively multilingual instruction corpora McEval-Instruct. In addition, we introduce an effective multilingual coder mCoder trained on McEval-Instruct to support multilingual programming language generation. Extensive experimental results on McEval show that there is still a difficult journey between open-source models and closed-source LLMs in numerous languages. The instruction corpora and evaluation benchmark are available at https://github.com/MCEVAL/McEval. Linzheng Chai, Jian Yang 0030, Yuwei Yin, Tao Sun 0016, Ge Zhang 0009, Changyu Ren, Hongcheng Guo, Noah Wang, Boyang Wang 0006, Xianjie Wu, Tongliang Li, Liqun Yang, Sufeng Duan, Zhaoxiang Zhang 0001, Zhoujun Li 0001 |
ICLR | 10 |
| 2025 | Revealing Weaknesses in Text Watermarking Through Self-Information Rewrite AttacksabstractText watermarking aims to subtly embeds statistical signals into text by controlling the Large Language Model (LLM)’s sampling process, enabling watermark detectors to verify that the output was generated by the specified model. The robustness of these watermarking algorithms has become a key factor in evaluating their effectiveness. Current text watermarking algorithms embed watermarks in high-entropy tokens to ensure text quality. In this paper, we reveal that this seemingly benign design can be exploited by attackers, posing a significant risk to the robustness of the watermark. We introduce a generic efficient paraphrasing attack, the Self-Information Rewrite Attack (SIRA), which leverages the vulnerability by calculating the self-information of each token to identify potential pattern tokens and perform targeted attack. Our work exposes a widely prevalent vulnerability in current watermarking algorithms. The experimental results show SIRA achieves nearly 100% attack success rates on seven recent watermarking methods with only $0.88 per million tokens cost. Our approach does not require any access to the watermark algorithms or the watermarked LLM and can seamlessly transfer to any LLM as the attack model even mobile-level models. Our findings highlight the urgent need for more robust watermarking. Yixin Cheng, Hongcheng Guo, Yangming Li, Leonid Sigal |
ICML | 2 |
| 2025 | SNS-Bench: Defining, Building, and Assessing Capabilities of Large Language Models in Social Networking ServicesabstractWith the rapid advancement of Social Networking Services (SNS), the need for intelligent and efficient interaction within diverse platforms has become more crucial. Large Language Models (LLMs) play an important role in SNS as they possess the potential to revolutionize user experience, content generation, and communication dynamics. However, recent studies focus on isolated SNS tasks rather than a comprehensive evaluation.
In this paper, we introduce SNS-Bench, specially constructed for assessing the abilities of large language models from different Social Networking Services, with a wide range of SNS-related information. SNS-Bench encompasses 8 different tasks such as note classification, query content relevance, and highlight words generation in comments. Finally, 6,658 questions of social media text, including subjective questions, single-choice, and multiple-choice questions, are concluded in SNS-Bench. Further, we evaluate the performance of over 25+ current diverse LLMs on our SNS-Bench. Models with different sizes exhibit performance variations, yet adhere to the scaling law.
Moreover, we hope provide more insights to revolutionize the techniques of social network services with LLMs. Hongcheng Guo, Shaosheng Cao, Fei Zhao 0012, Boyang Wang 0006, Lei Li 0039, Liang Chen 0024, Xinze Lyu, Yao Hu 0002, Zhoujun Li 0001 |
ICML | 1 |
| 2025 | H2HTalk: Evaluating Large Language Models as Emotional Companion
Boyang Wang 0006, Yalun Wu, Hongcheng Guo, Zhoujun Li 0001 |
NLPCC (1) | 3 |
| 2025 | Expanding Virtual Production Frontiers: AI-Driven Workflows for Enhanced Cinematic CreationabstractAlthough virtual production (VP) offers cinematic and immersive storytelling by aligning a high-end camera with multiple LED screens through a central server, the creation of high-quality 3D scenery with real-time interactions remains complex and resource-intensive. This paper explores the potential of expanding cinematic virtual production scenes through three innovative AI-driven approaches: (1) AI-generated 360° panoramas from text prompts to produce immersive backgrounds; (2) Direct text/AI-generated image-to-3D environment conversion using mesh generation; (3) An end-to-end AI pipeline for rapid stylized scene construction. We demonstrate each workflow through real-time avatar-interactive shooting scenarios. Our approach bridges technical and artistic domains and aims to show how these AI-driven workflows could accelerate scene creation, enable novel cinematic experiences, and reveal the future potential of visual content generation. Junrong Song, Hongcheng Guo, Lujin Zhang, Zeyu Wang 0003, David Kei-Man Yip |
VINCI | 2 |
| 2024 | LogFormer: A Pre-train and Tuning Pipeline for Log Anomaly DetectionabstractLog anomaly detection is a key component in the field of artificial intelligence for IT operations (AIOps). Considering log data of variant domains, retraining the whole network for unknown domains is inefficient in real industrial scenarios. However, previous deep models merely focused on extracting the semantics of log sequences in the same domain, leading to poor generalization on multi-domain logs. To alleviate this issue, we propose a unified Transformer-based framework for Log anomaly detection (LogFormer) to improve the generalization ability across different domains, where we establish a two-stage process including the pre-training and adapter-based tuning stage. Specifically, our model is first pre-trained on the source domain to obtain shared semantic knowledge of log data. Then, we transfer such knowledge to the target domain via shared parameters. Besides, the Log-Attention module is proposed to supplement the information ignored by the log-paring. The proposed method is evaluated on three public datasets and one real-world dataset. Experimental results on multiple benchmarks demonstrate the effectiveness of our LogFormer with fewer trainable parameters and lower training costs. Hongcheng Guo, Jian Yang 0030, Jiaqi Bai 0001, Boyang Wang 0006, Zhoujun Li 0001, Tieqiao Zheng, Bo Zhang 0096, Junran Peng |
AAAI | 1 |
| 2024 | UniCoder: Scaling Code Large Language Model via Universal CodeabstractTao Sun, Linzheng Chai, Jian Yang, Yuwei Yin, Hongcheng Guo, Jiaheng Liu, Bing Wang, Liqun Yang, Zhoujun Li. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Tao Sun 0016, Linzheng Chai, Jian Yang 0030, Yuwei Yin, Hongcheng Guo, Liqun Yang, Zhoujun Li 0001 |
ACL (1) | 5 |
| 2024 | m3P: Towards Multimodal Multilingual Translation with Multimodal PromptabstractMultilingual translation supports multiple translation directions by projecting all languages in a shared space, but the translation quality is undermined by the difference between languages in the text-only modality, especially when the number of languages is large. To bridge this gap, we introduce visual context as the universal language-independent representation to facilitate multilingual translation. In this paper, we propose a framework to leverage the multimodal prompt to guide the Multimodal Multilingual Neural Machine Translation (m3P), which aligns the representations of different languages with the same meaning and generates the conditional vision-language memory for translation. We construct a multilingual multimodal instruction dataset (InstrMulti102) to support 102 languages Our method aims to minimize the representation distance of different languages by regarding the image as a central language. Experimental results show that m3P outperforms previous text-only baselines and multilingual multimodal methods by a large margin. Furthermore, the probing experiments validate the effectiveness of our method in enhancing translation under the low-resource and massively multilingual scenario. Jian Yang 0030, Hongcheng Guo, Yuwei Yin, Jiaqi Bai 0001, Xinnian Liang, Linzheng Chai, Liqun Yang, Zhoujun Li 0001 |
LREC/COLING | 2 |
| 2024 | LTA-PCS: Learnable Task-Agnostic Point Cloud SamplingabstractRecently, many approaches directly operate on point clouds for different tasks. These approaches become more computation and storage demanding when point cloud size is large. To reduce the required computation and storage, one possible solution is to sample the point cloud. In this paper, we propose the first Learnable Task-Agnostic Point Cloud Sampling (LTA-PCS) framework. Existing task-agnostic point cloud sampling strategy (e.g., FPS) does not consider semantic information of point clouds, causing de-graded performance on downstream tasks. While learning-based point cloud sampling methods consider semantic in-formation, they are task-specific and require task-oriented ground-truth annotations. So they cannot generalize well on different downstream tasks. Our LTA-PCS achieves task-agnostic point cloud sampling without requiring task-oriented labels, in which both the geometric and semantic information of points is considered in sampling. Extensive experiments on multiple downstream tasks demonstrate the effectiveness of our LTA-PCS. Kaisiyuan Wang, Hongcheng Guo, Jian Yang 0030, Junran Peng, Ke Xu 0001, Xianglong Liu 0001, Jinyang Guo 0002 |
CVPR | 4 |
| 2024 | OWL: A Large Language Model for IT OperationsabstractWith the rapid advancement of IT operations, managing and analyzing large data volumes efficiently for practical applications has become increasingly critical. Natural Language Processing (NLP) techniques have demonstrated remarkable capabilities in various tasks, including named entity recognition, machine translation, and dialogue systems. Recently, Large Language Models (LLMs) have achieved significant improvements across various domain-specific areas. However, there is a noticeable gap in the development of specialized Large Language Models (LLMs) tailored for IT operations. In this paper, we introduce the OWL, a large language model trained on our constructed Owl-Instruct with a wide range of IT-related information. Specifically, limited by the maximum input length, we propose the \textbf{H}omogeneous \textbf{M}arkov \textbf{C}ontext \textbf{E}xtension method (HMCE). The mixture-of-adapter strategy is leveraged to improve the parameter-efficient tuning across different domains or tasks.
Further, we evaluate the performance of OWL on the Owl-Bench established by us and open IT-related benchmarks. OWL demonstrates superior performance results on IT tasks, which outperforms existing models by significant margins. Moreover, we hope that the findings of our work will provide more insights to revolutionize the techniques of IT operations with specialized LLMs. Hongcheng Guo, Jian Yang 0030, Liqun Yang, Linzheng Chai, Jiaqi Bai 0001, Junran Peng, Xiaorong Hu, Dongfeng Zhang, Xu Shi 0005, Tieqiao Zheng, Liangfan Zheng, Bo Zhang 0096, Ke Xu 0001, Zhoujun Li 0001 |
ICLR | 1 |
| 2024 | RoleAgent: Building, Interacting, and Benchmarking High-quality Role-Playing Agents from ScriptsabstractBelievable agents can empower interactive applications ranging from immersive environments to rehearsal spaces for interpersonal communication. Recently, generative agents have been proposed to simulate believable human behavior by using Large Language Models. However, the existing method heavily relies on human-annotated agent profiles (e.g., name, age, personality, relationships with others, and so on) for the initialization of each agent, which cannot be scaled up easily. In this paper, we propose a scalable RoleAgent framework to generate high-quality role-playing agents from raw scripts, which includes building and interacting stages. Specifically, in the building stage, we use a hierarchical memory system to extract and summarize the structure and high-level information of each agent for the raw script. In the interacting stage, we propose a novel innovative mechanism with four steps to achieve a high-quality interaction between agents. Finally, we introduce a systematic and comprehensive evaluation benchmark called RoleAgentBench to evaluate the effectiveness of our RoleAgent, which includes 100 and 28 roles for 20 English and 5 Chinese scripts, respectively. Extensive experimental results on RoleAgentBench demonstrate the effectiveness of RoleAgent. Zehao Ni, Haoran Que, Tao Sun 0016, Noah Wang, Jian Yang 0030, Jiakai Wang, Hongcheng Guo, Zhongyuan Peng, Ge Zhang 0009, Xingyuan Bu, Ke Xu 0001, Wenge Rong, Junran Peng, Zhaoxiang Zhang 0001 |
NeurIPS | 8 |
| 2024 | mt4CrossOIE: Multi-stage tuning for cross-lingual open information extraction
Tongliang Li, Linzheng Chai, Jian Yang 0030, Jiaqi Bai 0001, Yuwei Yin, Hongcheng Guo, Liqun Yang, Hebboul Zine El Abidine, Zhoujun Li 0001 |
Expert Syst. Appl. | 8 |
| 2024 | Infusing internalized knowledge of language models into hybrid prompts for knowledgeable dialogue generation
Jiaqi Bai 0001, Jian Yang 0030, Hongcheng Guo, Zhoujun Li 0001 |
Knowl. Based Syst. | 5 |
| 2023 | GripRank: Bridging the Gap between Retrieval and Generation via the Generative Knowledge Improved Passage RankingabstractRetrieval-enhanced text generation has shown remarkable progress on knowledge-intensive language tasks, such as open-domain question answering and knowledge-enhanced dialogue generation, by leveraging passages retrieved from a large passage corpus for delivering a proper answer given the input query. However, the retrieved passages are not ideal for guiding answer generation because of the discrepancy between retrieval and generation, i.e., the candidate passages are all treated equally during the retrieval procedure without considering their potential to generate a proper answer. This discrepancy makes a passage retriever deliver a sub-optimal collection of candidate passages to generate the answer. In this paper, we propose the GeneRative Knowledge Improved Passage Ranking (GripRank) approach, addressing the above challenge by distilling knowledge from a generative passage estimator (GPE) to a passage ranker, where the GPE is a generative language model used to measure how likely the candidate passages can generate the proper answer. We realize the distillation procedure by teaching the passage ranker learning to rank the passages ordered by the GPE. Furthermore, we improve the distillation quality by devising a curriculum knowledge distillation mechanism, which allows the knowledge provided by the GPE can be progressively distilled to the ranker through an easy-to-hard curriculum, enabling the passage ranker to correctly recognize the provenance of the answer from many plausible candidates. We conduct extensive experiments on four datasets across three knowledge-intensive language tasks. Experimental results show advantages over the state-of-the-art methods for both passage ranking and answer generation on the KILT benchmark. Jiaqi Bai 0001, Hongcheng Guo, Jian Yang 0030, Xinnian Liang, Zhoujun Li 0001 |
CIKM | 2 |
| 2023 | LogLG: Weakly Supervised Log Anomaly Detection via Log-Event Graph Construction
Hongcheng Guo, Yuhui Guo, Jian Yang 0030, Zhoujun Li 0001, Tieqiao Zheng, Liangfan Zheng, Weichao Hou, Bo Zhang 0096 |
DASFAA (4) | 1 |
| 2023 | HanoiT: Enhancing Context-aware Translation via Selective Context
Jian Yang 0030, Yuwei Yin, Shuming Ma, Liqun Yang, Hongcheng Guo, Haoyang Huang, Dongdong Zhang 0001, Yutao Zeng, Zhoujun Li 0001, Furu Wei |
DASFAA (3) | 5 |
| 2023 | KnowPrefix-Tuning: A Two-Stage Prefix-Tuning Framework for Knowledge-Grounded Dialogue Generation
Jiaqi Bai 0001, Ze Yang 0001, Jian Yang 0030, Xinnian Liang, Hongcheng Guo, Zhoujun Li 0001 |
ECML/PKDD (2) | 6 |
| 2023 | KINet: Incorporating Relevant Facts Into Knowledge-Grounded Dialog GenerationabstractKnowledge-grounded conversation has led to great progress in producing informative dialog responses by leveraging external knowledge. This work focuses on two affiliated knowledge grounded conversation tasks:Knowledge SelectionandResponse Generation. Previous work followed the paradigm of selecting the most optimal knowledge piece to guide the conversation towards generating the proper response. However, some knowledge pieces, which are not recognized as optimal, may still benefit response generation. How to effectively leverage these relevant knowledge pieces for response generation still remain a tricky issue. To address this problem, we proposeKINet, aKnowledgeIncorporationNetwork, which deals with the problem by boosting both the knowledge selection and the response generation. The proposed model contains a negative enhanced knowledge approximator which improves knowledge selection by enhancing the dense representation of knowledge pieces, and a curriculum knowledge sampler which improves generated responses by incorporating more relevant knowledge pieces in an easy-to-hard manner. We conduct the experiment on two datasets of knowledge-grounded conversations, the results show that the proposed model significantly outperforms state-of-the-art methods in terms of both automatic and human evaluations. Jiaqi Bai 0001, Ze Yang 0001, Jian Yang 0030, Hongcheng Guo, Zhoujun Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2022 | LVP-M3: Language-aware Visual Prompt for Multilingual Multimodal Machine TranslationabstractMultimodal Machine Translation (MMT) focuses on enhancing text-only translation with visual features, which has attracted considerable attention from both natural language processing and computer vision communities.Recent advances still struggle to train a separate model for each language pair, which is costly and unaffordable when the number of languages increases in the real world.In other words, the multilingual multimodal machine translation (Multilingual MMT) task has not been investigated, which aims to handle the aforementioned issues by providing a shared semantic space for multiple languages.Besides, the image modality has no language boundaries, which is superior to bridging the semantic gap between languages.To this end, we first propose the Multilingual MMT task by establishing two new Multilingual MMT benchmark datasets covering seven languages.Then, an effective baseline LVP-M 3 using visual prompts is proposed to support translations between different languages, which includes three stages (token encoding, language-aware visual prompt generation, and language translation).Extensive experimental results on our constructed benchmark datasets demonstrate the effectiveness of LVP-M 3 method for Multilingual MMT.* First two authors contributed equally. Hongcheng Guo, Haoyang Huang, Jian Yang 0030, Zhoujun Li 0001, Dongdong Zhang 0001 |
EMNLP | 1 |
| 2022 | UM4: Unified Multilingual Multiple Teacher-Student Model for Zero-Resource Neural Machine TranslationabstractMost translation tasks among languages belong to the zero-resource translation problem where parallel corpora are unavailable. Multilingual neural machine translation (MNMT) enables one-pass translation using shared semantic space for all languages compared to the two-pass pivot translation but often underperforms the pivot-based method. In this paper, we propose a novel method, named as Unified Multilingual Multiple teacher-student Model for NMT (UM4). Our method unifies source-teacher, target-teacher, and pivot-teacher models to guide the student model for the zero-resource translation. The source teacher and target teacher force the student to learn the direct source-target translation by the distilled knowledge on both source and target sides. The monolingual corpus is further leveraged by the pivot-teacher model to enhance the student model. Experimental results demonstrate that our model of 72 directions significantly outperforms previous methods on the WMT benchmark. Jian Yang 0030, Yuwei Yin, Shuming Ma, Dongdong Zhang 0001, Shuangzhi Wu, Hongcheng Guo, Zhoujun Li 0001, Furu Wei |
IJCAI | 6 |