Nayu Liu

dblp:278/8050 · DBLP profile ↗
← Back
24ranked-venue papers
9as first author
23since 2021 · last 2026
0000-0002-7664-9856ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 8 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Rectify Evaluation Preference: Improving LLMs' Critique on Math Reasoning via Perplexity-aware Reinforcement Learning
abstract
To improve Multi-step Mathematical Reasoning (MsMR) of Large Language Models (LLMs), it is crucial to obtain scalable supervision from the corpus by automatically critiquing mistakes in the reasoning process of MsMR and rendering a final verdict of the problem-solution. Most existing methods rely on crafting high-quality supervised fine-tuning demonstrations for critiquing capability enhancement and pay little attention to delving into the underlying reason for the poor critiquing performance of LLMs. In this paper, we orthogonally quantify and investigate the potential reason — imbalanced evaluation preference, and conduct a statistical preference analysis. Motivated by the analysis of the reason, a novel perplexity-aware reinforcement learning algorithm is proposed to rectify the evaluation preference, elevating the critiquing capability. Specifically, to probe into LLMs' critiquing characteristics, a One-to-many Problem-Solution (OPS) benchmark is meticulously constructed to quantify the behavior difference of LLMs when evaluating the problem solutions generated by itself and others. Then, to investigate the behavior difference in depth, we conduct a statistical preference analysis oriented on perplexity and find an intriguing phenomenon — "LLMs incline to judge solutions with lower perplexity as correct", which is dubbed as imbalanced evaluation preference. To rectify this preference, we regard perplexity as the baton in the algorithm of Group Relative Policy Optimization, supporting the LLMs to explore trajectories that judge lower perplexity as wrong and higher perplexity as correct. Extensive experimental results on our built OPS and existing available critic benchmarks demonstrate the validity of our method.
Changyuan Tian 0001, Zhicong Lu, Shuang Qian, Nayu Liu, Peiguang Li, Li Jin 0001, Leiyi Hu, Zhizhao Zeng, Guozhi Cas
AAAI4
2026 UMNet: Uncertainty-guided Memory Network for Hyperspectral Pansharpening
abstract
At present, most hyperspectral (HS) sharpening methods have not fully utilized the feature correlation between adjacent bands in HS images, nor have they explored the problem of feature uncertainty generated by the model during the fusion process. This may lead to inaccurate fusion features generated by the model, resulting in spatial and spectral distortions in the fusion results. To address these issues, we propose an uncertainty-guided memory network (UMNet) for HS pansharpening. A spatial-spectral recurrent fusion unit (SRFU) is designed based on the concept of temporal data modeling, which utilizes the correlation between adjacent bands to fuse spectral and spatial features from PAN and LRHS images. In SRFU, a state memory interaction unit (SMIU) is constructed based on non-negative matrix factorization (NMF) to learn the global spatial-spectral dependency of PAN and HS images in the recurrent state space. Moreover, based on uncertainty theory, we define two spatial-spectral uncertainty-guided loss functions for the HS pansharpening task to train the model step by step, ensuring that the network can reconstruct more accurate spectral and spatial features. Extensive experiments on three widely used datasets demonstrate that, compared with some state-of-the-art (SOTA) methods, the proposed UMNet has achieved significant improvements in both spatial and spectral quality metrics.
Yong Yang 0001, Shuying Huang, Nayu Liu
AAAI4
2026 HyCoRA: Hyper-Contrastive Role-Adaptive Learning for Role-Playing
abstract
Multi-character role-playing aims to equip models with the capability to simulate diverse roles. Existing methods either use one shared parameterized module across all roles or assign a separate parameterized module to each role. However, the role-shared module may ignore distinct traits of each role, weakening personality learning, while the role-specific module may overlook shared traits across multiple roles, hindering commonality modeling. In this paper, we propose a novel HyCoRA: Hyper-Contrastive Role-Adaptive learning framework, which efficiently improves multi-character role-playing agents' ability by balancing the learning of distinct and shared traits. Specifically, we propose a Hyper-Half Low-Rank Adaptation structure, where one half is a role-specific module generated by a lightweight hyper-network, and the other half is a trainable role-shared module. The role-specific module is devised to represent distinct persona signatures, while the role-shared module serves to capture common traits. Moreover, to better reflect distinct personalities across different roles, we design a hyper-contrastive learning mechanism to help the hyper-network distinguish their unique characteristics. Extensive experimental results on both English and Chinese available benchmarks demonstrate the superiority of our framework. Further GPT-4 evaluations and visual analyses also verify the capability of HyCoRA to capture role characteristics.
Zhicong Lu, Yong Yang 0001, Nayu Liu
AAAI6
2026 FocalOrder: Focal Preference Optimization for Reading Order Detection
abstract
Fuyuan Liu, Dianyu Yu, He Ren, Nayu Liu, Xiaomian Kang, Delai Qiu, Fa Zhang, Genpeng Zhen, Shengping Liu, Liang Jiaen, Weihuang, Yining Wang, Junnan Zhu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Fuyuan Liu, Dianyu Yu, Nayu Liu, Xiaomian Kang, Delai Qiu, Genpeng Zhen, Shengping Liu, Jiaen Liang, Junnan Zhu
ACL (1)4
2026 Beyond Query Memorization: Large Language Model Routing with Query Decomposition and Historical Matching
abstract
Bo Lv, Jingbo Sun, Jianwei Lv, Chen Tang, Shaojie Zhang, Nayu Liu, Guoxin Yu, Zihao Li, Qichao Zhang, Dongbin Zhao, Ping Luo, Yue Yu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jingbo Sun 0001, Jianwei Lv, Nayu Liu, Guoxin Yu, Dongbin Zhao, Ping Luo 0002, Yue Yu 0001
ACL (1)6
2026 MIRE: A medical information enhanced framework for long-tail medical dialogue synthesis
abstract
In recent years, deep-learning-based approaches for medical dialogue generation have become the predominant paradigm. However, real-world medical dialogues often face data imbalance issues, especially long-tail distribution problems. The scarcity of training samples for low-resource diseases makes it challenging for language models to provide accurate and comprehensive diagnostic support. In this paper, we propose MIRE, a novel framework that leverages external medical knowledge of tail diseases and dialogue data of common diseases to guide large language models (LLMs) in generating synthetic dialogues for tail diseases. Specifically, MIRE retrieves and crawls medical information about tail diseases from multiple online sources, enhancing subtype coverage in the generated synthetic dialogues. Moreover, we introduce a style transfer mechanism that utilises rich style templates extracted from common disease conversations to guide LLMs in augmenting dialogues in low-resource domains, thereby narrowing the gap between synthetic and real human dialogues. To evaluate the effectiveness of our method in addressing the long-tail disease problems, we construct a long-tail medical dialogue dataset, named TailMed. Experimental results show that training the model with a mixture of synthetic dialogues and the original dataset significantly improves both automatic metrics and human evaluations. Specifically, the model trained on the MIRE-enhanced dataset outperforms the original by over 20% in average metrics for tail diseases. These results demonstrate the potential of MIRE to enhance clinical dialogue systems, enabling more equitable diagnostic assistance for rare and underrepresented diseases, and contributing to improved accessibility in intelligent healthcare applications.
Nayu Liu, Guoxin Yu, Xin Liu 0039, Riyan Zhang, Yue Yu 0001
Expert Syst. Appl.3
2025 Language Constrained Multimodal Hyper Adapter For Many-to-Many Multimodal Summarization
abstract
Multimodal summarization (MS) combines text and visuals to generate summaries.Recently, many-to-many multimodal summarization (M3S) garnered interest as it enables a unified model for multilingual and cross-lingual MS.Existing methods have made progress by facilitating the transfer of common multimodal summarization knowledge.While, prior M3S models that fully share parameters neglect the language-specific knowledge learning, where potential interference between languages may limit the flexible adaptation of MS modes across different language combinations and hinder further collaborative improvements in joint M3S training.Based on this observation, we propose Language Constrained Multimodal Hyper Adapter (LCMHA) for M3S.LCMHA integrates language-specific multimodal adapters into multilingual pre-trained backbones via a language constrained hypernetwork, enabling relaxed parameter sharing that enhances language-specific learning while preserving shared MS knowledge learning.In addition, a language-regularized hypernetwork is designed to balance intra-and inter-language learning, generating language-specific adaptation weights and enhancing the retention of distinct language features through the regularization of generated parameters.Experimental results on the M3Sum benchmark show LCMHA's effectiveness and scalability across multiple multilingual pre-trained backbones.
Nayu Liu, Fanglong Yao, Yong Yang 0001
ACL (1)1
2025 SARA: Salience-Aware Reinforced Adaptive Decoding for Large Language Models in Abstractive Summarization
abstract
Nayu Liu, Junnan Zhu, Yiming Ma, Zhicong Lu, Wenlei Xu, Yong Yang, Jiang Zhong, Kaiwen Wei. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Nayu Liu, Junnan Zhu, Zhicong Lu, Wenlei Xu, Yong Yang 0001, Kaiwen Wei
ACL (1)1
2025 U-MERE: Unconstrained Multimodal Entity and Relation Extraction with Collaborative Modeling and Order-Sensitive Optimization
abstract
Existing multimodal entity and relation extraction tasks primarily focus on text-to-text or text-to-visual entity relations, overlooking real-world complexities involving visual-to-text and visual-to-visual cases, thus failing to capture the richer semantic structures in complex cross-modal interactions. To address the limitations, we propose a new task, Unconstrained Multimodal Entity and Relation Extraction (U-MERE), which jointly extracts arbitrary visual and textual entities, and their relations from image-text pairs. To accomplish U-MERE, we construct UMERE-Bench, a benchmark with over 9,000 samples that comprehensively covers four cross-modal entity relation directions and three task settings. Given the difficulty of jointly modeling diverse directions of cross-modal entity relations, we introduce Collaborative Modeling and Order-Sensitive (CMOS), which collaboratively guides large vision-language models (LVLMs) to decompose task complexity and mitigates generation order bias from fixed target relation sequences. CMOS employs small models to generate candidate entities, guiding LVLMs to capture key information and jointly optimizes multiple feasible relation orderings to reduce order dependency. Additionally, we design a Multimodal Order-aware Matching (MOM) evaluation method to align predictions with ground truth for precise assessment. Experimental results reveal that current LVLMs show limited performance on U-MERE, underscoring its inherent challenges, while CMOS consistently achieves superior performance across multiple advanced LVLMs, demonstrating its effectiveness and generalization capability. The dataset and code will be available in https://github.com/jiaweidoris/U-MERE.
Li Jin 0001, Kaiwen Wei, Yuying Shang, Nayu Liu, Zhicong Lu, Qing Liu 0021, Linhao Zhang, Yanfeng Hu
ACM Multimedia5
2025 SpecEM: Training-Free LLM Ensembling via Iterative Drafting, Verification, and Online Feedback
abstract
Ensembles of generative large language models (LLMs) are a promising way to compensate for individual model limitations, integrating the strengths of different LLMs. Existing LLM ensemble methods, however, face limitations such as first-token delay and challenges in long-range semantic collaboration between models, Moreover, they typically assume equal voting weights for all models during ensemble, ignoring performance differences between models for a given task. In this work, we propose SpecEM, a training-free, plug-and-play LLM ensemble framework that dynamically adjusts each model's model contribution in real time based on task performance. Inspired by speculative decoding, SpecFuse iteratively performs drafting and verification, allowing models to collaborate semantically at the segment level for integrated output. Furthermore, we introduce an online feedback mechanism with multiplicative weight updates, where each model's voting weight is adjusted on-the-fly according to how often it "outperforms" others during verification stage, ensuring that stronger models exert greater influence on the ensemble during generation. Experimental results on five popular LLMs (ranging from 7B to 72B parameters) and six benchmark tasks, spanning instruction following, reasoning, commonsense, and general instruction response, demonstrate consistent performance improvements compared to state-of-the-art LLM ensemble methods.
Nayu Liu, Xin Liu 0039, Yue Yu 0001, Ping Luo 0002
NeurIPS2
2024 Video Event Extraction with Multi-View Interaction Knowledge Distillation
abstract
Video event extraction (VEE) aims to extract key events and generate the event arguments for their semantic roles from the video. Despite promising results have been achieved by existing methods, they still lack an elaborate learning strategy to adequately consider: (1) inter-object interaction, which reflects the relation between objects; (2) inter-modality interaction, which aligns the features from text and video modality. In this paper, we propose a Multi-view Interaction with knowledge Distillation (MID) framework to solve the above problems with the Knowledge Distillation (KD) mechanism. Specifically, we propose the self-Relational KD (self-RKD) to enhance the inter-object interaction, where the relation between objects is measured by distance metric, and the high-level relational knowledge from the deeper layer is taken as the guidance for boosting the shallow layer in the video encoder. Meanwhile, to improve the inter-modality interaction, the Layer-to-layer KD (LKD) is proposed, which integrates additional cross-modal supervisions (i.e., the results of cross-attention) with the textual supervising signal for training each transformer decoder layer. Extensive experiments show that without any additional parameters, MID achieves the state-of-the-art performance compared to other strong methods in VEE.
Kaiwen Wei, Runyan Du, Li Jin 0001, Jian Liu 0032, Jianhua Yin 0001, Linhao Zhang, Nayu Liu, Zhi Guo
AAAI8
2024 CAMEL: Capturing Metaphorical Alignment with Context Disentangling for Multimodal Emotion Recognition
abstract
Understanding the emotional polarity of multimodal content with metaphorical characteristics, such as memes, poses a significant challenge in Multimodal Emotion Recognition (MER). Previous MER researches have overlooked the phenomenon of metaphorical alignment in multimedia content, which involves non-literal associations between concepts to convey implicit emotional tones. Metaphor-agnostic MER methods may be misinformed by the isolated unimodal emotions, which are distinct from the real emotions blended in multimodal metaphors. Moreover, contextual semantics can further affect the emotions associated with similar metaphors, leading to the challenge of maintaining contextual compatibility. To address the issue of metaphorical alignment in MER, we propose to leverage a conditional generative approach for capturing metaphorical analogies. Our approach formulates schematic prompts and corresponding references based on theoretical foundations, which allows the model to better grasp metaphorical nuances. In order to maintain contextual sensitivity, we incorporate a disentangled contrastive matching mechanism, which undergoes curricular adjustment to regulate its intensity during the learning process. The automatic and human evaluation experiments on two benchmarks prove that, our model provides considerable and stable improvements in recognizing multimodal emotion with metaphor attributes.
Linhao Zhang, Li Jin 0001, Guangluan Xu, Xiaoyu Li 0004, Kaiwen Wei, Nayu Liu
AAAI7
2024 DuaPIN: Auxiliary task enhanced dual path interaction network for civil court view generation
Nayu Liu, Yiquan Wu 0001, Kaiwen Wei, Cunhang Fan
Knowl. Based Syst.1
2024 Multimodal Cross-Lingual Summarization for Videos: A Revisit in Knowledge Distillation Induced Triple-Stage Training Method
abstract
Multimodal summarization (MS) for videos aims to generate summaries from multi-source information (e.g., video and text transcript), showing promising progress recently. However, existing works are limited to monolingual scenarios, neglecting non-native viewers' needs to understand videos in other languages. It stimulates us to introduce multimodal cross-lingual summarization for videos (MCLS), which aims to generate cross-lingual summaries from multimodal input of videos. Considering the challenge of high annotation cost and resource constraints in MCLS, we propose a knowledge distillation (KD) induced triple-stage training method to assist MCLS by transferring knowledge from abundant monolingual MS data to those data with insufficient volumes. In the triple-stage training method, a video-guided dual fusion network (VDF) is designed as the backbone network to integrate multimodal and cross-lingual information through diverse fusion strategies in the encoder and decoder; What's more, we propose two cross-lingual knowledge distillation strategies: adaptive pooling distillation and language-adaptive warping distillation (LAWD), designed for encoder-level and vocab-level distillation objects to facilitate effective knowledge transfer across cross-lingual sequences of varying lengths between MS and MCLS models. Specifically, to tackle lingual sequences of varying lengths between MS and MCLS models. Specifically, to tackle the challenge of unequal length of parallel cross-language sequences in KD, LAWD can directly conduct cross-language distillation while keeping the language feature shape unchanged to reduce potential information loss. We meticulously annotated the How2-MCLS dataset based on the How2 dataset to simulate MCLS scenarios. Experimental results show that the proposed method achieves competitive performance compared to strong baselines, and can bring substantial performance improvements to MCLS models by transferring knowledge from the MS model.
Nayu Liu, Kaiwen Wei, Yong Yang 0001, Jianhua Tao 0001, Xian Sun 0001, Fanglong Yao, Li Jin 0001, Zhao Lv, Cunhang Fan
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 M2DCapsN: Multimodal, Multichannel, and Dual-Step Capsule Network for Natural Language Moment Localization
abstract
Natural language moment localization aims to localize the target moment that matches a given natural language query in an untrimmed video. The key to this challenging task is to capture fine-grained video-language correlations to establish the alignment between the query and target moment. Most existing works establish a single-pass interaction schema to capture correlations between queries and moments. Considering the complex feature space of lengthy video and diverse information between frames, the weight distribution of information interaction flow is prone to dispersion or misalignment, which leads to redundant information flow affecting the final prediction. We address this issue by proposing a capsule-based approach to model the query-video interactions, termed the Multimodal, Multichannel, and Dual-step Capsule Network ( [Formula: see text]DCapsN), which is derived from the intuition that "multiple people viewing multiple times is better than one person viewing one time." First, we introduce a multimodal capsule network, replacing the single-pass interaction schema of "one person viewing one time" with the iterative interaction schema of "one person viewing multiple times," which cyclically updates cross-modal interactions and modifies potential redundant interactions via its routing-by-agreement. Then, considering that the conventional routing mechanism only learns a single iterative interaction schema, we further propose a multichannel dynamic routing mechanism to learn multiple iterative interaction schemas, where each channel performs independent routing iteration to collectively capture cross-modal correlations from multiple subspaces, that is, "multiple people viewing." Moreover, we design a dual-step capsule network structure based on the multimodal, multichannel capsule network, bringing together the query and query-guided key moments to jointly enhance the original video, so as to select the target moments according to the enhanced part. Experimental results on three public datasets demonstrate the superiority of our approach in comparison with state-of-the-art methods, and comprehensive ablation and visualization analysis validate the effectiveness of each component of the proposed model.
Nayu Liu, Xian Sun 0001, Fanglong Yao, Guangluan Xu, Kun Fu 0001
IEEE Trans. Neural Networks Learn. Syst.1
2023 TOT:Topology-Aware Optimal Transport for Multimodal Hate Detection
abstract
Multimodal hate detection, which aims to identify the harmful content online such as memes, is crucial for building a wholesome internet environment. Previous work has made enlightening exploration in detecting explicit hate remarks. However, most of their approaches neglect the analysis of implicit harm, which is particularly challenging as explicit text markers and demographic visual cues are often twisted or missing. The leveraged cross-modal attention mechanisms also suffer from the distributional modality gap and lack logical interpretability. To address these semantic gap issues, we propose TOT: a topology-aware optimal transport framework to decipher the implicit harm in memes scenario, which formulates the cross-modal aligning problem as solutions for optimal transportation plans. Specifically, we leverage an optimal transport kernel method to capture complementary information from multiple modalities. The kernel embedding provides a non-linear transformation ability to reproduce a kernel Hilbert space (RKHS), which reflects significance for eliminating the distributional modality gap. Moreover, we perceive the topology information based on aligned representations to conduct bipartite graph path reasoning. The newly achieved state-of-the-art performance on two publicly available benchmark datasets, together with further visual analysis, demonstrate the superiority of TOT in capturing implicit cross-modal alignment.
Linhao Zhang, Li Jin 0001, Xian Sun 0001, Guangluan Xu, Zequn Zhang, Xiaoyu Li 0004, Nayu Liu, Qing Liu 0021, Shiyao Yan
AAAI7
2023 RingMo-Sense: Remote Sensing Foundation Model for Spatiotemporal Prediction via Spatiotemporal Evolution Disentangling
abstract
Remote sensing spatiotemporal prediction aims to infer future trends from historical spatiotemporal data, e.g., videos and time series images, has a broad application prospect in many fields. The foundation model is a promising research direction for spatiotemporal information mining because of its robust feature extraction capability, and has made rapid progress in natural scenes. Nevertheless, due to the spatially multi-scale and temporally multi-scale properties in remote sensing data, these methods still encounter bottlenecks when applied to remote sensing. Therefore, we propose a foundation model for remote sensing spatiotemporal prediction via spatiotemporal evolution decoupling, abbreviated as RingMo-Sense. Considering spatial affinity, temporal continuity, and spatiotemporal interaction, we construct spatial, temporal, and spatiotemporal triple-branch prediction networks. Specifically, we use parameter-sharing and progressive joint training strategies to achieve stable long-range prediction and parameter reduction simultaneously. In addition, we build a remote sensing spatiotemporal dataset by collecting various remote sensing videos and time series images. The experimental results on six downstream spatiotemporal tasks demonstrate that the proposed model yields competitive performance.
Fanglong Yao, Wanxuan Lu, Heming Yang 0003, Liangyu Xu, Leiyi Hu, Nayu Liu, Chubo Deng, Deke Tang, Changshuo Chen, Xian Sun 0001, Kun Fu 0001
IEEE Trans. Geosci. Remote. Sens.8
2023 GAL: Graph-Induced Adaptive Learning for Weakly Supervised 3D Object Detection
abstract
Weakly Supervised 3D Object Detection (WS3DOD) aims to perform 3D object detection with little reliance on 3D labels, which greatly reduces the cost of 3D annotations. In recent literature, the pseudo-label-based approach brings impressive performance, which generates 3D pseudo-labels from 2D bounding boxes. Despite their success, two key issues remain unresolved that reduce the quality of 3D pseudo-labels: 1) the existing local object locating algorithm can not capture complete clusters of points globally, and 2) the existing algorithm can not capture sparse points caused by the unevenly distributed points obtained by LiDAR cameras. Hence, we propose GAL, a Graph-induced Adaptive Learning algorithm, to generate 3D pseudo-labels. First, we propose the Cluster Locating algorithm based on the Minimum Spanning Tree (MST) to globally locate the objects, which can leverage the characteristic that points inside an object are compact while points between objects are discrete. Second, we propose a density-guided adaptive learning algorithm to optimise the Cluster Locating algorithm, named Cuboid Drift. Cuboid Drift considers the inhomogeneous distribution of reflected points on different reflective surfaces of LiDAR imaging. Finally, 3D pseudo-labels generated by GAL are leveraged to train 3D detectors. Extensive experiments on the challenging KITTI and DAIR-V2X-V dataset demonstrate that GAL without 3D labels can be comparable with strongly supervised approaches and outperforms the previous state-of-the-art WS3DOD methods. Moreover, our method saves 88% of the time spent on pseudo-label generation.
Dongshuo Yin, Nayu Liu, Fanglong Yao, Qibin He 0001, Shiyao Yan, Xian Sun 0001
IEEE Trans. Intell. Transp. Syst.3
2023 Abstractive Summarization for Video: A Revisit in Multistage Fusion Network With Forget Gate
abstract
Multimodal abstractive summarization for videos is an emerging task that aims to generate a summary from multi-source information (i.e., video, audio transcript). The challenge is how to merge multimodal long sequences to capture rich semantic information without allowing possible noise from either lengthy modal sequence to degrade the other modality and thus hurt the entire model. To address the issues, we propose amultistagefusion network withforgetgate (MFFG), which selectively integrates multi-source information through the cross-fusion in encoding and hierarchical fusion in decoding between modalities, and design a fusion forget gate module to suppress the potential multimodal noise flow of multi-source long sequence. Meanwhile, considering that the source text in this task is lengthy and has the same distribution as the output summary text, we inherit the partial structure of the MFFG model and again propose its variant, single-stage fusion network with forget gate (SFFG), which simplifies the fusion schema, and leverages the long source text to enhance the representation of the target summary. Experimental results on How2 dataset and How2-300 dataset demonstrate the superiority of the two multimodal fusion methods. Further, we provide a version of ASR transcription data of How2 dataset to evaluate model performance under noisy scenarios, and experimental results show obvious advantages of our proposed models over prior systems.
Nayu Liu, Xian Sun 0001, Fanglong Yao, Guangluan Xu, Kun Fu 0001
IEEE Trans. Multim.1
2022 PolygonE: Modeling N-ary Relational Data as Gyro-Polygons in Hyperbolic Space
abstract
N-ary relational knowledge base (KBs) embedding aims to map binary and beyond-binary facts into low-dimensional vector space simultaneously. Existing approaches typically decompose n-ary relational facts into subtuples (entity pairs, triples or quintuples, etc.), and they generally model n-ary relational KBs in Euclidean space. However, n-ary relational facts are semantically and structurally intact, decomposition leads to the loss of global information and undermines the semantical and structural integrity. Moreover, compared to the binary relational KBs, n-ary ones are characterized by more abundant and complicated hierarchy structures, which could not be well expressed in Euclidean space. To address the issues, we propose a gyro-polygon embedding approach to realize n-ary fact integrity keeping and hierarchy capturing, termed as PolygonE. Specifically, n-ary relational facts are modeled as gyro-polygons in the hyperbolic space, where we denote entities in facts as vertexes of gyro-polygons and relations as entity translocation operations. Importantly, we design a fact plausibility measuring strategy based on the vertex-gyrocentroid geodesic to optimize the relation-adjusted gyro-polygon. Extensive experiments demonstrate that PolygonE shows SOTA performance on all benchmark datasets, generalizability to binary data, and applicability to arbitrary arity fact. Finally, we also visualize the embedding to help comprehend PolygonE's awareness of hierarchies.
Shiyao Yan, Zequn Zhang, Xian Sun 0001, Guangluan Xu, Shuchao Li, Qing Liu 0021, Nayu Liu, Shensi Wang
AAAI7
2022 Assist Non-native Viewers: Multimodal Cross-Lingual Summarization for How2 Videos
abstract
Multimodal summarization for videos aims to generate summaries from multi-source information (videos, audio transcripts), which has achieved promising progress.However, existing works are restricted to monolingual video scenarios, ignoring the demands of non-native video viewers to understand the cross-language videos in practical applications.It stimulates us to propose a new task, named Multimodal Cross-Lingual Summarization for videos (MCLS), which aims to generate cross-lingual summaries from multimodal inputs of videos.First, to make it applicable to MCLS scenarios, we conduct a Video-guided Dual Fusion network (VDF) that integrates multimodal and cross-lingual information via diverse fusion strategies at both encoder and decoder.Moreover, to alleviate the problem of high annotation costs and limited resources in MCLS, we propose a triple-stage training framework to assist MCLS by transferring the knowledge from monolingual multimodal summarization data, which includes: 1) multimodal summarization on sufficient prevalent language videos with a VDF model; 2) knowledge distillation (KD) guided adjustment on bilingual transcripts; 3) multimodal summarization for cross-lingual videos with a KD induced VDF model.Experiment results on the reorganized How2 dataset show that the VDF model alone outperforms previous methods for multimodal summarization, and the performance further improves by a large margin via the proposed triple-stage training framework. * Equal contribution. † Corresponding author.Portuguese (Pt) Transcript: vamos falar hoje sobre o solo.em primeiro lugar, precisamos de uma grande quan dade de solo bom para transplantes na primavera.ela vai adicionar partes iguais de musgo de turfa e composto de jardinagem que extraímos do nosso sistema interno de compostagem, e então um agregado orgânico, uma pedra chamada perlite, que serve para adicionar volume e aumentar a capacidade de retenção de água e de aeração de sua mistura... English (En) Summary: mix sterile soil for plan ng greens in trays to keep in a hoop house.learn to mix soil for growing greens from an organic farmer in this free gardening video.
Nayu Liu, Kaiwen Wei, Xian Sun 0001, Fanglong Yao, Li Jin 0001, Zhi Guo, Guangluan Xu
EMNLP1
2022 Cross-Modal Remote Sensing Image Retrieval Via Intra- and Inter-Modal Feature Matching
abstract
With the development of remote sensing (RS) acquisition technology, a mass of RS images have been produced, which brings challenges to the traditional manual retrieval methods and gives birth to the automatic RS image retrieval methods. Cross-modal RS image retrieval allows the usage of text and other modalities to retrieve RS images. For its flexible and convenient advantages, it has become a research hotspot. However, cross-modal RS image retrieval encounters the information asymmetry between modalities, i.e., RS images possess multi-scale, multi-objective properties and own rich information. At the same time, the query text is usually short and with less information. To solve the issues above, a cross-modal feature matching network is proposed to learn the feature fusion intra-modalities and the feature association inter-modalities to avoid the poor retrieval performance caused by the information asymmetry. Specifically, for the feature fusion intra-modalities, relying on the powerful feature representation ability of graph network, text and RS image graph modules are designed to fuse the intra-modal features. In terms of the feature correlation between modalities, RS image-text association module is created to attend the parts in text related to RS images and vice versa. Extended experiments on two public standard datasets verify the effectiveness of the proposed model.
Fanglong Yao, Nayu Liu, Peiguang Li, Dongshuo Yin, Xian Sun 0001
IGARSS2
2021 D-MmT: A concise decoder-only multi-modal transformer for abstractive summarization in videos
Nayu Liu, Xian Sun 0001, Wenkai Zhang 0002, Guangluan Xu
Neurocomputing1
2020 Multistage Fusion with Forget Gate for Multimodal Summarization in Open-Domain Videos
abstract
Multimodal summarization for open-domain videos is an emerging task, aiming to generate a summary from multisource information (video, audio, transcript).Despite the success of recent multiencoder-decoder frameworks on this task, existing methods lack finegrained multimodality interactions of multisource inputs.Besides, unlike other multimodal tasks, this task has longer multimodal sequences with more redundancy and noise.To address these two issues, we propose a multistage fusion network with the fusion forget gate module, which builds upon this approach by modeling fine-grained interactions between the multisource modalities through a multistep fusion schema and controlling the flow of redundant information between multimodal long sequences via a forgetting module.Experimental results on the How2 dataset show that our proposed model achieves a new state-of-the-art performance.Comprehensive analysis empirically verifies the effectiveness of our fusion schema and forgetting module on multiple encoder-decoder architectures.Specially, when using high noise ASR transcripts (W ER>30%), our model still achieves performance close to the ground-truth transcript model, which reduces manual annotation cost.
Nayu Liu, Xian Sun 0001, Wenkai Zhang 0002, Guangluan Xu
EMNLP (1)1