EDBT 2026 Demo / reviewers in the wild / expert
Zhiyuan Ma 0005
dblp:138/5978-5
· DBLP profile ↗
22ranked-venue papers
11as first author
22since 2021 · last 2026
0009-0006-3756-5621ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 8 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 6 first-author · 11 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OPERA: A Reinforcement Learning-Enhanced Orchestrated Planner-Executor Architecture for Reasoning-Oriented Multi-Hop RetrievalabstractRecent advances in large language models (LLMs) and dense retrievers have driven significant progress in retrieval-augmented generation (RAG). However, existing approaches face significant challenges in complex reasoning-oriented multi-hop retrieval tasks: 1) Ineffective reasoning-oriented planning: Prior methods struggle to generate robust multi-step plans for complex queries, as rule-based decomposers perform poorly on out-of-template questions. 2) Suboptimal reasoning-driven retrieval: Related methods employ limited query reformulation, leading to iterative retrieval loops that often fail to locate golden documents. 3) Insufficient reasoning-guided filtering: Prevailing methods lack the fine-grained reasoning to effectively filter salient information from noisy results, hindering utilization of retrieved knowledge. Fundamentally, these limitations all stem from the weak coupling between retrieval and reasoning in current RAG architectures. We introduce the Orchestrated Planner-Executor Reasoning Architecture (OPERA), a novel reasoning-driven retrieval framework. OPERA's Goal Planning Module (GPM) decomposes questions into sub-goals, which are executed by a Reason-Execute Module (REM) with specialized components for precise reasoning and effective retrieval. To train OPERA, we propose Multi-Agents Progressive Group Relative Policy Optimization (MAPGRPO), a novel variant of GRPO. Experiments on complex multi-hop benchmarks show OPERA's superior performance, validating both the MAPGRPO method and OPERA's design. Yanbing Liu 0007, Fangfang Yuan, Cong Cao 0001, Youbang Sun, Weizhuo Chen, Jianjun Li 0010, Zhiyuan Ma 0005 |
AAAI | 9 |
| 2026 | I2E: From Image Pixels to Actionable Interactive Environments for Text-Guided Image EditingabstractJinghan Yu, Junhao Xiao, Chenyu Zhu, Jiaming Li, Jia Li, HanMing Deng, Xirui Wang, Guoli Jia, Jianjun Li, Xiang Bai, Bowen Zhou, Zhiyuan Ma. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jinghan Yu, Chenyu Zhu, HanMing Deng, Xirui Wang, Guoli Jia, Jianjun Li 0010, Xiang Bai, Bowen Zhou 0002, Zhiyuan Ma 0005 |
ACL (1) | 12 |
| 2026 | UTAG: Leveraging LLM as a Unified Embedding Generator for Text-Attributed Graphs
Mingqian Ding, Jianjun Li 0010, Zhiyuan Ma 0005, Wenqi Yang |
WWW | 3 |
| 2026 | MIDE: Multimodal Dialogue Emotion Recognition via Mutual Information Enhancement and Dynamic Modality Selection
Zhibo Zhang 0009, Jianjun Li 0010, Zhiyuan Ma 0005 |
WWW | 3 |
| 2026 | VC-VTON: Toward Across-View and Multi-Posture-Driven Virtual Try-On via Spatiotemporal-Aware View-Consistency TrainingabstractVirtual try-on (VTON) aims to synthesize specific fashion images dressed in given garments, which possesses great potential in real-world scenarios. Existing methods generally stand on the shoulder of the single-view VTON to train a warping model and then fit the given garments onto the human body under a fixed posture and viewpoint, which often fails to preserve the consistent garment characteristics in across-view and multi-pose guided try-on scenarios due to the lack of both across-view data and effective view consistency training. To alleviate this dilemma, we propose a fresh view consistency-driven VTON task (VC-VTON) and release a multi-view virtual try-on dataset with complete annotation (e.g., viewpoint, text, posture, parsing maps, etc.) to encourage across-view training scenarios. Based on this hard-won dataset, we further propose VC-TwinNet, a Twin-UNet baseline based on spatiotemporal-aware View Consistency training, designed specifically for the challenging task. Specifically, to enable view-aware denoising and sparse-to-continuous view generalization, we introduce RoPE and circle embedding to represent the relative and continuous position relation across viewpoints, serving to distinguish their outfitting appearance and warping states. Afterwards, to implicitly learn the interactions across views under given multiple posture conditions, we further contribute a spatiotemporal-aware view attention module to capture the spatial and temporal details for across-view training. Moreover, we utilize an across-view consistency loss to supervise the model training, to ultimately improve the performance of our VC-VTON. Extensive experiments demonstrate the superiority of our approach and state-of-the-art results on various evaluations without declining single-view performance. And as for practicality and timeliness, our proposed components are essentially plug-and-play and remain effective in the new DiT-centered paradigm. Zhiyuan Ma 0005, Jiabao Wei, Zhihan Cai, Chundi Yang, Ermo Hua, Shulei Xie, Jianjun Li 0010, Bowen Zhou 0002 |
IEEE Trans. Image Process. | 1 |
| 2025 | DreamAlign: Dynamic Text-to-3D Optimization with Human Preference AlignmentabstractRecent years have witnessed the remarkable success of Text-to-3D generation, particularly with the rise of mainstream conditional diffusion models (DMs). Though achieving substantial progress, existing methods still face a knotty "human preference" dilemma, that is the 3D contents generated by the models often deviate greatly from the desired effects (e.g., perspective, aesthetics, shading, appearance, etc.) due to the lack of attention to human preferences. To mitigate the limitation of data deficiency and enable human preference learning, we first elaborately curate the HP3D, a text-to-3D dataset with expert preference annotations which is initally captioned by the multimodal large model LLava and then refined by human expert. Based on such a brand-new HP3D, we further propose DreamAlign, a reward-free method that does not require designing any complex reward models whereas only by introducing a light-weight lora adapter and then designing a novel direct 3D preference optimization (D-3DPO) algorithm for training. Moreover, in the stage of text-to-3D we design an additional Preference Contrastive Feedback training for score distillation sampling, which enables the generated 3D objects to align the human preferences (e.g., aesthetics, material, etc.). Extensive experiments demonstrate that DreamAlign consistently achieves state-of-the-art performance on generative effects and human preference alignment across various benchmark evaluations. Gaofeng Liu, Zhiyuan Ma 0005 |
AAAI | 2 |
| 2025 | Retrieval-Augmented Visual Question Answering via Built-in Autoregressive Search EnginesabstractRetrieval-augmented generation (RAG) has emerged to address the knowledge-intensive visual question answering (VQA) task. Current methods mainly employ separate retrieval and generation modules to acquire external knowledge and generate answers, respectively. We propose ReAuSE, an alternative to the previous RAG model for the knowledge-based VQA task, which seamlessly integrates knowledge retriever into the generative multi-modal large language model, serving as a built-in search engine. Specifically, our model functions both as a generative retriever and an accurate answer generator. It not only helps retrieve documents from the knowledge base by producing identifier for each document, but it also answers visual questions based on the retrieved documents. Furthermore, we also propose a reinforced retrieval calibration module from relevance feedback to improve retrieval performance and align with the preferences for accurate answer generation. Extensive experiments on two representative OKVQA and A-OKVQA datasets demonstrate significant improvements ranging from 2.9% to 9.6% across all evaluation metrics when compared to strong baselines. Xinwei Long, Zhiyuan Ma 0005, Ermo Hua, Biqing Qi, Bowen Zhou 0002 |
AAAI | 2 |
| 2025 | SAKR-Edit: Scene-Aware Knowledge Reasoning for Text-to-Image EditingabstractImage editing requires semantically modifying specific regions according to user instructions while preserving overall visual coherence. Although diffusion models have shown remarkable progress in image generation, their application to editing tasks faces two critical limitations: (1) insufficient understanding of editing objectives often leads to inconsistencies in style, attribute, or texture between generated content and background regions, and (2) over-reliance on ambiguous textual prompts that frequently lack crucial details, resulting in suboptimal edits. To address these challenges, we propose SAKR-Edit, a novel framework that enhances editing quality and controllability through Scene-Aware Knowledge Reasoning. Specifically, our approach introduces a scene-aware knowledge reasoning module that combines large language models (LLMs) with vision-language models (e.g., BLIP-2) to integrate global and local semantic information for improved instruction comprehension. The system employs chain-of-thought reasoning and contextual learning to parse instructions, infer implicit editing intentions, and supplement missing details, thereby improving editing precision. Additionally, we construct SSUD, a structured scene understanding dataset for evaluating editing models in real-world scenarios. Extensive experiments demonstrate that SAKR-Edit outperforms existing methods in image realism, style consistency, and structural integrity, while showing robust stability and adaptability in real-world applications. Our code and dataset are released at https://github.com/SAKR-Edit/sakr-edit.github.io. Jianjun Li 0010, Zhiyuan Ma 0005, Ruixia Bai |
ACM Multimedia | 3 |
| 2025 | TTRL: Test-Time Reinforcement LearningabstractThis paper investigates Reinforcement Learning (RL) on data without explicit labels for reasoning tasks in Large Language Models (LLMs). The core challenge of the problem is reward estimation during inference while not having access to ground-truth information. While this setting appears elusive, we find that common practices in Test-Time Scaling (TTS), such as majority voting, yield surprisingly effective rewards suitable for driving RL training. In this work, we introduce Test-Time Reinforcement Learning (TTRL), a novel method for training LLMs using RL on unlabeled data. TTRL enables self-evolution of LLMs by utilizing the priors in the pre-trained models. Our experiments demonstrate that TTRL consistently improves performance across a variety of tasks and models. Notably, TTRL boosts the pass@1 performance of Qwen-2.5-Math-7B by approximately 211% on the AIME 2024 with only unlabeled test data. Furthermore, although TTRL is only supervised by the Maj@N metric, TTRL has demonstrated performance to consistently surpass the upper limit of the initial model, and approach the performance of models trained directly on test data with ground-truth labels. Our experimental findings validate the general effectiveness of TTRL across various tasks and highlight TTRL's potential for broader tasks and domains. Yuxin Zuo, Shang Qu, Ganqu Cui, Xuekai Zhu, Haozhan Li, Xinwei Long, Ermo Hua, Biqing Qi, Youbang Sun, Zhiyuan Ma 0005, Lifan Yuan, Ning Ding 0002, Bowen Zhou 0002 |
NeurIPS | 13 |
| 2025 | Efficient Diffusion Models: A Comprehensive Survey From Principles to PracticesabstractAs one of the most popular and sought-after generative models in recent years, diffusion models have sparked the interests of many researchers and steadily shown excellent advantage in various generative tasks such as image synthesis, video generation, bioinformatics engineering, 3D scene rendering and multimodal generation, relying on their dense theoretical principles and reliable application practices. The remarkable success of these recent efforts on diffusion models comes largely from progressive design principles and efficient architecture, training, inference, and deployment methodologies. However, there has not been a comprehensive and in-depth review to summarize these principles and practices to help the rapid understanding and application of diffusion models. In this survey, we provide a new efficiency-oriented perspective on these existing efforts, which mainly focuses on the profound principles and efficient practices in architecture designs, model training, fast inference and reliable deployment, to guide further theoretical research, algorithm migration and model application for new scenarios in a reader-friendly way. Zhiyuan Ma 0005, Yuzhu Zhang, Guoli Jia, Yichao Ma, Gaofeng Liu, Ning Ding 0002, Jianjun Li 0010, Bowen Zhou 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Generative Multi-Modal Knowledge Retrieval with Large Language ModelsabstractKnowledge retrieval with multi-modal queries plays a crucial role in supporting knowledge-intensive multi-modal applications. However, existing methods face challenges in terms of their effectiveness and training efficiency, especially when it comes to training and integrating multiple retrievers to handle multi-modal queries. In this paper, we propose an innovative end-to-end generative framework for multi-modal knowledge retrieval. Our framework takes advantage of the fact that large language models (LLMs) can effectively serve as virtual knowledge bases, even when trained with limited data. We retrieve knowledge via a two-step process: 1) generating knowledge clues related to the queries, and 2) obtaining the relevant document by searching databases using the knowledge clue. In particular, we first introduce an object-aware prefix-tuning technique to guide multi-grained visual learning. Then, we align multi-grained visual features into the textual feature space of the LLM, employing the LLM to capture cross-modal interactions. Subsequently, we construct instruction data with a unified format for model training. Finally, we propose the knowledge-guided generation strategy to impose prior constraints in the decoding steps, thereby promoting the generation of distinctive knowledge clues. Through experiments conducted on three benchmarks, we demonstrate significant improvements ranging from 3.0% to 14.6% across all evaluation metrics when compared to strong baselines. Xinwei Long, Jiali Zeng, Fandong Meng, Zhiyuan Ma 0005, Bowen Zhou 0002, Jie Zhou 0016 |
AAAI | 4 |
| 2024 | AdapEdit: Spatio-Temporal Guided Adaptive Editing Algorithm for Text-Based Continuity-Sensitive Image EditingabstractWith the great success of text-conditioned diffusion models in creative text-to-image generation, various text-driven image editing approaches have attracted the attentions of many researchers. However, previous works mainly focus on discreteness-sensitive instructions such as adding, removing or replacing specific objects, background elements or global styles (i.e., “hard editing”), while generally ignoring subject-binding but semantically fine-changing continuity-sensitive instructions such as actions, poses or adjectives, and so on (i.e., “soft editing”), which hampers generative AI from generating user-customized visual contents. To mitigate this predicament, we propose a spatio-temporal guided adaptive editing algorithm AdapEdit, which realizes adaptive image editing by introducing a soft-attention strategy to dynamically vary the guiding degree from the editing conditions to visual pixels from both temporal and spatial perspectives. Note our approach has a significant advantage in preserving model priors and does not require model training, fine-tuning, extra data, or optimization. We present our results over a wide variety of raw images and editing instructions, demonstrating competitive performance and showing it significantly outperforms the previous approaches. Code is available: https://github.com/AnonymousPony/adap-edit. Zhiyuan Ma 0005, Guoli Jia, Bowen Zhou 0002 |
AAAI | 1 |
| 2024 | LMD: Faster Image Reconstruction with Latent Masking DiffusionabstractAs a class of fruitful approaches, diffusion probabilistic models (DPMs) have shown excellent advantages in high-resolution image reconstruction. On the other hand, masked autoencoders (MAEs), as popular self-supervised vision learners, have demonstrated simpler and more effective image reconstruction and transfer capabilities on downstream tasks. However, they all require extremely high training costs, either due to inherent high temporal-dependence (i.e., excessively long diffusion steps) or due to artificially low spatial-dependence (i.e., human-formulated high mask ratio, such as 0.75). To the end, this paper presents LMD, a faster image reconstruction framework with Latent Masking Diffusion. First, we propose to project and reconstruct images in latent space through a pre-trained variational autoencoder, which is theoretically more efficient than in the pixel-based space. Then, we combine the advantages of MAEs and DPMs to design a progressive masking diffusion model, which gradually increases the masking proportion by three different schedulers and reconstructs the latent features from simple to difficult, without sequentially performing denoising diffusion as in DPMs or using fixed high masking ratio as in MAEs, so as to alleviate the high training time-consumption predicament. Our approach allows for learning high-capacity models and accelerate their training (by 3x or more) and barely reduces the original accuracy. Inference speed in downstream tasks also significantly outperforms the previous approaches. Zhiyuan Ma 0005, Zhihuan Yu, Jianjun Li 0010, Bowen Zhou 0002 |
AAAI | 1 |
| 2024 | Safe-SD: Safe and Traceable Stable Diffusion with Text Prompt Trigger for Invisible Generative WatermarkingabstractRecently, stable diffusion (SD) models have typically flourished in the field of image synthesis and personalized editing, with a range of photorealistic and unprecedented images being successfully generated. As a result, widespread interest has been ignited to develop and use various SD-based tools for visual content creation. However, the exposure of AI-created content on public platforms could raise both legal and ethical risks. In this regard, the traditional methods of adding watermarks to the already generated images (i.e. post-processing) may face a dilemma (e.g., being erased or modified) in terms of copyright protection and content monitoring, since the powerful image inversion and text-to-image editing techniques have been widely explored in SD-based methods. In this work, we propose a Safe and high-traceable Stable Diffusion framework (namely Safe-SD) to adaptively implant the graphical watermarks (e.g., QR code) into the imperceptible structure-related pixels during the generative diffusion process for supporting text-driven invisible watermarking and detection. Different from the previous high-cost injection-then-detection training framework, we design a simple and unified architecture, which makes it possible to simultaneously train watermark injection and detection in a single network, greatly improving the efficiency and convenience of use. Moreover, to further support text-driven generative watermarking and deeply explore its robustness and high-traceability, we elaborately design a λ-sampling and λ-encryption algorithm to fine-tune a latent diffuser wrapped by a VAE for balancing high-fidelity image synthesis and high-traceable watermark detection. We present our quantitative and qualitative results on two representative datasets LSUN, COCO and FFHQ, demonstrating state-of-the-art performance of Safe-SD and showing it significantly outperforms the previous approaches. Zhiyuan Ma 0005, Guoli Jia, Biqing Qi, Bowen Zhou 0002 |
ACM Multimedia | 1 |
| 2024 | Neural Residual Diffusion Models for Deep Scalable Vision GenerationabstractThe most advanced diffusion models have recently adopted increasingly deep stacked networks (e.g., U-Net or Transformer) to promote the generative emergence capabilities of vision generation models similar to large language models (LLMs). However, progressively deeper stacked networks will intuitively cause numerical propagation errors and reduce noisy prediction capabilities on generative data, which hinders massively deep scalable training of vision generation models. In this paper, we first uncover the nature that neural networks being able to effectively perform generative denoising lies in the fact that the intrinsic residual unit has consistent dynamic property with the input signal's reverse diffusion process, thus supporting excellent generative abilities.
Afterwards, we stand on the shoulders of two common types of deep stacked networks to propose a unified and massively scalable Neural Residual Diffusion Models framework (Neural-RDM for short), which is a simple yet meaningful change to the common architecture of deep generative networks by introducing a series of learnable gated residual parameters that conform to the generative dynamics. Experimental results on various generative tasks show that the proposed neural residual models obtain state-of-the-art scores on image's and video's generative benchmarks. Rigorous theoretical proofs and extensive experiments also demonstrate the advantages of this simple gated residual mechanism consistent with dynamic modeling in improving the fidelity and consistency of generated content and supporting large-scale scalable training. Zhiyuan Ma 0005, Biqing Qi, Bowen Zhou 0002 |
NeurIPS | 1 |
| 2024 | Exploring Adversarial Robustness of Deep State Space ModelsabstractDeep State Space Models (SSMs) have proven effective in numerous task scenarios but face significant security challenges due to Adversarial Perturbations (APs) in real-world deployments. Adversarial Training (AT) is a mainstream approach to enhancing Adversarial Robustness (AR) and has been validated on various traditional DNN architectures. However, its effectiveness in improving the AR of SSMs remains unclear.
While many enhancements in SSM components, such as integrating Attention mechanisms and expanding to data-dependent SSM parameterizations, have brought significant gains in Standard Training (ST) settings, their potential benefits in AT remain unexplored. To investigate this, we evaluate existing structural variants of SSMs with AT to assess their AR performance. We observe that pure SSM structures struggle to benefit from AT, whereas incorporating Attention yields a markedly better trade-off between robustness and generalization for SSMs in AT compared to other components. Nonetheless, the integration of Attention also leads to Robust Overfitting (RO) issues.
To understand these phenomena, we empirically and theoretically analyze the output error of SSMs under AP. We find that fixed-parameterized SSMs have output error bounds strictly related to their parameters, limiting their AT benefits, while input-dependent SSMs may face the problem of error explosion. Furthermore, we show that the Attention component effectively scales the output error of SSMs during training, enabling them to benefit more from AT, but at the cost of introducing RO due to its high model complexity.
Inspired by this, we propose a simple and effective Adaptive Scaling (AdS) mechanism that brings AT performance close to Attention-integrated SSMs without introducing the issue of RO. Biqing Qi, Yiang Luo, Junqi Gao, Pengfei Li 0011, Zhiyuan Ma 0005, Bowen Zhou 0002 |
NeurIPS | 6 |
| 2024 | UltraMedical: Building Specialized Generalists in BiomedicineabstractLarge Language Models (LLMs) have demonstrated remarkable capabilities across various domains and are moving towards more specialized areas. Recent advanced proprietary models such as GPT-4 and Gemini have achieved significant advancements in biomedicine, which have also raised privacy and security challenges. The construction of specialized generalists hinges largely on high-quality datasets, enhanced by techniques like supervised fine-tuning and reinforcement learning from human or AI feedback, and direct preference optimization. However, these leading technologies (e.g., preference learning) are still significantly limited in the open source community due to the scarcity of specialized data. In this paper, we present the UltraMedical collections, which consist of high-quality manual and synthetic datasets in the biomedicine domain, featuring preference annotations across multiple advanced LLMs. By utilizing these datasets, we fine-tune a suite of specialized medical models based on Llama-3 series, demonstrating breathtaking capabilities across various medical benchmarks. Moreover, we develop powerful reward models skilled in biomedical and general reward benchmark, enhancing further online preference learning within the biomedical LLM community. Sihang Zeng, Ermo Hua, Ning Ding 0002, Zhang-Ren Chen, Zhiyuan Ma 0005, Haoxin Li, Ganqu Cui, Biqing Qi, Xuekai Zhu, Xingtai Lv, Jinfang Hu, Zhiyuan Liu 0001, Bowen Zhou 0002 |
NeurIPS | 6 |
| 2023 | HybridPrompt: Bridging Language Models and Human Priors in Prompt Tuning for Visual Question AnsweringabstractVisual Question Answering (VQA) aims to answer the natural language question about a given image by understanding multimodal content. However, the answer quality of most existing visual-language pre-training (VLP) methods is still limited, mainly due to: (1) Incompatibility. Upstream pre-training tasks are generally incompatible with downstream question answering tasks, which makes the knowledge from the language model not well transferable to downstream tasks, and greatly limits their performance in few-shot scenarios; (2) Under-fitting. They generally do not integrate human priors to compensate for universal knowledge from language models, so as to fit the challenging VQA problem and generate reliable answers. To address these issues, we propose HybridPrompt, a cloze- and verify-style hybrid prompt framework with bridging language models and human priors in prompt tuning for VQA. Specifically, we first modify the input questions into the cloze-style prompts to narrow the gap between upstream pre-training tasks and downstream VQA task, which ensures that the universal knowledge in the language model can be better transferred to subsequent human prior-guided prompt tuning. Then, we imitate the cognitive process of human brain to introduce topic and sample related priors to construct a dynamic learnable prompt template for human prior-guided prompt learning. Finally, we add fixed-length learnable free-parameters to further enhance the generalizability and scalability of prompt learning in the VQA model. Experimental results verify the effectiveness of HybridPrompt, showing that it achieves competitive performance against previous methods on widely-used VQAv2 dataset and obtains new state-of-the-art results. Our code is released at: https://github.com/zhizhi111/hybrid. Zhiyuan Ma 0005, Zhihuan Yu, Jianjun Li 0010, Guohui Li 0001 |
AAAI | 1 |
| 2022 | UniTranSeR: A Unified Transformer Semantic Representation Framework for Multimodal Task-Oriented Dialog SystemabstractAs a more natural and intelligent interaction manner, multimodal task-oriented dialog system recently has received great attention and many remarkable progresses have been achieved.Nevertheless, almost all existing studies follow the pipeline to first learn intra-modal features separately and then conduct simple feature concatenation or attention-based feature fusion to generate responses, which hampers them from learning inter-modal interactions and conducting crossmodal feature alignment for generating more intention-aware responses.To address these issues, we propose UniTranSeR, a Unified Transformer Semantic Representation framework with feature alignment and intention reasoning for multimodal dialog systems.Specifically, we first embed the multimodal features into a unified Transformer semantic space to prompt inter-modal interactions, and then devise a feature alignment and intention reasoning (FAIR) layer to perform cross-modal entity alignment and fine-grained key-value reasoning, so as to effectively identify user's intention for generating more accurate responses.Experimental results verify the effectiveness of UniTranSeR, showing that it significantly outperforms state-of-the-art approaches on the representative MMD dataset. Zhiyuan Ma 0005, Jianjun Li 0010, Guohui Li 0001, Yongjing Cheng |
ACL (1) | 1 |
| 2022 | GLAF: Global-to-Local Aggregation and Fission Network for Semantic Level Fact VerificationabstractAccurate fact verification depends on performing fine-grained reasoning over crucial entities by capturing their latent logical relations hidden in multiple evidence clues, which is generally lacking in existing fact verification models. In this work, we propose a novel Global-to-Local Aggregation and Fission network (GLAF) to fill this gap. Instead of treating entire sentences or all semantic elements within them as nodes to construct a coarse-grained or unstructured evidence graph as in previous methods, GLAF constructs a fine-grained and structured evidence graph by parsing the rambling sentences into structural triple-level reasoning clues and regarding them as graph nodes to achieve fine-grained and interpretable evidence graph reasoning. Specifically, to capture latent logical relations between the clues, GLAF first employs a local fission reasoning layer to conduct fine-grained multi-hop reasoning, and then uses a global evidence aggregation layer to achieve information sharing and the interchange of evidence clues for final claim label prediction. Experimental results on the FEVER dataset demonstrate the effectiveness of GLAF, showing that it achieves the state-of-the-art performance by obtaining a 77.62% FEVER score. Zhiyuan Ma 0005, Jianjun Li 0010, Guohui Li 0001, Yongjing Cheng |
COLING | 1 |
| 2022 | CMAL: A Novel Cross-Modal Associative Learning Framework for Vision-Language Pre-TrainingabstractWith the flourishing of social media platforms, vision-language pre-training (VLP) recently has received great attention and many remarkable progresses have been achieved. The success of VLP largely benefits from the information complementation and enhancement between different modalities. However, most of recent studies focus on cross-modal contrastive learning (CMCL) to promote image-text alignment by pulling embeddings of positive sample pairs together while pushing those of negative pairs apart, which ignores the natural asymmetry property between different modalities and requires large-scale image-text corpus to achieve arduous progress. To mitigate this predicament, we propose CMAL, a Cross-Modal Associative Learning framework with anchor points detection and cross-modal associative learning for VLP. Specifically, we first respectively embed visual objects and textual tokens into separate hypersphere spaces to learn intra-modal hidden features, and then design a cross-modal associative prompt layer to perform anchor point masking and swap feature filling for constructing a hybrid cross-modal associative prompt. Afterwards, we exploit a unified semantic encoder to learn their cross-modal interactive features for context adaptation. Finally, we design an associative mapping classification layer to learn potential associative mappings between modalities at anchor points, within which we develop a fresh self-supervised associative mapping classification task to boost CMAL's performance. Experimental results verify the effectiveness of CMAL, showing that it achieves competitive performance against previous CMCL-based methods on four common downstream vision-and-language tasks, with significantly fewer corpus. Noteably, CMAL obtains new state-of-the-art results on SNLI-VE and REC (testA). Zhiyuan Ma 0005, Jianjun Li 0010, Guohui Li 0001, Kaiyan Huang |
ACM Multimedia | 1 |
| 2021 | Intention Reasoning Network for Multi-Domain End-to-end Task-Oriented DialogueabstractRecent years has witnessed the remarkable success in end-to-end task-oriented dialog system, especially when incorporating external knowledge information.However, the quality of most existing models' generated response is still limited, mainly due to their lack of finegrained reasoning on deterministic knowledge (w.r.t.conceptual tokens), which makes them difficult to capture the concept shifts and identify user's real intention in cross-task scenarios.To address these issues, we propose a novel intention mechanism to better model deterministic entity knowledge.Based on such a mechanism, we further propose an intention reasoning network (IR-Net), which consists of joint and multi-hop reasoning, to obtain intention-aware representations of conceptual tokens that can be used to capture the concept shifts involved in task-oriented conversations, so as to effectively identify user's intention and generate more accurate responses.Experimental results verify the effectiveness of IR-Net, showing that it achieves the stateof-the-art performance on two representative multi-domain dialog datasets. Zhiyuan Ma 0005, Jianjun Li 0010, Zezheng Zhang, Guohui Li 0001, Yongjing Cheng |
EMNLP (1) | 1 |