Yunxin Li

dblp:11/2484 · DBLP profile ↗
← Back
37ranked-venue papers
16as first author
25since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 13 · 5 first-author · 4 since 2021Artificial intelligence and machine learning · 12 · 6 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2026 UniMoE-Audio: Unified Speech and Music Generation with Dynamic-Capacity Mixture-of-Experts
abstract
Zhenyu Liu, Yunxin li, Xuanyu Zhang, Qixun Teng, Shenyuan Jiang, Xinyu Chen, Haoyuan Shi, Haolan Chen, Fanbo Meng, Mingjun Zhao, Yu Xu, Yancheng He, Baotian Hu, Haizhou Li, Min Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yunxin Li, Xuanyu Zhang 0006, Qixun Teng, Shenyuan Jiang, Xinyu Chen 0003, Haolan Chen, Mingjun Zhao, Yancheng He, Baotian Hu, Haizhou Li 0001, Min Zhang 0005
ACL (1)2
2026 Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs
abstract
Zhenyu Liu, Xuanyu Zhang, Yunxin li, Qixun Teng, Shenyuan Jiang, Haolan Chen, Mingjun Zhao, Fanbo Meng, Yu Xu, Yancheng He, Baotian Hu, Haizhou Li, Min Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xuanyu Zhang 0006, Yunxin Li, Qixun Teng, Shenyuan Jiang, Haolan Chen, Mingjun Zhao, Yancheng He, Baotian Hu, Haizhou Li 0001, Min Zhang 0005
ACL (1)3
2026 Vision Enhancing LLMs: Empowering Multimodal Knowledge Storage and Sharing in LLMs
abstract
Recent advancements in multimodal large language models (MLLMs) have achieved significant multimodal generation capabilities, akin to GPT-4. These models predominantly map visual information into language representation space, leveraging the vast knowledge and powerful text generation abilities of LLMs to produce multimodal instruction-following responses. We could term this method as LLMs for Vision because of its employing LLMs for visual understanding and reasoning, yet observe that these MLLMs neglect the potential of harnessing visual knowledge to enhance the overall capabilities of LLMs, which could be regarded as Vision Enhancing LLMs. In this paper, we propose an approach called MKS2, aimed at enhancing LLMs through empowering Multimodal Knowledge Storage and Sharing in LLMs. Specifically, we introduce Modular Visual Memory (MVM), a component integrated into the internal blocks of LLMs, designed to store open-world visual information efficiently. Additionally, we present a soft Mixture of Multimodal Experts (MoMEs) architecture in LLMs to invoke multimodal knowledge collaboration during text generation. Our comprehensive experiments demonstrate that MKS2 substantially augments the reasoning capabilities of LLMs in contexts necessitating physical or commonsense knowledge. It also delivers competitive results on image-text understanding multimodal benchmarks. The codes will be available at: https://github.com/HITsz-TMG/MKS2-Multimodal-Knowledge-Storage-and-Sharing.
Yunxin Li, Baotian Hu, Wei Wang 0335, Xiaochun Cao, Min Zhang 0005
IEEE Trans. Image Process.1
2026 Look Before Switch: Sensing-Assisted Handover in 5G NR V2I Networks
abstract
Integrated Sensing and Communication (ISAC) has emerged as a promising solution in addressing the challenges of high-mobility scenarios in 5 G NR Vehicle-to-Infrastructure (V2I) communications. This paper proposes a novel sensing-assisted handover framework that leverages ISAC capabilities to enable precise beamforming and proactive handover decisions. Two sensing-enabled handover triggering algorithms are developed: a distance-based scheme that utilizes estimated spatial positioning, and a probability-based approach that predicts vehicle maneuvers using interacting multiple model extended Kalman filter (IMM-EKF) tracking. The proposed methods eliminate the need for uplink feedback and beam sweeping, thus significantly reducing signaling overhead and handover interruption time. A sensing-assisted NR frame structure and corresponding protocol design are also introduced to support rapid synchronization and access under vehicular mobility. Extensive link-level simulations using real-world map data demonstrate that the proposed framework reduces the average handover interruption time by over 50%, achieves lower handover rates, and enhances overall communication performance.
Yunxin Li, Fan Liu 0005, Haoqiu Xiong, Zhenkun Wang 0001, Narengerile, Christos Masouros
IEEE Trans. Mob. Comput.1
2025 VideoVista-CulturalLingo: 360° Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension
abstract
Xinyu Chen, Yunxin Li, Haoyuan Shi, Baotian Hu, Wenhan Luo, Yaowei Wang, Min Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Xinyu Chen 0003, Yunxin Li, Baotian Hu, Wenhan Luo, Yaowei Wang 0001, Min Zhang 0005
ACL (1)2
2025 Advancing Temporal Sensitive Question Answering through Progressive Multi-Step Reflection
abstract
Retrieval-augmented generation (RAG) has demonstrated strong potential in enhancing large language models (LLMs) for complex, real-world question answering. However, existing RAG frameworks remain inadequate for temporal scenarios, primarily due to their inability to jointly model temporal constraints in both retrieval and reasoning. On the retrieval side, traditional approaches focus on semantic similarity, often returning outdated or temporally misaligned evidence. On the generation side, these systems frequently produce factually incorrect or hallucinated answers when confronted with incomplete or temporally inconsistent information. Motivated by the observed limitations, we propose ChronoReflect+, a temporal logic-aware RAG framework that incorporates hybrid temporal-aware retrieval and progressive multi-step reflection. Our method iteratively refines both retrieval and reasoning, identifying and bridging information gaps as context accumulates. Extensive experiments demonstrate that ChronoReflect+ significantly outperforms state-of-the-art RAG baselines-improving end-to-end accuracy by 15.2%-particularly on questions involving implicit time expressions and multi-hop reasoning.
Erxue Min, Xiang Zhao 0002, Yunxin Li, Jinzhi Liao, Shuaiqiang Wang, Baotian Hu, Dawei Yin 0001
CIKM4
2025 Beyond RSRP: A Sensing-Assisted Handover Framework in V2I Networks
abstract
Given the challenges of high mobility and frequent handovers in 5G NR's Vehicle-to-Infrastructure (V2I) communications, the Integrated Sensing and Communications (ISAC) technique, now recognized in IMT-2030 as a usage scenario, is introduced to boost mobile network efficiency. Leveraging this emerging paradigm, our study delves into a V2I network within the 5G NR context, proposing a sensing-assisted handover mechanism and protocol that integrates ISAC to refine the intercell handover process. This mechanism uses precise beamforming and kinematic parameter estimation to reduce signaling and interruption time in inter-cell handover process. This approach enables a distance-based handover triggering mechanism that is both proactive and information-rich, allowing the serving gNB to inform the target gNB of the vehicle's location preemptively. Numerical results from link-level simulations at a crossroad scenario demonstrate that this sensing-assisted handover mechanism enables faster triggering and reduces the interruption time by 76.46%, while enhancing communication performance.
Yunxin Li, Fan Liu 0005, Christos Masouros
ICC1
2025 AniMaker: Multi-Agent Animated Storytelling with MCTS-Driven Clip Generation
abstract
Despite rapid advancements in video generation models, generating coherent, long-form storytelling videos that span multiple scenes and characters remains challenging. Current methods often rigidly convert pre-generated keyframes into fixed-length clips, resulting in disjointed narratives and pacing issues. Furthermore, the inherent instability of video generation models means that even a single low-quality clip can significantly degrade the entire output animation’s logical coherence and visual continuity. To overcome these obstacles, we introduce AniMaker, a multi-agent framework enabling efficient multi-candidate clip generation and storytelling-aware clip selection, thus creating globally consistent and story-coherent animation solely from text input. The framework is structured around specialized agents, including the Director Agent for storyboard generation, the Photography Agent for video clip generation, the Reviewer Agent for evaluation, and the Post-Production Agent for editing and voiceover, collectively realizing multi-character, multi-scene animation. Central to AniMaker’s approach are two key technical components: MCTS-Gen in Photography Agent, an efficient Monte Carlo Tree Search (MCTS)-inspired strategy that intelligently navigates the candidate space to generate high-potential clips while optimizing resource usage; and AniEval in Reviewer Agent, the first framework specifically designed for multi-shot animation evaluation, which assesses critical aspects such as story-level consistency, action completion, and animation-specific features by considering each clip in the context of its preceding and succeeding clips. Experiments demonstrate that AniMaker achieves superior quality as measured by popular metrics including VBench and our proposed AniEval framework, while significantly improving the efficiency of multi-candidate generation, pushing AI-generated storytelling animation closer to production standards. Code and data for this paper are at https://animaker-dev.github.io/
Yunxin Li, Xinyu Chen 0003, Longyue Wang, Baotian Hu, Min Zhang 0005
SIGGRAPH Asia2
2025 Uni-MoE: Scaling Unified Multimodal LLMs With Mixture of Experts
abstract
Recent advancements in Multimodal Large Language Models (MLLMs) underscore the significance of scalable models and data to boost performance, yet this often incurs substantial computational costs. Although the Mixture of Experts (MoE) architecture has been employed to scale large language or visual-language models efficiently, these efforts typically involve fewer experts and limited modalities. To address this, our work presents the pioneering attempt to develop a unified MLLM with the MoE architecture, named Uni-MoE that can handle a wide array of modalities. Specifically, it features modality-specific encoders with connectors for a unified multimodal representation. We also implement a sparse MoE architecture within the LLMs to enable efficient training and inference through modality-level data parallelism and expert-level model parallelism. To enhance the multi-expert collaboration and generalization, we present a progressive training strategy: 1) Cross-modality alignment using various connectors with different cross-modality data, 2) Training modality-specific experts with cross-modality instruction data to activate experts' preferences, and 3) Tuning the whole Uni-MoE framework utilizing Low-Rank Adaptation (LoRA) on mixed multimodal instruction data. We evaluate the instruction-tuned Uni-MoE on a comprehensive set of multimodal datasets. The extensive experimental results demonstrate Uni-MoE's principal advantage of significantly reducing performance bias in handling mixed multimodal datasets, alongside improved multi-expert collaboration and generalization.
Yunxin Li, Shenyuan Jiang, Baotian Hu, Longyue Wang, Wanqi Zhong, Wenhan Luo, Lin Ma 0002, Min Zhang 0005
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 TEFormer: Thermal Infrared Image Enhancement by Preserving Spatial Consistency and Details
abstract
Thermal infrared (TIR) images suffer from low contrast due to the atmospheric thermal radiation effect, especially under extreme conditions like low temperature. TIR image enhancement aims to improve image contrast, but previous enhancement approaches usually produce enhanced results with two limitations: spatial inconsistency and detail blurring. To deal with the limitations, we propose a novel TIR image enhancement method, named TEFormer, to preserve spatial consistency and restore fine-grained details during image enhancement. To preserve spatial consistency, we devise the global enhancement module (GEM) to enhance the low-resolution representation. The GEM performs long-range interactions across spatial dimensions and channel dimensions to condition the enhancement curve fitting. To keep details clear, we design the local enhancement module (LEM) as the decoding unit. The LEM injects additional detail structures into the enhanced low-resolution representation for high-resolution reconstruction. Besides, we further apply histogram-based supervision to facilitate learning in intensity distribution of clear images. Extensive experimental results on three challenging benchmarks demonstrate that the proposed method outperforms other state-of-the-art approaches.
Yunxin Li, Runmin Zhang, Si-Yuan Cao, Jiacheng Ying, Xiaokai Bai, Shujie Chen 0001, Bailin Yang
IEEE Trans. Geosci. Remote. Sens.2
2025 A vision-language model with multi-granular knowledge fusion in medical imaging
Kai Chen 0020, Yunxin Li, Xiwen Zhu, Wentai Zhang 0003, Baotian Hu
World Wide Web (WWW)2
2024 Cognitive Visual-Language Mapper: Advancing Multimodal Comprehension with Enhanced Visual Knowledge Alignment
abstract
Evaluating and Rethinking the current landscape of Large Multimodal Models (LMMs), we observe that widely-used visual-language projection approaches (e.g., Q-former or MLP) focus on the alignment of image-text descriptions yet ignore the visual knowledgedimension alignment, i.e., connecting visuals to their relevant knowledge.Visual knowledge plays a significant role in analyzing, inferring, and interpreting information from visuals, helping improve the accuracy of answers to knowledge-based visual questions.In this paper, we mainly explore improving LMMs with visual-language knowledge alignment, especially aimed at challenging knowledge-based visual question answering (VQA).To this end, we present a Cognitive Visual-Language Mapper (CVLM), which contains a pretrained Visual Knowledge Aligner (VKA) and a Finegrained Knowledge Adapter (FKA) used in the multimodal instruction tuning stage.Specifically, we design the VKA based on the interaction between a small language model and a visual encoder, training it on collected imageknowledge pairs to achieve visual knowledge acquisition and projection.FKA is employed to distill the fine-grained visual knowledge of an image and inject it into Large Language Models (LLMs).We conduct extensive experiments on knowledge-based VQA benchmarks and experimental results show that CVLM significantly improves the performance of LMMs on knowledge-based VQA (average gain by 5.0%).Ablation studies also verify the effectiveness of VKA and FKA, respectively.1
Yunxin Li, Xinyu Chen 0003, Baotian Hu, Min Zhang 0005
ACL (1)1
2024 A Multimodal In-Context Tuning Approach for E-Commerce Product Description Generation
abstract
In this paper, we propose a new setting for generating product descriptions from images, augmented by marketing keywords. It leverages the combined power of visual and textual information to create descriptions that are more tailored to the unique features of products. For this setting, previous methods utilize visual and textual encoders to encode the image and keywords and employ a language model-based decoder to generate the product description. However, the generated description is often inaccurate and generic since same-category products have similar copy-writings, and optimizing the overall framework on large-scale samples makes models concentrate on common words yet ignore the product features. To alleviate the issue, we present a simple and effective Multimodal In-Context Tuning approach, named ModICT, which introduces a similar product sample as the reference and utilizes the in-context learning capability of language models to produce the description. During training, we keep the visual encoder and language model frozen, focusing on optimizing the modules responsible for creating multimodal in-context references and dynamic prompts. This approach preserves the language generation prowess of large language models (LLMs), facilitating a substantial increase in description diversity. To assess the effectiveness of ModICT across various language model scales and types, we collect data from three distinct product categories within the E-commerce domain. Extensive experiments demonstrate that ModICT significantly improves the accuracy (by up to 3.3% on Rouge-L) and diversity (by up to 9.4% on D-5) of generated results compared to conventional methods. Our findings underscore the potential of ModICT as a valuable tool for enhancing the automatic generation of product descriptions in a wide range of applications. Data and code are at https://github.com/HITsz-TMG/Multimodal-In-Context-Tuning
Yunxin Li, Baotian Hu, Wenhan Luo, Lin Ma 0002, Min Zhang 0005
LREC/COLING1
2024 VisionGraph: Leveraging Large Multimodal Models for Graph Theory Problems in Visual Context
abstract
Large Multimodal Models (LMMs) have achieved impressive success in visual reasoning, particularly in visual mathematics. However, problem-solving capabilities in graph theory remain less explored for LMMs, despite being a crucial aspect of mathematical reasoning that requires an accurate understanding of graphical structures and multi-step reasoning on visual graphs. To step forward in this direction, we are the first to design a benchmark named VisionGraph, used to explore the capabilities of advanced LMMs in solving multimodal graph theory problems. It encompasses eight complex graph problem tasks, from connectivity to shortest path problems. Subsequently, we present a Description-Program-Reasoning (DPR) chain to enhance the logical accuracy of reasoning processes through graphical structure description generation and algorithm-aware multi-step reasoning. Our extensive study shows that 1) GPT-4V outperforms Gemini Pro in multi-step graph reasoning; 2) All LMMs exhibit inferior perception accuracy for graphical structures, whether in zero/few-shot settings or with supervised fine-tuning (SFT), which further affects problem-solving performance; 3) DPR significantly improves the multi-step graph reasoning capabilities of LMMs and the GPT-4V (DPR) agent achieves SOTA performance.
Yunxin Li, Baotian Hu, Wei Wang 0164, Longyue Wang, Min Zhang 0005
ICML1
2024 Anim-Director: A Large Multimodal Model Powered Agent for Controllable Animation Video Generation
abstract
demands substantial human effort and incurs high training
Yunxin Li, Baotian Hu, Longyue Wang, Jiashun Zhu, Jinyi Xu, Zhen Zhao 0001, Min Zhang 0005
SIGGRAPH Asia1
2024 Integrated Sensing and Communications: Recent Advances and Ten Open Challenges
abstract
It is anticipated that integrated sensing and communications (ISAC) would be one of the key enablers of next-generation wireless networks (such as beyond 5G (B5G) and 6G) for supporting a variety of emerging applications. In this paper, we provide a comprehensive review of the recent advances in ISAC systems, with a particular focus on their foundations, physical-layer system design, networking aspects and ISAC applications. Furthermore, we discuss the corresponding open questions of the above that emerged in each issue. Hence, we commence with the information theory of sensing and communications (S&C), followed by the information-theoretic limits of ISAC systems by shedding light on the fundamental performance metrics. Next, we discuss their clock synchronization and phase offset problems, the associated Pareto-optimal signaling strategies, as well as the associated super-resolution physical-layer ISAC system design. Moreover, we envision that ISAC ushers in a paradigm shift for the future cellular networks relying on network sensing, transforming the classic cellular architecture, cross-layer resource management methods, and transmission protocols. In ISAC applications, we further highlight the security and privacy issues of wireless sensing. Finally, we close by studying the recent advances in a representative ISAC use case, namely the multi-object multi-task (MOMT) recognition problem using wireless signals.
Shihang Lu, Fan Liu 0005, Yunxin Li, Kecheng Zhang, Hongjia Huang, Jiaqi Zou, Xinyu Li 0007, Yuxiang Dong, Fuwang Dong, Jia Zhu 0001, Yifeng Xiong, Weijie Yuan 0001, Yuanhao Cui, Lajos Hanzo
IEEE Internet Things J.3
2024 Frame Structure and Protocol Design for Sensing-Assisted NR-V2X Communications
abstract
The emergence of the fifth-generation (5G) New Radio (NR) technology has provided unprecedented opportunities for vehicle-to-everything (V2X) networks, enabling enhanced quality of services. However, high-mobility V2X networks require frequent handovers and acquiring accurate channel state information (CSI) necessitates the utilization of pilot signals, leading to increased overhead and reduced communication throughput. To address this challenge, integrated sensing and communications (ISAC) techniques have been employed at the base station (gNB) within vehicle-to-infrastructure (V2I) networks, aiming to minimize overhead and improve spectral efficiency. In this study, we propose novel frame structures that incorporate ISAC signals for three crucial stages in the NR-V2X system: initial access, connected mode, and beam failure and recovery. These new frame structures employ 75% fewer pilots and reduce reference signals by 43.24%, capitalizing on the sensing capability of ISAC signals. Through extensive link-level simulations, we demonstrate that our proposed approach enables faster beam establishment during initial access, higher throughput and more precise beam tracking in connected mode with reduced overhead, and expedited detection and recovery from beam failures. Furthermore, the numerical results obtained from our simulations showcase enhanced spectrum efficiency, improved communication performance and minimal overhead, validating the effectiveness of the proposed ISAC-based techniques in NR V2I networks.
Yunxin Li, Fan Liu 0005, Zhen Du, Weijie Yuan 0001, Qingjiang Shi, Christos Masouros
IEEE Trans. Mob. Comput.1
2024 LMEye: An Interactive Perception Network for Large Language Models
abstract
Current efficient approaches to building Multimodal Large Language Models (MLLMs) mainly incorporate visual information into LLMs with a simple visual mapping network such as a linear projection layer, a multilayer perceptron (MLP), or Q-former from BLIP-2. Such networks project the image feature once and do not consider the interaction between the image and the human inputs. Hence, the obtained visual information without being connected to human intention may be inadequate for LLMs to generate intention-following responses, which we refer to as static visual information. To alleviate this issue, our paper introduces LMEye, a human-like eye with a play-and-plug interactive perception network, designed to enable dynamic interaction between LLMs and external visual information. It can allow the LLM to request the desired visual information aligned with various human instructions, which we term dynamic visual information acquisition. Specifically, LMEye consists of a simple visual mapping network to provide the basic perception of an image for LLMs. It also contains additional modules responsible for acquiring requests from LLMs, performing request-based visual information seeking, and transmitting the resulting interacted visual information to LLMs, respectively. In this way, LLMs act to understand the human query, deliver the corresponding request to the request-based visual information interaction module, and generate the response based on the interleaved multimodal information. We evaluate LMEye through extensive experiments on multimodal benchmarks, demonstrating that it significantly improves zero-shot performances on various multimodal tasks compared to previous methods, with fewer parameters. Moreover, we also verify its effectiveness and scalability on various language models and video understanding, respectively.
Yunxin Li, Baotian Hu, Xinyu Chen 0003, Lin Ma 0002, Yong Xu 0001, Min Zhang 0005
IEEE Trans. Multim.1
2023 A Multi-Modal Context Reasoning Approach for Conditional Inference on Joint Textual and Visual Clues
abstract
Conditional inference on joint textual and visual clues is a multi-modal reasoning task that textual clues provide prior permutation or external knowledge, which are complementary with visual content and pivotal to deducing the correct option.Previous methods utilizing pretrained vision-language models (VLMs) have achieved impressive performances, yet they show a lack of multimodal context reasoning capability, especially for text-modal information.To address this issue, we propose a Multi-modal Context Reasoning approach, named ModCR.Compared to VLMs performing reasoning via cross modal semantic alignment, it regards the given textual abstract semantic and objective image information as the pre-context information and embeds them into the language model to perform context reasoning.Different from recent vision-aided language models used in natural language processing, ModCR incorporates the multi-view semantic alignment information between language and vision by introducing the learnable alignment prefix between image and text in the pretrained language model.This makes the language model well-suitable for such multi-modal reasoning scenario on joint textual and visual clues.We conduct extensive experiments on two corresponding data sets and experimental results show significantly improved performance (exact gain by 4.8% on PMR test set) compared to previous strong baselines.Code
Yunxin Li, Baotian Hu, Xinyu Chen 0003, Lin Ma 0002, Min Zhang 0005
ACL (1)1
2023 A Neural Divide-and-Conquer Reasoning Framework for Image Retrieval from Linguistically Complex Text
abstract
Pretrained Vision-Language Models (VLMs) have achieved remarkable performance in image retrieval from text.However, their performance drops drastically when confronted with linguistically complex texts that they struggle to comprehend.Inspired by the Divide-and-Conquer (Smith, 1985) algorithm and dualprocess theory (Groves and Thompson, 1970), in this paper, we regard linguistically complex texts as compound proposition texts composed of multiple simple proposition sentences and propose an end-to-end Neural Divide-and-Conquer Reasoning framework, dubbed NDCR.It contains three main components: 1) Divide: a proposition generator divides the compound proposition text into simple proposition sentences and produces their corresponding representations, 2) Conquer: a pretrained VLMsbased visual-linguistic interactor achieves the interaction between decomposed proposition sentences and images, 3) Combine: a neuralsymbolic reasoner combines the above reasoning states to obtain the final solution via a neural logic reasoning approach.According to the dual-process theory, the visual-linguistic interactor and neural-symbolic reasoner could be regarded as analogical reasoning System 1 and logical reasoning System 2. We conduct extensive experiments on a challenging image retrieval from contextual descriptions data set.Experimental results and analyses indicate NDCR significantly improves performance in the complex image-text reasoning problem.
Yunxin Li, Baotian Hu, Lin Ma 0002, Min Zhang 0005
ACL (1)1
2023 Training Multimedia Event Extraction With Generated Images and Captions
abstract
Contemporary news reporting increasingly features multimedia content, motivating research on multimedia event extraction. However, the task lacks annotated multimodal training data and artificially generated training data suffer from the distribution shift from the real-world data. In this paper, we propose Cross-modality Augmented Multimedia Event Learning (CAMEL), which successfully utilizes artificially generated multimodal training data and achieves state-of-the-art performance. Conditioned on unimodal training data, we generate multimodal training data using off-the-shelf image generators like Stable Diffusion [45] and image captioners like BLIP [24]. After that, we train the network on the resultant multimodal datasets. In order to learn robust features that are effective across domains, we devise an iterative and gradual training strategy. Substantial experiments show that CAMEL surpasses state-of-the-art (SOTA) baselines on the M2E2 benchmark. On multimedia events in particular, we outperform the prior SOTA by 4.2% F1 on event mention identification and by 9.8% F1 on argument identification, which demonstrates that CAMEL learns synergistic representations from the two modalities. Our work demonstrates a recipe to unleash the power of synthetic training data in structured prediction.
Zilin Du, Yunxin Li, Xu Guo 0002, Boyang Li 0001
ACM Multimedia2
2023 Fast and Robust Online Handwritten Chinese Character Recognition With Deep Spatial and Contextual Information Fusion Network
abstract
Deep convolutional neuralnetworks have achieved fairly high accuracy for single online handwritten Chinese character recognition (SOLHCCR). However, in real application scenarios, users always write multiple characters to form a complete sentence, and previous contextual information holds significant potential for improving the accuracy, robustness and efficiency of recognition. In this work, we first propose a simple and straightforward model named the vanilla compositional network (VCN) by coupling convolutional neural network with a sequence modeling architecture (i.e., a recurrent neural network or Transformer), which exploits the handwritten character’s previous contextual information. Although VCN performs much better than the previous state-of-the-art SOLHCCR models, it is a two-stage architecture in nature. It suffers from high fragility when confronting with poorly written characters such as sloppy writing, and missing or broken strokes, due to relying heavily on contextual information. To improve the robustness of the OLHCCR model, we further propose a novel deep spatial & contextual information fusion network (DSCIFN). It utilizes an autoregresssive framework pre-trained on a large-scale sentence corpora as the backbone component, and highly integrates the spatial features of handwritten characters and their previous contextual information in a multi-layer fusion module. To verify the effectiveness of models, we reorganize a new form of online Chinese handwritten character with its previous context dataset, named OHCCC. Extensive experimental results demonstrate that DSCIFN achieves state-of-the-art performance and has increased strong robustness compared to VCN and previous SOLHCCR models. The in-depth empirical analysis and case study indicate that DSCIFN can significantly improve the efficiency of handwriting input because it does not need complete strokes to recognize a handwritten Chinese character precisely.
Yunxin Li, Qian Yang 0007, Qingcai Chen, Baotian Hu, Xiaolong Wang 0001, Lin Ma 0002
IEEE Trans. Multim.1
2022 Medical Dialogue Response Generation with Pivotal Information Recalling
abstract
Medical dialogue generation is an important yet challenging task. Most previous works rely on the attention mechanism and large-scale pretrained language models. However, these methods often fail to acquire pivotal information from the long dialogue history to yield an accurate and informative response, due to the fact that the medical entities usually scatters throughout multiple utterances along with the complex relationships between them. To mitigate this problem, we propose a medical response generation model with Pivotal Information Recalling (MedPIR), which is built on two components, i.e., knowledge-aware dialogue graph encoder and recall-enhanced generator. The knowledge-aware dialogue graph encoder constructs a dialogue graph by exploiting the knowledge relationships between entities in the utterances, and encodes it with a graph attention network. Then, the recall-enhanced generator strengthens the usage of these pivotal information by generating a summary of the dialogue before producing the actual response. Experimental results on two large-scale medical dialogue datasets show that MedPIR outperforms the strong baselines in BLEU scores and medical entities F1 measure.
Yu Zhao 0043, Yunxin Li, Yuxiang Wu, Baotian Hu, Qingcai Chen, Xiaolong Wang 0001, Min Zhang 0005
KDD2
2022 Chunk-aware Alignment and Lexical Constraint for Visual Entailment with Natural Language Explanations
abstract
Visual Entailment with natural language explanations aims to infer the relationship between a text-image pair and generate a sentence to explain the decision-making process. Previous methods rely mainly on a pre-trained vision-language model to perform the relation inference and a language model to generate the corresponding explanation. However, the pre-trained vision-language models mainly build token-level alignment between text and image yet ignore the high-level semantic alignment between the phrases (chunks) and visual contents, which is critical for vision-language reasoning. Moreover, the explanation generator based only on the encoded joint representation does not explicitly consider the critical decision-making points of relation inference. Thus the generated explanations are less faithful to visual-language reasoning. To mitigate these problems, we propose a unified Chunk-aware Alignment and Lexical Constraint based method, dubbed as CALeC. It contains a Chunk-aware Semantic Interactor (arr. CSI), a relation inferrer, and a Lexical Constraint-aware Generator (arr. LeCG). Specifically, CSI exploits the sentence structure inherent in language and various image regions to build chunk-aware semantic alignment. Relation inferrer uses an attention-based reasoning network to incorporate the token-level and chunk-level vision-language representations. LeCG utilizes lexical constraints to expressly incorporate the words or chunks focused by the relation inferrer into explanation generation, improving the faithfulness and informativeness of the explanations. We conduct extensive experiments on three datasets, and experimental results indicate that CALeC significantly outperforms other competitor models on inference accuracy and quality of generated explanations.
Qian Yang 0007, Yunxin Li, Baotian Hu, Lin Ma 0002, Min Zhang 0005
ACM Multimedia2
2021 MSDF: A General Open-Domain Multi-skill Dialog Framework
Yu Zhao 0043, Xinshuo Hu, Yunxin Li, Baotian Hu, Dongfang Li 0002, Sichao Chen, Xiaolong Wang 0001
NLPCC (2)3
2011 A Simplified User Identification Approach for Multi-User Diversity with Enhanced Throughput
abstract
In a wireless network, multi-user diversity can be employed to improve system throughput performance by scheduling the channel to the user with the best instantaneous channel state information (CSI). However, the overhead induced by polling CSIs of a large number of users can overshadow the multi-user diversity gain. In our previous work, a user identification approach (UIDA) was proposed to reduce the system overhead. In this paper, by allowing a small degree of outage to occur, we simplify the UIDA to reduce the overhead further, and present the throughput and outage analysis of the simplified UIDA over Rayleigh fading channels. Computer simulations based on IEEE 802.11a systems show that the simplified UIDA achieves considerable throughput improvement.
Hang Li 0002, Qinghua Guo 0001, Yunxin Li, Defeng Huang
GLOBECOM3
2005 Simple CPM receivers based on a switched linear modulation model
abstract
Based on a switched linear modulation model recently developed for continuous-phase modulated (CPM) signal representation and approximation, and incorporated with new phase state symbol definitions, three simple CPM receivers are proposed in this letter. Their performance simulation results and complexity comparison are given using a quaternary 2RC (raised cosine frequency pulse) CPM scheme.
Xiaojing Huang 0001, Yunxin Li
IEEE Trans. Commun.2
2005 MMSE-optimal approximation of continuous-phase modulated signal as superposition of linearly modulated pulses
abstract
The optimal linear modulation approximation of any M-ary continuous-phase modulated (CPM) signal under the minimum mean-square error (MMSE) criterion is presented in this paper. With the introduction of the MMSE signal component, an M-ary CPM signal is exactly represented as the superposition of a finite number of MMSE incremental pulses, resulting in the novel switched linear modulation CPM signal models. Then, the MMSE incremental pulse is further decomposed into a finite number of MMSE pulse-amplitude modulated (PAM) pulses, so that an M-ary CPM signal is alternatively expressed as the superposition of a finite number of MMSE PAM components, similar to the Laurent representation. Advantageously, these MMSE PAM components are mutually independent for any modulation index. The optimal CPM signal approximation using lower order MMSE incremental pulses, or alternatively, using a small number of MMSE PAM pulses, is also made possible, since the approximation error is minimized in the MMSE sense. Finally, examples of the MMSE-optimal CPM signal approximation and its comparison with the Laurent approximation approach are given using raised-cosine frequency-pulse CPM schemes.
Xiaojing Huang 0001, Yunxin Li
IEEE Trans. Commun.2
2003 Simple noncoherent CPM receivers by PAM decomposition and MMSE equalization
abstract
Conventional differential detection and energy detection offer poor performances for demodulating continuous phase modulated (CPM) signal because of inherent inter-symbol interference (ISI). Based on the pulse amplitude modulation (PAM) decomposition of CPM signal, the binary CPM with any non-integer modulation index can be interpreted as a linear modulation with differentially encoded pseudo-symbols or equivalently as a modulation scheme with two overlapped waveforms to carry information bits. By combining the equalization of ISI with the design of the receiver filters under the minimum mean square error (MMSE) criterion, the equalized differential detection and equalized energy detection for binary CPM signal demonstrate significant performance improvements while maintaining the same simple receiver structures as the conventional ones. Through theoretical analysis, a closed-form formula is derived to efficiently evaluate the bit error probabilities of the proposed noncoherent receivers. Simulation results using minimum shift keying (MSK) and Gaussian frequency shift keying (GFSK) with BT = 0.5, h = 0.34 are also provided to conform with the performance improvements.
Xiaojing Huang 0001, Yunxin Li
PIMRC2
2003 Sample rate conversion by trapezoidal interpolation for software defined radio
abstract
A novel arbitrary conversion ratio sample rate conversion (SRC) structure for software defined radio (SDR) is proposed in this paper. Simplified implementation is also presented, which uses only one multiplication per sample for the trapezoidal interpolation. The passband droop of the trapezoidal interpolation is compensated by the reconstruction filter, which is designed using frequency sampling method. A root raised cosine filter design example is given to show that both anti-imaging and anti-aliasing requirements for SRC can be satisfied by proper selection of design parameters.
Xiaojing Huang 0001, Yunxin Li
PIMRC2
2003 The PAM decomposition of CPM signals with integer modulation index
abstract
This article presents the solution for decomposing continuous-phase modulated (CPM) signals with integer modulation index into pulse-amplitude modulated components. The notion of main complex pulse is also introduced. A simplified demodulator for CPM signals with integer modulation index is proposed as an application example and simulation results using a quaternary 2 raised cosine (RC) scheme are given.
Xiaojing Huang 0001, Yunxin Li
IEEE Trans. Commun.2
2002 Scalable complete complementary sets of sequences
abstract
A family of complete complementary sets of sequences with closed-form expression is presented in this paper. Firstly, by introducing the notion of Golay (1961)-paired matrix and providing its synthesis algorithms, an orthogonal Golay-paired matrix called Golay-paired Hadamard matrix is derived. Then, general procedures for constructing mutually orthogonal Golay-paired matrices are proposed. Finally, the complete complementary sets of sequences represented by Golay-paired Hadamard matrices are generated and their scalability is illustrated. The unique properties of this new family of scalable complete complementary sets of sequences make it an ideal candidate for applications in future advanced signal processing and communications systems.
Xiaojing Huang 0001, Yunxin Li
GLOBECOM2
2002 Performances of impulse train modulated ultra-wideband systems
abstract
This paper analyzes the performances of three impulse train modulated ultra-wideband (UWB) communications systems in an additive white Gaussian noise (AWGN) channel. First, the mathematical models for describing biphase, pulse position and hybrid modulated ultra-wideband signals are developed and the decision rules for detecting them with only AWGN interference are proposed. Then, the exact formulae of the bit error probabilities of these UWB systems and their closed-form approximations are derived. Finally, the derived formulae are applied to optimize the modulation parameter of a Gaussian monocycle UWB impulse radio.
Xiaojing Huang 0001, Yunxin Li
ICC2
2002 The simulation of independent Rayleigh faders
abstract
Multiple independent Rayleigh fading waveforms are often required for the simulation of wireless communications channels. Jakes (1974) Rayleigh fading model and its derivatives based on the sum-of-sinusoids provide simple simulators, but they have major shortcomings in their simulated correlation functions. A novel sum-of-sinusoids fading model is proposed and verified, which generates Rayleigh fading processes satisfying the theoretical independence requirements and providing desired power spectral densities with ideal second-order moment. The effects of replacing sinusoids in the proposed model by their approximate waveforms are also analyzed and tested. Performance evaluation and comparison are provided, using the quality measures of the mean-square-error of autocorrelation function and the second-order moment of the power spectral density.
Yunxin Li, Xiaojing Huang 0001
IEEE Trans. Commun.1
2001 The multicode interleaved DSSS system for high speed wireless digital communications
abstract
An intersymbol interference (ISI) free spread spectrum system, named multicode interleaved direct sequence spread spectrum (MCI-DSSS) system, is proposed in this paper. By combining the multicode direct sequence spreading and block interleaving together, this system is able to provide high-speed data transmission over severe multipath frequency-selective fading channel While the ISI is resolved by the orthogonality of the code set used in the multicode spreading, high degree path diversity is exploited to overcome the significant dispersion of the channel which may have a delay spread much longer than one bit duration, Another advantage of this MCI-DSSS system over conventional multicode and multicarrier systems is its higher power efficiency as constant envelope modulations can be applied to the binary MCI-DSSS baseband signal. The principle of the multicode interleaved direct sequence spreading and the MCI-DSSS system models are described. The selection of practical code sets in relation to the system structures is summarized and numerical results are given.
Xiaojing Huang 0001, Yunxin Li
ICC2
2000 The Generation of Independant Rayleigh Faders
abstract
The simulation of wireless communication channels often requires the generation of multiple complex waveforms satisfying the following conditions. (1) The real and imaginary parts are independent zero-mean random Gaussian processes with identical auto-correlation functions. (2) The complex waveforms are independent so that the cross-correlation function between any two waveforms is zero. We first analyze previously proposed methods based on Jakes fading model to demonstrate that the above two conditions have not been satisfied. We then propose a novel method for the generation of multiple Rayleigh faders, which uses improved arrival angle patterns and incident wave phases. The auto-correlation and cross-correlation functions are derived to prove theoretically that the above two conditions are satisfied. Finally the statistical properties are verified by numerical results.
Yunxin Li, Xiaojing Huang 0001
ICC (1)1
1995 Optimum soft-output detection for channels with intersymbol interference
abstract
In contrast to the conventional Viterbi algorithm (VA) which generates hard-outputs, an optimum soft-output algorithm (OSA) is derived under the constraint of fixed decision delay for detection of M-ary digital signals in the presence of intersymbol interference and additive white Gaussian noise. The OSA, a new type of the conventional symbol-by-symbol maximum a posteriori probability algorithm, requires only a forward recursion and the number of variables to be stored and recursively updated increases linearly, rather than exponentially, with the decision delay. Then, with little performance degradation, a suboptimum soft-output algorithm (SSA) is obtained by simplifying the OSA. The main computations in the SSA, as in the VA, are the add-compare-select operations. Simulation results of a convolutional-coded communication system are presented that demonstrate the superiority of the OSA and the SSA over the conventional VA when they are used as detectors. When the decision delay of the detectors equals the channel memory, a significant performance improvement is achieved with only a small increase in computational complexity.>
Yunxin Li, Branka Vucetic, Yoichi Sato 0004
IEEE Trans. Inf. Theory1