VLDB 2026 Research / reviewers in the wild / expert
Yangfan He
dblp:54/3082
· DBLP profile ↗
37ranked-venue papers
3as first author
32since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 1 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 14 since 2021Theory of computation · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ViG-RAG: Video-aware Graph Retrieval-Augmented Generation via Temporal and Semantic Hybrid ReasoningabstractRetrieval-augmented generation (RAG) has greatly improved Large Language Models (LLMs) by adding external knowledge. However, current RAG-based methods face difficulties with long-context video understanding due to two main challenges. First, Current RAG-based methods for long-context video understanding struggle to effectively integrate multimodal and long-range temporal information, resulting in fragmented and context-insensitive knowledge representations. Furthermore, their retrieval mechanisms often rely on static textual matching, failing to dynamically align user queries with the most relevant video segments and leading to suboptimal downstream performance. To overcome these issues, we introduce ViG-RAG, a new framework to enhance long-context video understanding through structured textual knowledge grounding and multi-modal retrieval. Specifically, we segment video transcripts into structured units, extract key entities, form temporal connections, and assign confidence for evidence, enabling coherent long-range reasoning. In this way, it utilizes a knowledge-aware grounding mechanism and a context-aware retrieval process that dynamically builds a probabilistic temporal knowledge graph to organize multi-video content. To improve retrieval accuracy, we propose a hybrid retrieval strategy for semantic and temporal features, with an adaptive distribution modeling the relevance. In this way, it achieves the optimal retrieval distribution for each query, enhancing generation efficiency by reducing unnecessary computations. On top of this, ViG-RAG uses a vision-language model to integrate semantic anchors, expanded contextual fields, and selected video frames, generating an accurate response. We evaluate ViG-RAG on several benchmarks, demonstrating that it significantly surpasses current RAG-based methods. Zongsheng Cao, Yangfan He, Jing Li 0114, Bo Zhang 0069, Zigan Wang |
AAAI | 3 |
| 2026 | GE-adapter: A general and efficient adapter for enhanced video editing with pretrained text-to-image diffusion models
Yangfan He, Kun Li 0014, Jianhui Wang 0001, Binxu Li, Tianyu Shi 0003, Miao Zhang 0010, Xueqian Wang 0001 |
Expert Syst. Appl. | 1 |
| 2026 | TCSTNet: A text-driven color style transfer network for low-light image enhancement
Tianyi Zeng, Miao Zhang 0010, Zimo Zeng, Junfeng Jiao, Yuantao Wang, Yangfan He, Junbo Tan, Christian G. Claudel, Xueqian Wang 0001 |
Expert Syst. Appl. | 10 |
| 2026 | Parameter-Efficient Fine-Tuning for Pre-Trained Vision Models: A Survey and Benchmark
Yi Xin 0003, Jianjiang Yang, Yuntao Du 0001, Haoxing Chen, Kangrui Cen, Yangfan He, Yuewen Cao, Junjun He, Xiaokang Yang 0001, Guangtao Zhai, Ming-Hsuan Yang 0001, Xiaohong Liu 0001 |
Int. J. Comput. Vis. | 8 |
| 2026 | Spectral-Mamba: Efficient Graph Signal Recovery via Spectral-Serialized State Space Models
Feiyue Zhao, Yangfan He |
IEEE Signal Process. Lett. | 3 |
| 2025 | Multi-Scale Volumetric Transformers with Adaptive Uncertainty Modeling for Robust Bacterial Flagellar Motor Localization in Cryo-Electron TomographyabstractAutomated identification of bacterial flagellar motors in cryo-electron tomography (cryo-ET) remains a significant challenge due to low signal-to-noise ratios, imaging artifacts, and the complexity of capturing both local and global structural contexts. In this work, we propose MSVT+AUM, a hybrid framework that combines a Multi-Scale Volumetric Transformer (MSVT) with Adaptive Uncertainty Modeling (AUM) to address these limitations. MSVT employs a 3D ResNet-50 backbone to extract hierarchical volumetric features, which are integrated through a volume-level self-attention mechanism and refined via 3D Vision Transformer blocks with positional encoding. To enhance discriminative capability in densely packed molecular environments, we introduce a Contextual Spatial Attention module that adaptively reweights spatial and channel-wise representations. To reduce false positives and provide calibrated confidence estimates, AUM jointly models aleatoric and epistemic uncertainties using heteroscedastic regression and Monte Carlo dropout, with detection thresholds dynamically adjusted based on uncertainty estimates. Evaluation on the BYU flagellar motor dataset and an extended subset of the CryoET Data Portal shows that our method achieves an F 2 score of 0.891, representing a 1.5 % improvement over existing approaches, while reducing in-ference time by 23%. MSVT+AUM exhibits robust generalization across diverse bacterial species and noise conditions, offering a scalable and reliable solution for high-throughput structural analysis in cryo-ET. Junqiao Wang, Yangfan He, Xinyuan Song 0002, Peilai Yu, Xunfei Zhu |
BIBM | 3 |
| 2025 | ArtFormer: Controllable Generation of Diverse 3D Articulated ObjectsabstractThis paper presents a novel framework for modeling and conditional generation of 3D articulated objects. Troubled by flexibility-quality tradeoffs, existing methods are often limited to using predefined structures or retrieving shapes from static datasets. To address these challenges, we parameterize an articulated object as a tree of tokens and employ a transformer to generate both the object’s high-level geometry code and its kinematic relations. Subsequently, each sub-part’s geometry is further decoded using a signed-distance-function (SDF) shape prior, facilitating the synthesis of high-quality 3D shapes. Our approach enables the generation of diverse objects with high-quality geometry and varying number of parts. Comprehensive experiments on conditional generation from text descriptions demonstrate the effectiveness and flexibility of our method. Jiayi Su, Youhe Feng, Jinhua Song, Yangfan He, Botao Ren, Botian Xu |
CVPR | 5 |
| 2025 | MaRI: Material Retrieval Integration across DomainsabstractAccurate material retrieval is critical for creating realistic 3D assets. Existing methods rely on datasets that capture shape-invariant and lighting-varied representations of materials, which are scarce and face challenges due to limited diversity and inadequate real-world generalization. Most current approaches adopt traditional image search techniques. They fall short in capturing the unique properties of material spaces, leading to suboptimal performance in retrieval tasks. Addressing these challenges, we introduce MaRI, a framework designed to bridge the feature space gap between synthetic and real-world materials. MaRI constructs a shared embedding space that harmonizes visual and material attributes through a contrastive learning strategy by jointly training an image and a material encoder, bringing similar materials and images closer while separating dissimilar pairs within the feature space. To support this, we construct a comprehensive dataset comprising high-quality synthetic materials rendered with controlled shape variations and diverse lighting conditions, along with real-world materials processed and standardized using material transfer techniques. Extensive experiments demonstrate the superior performance, accuracy, and generalization capabilities of MaRI across diverse and complex material retrieval tasks, outperforming existing methods. Zhifei Yang 0004, Yangfan He, Huixiong Zhang |
CVPR | 3 |
| 2025 | DocAgent: An Agentic Framework for Multi-Modal Long-Context Document UnderstandingabstractRecent advances in large language models (LLMs) have demonstrated significant promise in document understanding and questionanswering.Despite the progress, existing approaches can only process short documents due to limited context length or fail to fully leverage multi-modal information.In this work, we introduce DocAgent, a multi-agent framework for long-context document understanding that imitates the human reading practice.Specifically, we first extract a structured, tree-formatted outline from documents to help agents identify relevant sections efficiently.Further, we develop an interactive reading interface that enables agents to query and retrieve various types of content dynamically.To ensure answer reliability, we introduce a reviewer agent that cross-checks responses using complementary sources and maintains a task-agnostic memory bank to facilitate knowledge sharing across tasks.We evaluate our method on two long-context document understanding benchmarks, where it bridges the gap to human-level performance by surpassing competitive baselines, while maintaining a short context length.Our code is available at https://github.com/lisun-ai/DocAgent. Shuyue Jia, Yangfan He, Chenyu You |
EMNLP | 4 |
| 2025 | GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them?abstractYiyang Zhou, Linjie Li, Shi Qiu, Zhengyuan Yang, Yuyang Zhao, Siwei Han, Yangfan He, Kangqi Li, Haonian Ji, Zihao Zhao, Haibo Tong, Lijuan Wang, Huaxiu Yao. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yiyang Zhou, Shi Qiu 0016, Zhengyuan Yang, Siwei Han, Yangfan He, Kangqi Li, Haonian Ji, Haibo Tong, Huaxiu Yao |
EMNLP | 7 |
| 2025 | FALCON: Feedback-driven Adaptive Long/short-term memory reinforced Coding OptimizatioNabstractRecently, large language models (LLMs) have achieved significant progress in automated code generation. Despite their strong instruction-following capabilities, these models frequently struggled to align with user intent in the coding scenario. In particular, they were hampered by datasets that lacked diversity and failed to address specialized tasks or edge cases. Furthermore, challenges in supervised fine-tuning (SFT) and reinforcement learning from human feedback (RLHF) led to failures in generating precise, human-intent-aligned code. To tackle these challenges and improve the code generation performance for automated programming systems, we propose Feedback-driven Adaptive Long/short-term memory reinforced Coding OptimizatioN (i.e., FALCON). FALCON leverages long-term memory to retain and apply learned knowledge, short-term memory to incorporate immediate feedback, and meta-reinforcement learning with feedback rewards to address global-local bi-level optimization and enhance adaptability across diverse code generation tasks. Extensive experiments show that FALCON achieves state-of-the-art performance, outperforming other reinforcement learning methods by over 4.5% on MBPP and 6.1% on Humaneval, with the code publicly available. https://anonymous.4open.science/r/FALCON-3B64/README.md. Yangfan He, Lewei He, Jianhui Wang 0001, Tianyu Shi 0003, Yuchen Li 0015, Qiuwu Chen |
ICME | 2 |
| 2025 | Scene-Aware Explainable Multimodal Trajectory PredictionabstractAdvancements in intelligent technologies have significantly improved navigation in complex traffic environments by enhancing environment perception and trajectory prediction for automated vehicles. However, current research often overlooks the joint reasoning of scenario agents and lacks explainability in trajectory prediction models, limiting their practical use in real-world situations. To address this, we introduce the Explainable Conditional Diffusion-based Multimodal Trajectory Prediction (DMTP) model, which is designed to elucidate the environmental factors influencing predictions and reveal the underlying mechanisms. Our model integrates a modified conditional diffusion approach to capture multimodal trajectory patterns and employs a revised Shapley Value model to assess the significance of global and scenario-specific features. Experiments using the Waymo Open Motion Dataset demonstrate that our explainable model excels in identifying critical inputs and significantly outperforms baseline models in accuracy. Moreover, the factors identified align with the human driving experience, underscoring the model's effectiveness in learning accurate predictions. Code is available in our open-source repository: https://github.com/ocean-luna/Explainable-Prediction. Junlan Chen, Yangfan He, Jun Ma 0008 |
ICRA | 6 |
| 2025 | Wcdt: World-Centric Diffusion Transformer for Traffic Scene GenerationabstractIn this paper, we introduce a novel approach for autonomous driving trajectory generation by harnessing the complementary strengths of diffusion probabilistic models (a.k.a., diffusion models) and transformers. Our proposed framework, termed the “World-centric Diffusion Transformer” (WcDT), optimizes the entire trajectory generation process, from feature extraction to model inference. To enhance the scene diversity and stochasticity, the historical trajectory data is first preprocessed into “Agent Move Statement” and encoded into latent space using Denoising Diffusion Probabilistic Models (DDPM) enhanced with Diffusion with Transformer (DiT) blocks. Then, the latent features, historical trajectories, HD map features, and historical traffic signal information are fused with various transformer-based encoders that is used to enhance the interaction of agents with other elements in the traffic scene. The encoded traffic scenes are then decoded by a trajectory decoder to generate multimodal future trajectories. Comprehensive experimental results show that the proposed approach exhibits superior performance in generating both realistic and diverse trajectories, showing its potential for integration into automatic driving simulation systems. Our code is available at https://github.com/yangchen1997/WcDT. Yangfan He, Aaron Xuxiang Tian, Dong Chen 0016, Arsalan Heydarian |
ICRA | 2 |
| 2025 | Free-Mask: A Novel Paradigm of Integration Between the Segmentation Diffusion Model and Image Editing
Bo Gao 0004, Jianhui Wang 0001, Xinyuan Song 0002, Yangfan He, Fangxu Xing, Tianyu Shi 0003 |
ACM Multimedia | 4 |
| 2025 | PurifyGen: A Risk-Discrimination and Semantic-Purification Model for Safe Text-to-Image GenerationabstractRecent advances in diffusion models have notably enhanced text-to-image (T2I) generation quality, but they also raise the risk of generating unsafe content. Traditional safety methods like text blacklisting or harmful content classification have significant drawbacks: they can be easily circumvented or require extensive datasets and extra training. To overcome these challenges, we introduce PurifyGen, a novel, training-free approach for safe T2I generation that retains the model's original weights. PurifyGen introduces a dual-stage strategy for prompt purification. First, we evaluate the safety of each token in a prompt by computing its complementary semantic distance, which measures the semantic proximity between the prompt tokens and concept embeddings from predefined toxic and clean lists. This enables fine-grained prompt classification without explicit keyword matching or retraining. Tokens closer to toxic concepts are flagged as risky. Second, for risky prompts, we apply a dual-space transformation: we project toxic-aligned embeddings into the null space of the toxic concept matrix, effectively removing harmful semantic components, and simultaneously align them into the range space of clean concepts. This dual alignment purifies risky prompts by both subtracting unsafe semantics and reinforcing safe ones, while retaining the original intent and coherence. We further define a token-wise strategy to selectively replace only risky token embeddings, ensuring minimal disruption to safe content. PurifyGen offers a plug-and-play solution with theoretical grounding and strong generalization to unseen prompts and models. Extensive testing shows that PurifyGen surpasses current methods in reducing unsafe content across five datasets and competes well with training-dependent approaches. Zongsheng Cao, Yangfan He, Jun Xie 0003, Zhepeng Wang 0002, Feng Chen 0044 |
ACM Multimedia | 2 |
| 2025 | CoFi-Dec: Hallucination-Resistant Decoding via Coarse-to-Fine Generative Feedback in Large Vision-Language ModelsabstractLarge Vision-Language Models (LVLMs) have achieved impressive progress in multi-modal understanding and generation. However, they still tend to produce hallucinated content that is inconsistent with the visual input, which limits their reliability in real-world applications. We propose CoFi-Dec, a training-free decoding framework that mitigates hallucinations by integrating generative self-feedback with coarse-to-fine visual conditioning. Inspired by the human visual process from global scene perception to detailed inspection, CoFi-Dec first generates two intermediate textual responses conditioned on coarse- and fine-grained views of the original image. These responses are then transformed into synthetic images using a text-to-image model, forming multi-level visual hypotheses that enrich grounding cues. To unify the predictions from these multiple visual conditions, we introduce a Wasserstein-based fusion mechanism that aligns their predictive distributions into a geometrically consistent decoding trajectory. This principled fusion reconciles high-level semantic consistency with fine-grained visual grounding, leading to more robust and faithful outputs. Extensive experiments on six hallucination-focused benchmarks show that CoFi-Dec substantially reduces both entity-level and semantic-level hallucinations, outperforming existing decoding strategies. The framework is model-agnostic, requires no additional training, and can be seamlessly applied to a wide range of LVLMs. Zongsheng Cao, Yangfan He, Jun Xie 0003, Zhepeng Wang 0002, Feng Chen 0044 |
ACM Multimedia | 2 |
| 2025 | TV-RAG: A Temporal-aware and Semantic Entropy-Weighted Framework for Long Video Retrieval and UnderstandingabstractLarge Video Language Models (LVLMs) have rapidly emerged as the focus of multimedia AI research. Nonetheless, when confronted with lengthy videos, these models struggle: their temporal windows are narrow, and they fail to notice fine-grained semantic shifts that unfold over extended durations. Moreover, mainstream text-based retrieval pipelines, which rely chiefly on surface-level lexical overlap, ignore the rich temporal interdependence among visual, audio, and subtitle channels. To mitigate these limitations, we propose TV-RAG, a training-free architecture that couples temporal alignment with entropy-guided semantics to improve long-video reasoning. The framework contributes two main mechanisms: (i) a time-decay retrieval module that injects explicit temporal offsets into the similarity computation, thereby ranking text queries according to their true multimedia context; and (ii) an entropy-weighted key-frame sampler that selects evenly spaced, information-dense frames, reducing redundancy while preserving representativeness. By weaving these temporal and semantic signals together, TV-RAG realises a dual-level reasoning routine that can be grafted onto any LVLM without re-training or fine-tuning. The resulting system offers a lightweight, budget-friendly upgrade path and consistently surpasses most leading baselines across established long-video benchmarks such as Video-MME, MLVU, and LongVideoBench, confirming the effectiveness of our model. Zongsheng Cao, Yangfan He, Jun Xie 0003, Feng Chen 0044, Zhepeng Wang 0002 |
ACM Multimedia | 2 |
| 2025 | Twin Co-Adaptive Dialogue for Progressive Image GenerationabstractModern text-to-image generation systems have enabled the creation of remarkably realistic and high-quality visuals, yet they often falter when handling the inherent ambiguities in user prompts. In this work, we present Twin-Co, a framework that leverages synchronized, co-adaptive dialogue to progressively refine image generation. Instead of a static generation process, Twin-Co employs a dynamic, iterative workflow where an intelligent dialogue agent continuously interacts with the user. Initially, a base image is generated from the user's prompt. Then, through a series of synchronized dialogue exchanges, the system adapts and optimizes the image according to evolving user feedback. The co-adaptive process allows the system to progressively narrow down ambiguities and better align with user intent. Experiments demonstrate that Twin-Co not only enhances user experience by reducing trial-and-error iterations but also improves the quality of the generated images, streamlining creative process across various applications. Jianhui Wang 0001, Yangfan He, Yan Zhong 0001, Xinyuan Song 0002, Jiayi Su, Yuheng Feng, Hongyang He, Wenyu Zhu, Xinhang Yuan, Miao Zhang 0010, Tianyu Shi 0003, Xueqian Wang 0001 |
ACM Multimedia | 2 |
| 2025 | TRiCo: Triadic Game-Theoretic Co-Training for Robust Semi-Supervised LearningabstractWe introduce TRiCo, a novel triadic game-theoretic co-training framework that rethinks the structure of semi-supervised learning by incorporating a teacher, two students, and an adversarial generator into a unified training paradigm. Unlike existing co-training or teacher-student approaches, TRiCo formulates SSL as a structured interaction among three roles: (i) two student classifiers trained on frozen, complementary representations, (ii) a meta-learned teacher that adaptively regulates pseudo-label selection and loss balancing via validation-based feedback, and (iii) a non-parametric generator that perturbs embeddings to uncover decision boundary weaknesses. Pseudo-labels are selected based on mutual information rather than confidence, providing a more robust measure of epistemic uncertainty. This triadic interaction is formalized as a Stackelberg game, where the teacher leads strategy optimization and students follow under adversarial perturbations. By addressing key limitations in existing SSL frameworks—such as static view interactions, unreliable pseudo-labels, and lack of hard sample modeling—TRiCo provides a principled and generalizable solution. Extensive experiments on CIFAR-10, SVHN, STL-10, and ImageNet demonstrate that TRiCo consistently achieves state-of-the-art performance in low-label regimes, while remaining architecture-agnostic and compatible with frozen vision backbones. Hongyang He, Xinyuan Song 0002, Yangfan He, Yanshu Li, Haochen You, Lifan Sun, Wenqiao Zhang |
NeurIPS | 3 |
| 2025 | SRPO: Enhancing Multimodal LLM Reasoning via Reflection-Aware Reinforcement LearningabstractMultimodal large language models (MLLMs) have shown promising capabilities in reasoning tasks, yet still struggle significantly with complex problems requiring explicit self-reflection and self-correction, especially compared to their unimodal text-based counterparts. Existing reflection methods are simplistic and struggle to generate meaningful, instructive feedback, as the reasoning ability and knowledge limits of pre-trained models are largely fixed during initial training. To overcome these challenges, we propose \textit{multimodal \textbf{S}elf-\textbf{R}eflection enhanced reasoning with Group Relative \textbf{P}olicy \textbf{O}ptimization} \textbf{SRPO}, a two-stage reflection-aware reinforcement learning (RL) framework explicitly designed to enhance multimodal LLM reasoning. In the first stage, we construct a high-quality, reflection-focused dataset under the guidance of an advanced MLLM, which generates reflections based on initial responses to help the policy model to learn both reasoning and self-reflection. In the second stage, we introduce a novel reward mechanism within the GRPO framework that encourages concise and cognitively meaningful reflection while avoiding redundancy. Extensive experiments across multiple multimodal reasoning benchmarks—including MathVista, MathVision, Mathverse, and MMMU-Pro—using Qwen-2.5-VL-7B and Qwen-2.5-VL-32B demonstrate that SRPO significantly outperforms state-of-the-art models, achieving notable improvements in both reasoning accuracy and reflection quality. Zhongwei Wan, Zhihao Dou, Che Liu 0002, Yu Zhang 0133, Dongfei Cui, Qinjian Zhao, Hui Shen 0008, Yi Xin 0003, Chaofan Tao, Yangfan He, Mi Zhang 0002, Shen Yan 0008 |
NeurIPS | 12 |
| 2025 | ReAgent-V: A Reward-Driven Multi-Agent Framework for Video UnderstandingabstractVideo understanding is fundamental to tasks such as action recognition, video reasoning, and robotic control. Early video understanding methods based on large vision-language models (LVLMs) typically adopt a single-pass reasoning paradigm without dynamic feedback, limiting the model’s capacity to self-correct and adapt in complex scenarios. Recent efforts have attempted to address this limitation by incorporating reward models and reinforcement learning to enhance reasoning, or by employing tool-agent frameworks. However, these approaches face several challenges, including high annotation costs, reward signals that fail to capture real-time reasoning states, and low inference efficiency. To overcome these issues, we propose ReAgent-V, a novel agentic video understanding framework that integrates efficient frame selection with real-time reward generation during inference. These reward signals not only guide iterative answer refinement through a multi-perspective reflection mechanism—adjusting predictions from conservative, neutral, and aggressive viewpoints—but also enable automatic filtering of high-quality data for supervised fine-tuning (SFT), direct preference optimization (DPO), and group relative policy optimization (GRPO). ReAgent-V is lightweight, modular, and extensible, supporting flexible tool integration tailored to diverse tasks. Extensive experiments on 12 datasets across three core applications—video understanding, video reasoning enhancement, and vision-language-action model alignment—demonstrate significant gains in generalization and reasoning, with improvements of up to 6.9%, 2.1%, and 9.8%, respectively, highlighting the effectiveness and versatility of the proposed framework. Yiyang Zhou, Yangfan He, Yaofeng Su, Siwei Han, Joel Jang, Gedas Bertasius, Mohit Bansal, Huaxiu Yao |
NeurIPS | 2 |
| 2025 | Multi-Modal Large Language Model with RAG Strategies in Soccer Commentary GenerationabstractAs a globally celebrated sport, soccer has seen its appeal greatly amplified by engaging and vivid commentary. Recently, Multi-Modal Large Language Models (MLLMs) have attracted attention in generating soccer commentaries due to their remarkable capacities of understanding different modalities of the input videos. Most of these methods have shown that the use of multiple modalities can enhance the commentary quality, which includes video, audio, and structured meta-data. However, delivering precise and rich commentary requires the ability to accurately discern sub-tle differences in similar backgrounds, events, and players. This presents a significant challenge for existing MLLMs. So we propose SoccerComment, a framework for generating soccer commentary that integrates MLLMs with Retrieval-Augmented Generation (RAG) strategies. This framework enhances inference efficiency and reduces the need for continuous training through a multimodal clustering memory unit and retrieval-augmented in-context learning mechanisms, ultimately improving the accuracy and diversity of the commentary. Based on similar retrieved scenarios, SoccerComment demonstrates outstanding zero-shot performance, offering a new direction and scalable solution for future research in soccer commentary generation. Yangfan He, Shuaishuai Zu, Yiting Xie |
WACV | 2 |
| 2025 | UniTMGE: Uniform Text-Motion Generation and Editing Model via DiffusionabstractCurrent methods have shown promising results in applying diffusion models to motion generation given text input. However, these methods are limited to unimodal inputs and outputs, restricted to motion generation alone, and lacking multimodal control capabilities. To address these issues, we introduce UniTMGE, a text-motion multimodal generation and editing framework based on diffusion. UniTMGE overcomes single-modality limitations, enabling exceptional performance and strong generalization across multiple tasks like text-driven motion generation, motion captioning, motion completion, and multimodal motion editing. UniTMGE comprises three components: UTMV for mapping text and motion into a shared latent space using contrastive learning, a controllable diffusion model customized for the UTMV space, and MCRE for unifying multimodal conditions into CLIP representations, enabling precise multimodal control and flexible motion editing through simple linear operations. We conducted both closed-world experiments and open-world experiments using the Motion-X dataset with detailed text descriptions, with results demonstrating our model's effectiveness and generalizability across multiple tasks. Yangfan He, Tengjiao Sun |
WACV | 2 |
| 2025 | OMR-diffusion: Optimizing multi-round enhanced training in diffusion models for improved intent understanding
Kun Li 0014, Jianhui Wang 0001, Yangfan He, Miao Zhang 0010, Xueqian Wang 0001 |
Neurocomputing | 3 |
| 2025 | SAGE: Self-evolving Agents with Reflective and Memory-augmented Abilities
Xuechen Liang, Meiling Tao, Yinghui Xia, Jianhui Wang 0001, Kun Li 0014, Yangfan He, Jingsong Yang, Tianyu Shi 0003, Yuantao Wang, Miao Zhang 0010, Xueqian Wang 0001 |
Neurocomputing | 7 |
| 2025 | MDANet: A multi-stage domain adaptation framework for generalizable low-light image enhancement
Jianhui Wang 0001, Yangfan He, Kun Li 0014, Miao Zhang 0010, Tianyu Shi 0003, Xueqian Wang 0001 |
Neurocomputing | 2 |
| 2025 | Enhancing intent understanding for ambiguous prompt: A human-machine co-adaption strategy
Yangfan He, Jianhui Wang 0001, Kun Li 0014, Li Sun 0010, Miao Zhang 0010, Xueqian Wang 0001 |
Neurocomputing | 2 |
| 2024 | Cross Metaplectic Wigner Distribution: Definition, Properties, Relation to Short-Time Metaplectic Transform, and Uncertainty PrinciplesabstractThe metaplectic operator has shown to be a valid technique for generalizing the notion of cross Wigner distribution to achieve time-frequency superresolution. Inspired by the latest work (Cordero and Rodino, 2022), we revisit the notion of cross Wigner distribution in metaplectic transform domains by introducing two$2N\times 2N$symplectic matrices rather than integrating them into one$4N\times 4N$symplectic matrix. We style the derived general formulation as the cross metaplectic Wigner distribution and obtain its basic properties including time translation property, frequency modulation property, time translation and frequency modulation property, Moyal formula, complex conjugate symmetry, time reversal symmetry, and scaling property. We use it to define the so-called short-time metaplectic transform and clarify the equivalence between them. We establish the standard Heisenberg’s uncertainty principles for the cross metaplectic Wigner distribution, i.e., an attainable lower bound for two real-valued functions in cross metaplectic Wigner distribution domains and a sequence of attainable (unattainable) lower bounds for two complex-valued (one real-valued and the other complex-valued) functions in orthogonal, orthonormal, the minimum eigenvalue commutative and the maximum eigenvalue commutative cross metaplectic Wigner distribution domains. We further demonstrate the time-frequency superresolution superiority of the derived results over the conventional one through theoretical analyses and numerical experiments. Dong Li 0009, Yangfan He, Weiguo Huang |
IEEE Trans. Inf. Theory | 4 |
| 2023 | Wigner distribution associated with the symplectic coordinates transformation
Yangfan He |
Signal Process. | 2 |
| 2023 | K-Wigner Distribution: Definition, Uncertainty Principles and Time-Frequency AnalysisabstractTo tackle a challenge in high-dimensional complex features information processing, this study extends the permanent scale Wigner distribution and the single scale$k$-Wigner distribution (kWD, formerly known as$\tau $-Wigner distribution) to a novel multiscale parameterized Wigner distribution. That is the so-called$\mathbf {K}$-Wigner distribution (KWD) which is able to use different scales to extract different types of features at different dimensions. Heisenberg-type uncertainty inequalities of the KWD are established, giving rise to the tightest universal attainable lower bound for all functions on the uncertainty product in time-KWD and Fouier transform-KWD domains, and two versions of attainable lower bounds for complex-valued functions. The obtained results solve an important concern regarding the limit of the KWD’s time-frequency resolution influenced by the parameter matrix. As an application, the derived uncertainty inequalities are applied to estimate the bandwidth in KWD domains. The time-frequency resolution performance of the multiscale KWD, as compared with that of the single scale kWD, is investigated in details. The optimal parameter matrix of the KWD achieving the best performance is then generated, which solves an important concern regarding the KWD’s parameter matrix selection. Examples are also carried out to demonstrate the usefulness and effectiveness of the proposed technique. Dong Li 0009, Yangfan He, Jianwei Zhang 0005, Chengxi Zhou |
IEEE Trans. Inf. Theory | 3 |
| 2023 | Free Metaplectic Wigner Distribution: Definition and Heisenberg's Uncertainty PrinciplesabstractInspired by a definition of the closed-form instantaneous cross-correlation Wigner distribution (Zhang, 2019), we generalize the notion of Wigner distribution to the so-called free metaplectic Wigner distribution (FMWD) through three free metaplectic transforms, in order to tackle a challenge in high-dimensional complex features information processing. We provide some representative special cases for this general form, including the$N$-dimensional nonseparable affine characteristic Wigner distribution, kernel function Wigner distribution, convolution representation Wigner distribution and instantaneous cross-correlation Wigner distribution. We establish the standard Heisenberg’s uncertainty principles (HUPs) of the real-valued function for the FMWD. We also establish the standard HUPs of the complex-valued function for some specific (i.e., the orthogonal, the orthonormal, the minimum eigenvalue commutative and the maximum eigenvalue commutative) FMWDs. In view of the one-dimensional case of our results, we solve a burning question regarding the limit of time-frequency superresolution triggered by the linear canonical transform free parameters. Zhicheng Zhu, Dong Li 0009, Yangfan He |
IEEE Trans. Inf. Theory | 4 |
| 2022 | The optimal k-Wigner distribution
Yangfan He, Jianwei Zhang 0005, Chengxi Zhou |
Signal Process. | 2 |
| 2020 | COSINE: a software development model integrating collective intelligence, service and ecosystemabstractWith the development of the internet technology, a large amount of softwares have emerged to meet users' increasing needs. At the mean time, software systems have been faced with a problem that they must adapt to the dynamic network environment. It is obvious that a variety of software development models have been proposed in the past few decades. However, the majority of these methods are gradually unadaptable to new circumstances. In this paper, we proposed a new software development model integrating collective intelligence, service and ecosystem. On the one hand, we have introduced the model in detail. On the other hand, We took a practical example to demonstrate the effectiveness of the proposed model. Tianjing Hong, Jian Cao 0001, Haijun Zhang 0002, Changhai Nie, Bo Cheng 0001, Yangfan He, Li Kuang, Dun-Wei Gong, Wuhui Chen, Yuliang Shi, Deyi Huang |
SERVICES | 7 |
| 2020 | Software Construction Oriented Multi-agent Collaborative Modeling and SimulationabstractImproper collaboration frequency causes low efficiency in software development. To improve collaboration efficiency in software construction, we propose a method of multi-agent collaborative modeling and simulation. First, we propose collective action model and individual action model to describe the whole construction environment. Then, we design agents in our simulation, establish a mapping relationship between artificial bee colonies and developers group, and propose buzz factor as a parameter to describe the collaboration frequency. In particular, we introduce Q-learning algorithm to obtain the dynamic buzz factor that maximizes construction efficiency. Finally, we perform experiments on groups of agents based on our model to observe the effect of buzz factor. Comparing to the classic feature-oriented stigmergy-based model, our experiment results show a better construction efficiency and provide a reference for controlling the collaboration. Through our work, we can find a suitable collaboration frequency through the simulation, to guide the practice of software development. Yangfan He, Bing Li 0010 |
SERVICES | 3 |
| 2016 | RE_PROV: Modeling Requirement Provenance with PROVabstractRequirements are complex by nature. When describing something that has not been realized, users may find it difficult to interpret it accurately. Many problems cannot be discovered until some abstract concepts become concrete and some details have been confirmed. So checking the origination and processing history of the requirements and making adjustments are very normal practice. Requirement provenance records which have well-defined syntax and explicit semantics can provide solid support for requirement traceability analysis. By setting up the mapping between PROV and a typical requirement framework, this paper proposes RE_PROV, a novel model for the description of requirement provenance. The rationales of the mapping are explained and a comprehend example of an embedded software design is provided to show possible usage. Yangfan He |
APSEC | 1 |
| 2009 | A Contextual Information Acquisition Approach Based on Semantics and Mashup Technology
Yangfan He, Keqing He 0002, Xiuhong Chen |
CloudCom | 1 |
| 2008 | A new Evolutionary Algorithm based on quantum statistical mechanicsabstractA new evolutionary algorithm based on quantum statistical mechanics (QSEA) is raised in this paper. In the algorithm, the whole evolutionary system is treated as a quantum statistical system, where quantum coding is adopted to express chromosomes, and superposition of quantum bits is used to simulate the linear superposition state of the system. Quantum system entropy and statistical energy have been defined by analogy with corresponding concepts in quantum statistical mechanics. And the competition between quantum statistical energy and entropy of the system is used to simulate the conflict between dasiaselection pressurepsila and dasiadiversity of populationpsila, which helps the algorithm to keep a delicate balance between these two issues, and obtain optimal solution rapidly. Numerical experiments show that this new algorithm has high efficiency and strong ability to get global optimal solution. Dazhi Jiang, Yangfan He, Xingyan Huang |
IEEE Congress on Evolutionary Computation | 4 |