VLDB 2026 Research / reviewers in the wild / expert
Zhongjiang He
dblp:348/6925
· DBLP profile ↗
34ranked-venue papers
1as first author
34since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 1 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 18 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adaptive Evidential Learning for Temporal-Semantic Robustness in Moment RetrievalabstractIn the domain of moment retrieval, accurately identifying temporal segments within videos based on natural language queries remains challenging. Traditional methods often employ pre-trained models that struggle with fine-grained information and deterministic reasoning, leading to difficulties in aligning with complex or ambiguous moments. To overcome these limitations, we explore Deep Evidential Regression (DER) to construct a vanilla Evidential baseline. However, this approach encounters two major issues: the inability to effectively handle modality imbalance and the structural differences in DER's heuristic uncertainty regularizer, which adversely affect uncertainty estimation. This misalignment results in high uncertainty being incorrectly associated with accurate samples rather than challenging ones. Our observations indicate that existing methods lack the adaptability required for complex video scenarios. In response, we propose Debiased Evidential Learning for Moment Retrieval (DEMR), a novel framework that incorporates a Reflective Flipped Fusion (RFF) block for cross-modal alignment and a query reconstruction task to enhance text sensitivity, thereby reducing bias in uncertainty estimation. Additionally, we introduce a Geom-regularizer to refine uncertainty predictions, enabling adaptive alignment with difficult moments and improving retrieval accuracy. Extensive testing on standard datasets and debiased datasets ActivityNet-CD and Charades-CD demonstrates significant enhancements in effectiveness, robustness, and interpretability, positioning our approach as a promising solution for temporal-semantic robustness in moment retrieval. Haojian Huang, Kaijing Ma, Xianghao Zang, Han Fang 0002, Chao Ban, Hao Sun 0038, Mulin Chen, Zhongjiang He |
AAAI | 11 |
| 2026 | Introducing Visual Scenes and Reasoning: A More Realistic Benchmark for Spoken Language UnderstandingabstractSpoken Language Understanding (SLU) consists of two sub-tasks: intent detection (ID) and slot filling (SF). Given its broad range of real-world applications, enhancing SLU for practical deployment is increasingly critical. Profile-based SLU addresses ambiguous user utterances by incorporating context awareness (CA), user profiles (UP), and knowledge graphs (KG) to support disambiguation, thereby advancing SLU research toward real-world applicability. However, existing SLU datasets still fall short in representing real-world scenarios. Specifically, (1) CA uses one-hot vectors for representation, which is overly idealized, and (2) models typically focuses solely on predicting intents and slot labels, neglecting the reasoning process that could enhance performance and interpretability. To overcome these limitations, we introduce VRSLU, a novel SLU dataset that integrates both Visual images and explicit Reasoning. For over-idealized CA, we use GPT-4o and FLUX.1-dev to generate images reflecting users’ environments and statuses, followed by human verification to ensure quality. For reasoning, GPT-4o is employed to generate explanations for predicted labels, which are then refined by human annotators to ensure accuracy and coherence. Additionally, we propose an instructional template, LR-Instruct, which first predicts labels and then generates corresponding reasoning. This two-step approach helps mitigate the influence of reasoning bias on label prediction. Experimental results confirm the effectiveness of incorporating visual information and highlight the promise of explicit reasoning in advancing SLU. Di Wu 0088, Liting Jiang, Ruiyu Fang, Bianjing, Hongyan Xie, Haoxiang Su, Hao Huang 0009, Zhongjiang He, Shuangyong Song, Xuelong Li 0001 |
AAAI | 8 |
| 2026 | VAV-R1: Difficulty-Aware Multimodal Reasoning for Video Anomaly ValidationabstractVideo anomaly validation (VAV) serves as the final alarm validation step for false positive filtering in video anomaly detection (VAD), requiring both higher accuracy in anomaly identification and stronger justifications. Despite rapid VAD advancements, existing methods still lack sufficient interpretability and struggle with challenging cases. To address this, we propose VAV-R1, a multimodal reasoning model tailored for VAV. We first construct ThinkVAV, a dedicated benchmark with fine-grained reasoning annotations across diverse anomaly types. Furthermore, we introduce DA-GRPO, a difficulty-aware reinforcement learning strategy that prioritizes learning from more challenging cases. Extensive experiments demonstrate that VAV-R1 achieves SOTA performance across multiple tasks. Qianhao Ren, Yutong Wang 0001, Zhongjiang He, Jingmin Xin, Hao Sun 0038 |
ICMR | 6 |
| 2026 | Unifying Granularity and Reliability: A Robust and Efficient Framework for Text-based Person RetrievalabstractText-based person retrieval (TPR) has become a crucial task in cross-modal retrieval due to its broad application in fields such as public safety and criminal investigation. Existing TPR methods typically rely on fully fine-tuning large-scale pretrained vision-language models like CLIP, which incurs high computational costs and tends to exhibit poor generalization in unseen domains due to overfitting. Fortunately, Parameter-Efficient Transfer Learning (PETL) has emerged as a lightweight alternative. However, applying PETL to TPR remains challenging, as its limited adaptation capacity struggles to capture intricate identity cues and becomes highly susceptible to gradient interference from unreliable image-text pairs. To address these challenges, we present a PETL-based framework named UniGR that unifies granularity and reliability for robust and efficient TPR. Specifically, we design a multi-granularity relational adapter (MRA) to capture both coarse-grained global and fine-grained local relational features among tokens, equipping the generic backbone with the task-specific, precise understanding needed for TPR. To combat the noise sensitivity of PETL, a reliability-aware reweighting strategy (RRS) is introduced to adaptively down-weight unreliable samples during training. Furthermore, we propose a parameter-free cross-modal cyclic verification (CMCV) module to mitigate ambiguities in cross-modal matching computations and refine retrieval ranking further. Experiments on benchmarks corroborate the superiority of UniGR among parameter-efficient methods. Remarkably, with only 4.5% of trainable parameters, UniGR outperforms most fully fine-tuned methods while maintaining strong generalization. Jingchen Hao, Zhen Peng 0005, Yuting Zhang 0007, Zhongjiang He, Weizhan Zhang, Hao Sun 0038 |
SIGIR | 5 |
| 2026 | MEGE: A mixed emotion graph model for empathetic dialogue generation
Deji Zhao, Donghong Han, Ye Yuan 0001, Bo Ning 0002, Zhongjiang He, Chao Wang 0057, Shuangyong Song |
Neural Networks | 6 |
| 2026 | Benchmarking Semantic Segmentation Models via Appearance and Geometry Attribute EditingabstractSemantic segmentation takes a pivotal role in various applications such as autonomous driving and medical image analysis. When deploying segmentation models in practice, it is critical to test their behaviors in varied and complex scenes in advance. In this paper, we construct an automatic data generation pipeline Gen4Seg to stress-test semantic segmentation models by generating various challenging samples with different attribute changes. Beyond previous evaluation paradigms focusing solely on global weather and style transfer, we investigate variations in both appearance and geometry attributes at the object and image level. These include object color, material, size, and position, as well as image-level variations such as weather and style. To achieve this, we propose to edit visual attributes of existing real images with precise control of structural information, empowered by diffusion models. In this way, the existing segmentation labels can be reused for the edited images, which greatly reduces the labor costs of constructing datasets. Using our pipeline, we construct two new benchmarks, Pascal-EA and COCO-EA. We benchmark a broad variety of semantic segmentation models, spanning from conventional close-set models to recent open-vocabulary large models. We have several key findings: 1) advanced open-vocabulary models do not exhibit greater robustness compared to closed-set methods under geometric variations; 2) traditional data augmentation techniques, such as CutOut and CutMix, are limited in enhancing robustness against appearance variations; 3) our generation pipeline can also be employed as a data augmentation tool and improve both in-distribution and out-of-distribution performances. Our work suggests the potential of generative models as effective tools for automatically analyzing segmentation models, and we hope our findings will assist practitioners and researchers in developing more robust and reliable segmentation models. Zijin Yin, Bing Li 0015, Kongming Liang, Hao Sun 0015, Zhongjiang He, Zhanyu Ma, Jun Guo 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Trusted Unified Feature-Neighborhood Dynamics for Multi-View ClassificationabstractMulti-view classification (MVC) faces inherent challenges due to domain gaps and inconsistencies across different views, often resulting in uncertainties during the fusion process. While Evidential Deep Learning (EDL) has been effective in addressing view uncertainty, existing methods predominantly rely on the Dempster-Shafer combination rule, which is sensitive to conflicting evidence and often neglects the critical role of neighborhood structures within multi-view data. To address these limitations, we propose a Trusted Unified Feature-NEighborhood Dynamics (TUNED) model for robust MVC. This method effectively integrates local and global feature-neighborhood (F-N) structures for robust decision-making. Specifically, we begin by extracting local F-N structures within each view. To further mitigate potential uncertainties and conflicts in multi-view fusion, we employ a selective Markov random field that adaptively manages cross-view neighborhood dependencies. Additionally, we employ a shared parameterized evidence extractor that learns global consensus conditioned on local F-N structures, thereby enhancing the global integration of multi-view features. Experiments on benchmark datasets show that our method improves accuracy and robustness over existing approaches, particularly in scenarios with high uncertainty and conflicting views. Haojian Huang, Chuanyu Qin, Zhe Liu 0041, Kaijing Ma, Han Fang 0002, Chao Ban, Hao Sun 0038, Zhongjiang He |
AAAI | 9 |
| 2025 | MR-SQL: Multi-level Retrieval Enhances Inference for LLM in Text-to-SQL
Zhenhe Wu, Zhongqiu Li, Mengxiang Li, Zhongjiang He, Jian Yang 0003, Yu Zhao 0007, Ruiyu Fang, Zhoujun Li 0001, Shuangyong Song |
DASFAA (2) | 5 |
| 2025 | ViCo: A Multitask Video-enhanced and Cognition-preserving Modality Alignment Training FrameworkabstractThe rapid development of multimodal large language models (MLLMs) has brought significant breakthroughs to this field. However, current MLLMs typically rely on vision instruction tuning based on large language models (LLMs) to endow them with multimodal capabilities, which may lead to low video utilization and catastrophic forgetting due to the task-specific nature of vision instruction tuning. Take case of multiple-choice Video Question Answering (VideoQA) as an example, requiring only selecting correct answers makes it easier for LLMs to shortcut rather than fully comprehend video content. Moreover, the fixed task format of ’’predicting the correct option" risks catastrophic forgetting, which can cause LLMs to lose their origin cognitive abilities. To address this, we propose ViCo, a training framework that enables LLMs to perform a specific auxiliary tasks during vision instruction tuning, which helps the model leverage video information more thoroughly while mitigating catastrophic forgetting, thereby improving modality alignment and enhancing performance. We evaluate our method on several strong VideoQA benchmarks, where it outperforms all state-of-the-art methods, demonstrating its effectiveness and generality. Additionally, we also provide extensive analysis on the video information utilization and catastrophic forgetting resulting. The code will be made available at https://github.com/dunknsabsw/ViCo. Zhenda Yu, Jiayu Shen, Lanxiang Zhou, Han Fang 0002, Xianghao Zang, Chao Ban, Jingfeng Chen, Zhongjiang He, Hao Sun 0038, Yanmei Kang |
ICASSP | 9 |
| 2025 | FairHuman: Boosting Hand and Face Quality in Human Image Generation with Minimum Potential Delay Fairness in Diffusion Models
Tianwei Cao, Zhongjiang He, Kongming Liang, Zhanyu Ma |
ICCV | 4 |
| 2025 | SelectVision: Adaptive Vision Resolution Selection for Visual Document Understanding
Zhongjiang He, Han Fang 0002, Hao Sun 0015, Kongming Liang, Zhanyu Ma |
ICDAR (4) | 1 |
| 2025 | When Less is More: Minimal Prompts with LoRA for LLM Text Detection
Shiquan Wang, Ruiyu Fang, Mengxiang Li, Zhongjiang He, Shuangyong Song |
NLPCC (4) | 4 |
| 2025 | Empathetic Dialogue Generation with LLMs for Emotional Support
Shiquan Wang, Ruiyu Fang, Mengxiang Li, Zhongjiang He, Shuangyong Song |
NLPCC (4) | 4 |
| 2025 | Animal-CLIP: A Dual-Prompt Enhanced Vision-Language Model for Animal Action Recognition
Yinuo Jing, Kongming Liang, Ruxu Zhang, Hao Sun 0015, Zhongjiang He, Zhanyu Ma |
Int. J. Comput. Vis. | 6 |
| 2025 | RAICL-DSC: Retrieval-Augmented In-Context Learning for Dialogue State Correction
Haoxiang Su, Hongyan Xie, Di Wu 0088, Liting Jiang, Hao Huang 0009, Zhongjiang He, Ruiyu Fang, Shuangyong Song |
Knowl. Based Syst. | 7 |
| 2025 | Enhancing math reasoning ability of large language models via computation logic graphs
Deji Zhao, Donghong Han, Jia Wu 0001, Zhongjiang He, Bo Ning 0002, Ye Yuan 0001, Chao Wang 0057, Shuangyong Song |
Knowl. Based Syst. | 4 |
| 2025 | Reserve to Adapt: Mining Inter-Class Relations for Open-Set Domain AdaptationabstractOpen-Set Domain Adaptation (OSDA) aims at adapting a model trained on a labelled source domain, to an unlabeled target domain that is corrupted with unknown classes. The key challenge inherent to this open-set setting is therefore how best to avoid the negative transfer incurred by unknown classes during model adaptation. Most existing works tackle this challenge by simply pushing the entire unknown classes away. In this paper, we take a different stance - instead of addressing these unknown classes as a single entity, we "reserve" in-between spaces for their subsets in the learned embedding. Our key finding is that the inter-class relations learned off the source domain, can help to enforce class separations in the target domain - thereby reserving spaces for unknown classes. More specifically, we first prep the "reservation" by tightening the known-class representations while enlarging their inter-class margin. We then learn soft-label prototypes in the source domain to facilitate the discrimination of known and unknown samples in the target domain. It follows that these two steps are iterated at each epoch in a mutually beneficial manner - better discrimination of unknown samples helps with space reservation, and vice versa. We show state-of-the-art results on four standard OSDA datasets, Office-31, Office-Home, VisDA and ImageCLEF, and conduct further analysis to help understand our method. Codes are available at: https://github.com/PRIS-CV/Reserve_to_Adapt. Yujun Tong, Dongliang Chang, Da Li 0001, Kongming Liang, Zhongjiang He, Yi-Zhe Song, Zhanyu Ma |
IEEE Trans. Image Process. | 6 |
| 2025 | Detailed Object Description With Controllable DimensionsabstractObject description plays an important role for visually impaired individuals to understand and compare the differences between objects. Recent multimodal large language models (MLLMs) exhibit powerful perceptual abilities and demonstrate impressive potential for generating object-centric descriptions. However, the descriptions generated by such models may still usually contain a lot of content that is not relevant to the user intent or miss some important object dimension details. Under special scenarios, users may only need the details of certain dimensions of an object. In this paper, we propose a training-free object description refinement pipeline,Dimension Tailor, designed to enhance user-specified details in object descriptions. This pipeline includes three steps: dimension extracting, erasing, and supplementing, which decompose the description into user-specified dimensions. Dimension Tailor can not only improve the quality of object details but also offer flexibility in including or excluding specific dimensions based on user preferences. We conducted extensive experiments to demonstrate the effectiveness of Dimension Tailor on controllable object descriptions. Notably, the proposed pipeline can consistently improve the performance of the recent MLLMs. The code is currently accessible athttps://github.com/PRIS-CV/ControllableObjectDescription. Haiwen Zhang, Baoteng Li, Kongming Liang, Hao Sun 0015, Zhongjiang He, Zhanyu Ma, Jun Guo 0002 |
IEEE Trans. Multim. | 6 |
| 2024 | Referred by Multi-Modality: A Unified Temporal Transformer for Video Object SegmentationabstractRecently, video object segmentation (VOS) referred by multi-modal signals, e.g., language and audio, has evoked increasing attention in both industry and academia. It is challenging for exploring the semantic alignment within modalities and the visual correspondence across frames. However, existing methods adopt separate network architectures for different modalities, and neglect the inter-frame temporal interaction with references. In this paper, we propose MUTR, a Multi-modal Unified Temporal transformer for Referring video object segmentation. With a unified framework for the first time, MUTR adopts a DETR-style transformer and is capable of segmenting video objects designated by either text or audio reference. Specifically, we introduce two strategies to fully explore the temporal relations between videos and multi-modal signals. Firstly, for low-level temporal aggregation before the transformer, we enable the multi-modal references to capture multi-scale visual cues from consecutive video frames. This effectively endows the text or audio signals with temporal knowledge and boosts the semantic alignment between modalities. Secondly, for high-level temporal interaction after the transformer, we conduct inter-frame feature communication for different object embeddings, contributing to better object-wise correspondence for tracking along the video. On Ref-YouTube-VOS and AVSBench datasets with respective text and audio references, MUTR achieves +4.2% and +8.7% J&F improvements to state-of-the-art methods, demonstrating our significance for unified multi-modal VOS. Code is released at https://github.com/OpenGVLab/MUTR. Shilin Yan, Renrui Zhang, Wei Zhang 0016, Hongyang Li 0001, Yu Qiao 0001, Hao Dong 0003, Zhongjiang He, Peng Gao 0007 |
AAAI | 9 |
| 2024 | Class-Aware Contrastive Learning for Fine-Grained Skeleton-Based Action Recognition
Xinyu Bian, Dongliang Chang, Zhongjiang He, Kongming Liang, Zhanyu Ma |
ACCV (1) | 4 |
| 2024 | Hierarchical Prompting for Diffusion Classifiers
Wenxin Ning, Dongliang Chang, Yujun Tong, Zhongjiang He, Kongming Liang, Zhanyu Ma |
ACCV (8) | 4 |
| 2024 | Domain-Slot Aware Contrastive Learning for Improved Dialogue State TrackingabstractLarge-scale pre-trained neural language model has facilitated to achieve the state-of-the-art performance on Dialogue State Tracking (DST) tasks. One of the existing works models the semantic correlation between the dialogue context and (domain, slot) pair encoded by BERT and make the prediction. Despite the effectiveness, they ignore the fact that there is no perfect semantic correspondence between (domain, slot) pair and the dialogue context. In this paper, we propose a domain-slot aware contrastive learning framework to solve this problem, which proposes three methods to bridge the semantic gap between the dialogue context and the (domain, slot) by constructing training sample pairs to fine-tune the BERT model and use it for base DST model. The experiments demonstrate that our proposed method has improved the performance of the baseline model on the MultiWOZ2.1 and MultiWOZ2.4 datasets, yielding competitive results. Haoxiang Su, Sijie Feng, Hongyan Xie, Di Wu 0088, Hao Huang 0009, Zhongjiang He, Shuangyong Song, Ruiyu Fang, Xiaomeng Huang, Wushour Slamu |
ICASSP | 6 |
| 2024 | ProTA: Probabilistic Token Aggregation for Text-Video RetrievalabstractText-video retrieval aims to find the most relevant cross-modal samples for a given query. Recent methods focus on modeling the whole spatial-temporal relations. However, since video clips contain more diverse content than captions, the model aligning these asymmetric video-text pairs has a high risk of retrieving many false positive results. In this paper, we propose Probabilistic Token Aggregation (ProTA) to handle cross-modal interaction with content asymmetry. Specifically, we propose dual partial-related aggregation to disentangle and re-aggregate token representations in both low-dimension and high-dimension spaces. We propose token-based probabilistic alignment to generate token-level probabilistic representation and maintain the feature representation diversity. In addition, an adaptive contrastive loss is proposed to learn compact cross-modal distribution space. Based on extensive experiments, ProTA achieves significant improvements on MSR-VTT (50.9%), LSMDC (25.8%), and DiDeMo (47.2%). Han Fang 0002, Xianghao Zang, Chao Ban, Zerun Feng, Lanxiang Zhou, Zhongjiang He, Hao Sun 0038 |
ICME | 6 |
| 2024 | Towards Generalization beyond Pointwise Learning: A Unified Information-theoretic PerspectiveabstractThe recent surge in contrastive learning has intensified the interest in understanding the generalization of non-pointwise learning paradigms. While information-theoretic analysis achieves remarkable success in characterizing the generalization behavior of learning algorithms, its applicability is largely confined to pointwise learning, with extensions to the simplest pairwise settings remaining unexplored due to the challenges of non-i.i.d losses and dimensionality explosion. In this paper, we develop the first series of information-theoretic bounds extending beyond pointwise scenarios, encompassing pointwise, pairwise, triplet, quadruplet, and higher-order scenarios, all within a unified framework. Specifically, our hypothesis-based bounds elucidate the generalization behavior of iterative and noisy learning algorithms via gradient covariance analysis, and our prediction-based bounds accurately estimate the generalization gap with computationally tractable low-dimensional information metrics. Comprehensive numerical studies then demonstrate the effectiveness of our bounds in capturing the generalization dynamics across diverse learning scenarios. Yuxin Dong 0003, Tieliang Gong, Hong Chen 0004, Zhongjiang He, Mengxiang Li, Shuangyong Song, Chen Li 0011 |
ICML | 4 |
| 2024 | Improving Pointer Network based Dialogue State Tracking via Dual Hierarchical Selective AugmentationabstractDialogue state tracking is responsible for predicting the user’s dialogue state during the whole dialogue process. In practical applications, values for different slots exist in individual utterances of the dialog history. With the accumulation of the dialogue history, it becomes extremely difficult to accurately predict slots and corresponding values from the lengthy dialogue history. To solve the problem of the interference caused by lengthy dialogue history, we propose a dual hierarchical selective augmentation method, which makes use of two hierarchical level information selection strategy to generate slot values. In the encoding phase, we first extract word-level matching features between the slot and each dialogue turn, and then build turn-level context relevance. In the decoding phase, first of all, from a global perspective, the dialogue turn information is selected multiple according to the dialogue context and slot, so that the model focuses more on the turn containing slot value. Secondly, our model performs weighted context attention to capture the critical words of dialogue turn from the local view. This dual hierarchical context selection alleviates the interference caused by excessive redundant information in the dialogue history and enhances the judgment ability of the model for vital turns and words. Furthermore, to enhance the copying ability of the model, we use the turn selection-guided pointer network to copy slot values from the dialogue. Experimental results show that our model significantly outperforms multiple baselines on the released MultiWOZ benchmark. Shuangyong Song, Hongyan Xie, Haoxiang Su, Hao Huang 0009, Mengxiang Li, Zhongjiang He, Ruiyu Fang |
IJCNN | 6 |
| 2024 | Graph-based Dynamic Domain Selection for Dialogue State TrackingabstractThe Dialogue State Tracking (DST) module tracks the user’s intent by populating multiple predefined slots related to the dialogue task. In recent years, various graph neural network-based DST methods have been proposed to establish graph structures capturing the correlations between domains and slots, thereby enhancing model performance. However, these methods may involve redundant connections in the graph structure. To better construct relationships between domains and slots, we introduce a graph neural network-based dialogue state tracking method called Dynamic Domain Selection Graph DST (DDSG-DST). Specifically, (1) we employ Graphormer to establish hierarchical relationships between domains and slots; (2) we propose an additional domain prediction auxiliary task to predict the domain relevant to the dialogue context; (3) based on the predicted relevant domain from the auxiliary task, we dynamically select domain node information in the graph and perform dialogue state prediction. Experimental results demonstrate that we effectively establish hierarchical relationships between domains and slots, mitigate the negative impact of redundant connections in the graph structure, and enhance model performance. Shuangyong Song, Hao Huang 0009, Hongyan Xie, Haoxiang Su, Mengxiang Li, Zhongjiang He, Ruiyu Fang |
IJCNN | 7 |
| 2024 | Towards Robustness and Diversity: Continual Learning in Dialog Generation with Text-Mixup and Batch Nuclear-Norm MaximizationabstractIn our dynamic world where data arrives in a continuous stream, continual learning enables us to incrementally add new tasks/domains without the need to retrain from scratch. A major challenge in continual learning of language model is catastrophic forgetting, the tendency of models to forget knowledge from previously trained tasks/domains when training on new ones. This paper studies dialog generation under the continual learning setting. We propose a novel method that 1) uses Text-Mixup as data augmentation to avoid model overfitting on replay memory and 2) leverages Batch-Nuclear Norm Maximization (BNNM) to alleviate the problem of mode collapse. Experiments on a 37-domain task-oriented dialog dataset and DailyDialog (a 10-domain chitchat dataset) demonstrate that our proposed approach outperforms the state-of-the-art in continual learning. Jiayu Xiao, Mengxiang Li, Zhongjiang He, Shuangyong Song |
IJCNN | 4 |
| 2024 | GOAL: Grounded text-to-image Synthesis with Joint Layout Alignment TuningabstractRecent text-to-image (T2I) synthesis models have demonstrated intriguing abilities to produce high-quality images based on text prompts. However, current models still face Text-Image Misalignment problem (e.g., attribute errors and relation mistakes) for compositional generation. Existing models attempted to condition T2I models on grounding inputs to improve controllability while ignoring the explicit supervision from the layout conditions. To tackle this issue, we propose Grounded jOint lAyout aLignment (GOAL), an effective framework for T2I synthesis. Two novel modules, discriminative semantic alignment (DSAlign) and masked attention alignment (MAAlign), are proposed and incorporated in this framework to improve the text-image alignment. DSAlign leverages discriminative tasks at the region-wise level to ensure low-level semantic alignment. MAAlign provides high-level attention alignment by guiding the model to focus on the target object. We also build a dataset GOAL2K for model fine-tuning, which composes 2000 semantically accurate image-text pairs and their layout annotations. Comprehensive evaluations on T2I-Compbench, NSR-1K, and Drawbench demonstrate the superior generation performance of our method. Especially, there are improvements of 19%, 13%, and 12% in color, shape, and texture metrics for T2I-Compbench. Additionally, Q-Align metrics demonstrate that our method can generate images of higher quality. Han Fang 0002, Zerun Feng, Kaijing Ma, Chao Ban, Xianghao Zang, Lanxiang Zhou, Zhongjiang He, Jingyan Chen, Jiani Hu, Hao Sun 0038 |
ACM Multimedia | 8 |
| 2024 | AutoGraph: Enabling Visual Context via Graph Alignment in Open Domain Multi-Modal Dialogue GenerationabstractOpen-domain multi-modal dialogue system heavily relies on visual information to generate contextually relevant responses. The existing open-domain multi-modal dialog generation methods ignore the complementary relationship between multiple modalities, and are difficult to integrate with LLMs. To tackle these challenges, we introduce AutoGraph, an innovative method for constructing visual context graphs automatically. We aim to structure complex information and seamlessly integrate it with large language models (LLMs), aligning information from multiple modalities at both semantic and structural levels. Specifically, we fully connect the text graphs and scene graphs, and then trim unnecessary edges via LLMs to automatically construct a visual context graph. Next, we design several graph sampling grammar for the first time to convert graph structures into sequence which is suitable for LLMs. Finally, we propose a two-stage fine-tuning strategy to allow LLMs to understand graph sampling grammar and generate responses. We validate our proposed method on text-based LLMs, and visual-based LLMs, respectively. Experimental results show that our proposed method achieves state-of-the-art performance on multiple public datasets. Deji Zhao, Donghong Han, Ye Yuan 0001, Bo Ning 0002, Mengxiang Li, Zhongjiang He, Shuangyong Song |
ACM Multimedia | 6 |
| 2024 | Animal-Bench: Benchmarking Multimodal Video Models for Animal-centric Video UnderstandingabstractWith the emergence of large pre-trained multimodal video models, multiple benchmarks have been proposed to evaluate model capabilities. However, most of the benchmarks are human-centric, with evaluation data and tasks centered around human applications. Animals are an integral part of the natural world, and animal-centric video understanding is crucial for animal welfare and conservation efforts. Yet, existing benchmarks overlook evaluations focused on animals, limiting the application of the models. To address this limitation, our work established an animal-centric benchmark, namely Animal-Bench, to allow for a comprehensive evaluation of model capabilities in real-world contexts, overcoming agent-bias in previous benchmarks. Animal-Bench includes 13 tasks encompassing both common tasks shared with humans and special tasks relevant to animal conservation, spanning 7 major animal categories and 819 species, comprising a total of 41,839 data entries. To generate this benchmark, we defined a task system centered on animals and proposed an automated pipeline for animal-centric data processing. To further validate the robustness of models against real-world challenges, we utilized a video editing approach to simulate realistic scenarios like weather changes and shooting parameters due to animal movements. We evaluated 8 current multimodal video models on our benchmark and found considerable room for improvement. We hope our work provides insights for the community and opens up new avenues for research in multimodal video models. Our data and code will be released at https://github.com/PRIS-CV/Animal-Bench. Yinuo Jing, Ruxu Zhang, Kongming Liang, Zhongjiang He, Zhanyu Ma, Jun Guo 0002 |
NeurIPS | 5 |
| 2024 | Enhancing Chinese Argument Mining with Large Language Model
Shiquan Wang, Ruiyu Fang, Mengxiang Li, Zhongjiang He, Shuangyong Song |
NLPCC (5) | 4 |
| 2024 | Mixture-of-Hand-Experts: Repainting the Deformed Hand Images Generated by Diffusion Models
Tianwei Cao, Kongming Liang, Zhongjiang He, Hao Sun 0015, Zhanyu Ma |
PRCV (5) | 4 |
| 2023 | Mask to Reconstruct: Cooperative Semantics Completion for Video-text RetrievalabstractRecently, masked video modeling has been widely explored and improved the model's understanding ability of visual regions at a local level. However, existing methods usually adopt random masking and follow the same reconstruction paradigm to complete the masked regions, which do not leverage the correlations between cross-modal content. In this paper, we present MAsk for Semantics COmpleTion (MASCOT) based on semantic-based masked modeling. Specifically, after applying attention-based video masking to generate high-informed and low-informed masks, we propose Informed Semantics Completion to recover masked semantics information. The recovery mechanism is achieved by aligning the masked content with the unmasked visual regions and corresponding textual context, which makes the model capture more text-related details at a patch level. Additionally, we shift the emphasis of reconstruction from irrelevant backgrounds to discriminative parts to ignore regions with low-informed masks. Furthermore, we design co-learning to incorporate video cues under different masks and learn more aligned representation. Our MASCOT performs state-of-the-art performance on four text-video retrieval benchmarks, including MSR-VTT, LSMDC, ActivityNet, and DiDeMo. Han Fang 0002, Zhifei Yang 0004, Xianghao Zang, Chao Ban, Zhongjiang He, Hao Sun 0038, Lanxiang Zhou |
ACM Multimedia | 5 |
| 2023 | A Baseline Investigation: Transformer-based Cross-view Baseline for Text-based Person SearchabstractThis paper investigates a baseline approach for text-based person search by using a transformer-based framework. Existing methods usually treat the visual and textual features as independent entities for speeding up the model inference process. However, the attention to the same images should be changed according to different texts. In this paper, we use a commonly employed framework with a fused feature as the baseline, which overcomes the misalignment problem introduced by fixed features. A thorough investigation is conducted in this paper. Moreover, we propose Cross-View Matching (CVM) to provide challenging, positive text-image pairs that enable the model to learn cross-view meta-information. Furthermore, we suggest a novel evaluation process to reduce the inference time and GPU memory demand. The experiments are conducted on CUHK-PEDES, ICFG-PEDES, and RSTPReid benchmarks. Through extensive parameter analysis, the potentials of a transformer-based framework are fully explored. Although the proposed scheme is a simple framework, it achieves significant performance improvements compared with other state-of-the-art methods. Xianghao Zang, Wei Gao 0003, Ge Li 0002, Han Fang 0002, Chao Ban, Zhongjiang He, Hao Sun 0038 |
ACM Multimedia | 6 |