EDBT 2026 Demo / reviewers in the wild / expert
Shaozu Yuan
dblp:254/1904
· DBLP profile ↗
25ranked-venue papers
5as first author
23since 2021 · last 2026
0000-0001-5084-7064ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 2 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 13 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TeCES: Collaborative Geometric Knowledge Representation Framework under Evolving Fact SnapshotsabstractJiujiang Guo, Zhengliang Guo, Kai Wang, Meiyang Wang, Dehua Peng, Shaozu Yuan, Chengyin Hu, Shuan Ai, Yiwei Wei. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jiujiang Guo, Zhengliang Guo, Meiyang Wang, Dehua Peng, Shaozu Yuan, Chengyin Hu, Shuan Ai |
ACL (1) | 6 |
| 2026 | Diverse Sign Language TranslationabstractAbstract Like spoken languages, a single sign language expression could correspond to multiple valid textual interpretations. Hence, learning a rigid one-to-one mapping for sign language translation (SLT) models might be inadequate, particularly in the case of limited data. In this work, we introduce a Diverse Sign Language Translation (DivSLT) task, aiming to generate diverse yet accurate translations for sign language videos. Firstly, we employ large language models (LLM) to generate multiple references for the widely-used CSL-Daily and PHOENIX14T SLT datasets. Here, native speakers are only invited to touch up inaccurate references, thus significantly improving the annotation efficiency. Secondly, we provide a benchmark model to spur research in this task. Specifically, we investigate multi-reference training strategies enabling our DivSLT model to achieve diverse translations. Then, to enhance translation accuracy, we employ the max-reward-driven reinforcement learning objective that maximizes the reward of the translated result. Additionally, we utilize multiple metrics to assess the accuracy, diversity, and semantic precision of the DivSLT task. Experimental results on the enriched datasets demonstrate that our DivSLT method achieves not only better translation performance but also diverse translation results. Shaozu Yuan, Heming Du, Xin Yu 0002 |
Int. J. Comput. Vis. | 3 |
| 2025 | AutoMV: An Autonomous Agent Framework for Real Estate Marketing Video GenerationabstractIn this paper, we introduce AutoMV, an autonomous agent framework designed for generating real estate marketing videos. The framework integrates a diverse set of existing models into a tool library, allowing the agent to intelligently select and execute the appropriate tools. Given property images and text, the agent decomposes the task into manageable subtasks, generating storyline directives and corresponding camera movement trajectories to guide the video production process. By automatically applying video synthesis techniques and incorporating multimedia elements such as subtitles and background music, the agent transforms static real estate images into dynamic, visually appealing videos, thereby optimizing their impact for digital marketing purposes. Kuizong Wu, Shaozu Yuan, Chang Shen, Meng Chen 0006 |
AAAI | 2 |
| 2025 | Towards Multimodal Sentiment Analysis via Hierarchical Correlation Modeling with Semantic Distribution ConstraintsabstractSentiment analysis is rapidly advancing by utilizing various data modalities (e.g., text, video, and audio). However, most existing techniques only learn the atomic-level features that reflect strong correlations, while ignoring more complex compositions in multimodal data. Moreover, they also neglected the incongruity in semantic distribution among modalities. In light of this, we introduce a novel Hierarchical Correlation Modeling Network (HCMNet), which enhances the multimodal sentiment analysis by exploring both the atomic-level correlations based on dynamic attention reasoning and the composition-level correlations through topological graph reasoning. In addition, we also alleviate the impact of distributional inconsistencies between modalities from both atomic-level and composition-level perspectives. Specifically, we first design an atomic-level contrastive loss that constrains the semantic distribution across modalities to mitigate the atomic-level inconsistency. Then, we design a graph optimal transport module that integrates transport flows with different graphs to constrain the composition-level semantic distribution, thus reducing the inconsistency of compositional nodes. Experiments on three public benchmark datasets have demonstrated the superiority of the proposed model over the state-of-the-art methods. Qinfu Xu, Chunlei Wu, Leiquan Wang, Shaozu Yuan, Jie Wu 0033, Jing Lu 0013, Hengyang Zhou |
AAAI | 5 |
| 2025 | Multiple Feature Refining Network for Visual Emotion Distribution LearningabstractThe significance of visual emotion distribution learning (VEDL) has surged, particularly with the growing inclination to convey emotions through images. The key of VEDL lies in capturing both low- and high-level features within the same visual content, thus promoting the model for salient and subtle emotion awareness. To learn the distribution of emotions involved in images, most previous works learn coarse semantic knowledge with unbiased filtering. Consequently, they focus on the entire scene and suffer from the redundancy of semantic-irrelevant information, which diminishes the affective coherence, impeding the comprehension of emotional attributes within the treated features. In light of this, we reanalyze from the perspective of information filtering and propose a novel method called Multiple Feature Refining Network (MFRN). To minimize low-level feature redundancy, we design a wavelet-based separated frequency modeling, named Spectral Mixer, to learn invariant representations and enhance emotion saliency in low-level image features. At the higher semantic level, we design a Semantic Graph Prompt Learning for emotional semantic filtering, ensuring the purity of emotional information and providing the model with richer content semantics. Experiments conducted on three commonly used datasets have demonstrated the superiority of our MFRN model over cutting-edge methods. Qinfu Xu, Shaozu Yuan, Jie Wu 0033, Leiquan Wang, Chunlei Wu |
AAAI | 2 |
| 2025 | From Subtle Hints to Grand Expressions - Mastering Fine-grained Emotions with Dynamic Multimodal AnalysisabstractMultimodal Emotion Analysis (MEA) plays a crucial role in extracting and understanding emotional insights from diverse data sources, including text, video, and audio. However, existing methods may overlook the key issue that multimodal components exhibit asynchronism temporally and they obtain insufficient representation of fine-grained emotional expressions. In light of this, we propose a unified emotion reasoning model, EmoChat, which enhances multimodal emotion analysis by dynamically generating emotion-related tokens and fine-grained expression information through facial action modeling. To incorporate expression semantics, we design the AU Agent, a lightweight facial expression extractor, to provide LLMs with fine-grained facial knowledge for reasoning. In addition, we propose the Correlation Aggregator to alleviate the correlation differences between acoustic features and textual content. Therefore, our method decouples both the audio and vision modalities, allowing for efficient token-level emotion cues mining in misaligned multimodal input, while maintaining semantic consistency across different languages. Experiments on public benchmark datasets have demonstrated the superiority of our proposed EmoChat over the state-of-the-art methods. Qinfu Xu, Liyuan Pan, Shaozu Yuan, Chunlei Wu |
ACM Multimedia | 3 |
| 2025 | DeepMSD: Advancing Multimodal Sarcasm Detection Through Knowledge-Augmented Graph ReasoningabstractMultimodal sarcasm detection (MSD) requires predicting the sarcastic sentiment by understanding diverse modalities of data (e.g., text, image). Beyond the surface-level information conveyed in the post data, understanding the underlying deep-level knowledge-such as the background and intent behind the data-is crucial for understanding the sarcastic sentiment. However, previous works have often overlooked this aspect, limiting their potential to achieve superior performance. To tackle this challenge, we propose DeepMSD, a novel framework that generates supplemental deep-level knowledge to enhance the understanding of sarcastic content. Specifically, we first devise a Deep-level Knowledge Extraction Module that leverages large vision-language models to generate deep-level information behind the text-image pairs. Additionally, we devise a Cross-knowledge Graph Reasoning Module to model how humans use prior knowledge to identify sarcastic cues in multimodal posts. This module constructs cross-knowledge graphs that connect deep-level knowledge with surface-level knowledge. As such, it enables a more profound exploration of the cues underlying sarcasm. Experiments on the public MSD dataset demonstrate that our approach significantly surpasses previous state-of-the-art methods. Hengyang Zhou, Shaozu Yuan, Meng Chen 0006, Zhiyang Jia, Longbiao Wang, Xiaodong He 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Enhancing Semantic Awareness by Sentimental Constraint With Automatic Outlier Masking for Multimodal Sarcasm DetectionabstractMultimodal sarcasm detection, aiming to uncover sarcastic sentiment behind multimodal data, has gained substantial attention in multimodal communities. Recent advancements in multimodal sarcasm detection (MSD) methods have primarily focused on modality alignment with pre-trained vision-language (V-L) model. However, text-image pairs often exhibit weak or even opposite semantic correlations in MSD tasks. Consequently, directly aligning these modalities can potentially result in feature shift and inter-class confusion, ultimately hindering the model's ability. To alleviate this issue, we propose the Enhancing Semantic Awareness Model (ESAM) for multimodal sarcasm detection. Specifically, we first devise a Modality-decoupled Framework (MDF) to separate the textual and visual features from the fused multimodal representation. This decoupling enables the parallel integration of the Sentimental Congruity Constraint (SCC) within both visual and textual latent spaces, thereby enhancing the semantic awareness of different modalities. Furthermore, given that certain outlier samples with ambiguous sentiments can mislead the training and weaken the performance of SCC, we further incorporate Automatic Outlier Masking. This mechanism automatically detects and masks the outliers, guiding the model to focus on more informative samples during training. Experimental results on two public MSD datasets validate the robustness and superiority of our proposed ESAM model. Shaozu Yuan, Hengyang Zhou, Qinfu Xu, Meng Chen 0006, Xiaodong He 0001 |
IEEE Trans. Multim. | 1 |
| 2024 | G^2SAM: Graph-Based Global Semantic Awareness Method for Multimodal Sarcasm DetectionabstractMultimodal sarcasm detection, aiming to detect the ironic sentiment within multimodal social data, has gained substantial popularity in both the natural language processing and computer vision communities. Recently, graph-based studies by drawing sentimental relations to detect multimodal sarcasm have made notable advancements. However, they have neglected exploiting graph-based global semantic congruity from existing instances to facilitate the prediction, which ultimately hinders the model's performance. In this paper, we introduce a new inference paradigm that leverages global graph-based semantic awareness to handle this task. Firstly, we construct fine-grained multimodal graphs for each instance and integrate them into semantic space to draw graph-based relations. During inference, we leverage global semantic congruity to retrieve k-nearest neighbor instances in semantic space as references for voting on the final prediction. To enhance the semantic correlation of representation in semantic space, we also introduce label-aware graph contrastive learning to further improve the performance. Experimental results demonstrate that our model achieves state-of-the-art (SOTA) performance in multimodal sarcasm detection. The code will be available at https://github.com/upccpu/G2SAM. Shaozu Yuan, Hengyang Zhou, Longbiao Wang, Zhiling Yan, Ruosong Yang, Meng Chen 0006 |
AAAI | 2 |
| 2024 | Multi-receptive Field Distillation Network for seismic velocity model building
Jing Lu 0013, Chunlei Wu, Guolong Li, Shaozu Yuan |
Eng. Appl. Artif. Intell. | 5 |
| 2024 | Towards visual emotion analysis via Multi-Perspective Prompt Learning with Residual-Enhanced Adapter
Chunlei Wu, Qinfu Xu, Shaozu Yuan, Jie Wu 0033, Leiquan Wang |
Knowl. Based Syst. | 4 |
| 2024 | MuJo-SF: Multimodal Joint Slot Filling for Attribute Value Prediction of E-Commerce CommoditiesabstractSupplementing product attribute information is a critical step for E-commerce platforms, which further benefits various downstream tasks, including product recommendation, product search, and product knowledge graph construction. Intuitively, the visual information available on e-commerce platforms can effectively function as a primary source for certain product attributes. However, existing works either extract attribute values solely from textual product descriptions or leverage limited visual information (e.g., image features or optical character recognition tokens) to assist extraction, without mining the fine-grained visual cues linked with the products effectively. In this paper, we propose a novel task -Multimodal Joint Slot Filling(MuJo-SF) - that aims to combine multimodal information from both product descriptions and their corresponding product images to jointly fill values into the pre-defined product attribute set. To this end, we develop MAVP, a new dataset with 79 k instances of product description-image pairs. Specifically, we present a strategy to fulfill visualized saliency ascription, which aims to distinguish between text-dependent and image-dependent attributes. For those image-dependent attributes, we annotate the corresponding values from images using distant supervision. Then, we design a model for MuJo-SF, which combines multimodal representations and fills image-dependent and text-dependent attributes separately. Finally, we conduct extensive experiments on MAVP and provide rich results for MuJo-SF, which can be used as baselines to facilitate future research. Meihuizi Jia, Lei Shen 0001, Anh Tuan Luu, Meng Chen 0006, Lejian Liao, Shaozu Yuan, Xiaodong He 0001 |
IEEE Trans. Multim. | 7 |
| 2023 | Tackling Modality Heterogeneity with Multi-View Calibration Network for Multimodal Sentiment DetectionabstractYiwei Wei, Shaozu Yuan, Ruosong Yang, Lei Shen, Zhangmeizhi Li, Longbiao Wang, Meng Chen. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Shaozu Yuan, Ruosong Yang, Lei Shen 0001, Zhangmeizhi Li, Longbiao Wang, Meng Chen 0006 |
ACL (1) | 2 |
| 2023 | Nested Attention Network with Graph Filtering for Visual Question and AnsweringabstractRecently, Visual Question Answering(VQA), which is required to generate the answer by understanding both visual and textual content, has attracted considerable research interest. Most existing works extract visual features with the CNN network and learn its feature embedding with an attention mechanism. However, this mechanism may ignore the interaction between entities in the image, which has a fuzzy impact on the answer generation. To better explore the relationship between different entities in the image, a novel Nested Attention Network with Graph Filtering (NANGF) is proposed. It composes of two novel designed modules: a graph filtering mechanism to mine more precise visual semantics and avoid understanding deviation and nested attention to effectively guide the integration of visual features and question features. Extensive experiments conducted on the VQA2.0 datasets demonstrate the effectiveness of the proposed method. Jing Lu 0013, Chunlei Wu, Leiquan Wang, Shaozu Yuan, Jie Wu 0033 |
ICASSP | 4 |
| 2023 | Enhancing Multimodal Alignment with Momentum Augmentation for Dense Video CaptioningabstractDense video captioning aims to localize multiple events from an untrimmed video and generate corresponding captions for each event. Fusing different modalities(e.g. rgb, flow, audio) via transformer structure is a promising way to improve the caption performance. However, it is challenging for the cross-modal encoder to learn multimodal interactions due to their inherent disparities of distribution. In this paper, we propose a novel transformer structure with contrastive learning to align different modalities. Specifically, to avoid the limitation of small batch size and false contrastive targets, we design an event-aligned momentum augmentation strategy to apply contrast learning for dense video captioning. The experimental result shows that our proposals outperform all existing multimodal fusion methods for dense video captioning. Shaozu Yuan, Meng Chen 0006, Longbiao Wang |
ICASSP | 2 |
| 2023 | Auslan-Daily: Australian Sign Language Translation for Daily Communication and NewsabstractSign language translation (SLT) aims to convert a continuous sign language video clip into a spoken language. Considering different geographic regions generally have their own native sign languages, it is valuable to establish corresponding SLT datasets to support related communication and research. Auslan, as a sign language specific to Australia, still lacks a dedicated large-scale dataset for SLT.To fill this gap, we curate an Australian Sign Language translation dataset, dubbed Auslan-Daily, which is collected from the Auslan educational TV series and Auslan TV programs. The former involves daily communications among multiple signers in the wild, while the latter comprises sign language videos for up-to-date news, weather forecasts, and documentaries. In particular, Auslan-Daily has two main features: (1) the topics are diverse and signed by multiple signers, and (2) the scenes in our dataset are more complex, e.g., captured in various environments, gesture interference during multi-signers' interactions and various camera positions. With a collection of more than 45 hours of high-quality Auslan video materials, we invite Auslan experts to align different fine-grained visual and language pairs, including video $\leftrightarrow$ fingerspelling, video $\leftrightarrow$ gloss, and video $\leftrightarrow$ sentence. As a result, Auslan-Daily contains multi-grained annotations that can be utilized to accomplish various fundamental sign language tasks, such as signer detection, sign spotting, fingerspelling detection, isolated sign language recognition, sign language translation and alignment. Moreover, we benchmark results with state-of-the-art models for each task in Auslan-Daily. Experiments indicate that Auslan-Daily is a highly challenging SLT dataset, and we hope this dataset will contribute to the development of Auslan and the advancement of sign languages worldwide in a broader context. All datasets and benchmarks are available at Auslan-Daily. Shaozu Yuan, Hongwei Sheng, Heming Du, Xin Yu 0002 |
NeurIPS | 2 |
| 2023 | MPP-net: Multi-perspective perception network for dense video captioning
Shaozu Yuan, Meng Chen 0006, Longbiao Wang, Lei Shen 0001, Zhiling Yan |
Neurocomputing | 2 |
| 2022 | Few-Shot Table Understanding: A Benchmark Dataset and Pre-Training BaselineabstractFew-shot table understanding is a critical and challenging problem in real-world scenario as annotations over large amount of tables are usually costly. Pre-trained language models (PLMs), which have recently flourished on tabular data, have demonstrated their effectiveness for table understanding tasks. However, few-shot table understanding is rarely explored due to the deficiency of public table pre-training corpus and well-defined downstream benchmark tasks, especially in Chinese. In this paper, we establish a benchmark dataset, FewTUD, which consists of 5 different tasks with human annotations to systematically explore the few-shot table understanding in depth. Since there is no large number of public Chinese tables, we also collect a large-scale, multi-domain tabular corpus to facilitate future Chinese table pre-training, which includes one million tables and related natural language text with auxiliary supervised interaction signals. Finally, we present FewTPT, a novel table PLM with rich interactions over tabular data, and evaluate its performance comprehensively on the benchmark. Our dataset and model will be released to the public soon. Ruixue Liu, Shaozu Yuan, Aijun Dai, Lei Shen 0001, Tiangang Zhu, Meng Chen 0006, Xiaodong He 0001 |
COLING | 2 |
| 2022 | SE-GAN: Skeleton Enhanced Gan-Based Model for Brush Handwriting Font GenerationabstractPrevious works on font generation mainly focus on the standard print fonts where character's shape is stable and strokes are clearly separated. There is rare research on brush hand-writing font generation, which involves holistic structure changes and complex strokes transfer. To address this issue, we propose a novel GAN-based image translation model by integrating the skeleton information. We first extract the skeleton from training images, then design an image encoder and a skeleton encoder to extract corresponding features. A self-attentive refined attention module is devised to guide the model to learn distinctive features between different domains. A skeleton discriminator is involved to first synthesize the skeleton image from the generated image with a pre-trained generator, then to judge its realness to the target one. We also contribute a large-scale brush handwriting font image dataset with six styles and 15,000 high-resolution images. Both quantitative and qualitative experimental results demonstrate the competitiveness of our proposed model. Shaozu Yuan, Ruixue Liu, Meng Chen 0006, Baoyang Chen, Zhijie Qiu, Xiaodong He 0001 |
ICME | 1 |
| 2022 | Learning to Generate Poetic Chinese Landscape Painting with CalligraphyabstractIn this paper, we present a novel system (denoted as Polaca) to generate poetic Chinese landscape painting with calligraphy. Unlike previous single image-to-image painting generation, Polaca takes the classic poetry as input and outputs the artistic landscape painting image with the corresponding calligraphy. It is equipped with three different modules to complete the whole piece of landscape painting artwork: the first one is a text-to-image module to generate landscape painting image, the second one is an image-to-image module to generate stylistic calligraphy image, and the third one is an image fusion module to fuse the two images into a whole piece of aesthetic artwork. Shaozu Yuan, Aijun Dai, Zhiling Yan, Ruixue Liu, Meng Chen 0006, Baoyang Chen, Zhijie Qiu, Xiaodong He 0001 |
IJCAI | 1 |
| 2022 | MCIC: Multimodal Conversational Intent Classification for E-commerce Customer Service
Shaozu Yuan, Hang Liu 0005, Zhiling Yan, Ruixue Liu, Meng Chen 0006 |
NLPCC (1) | 1 |
| 2021 | Learning to Compose Stylistic Calligraphy Artwork with EmotionsabstractEmotion plays a critical role in calligraphy composition, which makes the calligraphy artwork impressive and have a soul. However, previous research on calligraphy generation all neglected the emotion as a major contributor to the artistry of calligraphy. Such defects prevent them from generating aesthetic, stylistic, and diverse calligraphy artworks, but only static handwriting font library instead. To address this problem, we propose a novel cross-modal approach to generate stylistic and diverse Chinese calligraphy artwork driven by different emotions automatically. We firstly detect the emotions in the text by a classifier, then generate the emotional Chinese character images via a novel modified Generative Adversarial Network (GAN) structure, finally we predict the layout for all character images with a recurrent neural network. We also collect a large-scale stylistic Chinese calligraphy image dataset with rich emotions. Experimental results demonstrate that our model outperforms all baseline image translation models significantly for different emotional styles in terms of content accuracy and style discrepancy. Besides, our layout algorithm can also learn the patterns and habits of calligrapher, and makes the generated calligraphy more artistic. To the best of our knowledge, we are the first to work on emotion-driven discourse-level Chinese calligraphy artwork composition. Shaozu Yuan, Ruixue Liu, Meng Chen 0006, Baoyang Chen, Zhijie Qiu, Xiaodong He 0001 |
ACM Multimedia | 1 |
| 2021 | Generate classical Chinese poems with theme-style from images
Chunlei Wu, Jiangnan Wang, Shaozu Yuan, Leiquan Wang, Weishan Zhang |
Pattern Recognit. Lett. | 3 |
| 2020 | The JDDC Corpus: A Large-Scale Multi-Turn Chinese Dialogue Dataset for E-commerce Customer ServiceabstractHuman conversations are complicated and building a human-like dialogue agent is an extremely challenging task. With the rapid development of deep learning techniques, data-driven models become more and more prevalent which need a huge amount of real conversation data. In this paper, we construct a large-scale real scenario Chinese E-commerce conversation corpus, JDDC, with more than 1 million multi-turn dialogues, 20 million utterances, and 150 million words. The dataset reflects several characteristics of human-human conversations, e.g., goal-driven, and long-term dependency among the context. It also covers various dialogue types including task-oriented, chitchat and question-answering. Extra intent information and three well-annotated challenge sets are also provided. Then, we evaluate several retrieval-based and generative models to provide basic benchmark performance on the JDDC corpus. And we hope JDDC can serve as an effective testbed and benefit the development of fundamental research in dialogue task. Meng Chen 0006, Ruixue Liu, Lei Shen 0001, Shaozu Yuan, Jingyan Zhou, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001 |
LREC | 4 |
| 2020 | MaLiang: An Emotion-driven Chinese Calligraphy Artwork Composition SystemabstractWe present a novel Chinese calligraphy artwork composition system (MaLiang) which can generate aesthetic, stylistic and diverse calligraphy images based on the emotion status from the input text. Different from previous research, it's the first work to endow the calligraphy synthesis with the ability to express fickle emotions and composite a whole piece of discourse-level calligraphy artwork instead of single character images. The system consists of three modules: emotion detection, character image generation, and layout prediction. As a creative form of interactive art, MaLiang has been exhibited in several famous international art festivals. Ruixue Liu, Shaozu Yuan, Meng Chen 0006, Baoyang Chen, Zhijie Qiu, Xiaodong He 0001 |
ACM Multimedia | 2 |