EDBT 2026 Demo / reviewers in the wild / expert
Peipei Song
dblp:154/3565
· DBLP profile ↗
27ranked-venue papers
10as first author
22since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 7 first-author · 15 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Computer networks · 2 · 1 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Social Media to Psychological Scale: An Adaptive Framework with Two-Hop Retrieval for Depression ScreeningabstractDepressive disorders represent a major global public health challenge. As an increasing number of individuals share their emotional experiences and concerns on social media, researchers have shown growing interest in leveraging such data for early depression screening. However, most existing methods rely on a fixed model and a singular reasoning paradigm, which constrains their adaptability to depression detection. The limited availability of mental health-related data and variability in training data distributions across different LLMs hinder their consistent and comprehensive understanding of diverse psychological symptoms. In this paper, we propose AdaDepression, a framework that enables explainable depression screening through a two-hop retrieval algorithm to identify symptom-relevant posts and a two-stage adaptive routing mechanism for selecting appropriate reasoning strategies and LLMs. Specifically, we first collect representative posts from the training dataset to capture the real-world symptom expressions, and then utilize these posts to retrieve symptom-relevant posts from the user's posting history. Subsequently, we employ the Mixture of Routers (MoR), which integrates the Mixture of Experts (MoE) into the routing mechanism to select the optimal reasoning strategies and LLMs in a cascaded manner. Finally, we complete the standardized psychological questionnaire using the selected LLMs and reasoning strategies. Experimental results on the Reddit-based benchmarks demonstrate the effectiveness of the proposed method, outperforming existing studies on various metrics. Our code is released at https://github.com/MindIntLab-HFUT/AdaDepression. Yangyang Xu 0002, Jinpeng Hu, Peipei Song, Zhangling Duan, Xun Yang 0001 |
WWW | 3 |
| 2026 | Updatable one-stream Vision-Language tracking via multilayer perceptual memory network
Peipei Song, Huanlong Zhang, Bin Jiang 0007, Bineng Zhong 0001 |
Pattern Recognit. | 3 |
| 2026 | Aleatoric-Epistemic Joint Uncertainty Modeling for Cross-Modal RetrievalabstractRecently, the cross-modal retrieval task has gained significant attention with the advent of large-scale vision-language pretraining models, e.g., CLIP. These methods typically map the vision and language modalities into a shared embedding space and then build similarity relations based on the joint feature representations. Despite tremendous progress in this field, most existing methods still suffer from unreliable retrieval results caused by data and model uncertainties, which can arise from inherent data ambiguity or noisy pairs. In this article, we propose a novel cross-modal retrieval framework with aleatoric-epistemic joint uncertainty modeling (AEUM). AEUM is committed to providing reliable uncertainty estimation for both data (aleatoric uncertainty, AU) and model (epistemic uncertainty, EU), which are then used to correct the initial cross-modal similarity to yield more accurate retrieval results. Specifically, for AU, we introduce learnable semantic tokens for each modality to estimate the data-induced uncertainty in another modality, offering guidance on data complexity or ambiguity. For the EU, we leverage the efficient evidential learning paradigm to estimate model-induced uncertainty and incorporate it into the model's predictions, thereby enhancing robustness against noisy data. Extensive experiments demonstrate the effectiveness and generalization of our method on multiple cross-modal retrieval benchmarks, including five video-text retrieval datasets (MSRVTT, LSMDC, MSVD, VATEX, and DiDeMo) and two image-text retrieval datasets (MSCOCO and Flickr30K). Our code is publicly available at https://github.com/cty8998/AEUM. Peipei Song, Xun Yang 0001, Dan Guo 0001, Xiaojun Chang |
IEEE Trans. Cybern. | 2 |
| 2026 | Subjective-Objective Emotion-Correlated Generation Network for Subjective Video CaptioningabstractThe emotional video captioning (EVC) task, which aims to generate factual descriptions based on the perceived subtle visual emotion cues, has received more and more attention and research. However, EVC is essentially an objective video captioning task, and ignores the subjective emotional reactions of video viewers, which cannot reflect personalized affective understandings of different viewers on the same video. To fill the research gap, we investigate the subjective video captioning (SVC) task in this paper, which aims to generate emotional captions by incorporating viewers' personalized emotional reactions upon the EVC task. SVC is extremely challenging, which lies in two aspects: 1) the correlative emotion perception between subjective and objective emotions and 2) the collaborative generation between emotional and factual information. To this end, we propose the Subjective-Objective Emotion-Correlated Generation Network (SO-ECGN) in this paper. Specifically, our SO-ECGN leverages the proposed dynamic mask attention and emotion domain shifting module to achieve the objective emotion incremental learning, and then, a subjective-objective emotions correlation module is proposed to adaptively combine two perspective emotions to provide accurate emotion guidance (i.e., emotional polarity and intensity) for each generation step. Furthermore, an emotion-correlated decoder is proposed to generate subjective captions by adaptively referring to factual information and emotional information. Extensive experiments on three challenging datasets demonstrate the superiority of our approach and each proposed module, i.e., reaching 79.2%, 45.1% on BLEU-1, CIDEr metrics on EmVidCap-L dataset. Weidong Chen 0013, Cheng Ye 0004, Peipei Song, Lei Zhang 0119, Yongdong Zhang 0001, Zhendong Mao 0001 |
IEEE Trans. Image Process. | 3 |
| 2026 | VMFormer: Visual Clues-Guided Multi-Stage Transformer for Image CaptioningabstractTransformer-based models have significantly advanced image captioning through self-attention mechanisms and parallel computation. However, existing methods typically adopt teacher-forcing strategies during training by conditioning the decoder exclusively on ground-truth tokens, whereas at inference, captions are generated autoregressively based solely on previously predicted tokens. Such discrepancy between training and inference conditions leads to a progressive accumulation of prediction errors, resulting in captions that deviate significantly from the visual content. While scheduled sampling strategies mitigate this issue, directly integrating them disrupts Transformer parallelism and overlooks visual saliency differences. To tackle these limitations, we present VMFormer, a multi-stage decoding Transformer featuring a Visual-Aware Scheduled Sampling (VASS) module that bridges this gap through two key innovations: 1) A two-stage decoding scheme where an initial self-study stage generates candidate tokens, followed by a hybrid stage dynamically blending ground-truth references with predictions via a visual clues controlled gate. 2) A cognitive-inspired prioritization mechanism that retains visual keywords (nouns/verbs/attributes) in early training phases before transitioning to linguistic refinements, mirroring human captioning patterns. Crucially, the VASS module preserves the parallel computational strengths of Transformer architectures and is designed as a plug-and-play component, readily adaptable to Transformer-based captioning models. Experiments on challenging MSCOCO achieve 142.2 CIDEr. We extend VMFormer to video captioning and demonstrate consistent improvements on MSRVTT and MSVD datasets. Yuchen Ren 0001, Xin Chen 0032, Hongrui Yuan, Peipei Song, Wanli Ouyang, Lan Zhang 0002, Jinyang Guo 0002 |
IEEE Trans. Image Process. | 5 |
| 2026 | Bridging Subjectivity in Affective Explanation Captioning via Consensus-Prompted Emotion ReasoningabstractAffective Explanation Captioning (AEC) aims to perform viewer-centered visual emotion analysis by not only identifying the emotions evoked by an image but also explaining their underlying causes. Prior efforts have achieved promising results by fine-tuning LLMs on affective data; however, two key challenges remain: 1) the inherent subjectivity of human emotion leads to diverse interpretations of the same image, making it difficult for models to catch dominant emotions; and 2) the affective gap between abstract emotions and concrete visual content hinders models from capturing both semantic and emotional aspects effectively. To tackle these challenges, we propose Consensus-Prompted Emotion Reasoning (CPER), a new framework that explicitly models emotional diversity and enforces emotional-semantic alignment. Inspired by psychological studies, we observe that common emotional patterns often emerge within certain groups, which we refer to as affective consensus. Capturing this consensus across varying levels is helpful for bridging the subjectivity in AEC. Specifically, we introduce a consensus-based bucket prompt, which depicts the consensus level of each emotional perspective, serving as a control signal to adjust the emotion reasoning. To reconcile abstract emotion understanding and concrete visual grounding, we design a dual-space representation, where a CLIP encoder extracts objective semantic evidence and an emotion encoder captures abstract affective cues for AEC. Furthermore, an emotion consistency learning strategy is devised, which explicitly aligns the generated explanation with the input image and the emotion label, ensuring both emotionally and semantically grounded explanations. Extensive experiments on three benchmark datasets, ranging from visual arts (ArtEmis v1.0 and ArtEmis v2.0) and real-world images (Affection), demonstrate the effectiveness of our CPER in terms of emotional diversity and semantic coherence compared to state-of-the-art methods. Our code is publicly available at https://github.com/songpipi/CPER. Peipei Song, Zhiyan Zhang, Weidong Chen 0013, Jinpeng Hu, Xun Yang 0001, Xiaojun Chang |
IEEE Trans. Image Process. | 1 |
| 2026 | Scene-Text Grounding for Text-Based Video Question AnsweringabstractExisting efforts in text-based video question answering (TextVideoQA) are criticized for their opaque decision-making and heavy reliance on scene-text recognition. In this paper, we propose to studyGrounded TextVideoQAby forcing models to answer questions and spatio-temporally localize the relevant scene-text regions, thus decoupling QA from scene-text recognition and promoting research towards interpretable QA. The task has three-fold significance. First, it encourages scene-text evidence versus other short-cuts for answer predictions. Second, it directly accepts scene-text regions as visual answers, thus circumventing the problem of ineffective answer evaluation by stringent string matching. Third, it isolates the challenges inherited in VideoQA and scene-text recognition. This enables the diagnosis of the root causes for failure predictions,e.g., wrong QA or wrong scene-text recognition? To achieve Grounded TextVideoQA, we propose the T2S-QA model that highlights a disentangled temporal-to-spatial contrastive learning strategy for weakly-supervised scene-text grounding and grounded TextVideoQA. To facilitate evaluation, we construct a new datasetViTXT-GQAwhich features 52Kscene-text bounding boxes within 2.2Ktemporal segments related to 2Kquestions and 729 videos. With ViTXT-GQA, we perform extensive experiments and demonstrate the severe limitations of existing techniques in Grounded TextVideoQA. While T2S-QA achieves superior results, the large performance gap with human leaves ample space for improvement. Our further analysis of oracle scene-text inputs posits that the major challenge is scene-text recognition. To advance the research of Grounded TextVideoQA, our dataset and code are athttps://github.com/zhousheng97/ViTXT-GQA.git Junbin Xiao, Xun Yang 0001, Peipei Song, Dan Guo 0001, Angela Yao, Meng Wang 0001, Tat-Seng Chua |
IEEE Trans. Multim. | 4 |
| 2025 | Linguistics-Vision Monotonic Consistent Network for Sign Language ProductionabstractSign Language Production (SLP) aims to generate sign videos corresponding to spoken language sentences, where the conversion of sign Glosses to Poses (G2P) is the key step. Due to the cross-modal semantic gap and the lack of word-action correspondence labels for strong supervision alignment, the SLP suffers huge challenges in linguistics-vision consistency. In this work, we propose a Transformer-based Linguistics-Vision Monotonic Consistent Network (LVMCN) for SLP, which constrains fine-grained cross-modal monotonic alignment and coarse-grained multimodal semantic consistency in language-visual cues through Cross-modal Semantic Aligner (CSA) and Multimodal Semantic Comparator (MSC). In the CSA, we constrain the implicit alignment between corresponding gloss and pose sequences by computing the cosine similarity association matrix between cross-modal feature sequences (i.e., the order consistency of fine-grained sign glosses and actions). As for MSC, we construct multimodal triplets based on paired and unpaired samples in batch data. By pulling closer the corresponding text-visual pairs and pushing apart the non-corresponding text-visual pairs, we constrain the semantic co-occurrence degree between corresponding gloss and pose sequences (i.e., the semantic consistency of coarse-grained textual sentences and sign videos). Extensive experiments on the popular PHOENIX14T benchmark show that the LVMCN outperforms the state-of-the-art. Shengeng Tang, Peipei Song, Shuo Wang 0008, Dan Guo 0001, Richang Hong |
ICASSP | 3 |
| 2025 | Benchmarking and Bridging Emotion Conflicts for Multimodal Emotion ReasoningabstractDespite their strong performance in multimodal emotion reasoning, existing Multimodal Large Language Models (MLLMs) often overlook the scenarios involving emotion conflicts, where emotional cues from different modalities are inconsistent. To fill this gap, we first introduce CA-MER, a new benchmark designed to examine MLLMs under realistic emotion conflicts. It consists of three subsets: video-aligned, audio-aligned, and consistent, where only one or all modalities reflect the true emotion. However, evaluations on our CA-MER reveal that current state-of-the-art emotion MLLMs systematically over-rely on audio signal during emotion conflicts, neglecting critical cues from visual modality. To mitigate this bias, we propose MoSEAR, a parameter-efficient framework that promotes balanced modality integration. MoSEAR consists of two modules: (1)MoSE, modality-specific experts with a regularized gating mechanism that reduces modality bias in the fine-tuning heads; and (2)AR, an attention reallocation mechanism that rebalances modality contributions in frozen backbones during inference. Our framework offers two key advantages: it mitigates emotion conflicts and improves performance on consistent samples-without incurring a trade-off between audio and visual modalities. Experiments on multiple benchmarks-including MER2023, EMER, DFEW, and our CA-MER-demonstrate that MoSEAR achieves state-of-the-art performance, particularly under modality conflict conditions. Zhiyuan Han, Beier Zhu, Yanlong Xu, Peipei Song, Xun Yang 0001 |
ACM Multimedia | 4 |
| 2025 | Beyond Emotion Recognition: A Multi-Turn Multimodal Emotion Understanding and Reasoning BenchmarkabstractMultimodal large language models (MLLMs) have been widely applied across various fields due to their powerful perceptual and reasoning capabilities. In the realm of psychology, these models hold promise for a deeper understanding of human emotions and behaviors. However, recent research primarily focuses on enhancing their emotion recognition abilities, leaving the substantial potential in emotion reasoning, which is crucial for improving the naturalness and effectiveness of human-machine interactions. Therefore, in this paper, we introduce a multi-turn multimodal emotion understanding and reasoning (MTMEUR) benchmark, which encompasses 1,451 video data from real-life scenarios, along with 5,101 progressive questions. These questions cover various aspects, including emotion recognition, potential causes of emotions, future action prediction, etc. Besides, we propose a multi-agent framework, where each agent specializes in a specific aspect, such as background context, character dynamics, and event details, to improve the system's reasoning capabilities. Furthermore, we conduct experiments with existing MLLMs and our agent-based method on the proposed benchmark, revealing that most models face significant challenges with this task. Jinpeng Hu, Hongchang Shi, Chongyuan Dai, Peipei Song, Meng Wang 0001 |
ACM Multimedia | 5 |
| 2025 | Multi-round Mutual Emotion-Cause Pair Extraction for Emotion-Attributed Video CaptioningabstractEmotional Video Captioning (EVC) is an emerging task that aims to describe factual content with the intrinsic emotions expressed in videos. Existing EVC methods perceive global emotional cues through visual features at first, and then combine them with the video features to guide the emotional caption generation, which ignores the critical characteristic of the EVC task that emotional cues have intrinsic motivational causes reflected in the video content. Such video causes have a facilitative effect on both emotion perception and emotion-attributed caption generation. To this end, a multi-round mutual emotion-cause pair extraction network (MM-ECPE) is proposed in this paper for the joint extraction of emotional cues and visual causes through iterative mutual refinement. Specifically, in the 1st-round mutual learning, we propose a spatio-temporal disentangled visual adaptive refinement (ST-DVAR) and a multi-level video-guided emotion affine transformation (MV-EAT) to achieve preliminary refinement on video features and emotion lexicon to eliminate the noise caused by emotion-irrelevant visual information and video-irrelevant emotional information. Then, in the 2nd-round mutual learning, we exploit the cross-attention of the preliminary refined features and the original features to obtain the ultimate emotional cues and visual causes, and couple them in pair-wise extraction through contrastive loss. Overall, our approach optimizes complex semantic understanding and emotion perception of videos, leading to a promising performance in emotional captioning. Extensive experiments on three challenging datasets demonstrate the superiority of our approach and each proposed module, e.g., improving the latest records by +97.5% and +76.2% w.r.t. CIDEr and CFS, respectively, on the EVC-MSVD dataset. Cheng Ye 0004, Weidong Chen 0013, Peipei Song, Xinyan Liu 0008, Lei Zhang 0119, Zhendong Mao 0001 |
ACM Multimedia | 3 |
| 2025 | Video Corpus Moment Retrieval With Query-Specific Context Learning and Progressive LocalizationabstractVideo corpus moment retrieval (VCMR) aims to retrieve a moment from a large corpus of untrimmed videos corresponding to a given language query. However, existing methods often fall short due to their reliance on simple cross-modal attention mechanisms and one-stop localization, which fail to handle the complex multimodal information and large search space effectively. To address these challenges, we propose a novel VCMR method with Query-specific Context Learning and Progressive Localization (QCLPL). First, we construct query-specific multimodal contexts that capture complementary and consistent semantics across subtitles and frames, ensuring informative and efficient context building. We further introduce a semantic contrastive loss to refine these multimodal contexts, filtering out query-irrelevant information. Additionally, we introduce a progressive localization strategy that transforms the moment localization task into a two-stage process. By classifying frames into foreground and background regions, we present a simplified binary classification problem before boundary prediction, constrained by a region-aware loss. This progressive approach leverages region priors to improve subsequent moment localization. Extensive experiments on the TVR and DiDeMo datasets demonstrate that our method significantly outperforms existing approaches, setting a new state of the art for VCMR. Peipei Song, Zhangling Duan, Shuo Wang 0008, Xiaojun Chang, Xun Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Alleviating Confirmation Bias in Learning with Noisy Labels via Two-Network CollaborationabstractDeep neural networks (DNNs) have achieved remarkable success in various computer vision tasks, e.g., image classification. However, most of the existing models depend heavily on annotated data, where label noise is inevitable. Training with such noisy data negatively impacts the generalization performance of DNNs. To this end, recent advances in learning with noisy labels (LNL) adopt the sample selection strategy that identifies clean samples from the noisy dataset to update DNNs, using semi-supervised learning where rejected samples are treated as unlabeled data. However, existing LNL methods often overlook the varying fitting difficulties of different classes, resulting in suboptimal sample selection and confirmation bias, and consequently, the errors accumulate during semi-supervised training. In this article, we propose a novel method, TNCollab, which aims at alleviating confirmation bias in both sample selection and semi-supervised training stages via two-network collaboration. Specifically, we introduce a class-adaptive threshold for sample selection to address the varying fitting difficulties across different classes. Additionally, we construct a hard set consisting of samples where the two networks disagree and introduce a noise-robust loss to extract potentially useful information while maintaining robustness against label noise. Furthermore, we propose a dual consistency loss to ensure consistent predictions between the networks across different augmented views of the same sample, facilitating mutual learning. Extensive experiments demonstrate that TNCollab achieves state-of-the-art performance on image classification and facial expression recognition tasks, particularly on CIFAR-10, CIFAR-100, WebVision, Clothing1M, Tiny-ImageNet, and RAF-DB datasets, showing improved visual understanding and generalization capabilities. Our codes are available at https://github.com/Delete12137/TNCollab . Peipei Song, Shengeng Tang, Dan Guo 0001, Xun Yang 0001 |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2025 | Towards Efficient Partially Relevant Video Retrieval With Active Moment DiscoveringabstractPartially relevant video retrieval (PRVR) is a practical yet challenging task in text-to-video retrieval, where videos are untrimmed and contain much background content. The pursuit here is of both effective and efficient solutions to capture the partial correspondence between text queries and untrimmed videos. Existing PRVR methods, which typically focus on modeling multi-scale clip representations, however, suffer from content independence and information redundancy, impairing retrieval performance. To overcome these limitations, we propose a simple yet effective approach with active moment discovering (AMDNet). We are committed to discovering video moments that are semantically consistent with their queries. By using learnable span anchors to capture distinct moments and applying masked multi-moment attention to emphasize salient moments while suppressing redundant backgrounds, we achieve more compact and informative video representations. To further enhance moment modeling, we introduce a moment diversity loss to encourage different moments of distinct regions and a moment relevance loss to promote semantically query-relevant moments, which cooperate with a partially relevant retrieval loss for end-to-end optimization. Extensive experiments on two large-scale video datasets (i.e., TVR and ActivityNet Captions) demonstrate the superiority and efficiency of our AMDNet. In particular, AMDNet is about 15.5 times smaller (#parameters) while 6.0 points higher (SumR) than the up-to-date method GMMFormer on TVR. Peipei Song, Long Lan, Weidong Chen 0013, Dan Guo 0001, Xun Yang 0001, Meng Wang 0001 |
IEEE Trans. Multim. | 1 |
| 2024 | Synergizing triple attention with depth quality for RGB-D salient object detection
Peipei Song, Peiyan Zhong, Jing Zhang 0052, Piotr Koniusz, Feng Duan 0006, Nick Barnes |
Neurocomputing | 1 |
| 2024 | Emotional Video Captioning With Vision-Based Emotion Interpretation NetworkabstractEffectively summarizing and re-expressing video content by natural languages in a more human-like fashion is one of the key topics in the field of multimedia content understanding. Despite good progress made in recent years, existing efforts usually overlooked the emotions in user-generated videos, thus making the generated sentence a bit boring and soulless. To fill the research gap, this paper presents a novel emotional video captioning framework in which we design a Vision-based Emotion Interpretation Network to effectively capture the emotions conveyed in videos and describe the visual content in both factual and emotional languages. Specifically, we first model the emotion distribution over an open psychological vocabulary to predict the emotional state of videos. Then, guided by the discovered emotional state, we incorporate visual context, textual context, and visual-textual relevance into an aggregated multimodal contextual vector to enhance video captioning. Furthermore, we optimize the network in a new emotion-fact coordinated way that involves two losses- Emotional Indication Loss and Factual Contrastive Loss, which penalize the error of emotion prediction and visual-textual factual relevance, respectively. In other words, we innovatively introduce emotional representation learning into an end-to-end video captioning network. Extensive experiments on public benchmark datasets, EmVidCap and EmVidCap-S, demonstrate that our method can significantly outperform the state-of-the-art methods by a large margin. Quantitative ablation studies and qualitative analyses clearly show that our method is able to effectively capture the emotions in videos and thus generate emotional language sentences to interpret the video content. Peipei Song, Dan Guo 0001, Xun Yang 0001, Shengeng Tang, Meng Wang 0001 |
IEEE Trans. Image Process. | 1 |
| 2024 | Efficiently Gluing Pre-Trained Language and Vision Models for Image CaptioningabstractVision-and-language pre-training models have achieved impressive performance for image captioning. But most of them are trained with millions of paired image-text data and require huge memory and computing overhead. To alleviate this, we try to stand on the shoulders of large-scale pre-trained language models (PLM) and pre-trained vision models (PVM) and efficiently connect them for image captioning. There are two major challenges: one is that language and vision modalities have different semantic granularity (e.g., a noun may cover many pixels), and the other is that the semantic gap still exists between the pre-trained language and vision models. To this end, we design a lightweight and efficient connector to glue PVM and PLM, which holds a criterion of selection-then-transformation . Specifically, in the selection phase, we treat each image as a set of patches instead of pixels. We select salient image patches and cluster them into visual regions to align with text. Then, to effectively reduce the semantic gap, we propose to map the selected image patches into text space through spatial and channel transformations. With training on image captioning datasets, the connector learns to bridge the semantic granularity and semantic gap via backpropagation, preparing for the PLM to generate descriptions. Experimental results on the MSCOCO and Flickr30k datasets demonstrate that our method yields comparable performance to existing works. By solely training the small connector, we achieve a CIDEr performance of 132.2% on the MSCOCO Karpathy test split. Moreover, our findings reveal that fine-tuning the PLM can further enhance performance potential, resulting in a CIDEr score of 140.6%. Code and models are available at https://github.com/YuanEZhou/PrefixCap . Peipei Song, Yuanen Zhou, Xun Yang 0001, Daqing Liu, Zhenzhen Hu 0004, Depeng Wang, Meng Wang 0001 |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2024 | Visual-linguistic-stylistic Triple Reward for Cross-lingual Image CaptioningabstractGenerating image captions in different languages is worth exploring and essential for non-native speakers. Nevertheless, collecting paired annotation for every language is time-consuming and impractical, particularly for minor languages. To this end, the cross-lingual image captioning task is proposed, which leverages existing image-source caption annotation data and wild unrelated target corpus to generate satisfactory caption in the target language. Current methods perform a two-step translation process of image-to-pivot (source) and pivot-to-target. The distinct two-step process comes with certain caption issues, such as the weak semantic alignment between the image and the generated caption and the generated caption’s non-target language style. To address these issues, we propose an end-to-end reinforce learning framework with Visual-linguistic-stylistic Triple Reward named TriR. In TriR, we jointly consider the visual, linguistic, and stylistic alignments to generate factual, fluent, and natural caption in the target language. To be specific, the image-source caption annotation provides factual semantic guidance, whereas the unrelated target corpus guides the language style of generated caption. To achieve this, we construct a visual reward module to measure the cross-modal semantic embedding of image and target caption, a linguistic reward module to measure the cross-linguistic embedding of source and target captions, and a stylistic reward module to imitate the presentation style of target corpus. The TriR can be implemented with either classical CNN-LSTM or prevalent Transformer architecture. Extensive experiments are conducted with four cross-lingual settings, i.e., Chinese-to-English, English-to-Chinese, English-to-German, and English-to-French. Experimental results demonstrate the remarkable superiority of our method, and sufficient ablation experiments validate the beneficial impact of every reward. Jing Zhang 0089, Dan Guo 0001, Xun Yang 0001, Peipei Song, Meng Wang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | Emotion-Prior Awareness Network for Emotional Video CaptioningabstractEmotional video captioning (EVC) is an emerging task to describe the factual content with the inherent emotion expressed in a video. It is crucial for the EVC task to effectively perceive subtle and ambiguous visual emotion cues in the stage of caption generation. However, existing captioning methods usually overlooked the learning of emotions in user-generated videos, thus making the generated sentence a bit boring and soulless. Peipei Song, Dan Guo 0001, Xun Yang 0001, Shengeng Tang, Erkun Yang, Meng Wang 0001 |
ACM Multimedia | 1 |
| 2023 | Memorial GAN With Joint Semantic Optimization for Unpaired Image CaptioningabstractMost works of image captioning are implemented under the full supervision of paired image-caption data. Limited to expensive cost of data collection, the task of unpaired image captioning has attracted researchers' attention. In this article, we propose a novel memorial GAN (MemGAN) with the joint semantic optimization for unpaired image captioning. The core idea is to explore implicit semantic correlation between disjointed images and sentences through building a multimodal semantic-aware space (SAS). Concretely, each modality is mapped into a unified multimodal SAS, where SAS includes the semantic vectors of image I , visual concepts O , unpaired sentence S , and the generated caption C . We adopt the memory unit based on multihead attention and relational gate as a backbone to preserve and transit crucial multimodal semantics in the SAS for image caption generation and sentence reconstruction. Then, the memory unit is embedded into a GAN framework to exploit the semantic similarity and relevance in SAS, that is, imposing a joint semantic-aware optimization on SAS without supervision clues. To summarize, the proposed MemGAN learns the latent semantic relevance of SAS's multimodalities in an adversarial manner. Extensive experiments and qualitative results demonstrate the effectiveness of MemGAN, achieving improvements over state of the arts on unpaired image captioning benchmarks. Peipei Song, Dan Guo 0001, Jinxing Zhou, Mingliang Xu 0001, Meng Wang 0001 |
IEEE Trans. Cybern. | 1 |
| 2023 | Contextual Attention Network for Emotional Video CaptioningabstractThis paper investigates an emerging and challenging task—emotional video captioning. Formally, given a video, the task aims to not only describe the factual content of the video, but also discover the emotional clues in the video. We propose a novel Contextual Attention Network (CANet), which recognizes and describes the fact and emotion in the video by semantic-rich context learning. To be specific, at each time step, we first extract visual and textual features from both input video and previously generated words. Then, we apply the attention mechanism to these features to capture informative contexts for captioning. We train the CANet model with the joint optimization of cross-entropy loss$\mathcal {L}_{CE}$and contrastive loss$\mathcal {L}_{CL}$, where$\mathcal {L}_{CE}$constrains the semantics of the generated sentence to be close to human annotation and$\mathcal {L}_{CL}$encourages discriminative representation learning from positive and negative pairs of video and caption. Experiments on two emotional video captioning datasets (i.e., EmVidCap and EmVidCap-S) demonstrate the superiority of CANet compared to the state-of-the-art approaches. Peipei Song, Dan Guo 0001, Jun Cheng 0002, Meng Wang 0001 |
IEEE Trans. Multim. | 1 |
| 2022 | Multi-Modal Transformer for RGB-D Salient Object DetectionabstractThe main focus of existing RGB-D salient object detection models is achieving effective multi-modal fusion. Due to the limited receptive field of conventional convolutional neural networks (CNNs), CNN-based multi-modal fusion strategies fail to extensively model the correlation between the two modalities (appearance information from the RGB image and geometric information from the depth data). Given the success of transformer networks for long-range dependency modeling, we investigate multi-modal transformer networks for RGB-D salient object detection. Specifically, a transformer-based multi-modal fusion module is presented to effectively fuse appearance features and geometric features. Experimental results on six challenging benchmark RGB-D salient object detection datasets demonstrate the effectiveness of our approach. Peipei Song, Jing Zhang 0052, Piotr Koniusz, Nick Barnes |
ICIP | 1 |
| 2020 | Weakly-Supervised Salient Object Detection via Scribble AnnotationsabstractCompared with laborious pixel-wise dense labeling, it is much easier to label data by scribbles, which only costs 1~2 seconds to label one image. However, using scribble labels to learn salient object detection has not been explored. In this paper, we propose a weakly-supervised salient object detection model to learn saliency from such annotations. In doing so, we first relabel an existing large-scale salient object detection dataset with scribbles, namely S-DUTS dataset. Since object structure and detail information is not identified by scribbles, directly training with scribble labels will lead to saliency maps of poor boundary localization. To mitigate this problem, we propose an auxiliary edge detection task to localize object edges explicitly, and a gated structure-aware loss to place constraints on the scope of structure to be recovered. Moreover, we design a scribble boosting scheme to iteratively consolidate our scribble annotations, which are then employed as supervision to learn high-quality saliency maps. As existing saliency evaluation metrics neglect to measure structure alignment of the predictions, the saliency map ranking may not comply with human perception. We present a new metric, termed saliency structure measure, as a complementary metric to evaluate sharpness of the prediction. Extensive experiments on six benchmark datasets demonstrate that our method not only outperforms existing weakly-supervised/unsupervised methods, but also is on par with several fully-supervised state-of-the-art models (Our code and data is publicly available at: https://github.com/JingZhang617/Scribble_Saliency). Jing Zhang 0052, Xin Yu 0002, Aixuan Li, Peipei Song, Bowen Liu 0012, Yuchao Dai |
CVPR | 4 |
| 2020 | Recurrent Relational Memory Network for Unsupervised Image CaptioningabstractUnsupervised image captioning with no annotations is an emerging challenge in computer vision, where the existing arts usually adopt GAN (Generative Adversarial Networks) models. In this paper, we propose a novel memory-based network rather than GAN, named Recurrent Relational Memory Network (R2M). Unlike complicated and sensitive adversarial learning that non-ideally performs for long sentence generation, R2M implements a concepts-to-sentence memory translator through two-stage memory mechanisms: fusion and recurrent memories, correlating the relational reasoning between common visual concepts and the generated words for long periods. R2M encodes visual context through unsupervised training on images, while enabling the memory to learn from irrelevant textual corpus via supervised fashion. Our solution enjoys less learnable parameters and higher computational efficiency than GAN-based methods, which heavily bear parameter sensitivity. We experimentally validate the superiority of R2M than state-of-the-arts on all benchmark datasets. Dan Guo 0001, Yang Wang 0023, Peipei Song, Meng Wang 0001 |
IJCAI | 3 |
| 2019 | Parallel Temporal Encoder For Sign Language TranslationabstractThis paper addresses the sign video interpretation which is a weakly supervised task. Each sign action in videos lacks exact boundaries or labels. We design a Parallel Temporal Encoder (PTEnc) to learn the temporal relation of a sign video from local and global sequential learning views in parallel. PTEnc utilizes the complementarity between the local and global temporal cues. Then, fused encoded feature sequence is fed into a Connectionist Temporal Classification (CTC) based sentence decoder. In addition, in order to enhance the temporal cues in each video, we introduce a reconstruction loss, which performs in an unsupervised way without additional labels. The CTC loss cooperates with the reconstruction loss in an end-to-end training manner. Experimental results on a benchmark dataset demonstrate the effectiveness of the proposed method. Peipei Song, Dan Guo 0001, Meng Wang 0001 |
ICIP | 1 |
| 2017 | Design of an SSVEP-based BCI system with visual servo module for a service robot to execute multiple tasksabstractBrain-computer interface (BCI) systems can translate the human mind into control commands, which makes it feasible to improve the life quality of physically challenged people. However, in real-life situations, it is still difficult for users to utilize robots to provide basic services with BCI systems. We aimed to propose a BCI-based system with a visual servo module to operate a service robot. We recorded single-channel steady-state visual evoked potentials (SSVEP) as input signals for the BCI system of this study. The visual stimuli for inducing SSVEP were modulated at seven different frequencies with the sampled sinusoidal method. Correspondingly, this SSVEP-based BCI system can generate seven control commands for the operation of the service robot, which can provide three fundamental services: mobility, manipulation, and delivery. The visual servo module was established to reduce the burden of users and accelerate service procedures. To evaluate the performance of this system, subjects were recruited to participate in the experiments. All the participants succeed in operating the robot to provide the basic services. According to the experimental results, this SSVEP-based BCI system that incorporates the visual servo module can be effectively used to operate service robots with reduced number of channels and increased ability to perform multiple tasks. Shili Sheng, Peipei Song, Lingyue Xie, Zhendong Luo, Wennan Chang, Shurui Jiang, Haoyong Yu, Chi Zhu 0001, Jeffrey Too Chuan Tan, Feng Duan 0006 |
ICRA | 2 |
| 2014 | Power consumption evaluation in random cellular networksabstractRecently, issues of power consumption at base stations (BSs) in wireless cellular networks have attracted great interest in both research communities and the industry. In this paper, we investigate the BS power consumption in multiple-input single-output (MISO) Poisson-Voronoi tessellation (PVT) random cellular networks. Taking into account the inter-cell interference, the impact of the traffic demands of users and the spatial traffic intensity on the BS power consumption are jointly considered for MISO PVT random cellular networks. Furthermore, the power consumption required at the BS in a typical PVT cell is modeled through characteristic functions. Simulation results are employed to evaluate the BS power consumption and the performance of the random cellular networks. Xiaohu Ge, Peipei Song, Tarik Taleb, Tao Han 0001, Jing Zhang 0025, Qiang Li 0009 |
WCNC | 2 |