Shuqin Chen

dblp:198/7684 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
12since 2021 · last 2027
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 5 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2027 Rethinking 3D point cloud adversarial attacks from models' inherent focus
Xiaowen Cai 0001, Shuqin Chen, Junhao Dong 0001, Keke Tang, Zhongliang Guo 0001, Daizong Liu
Expert Syst. Appl.3
2026 Ask and focus more: Question-prompt uncertainty allocation for dual-controllable video captioning
Shuqin Chen, Xingrui Yang 0003, Xiaohan Yu 0001, Xian Zhong
Pattern Recognit.1
2026 Fine-Grained Lexical-Centric Semantic Network for Coherent Video Paragraph Captioning
abstract
Video paragraph captioning (VPC) aims to generate coherent, detailed narratives that accurately reflect a video's content. However, existing methods typically depend on coarse-grained event correlations and neglect the nuanced spatio-temporal interactions critical for comprehensive understanding. Refined verbs and prepositions, encoding actions and spatial relations, are essential for clear, consistent descriptions. To address these issues, we propose the Fine-Grained Lexical-Centric Semantic Network (FLS-Net), which emphasizes verbs and prepositions linked to salient objects to improve spatio-temporal coherence across events. FLS-Net integrates a multi-lexical synergy mechanism, leveraging nouns obtained via multi-modal matching, and employs a Verb-Guided Event Consistency Module (VECM) alongside a Preposition-Driven Relation Representation Module (PRRM). A cyclic encoder-decoder architecture further enforces event consistency, significantly boosting VPC performance. Extensive experiments onActivityNet CaptionsandYouCook2demonstrate FLS-Net's superiority over state-of-the-art approaches. The source code is available athttps://github.com/yangxingrui/FLS.
Shuqin Chen, Xian Zhong, Xingrui Yang 0003, Bin Sheng 0001, Alex Chichung Kot
IEEE Trans. Multim.1
2025 Bridging the One-to-Many Gap: Multi-label Semantic Learning and Relay for Video Captioning
abstract
Many commonly used video captioning datasets contain multiple caption annotations per video. When training with cross-entropy loss, the model encounters ambiguity because the same input is mapped to different targets, leading to confusion. To address this issue, we propose the Multi-label Semantic Learning and Relay (MSLR) framework, which transforms the one-to-many relation between videos and their descriptions into a one-to-one mapping. Specifically, MSLR introduces two modules after the decoder. The Multi-label Multi-granularity Learning (MML) module integrates sentence-level granularity from multiple descriptions and employs attention mechanisms to extract word-level granularity weights, thereby capturing complementary semantic information from diverse perspectives. The Multi-label Semantic Relay (MSR) module subsequently leverages a parameter-sharing mechanism to feed complementary semantics back into the decoder, thus preventing the generation of overly generic descriptions during inference. Compared to state-of-the-art lightweight methods, MSLR achieves highly competitive results. The code is available at https://github.com/hyk0320/MSLR.
Shuqin Chen, Yikang Hu, Zhixin Sun, Liangjun Yu, Xian Zhong
ICME1
2025 Refined linguistic deliberation for video captioning via cascade transformer and LSTM
Shuqin Chen, Zhixin Sun, Yikang Hu, Shifeng Wu
Multim. Syst.1
2024 Action-aware Linguistic Skeleton Optimization Network for Non-autoregressive Video Captioning
abstract
Non-autoregressive video captioning methods generate visual words in parallel but often overlook semantic correlations among them, especially regarding verbs, leading to lower caption quality. To address this, we integrate action information of highlighted objects to enhance semantic connections among visual words. Our proposed Action-aware Language Skeleton Optimization Network (ALSO-Net) tackles the challenge of extracting action information across frames, improving understanding of complex context-dependent video actions and reducing sentence inconsistencies. ALSO-Net incorporates a linguistic skeleton tag generator to refine semantic correlations and a video action predictor to enhance verb prediction accuracy in video captions. We also address issues of unsatisfactory caption length and quality by jointly optimizing different levels of motion prediction loss. Experimental evaluation on prominent video captioning datasets demonstrates that ALSO-Net outperforms baseline methods by a significant margin and achieves competitive performance compared to state-of-the-art autoregressive methods with smaller model complexity and faster inference time.
Shuqin Chen, Xian Zhong, Lei Zhu 0003, Ping Li 0016, Xiaokang Yang 0001, Bin Sheng 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2023 Refined Semantic Enhancement towards Frequency Diffusion for Video Captioning
abstract
Video captioning aims to generate natural language sentences that describe the given video accurately. Existing methods obtain favorable generation by exploring richer visual representations in encode phase or improving the decoding ability. However, the long-tailed problem hinders these attempts at low-frequency tokens, which rarely occur but carry critical semantics, playing a vital role in the detailed generation. In this paper, we introduce a novel Refined Semantic enhancement method towards Frequency Diffusion (RSFD), a captioning model that constantly perceives the linguistic representation of the infrequent tokens. Concretely, a Frequency-Aware Diffusion (FAD) module is proposed to comprehend the semantics of low-frequency tokens to break through generation limitations. In this way, the caption is refined by promoting the absorption of tokens with insufficient occurrence. Based on FAD, we design a Divergent Semantic Supervisor (DSS) module to compensate for the information loss of high-frequency tokens brought by the diffusion process, where the semantics of low-frequency tokens is further emphasized to alleviate the long-tailed problem. Extensive experiments indicate that RSFD outperforms the state-of-the-art methods on two benchmark datasets, i.e., MSR-VTT and MSVD, demonstrate that the enhancement of low-frequency tokens semantics can obtain a competitive generation effect. Code is available at https://github.com/lzp870/RSFD.
Xian Zhong, Shuqin Chen, Kui Jiang, Chen Chen 0001, Mang Ye
AAAI3
2023 Background Disturbance Mitigation for Video Captioning Via Entity-Action Relocation
abstract
Video captioning aims to generate sentences to accurately describe the video content, in which video background plays the role of prompts. State-of-the-art methods tend to explore richer video representations adequately, fusing with language to improve caption quality, which has shown great success. However, they focus on exploiting foreground semantics, ignoring the potential negative impact of video background disturbance to caption generation, i.e., the entities and the actions are misjudged by a similar video background. To ameliorate this issue, we propose Entity-Action Relocation (EAR) to enhance the adaptability of entities and actions to various backgrounds by giving them the background. Specifically, for an extracted original video feature, we construct a mixed background for all entities and actions to form a distracting video feature sample. After that, contrastive learning is applied to pull the generated caption of the original representations and of the distracting representations closer, and to push the former away from the generated caption of other videos, explicitly concentrating on the entities and actions of the current video scene. Extensive experiments on two public datasets (MSR-VTT and MSVD) demonstrate that dealing with background disturbance for video can obtain a competitive caption generation effect.
Xian Zhong, Shuqin Chen, Wenxin Huang, Lin Li 0001
ICASSP3
2023 Video Captioning Based on Cascaded Attention-Guided Visual Feature Fusion
Shuqin Chen, Yikang Hu
Neural Process. Lett.1
2022 Visual-Aware Attention Dual-Stream Decoder for Video Captioning
abstract
Video captioning is a challenging task that captures different visual parts and describes them in sentences, for it requires visual and linguistic coherence. The attention mechanism in the current video captioning method learns to assign weight to each frame, promoting the decoder dynamically. This may not explicitly model the correlation and the temporal coherence of the visual features extracted in the sequence frames. To generate semantically coherent sentences, we propose a new Visual-aware Attention (VA) model, which concatenates dynamic changes of temporal sequence frames with the words at the previous moment, as the input of attention mechanism to extract sequence features. In addition, the prevalent approaches widely use the Teacher-forcing (TF) learning during training, where the next token is generated conditioned on the previous ground-truth tokens. The semantic information in the previously generated tokens is lost. Therefore, we design a Self-forcing (SF) stream that takes the semantic information in the probability distribution of the previous token as input to enhance the current token. The Dual-stream Decoder (DD) architecture unifies the TF and SF streams, generating sen-tences to promote the annotated captioning for both streams. Meanwhile, with the Dual-stream Decoder utilized, the ex-posure bias problem is alleviated, caused by the discrepancy between the training and testing in the TF learning. The effectiveness of the proposed Visual-aware Attention Dual-stream Decoder (VADD) is demonstrated through the result of ex-perimental studies on Microsoft video description (MSVD) corpus and MSR-Video to text (MSR-VTT) datasets.
Zhixin Sun, Shuqin Chen, Luo Zhong
ICME2
2022 Dual-Scale Alignment-Based Transformer on Linguistic Skeleton Tags for Non-Autoregressive Video Captioning
abstract
Due to the characteristic of one-time parallel generation of a caption, non-autoregressive video captioning lacks strong dependencies between words. Although using guideline of scene-related visual words can promote caption generation, the semantic relations among visual words are barely explored, limiting the accurate representation. To this end, we propose a Dual-Scale Alignment-based transformer on Linguistic Skeleton Tags (DSA-LST), which alleviates the defect above in the form of visual words group (several words representing a video frame). Different groups represent different semantic dependencies by attention. We utilize linguistic skeleton tags (i.e., several groups) as sentence-level supervision for visual words sequence. For visual words group to accurately express a specific frame, we further design dual scales of visual-language bi-direction alignment to achieve internal relevance of the tags. Extensive experiments conducted on widely used datasets: MSVD and MSR-VTT demonstrate the effectiveness of our method when compared with existing approaches.
Xian Zhong, Shuqin Chen, Zhixin Sun, Huantao Zheng, Kui Jiang
ICME3
2021 Modeling Context-Guided Visual and Linguistic Semantic Feature for Video Captioning
Zhixin Sun, Xian Zhong, Shuqin Chen, Duxiu Feng
ICANN (5)3
2020 Complementing Representation Deficiency in Few-shot Image Classification: A Meta-Learning Approach
abstract
Few-shot learning is a challenging problem that has attracted more and more attention recently since abundant training samples are difficult to obtain in practical applications. Meta-learning has been proposed to address this issue, which focuses on quickly adapting a predictor as a base-learner to new tasks, given limited labeled samples. However, a critical challenge for meta-learning is the representation deficiency since it is hard to discover common information from a small number of training samples or even one, as is the representation of key features from such little information. As a result, a meta-learner cannot be trained well in a high-dimensional parameter space to generalize to new tasks. Existing methods mostly resort to extracting less expressive features so as to avoid the representation deficiency. Aiming at learning better representations, we propose a meta-learning approach with complemented representations network (MCRNet) for few-shot image classification. In particular, we embed a latent space, where latent codes are reconstructed with extra representation information to complement the representation deficiency. Furthermore, the latent space is established with variational inference, collaborating well with different base-learners, and can be extended to other models. Finally, our end-to-end framework achieves the state-of-the-art performance in image classification on three standard few-shot learning datasets.
Xian Zhong, Wenxin Huang, Lin Li 0001, Shuqin Chen, Chia-Wen Lin
ICPR5
2020 An emotion classification algorithm based on SPT-CapsNet
Xian Zhong, Jinhang Liu, Lin Li 0001, Shuqin Chen, Yuyu Dong, Bingqing Wu, Luo Zhong
Neural Comput. Appl.4
2020 Adaptively Converting Auxiliary Attributes and Textual Embedding for Video Captioning Based on BiLSTM
Shuqin Chen, Xian Zhong, Lin Li 0001, Luo Zhong
Neural Process. Lett.1