EDBT 2026 Demo / reviewers in the wild / expert
Jinjing Gu
dblp:214/8887
· DBLP profile ↗
19ranked-venue papers
6as first author
18since 2021 · last 2026
0000-0002-7669-8201ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 4 first-author · 13 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Retrieval-based objects and relations prompt for image captioning
Jinjing Gu, Tianbao Qin, Zhengpeng Zhao |
Eng. Appl. Artif. Intell. | 1 |
| 2026 | Adverse multi-weather image restoration for boosting downstream object detection
Jinjing Gu, Chenggang Yang, Zhengpeng Zhao |
Eng. Appl. Artif. Intell. | 1 |
| 2026 | StrCCL: Structure-aware Contrastive Consistency Loss for Artistic Style Transfer
Shuyu Pan, Zhengpeng Zhao, Qiuxia Yang, Jinjing Gu, Dan Xu 0001 |
Expert Syst. Appl. | 5 |
| 2026 | Multimodal progressive contrastive learning for sentiment analysis
Lianmin Zhou, Zhengpeng Zhao, Jue Feng, Dan Xu 0001, Jinjing Gu |
Neurocomputing | 6 |
| 2026 | CoMPLe: Cross-Modal Hybrid Prompt Learning for End-to-End Multimodal Emotion RecognitionabstractThe quality of features directly affects the accuracy of Multimodal Emotion Recognition (MER). A key challenge in this context is the effective extraction of dynamically interactive multimodal features to enrich conversational emotion representations. However, existing approaches are often constrained by non-end-to-end architectures, overlooking the significance of feature extraction in MER. To address the problem of dynamic interaction in emotion feature extraction, this paper introduces an end-to-end network based on Cross-Modal Hybrid Prompt Learning (CoMPLe). The model takes raw video as input and leverages three prompt mechanisms to guide large-scale pre-trained encoders in extracting emotionally salient features with latent correlations. Specifically, we design a cross-modal soft prompt learning strategy to mine complementary information across modalities and dynamically adjust the cross-modal semantic space. To capture stage-dependent characteristics, deep feature prompts are incorporated to progressively learn intra-modal contextual representations. Furthermore, a label prompt mechanism is proposed to construct hard prompt templates from emotion labels. Finally, the highest cosine similarity is computed between each unimodal feature and the label prompt templates to activate factual knowledge relevant to emotion recognition. Experiments on three public datasets show that the end-to-end network proposed in this paper surpasses the existing State-Of-The-Art baselines. Jue Feng, Zhengpeng Zhao, Lianmin Zhou, Jiale Ye, Dan Xu 0001, Jinjing Gu |
IEEE Trans. Affect. Comput. | 7 |
| 2025 | FNContra: Frequency-domain Negative Sample Mining in Contrastive Learning for limited-data image generation
Qiuxia Yang, Zhengpeng Zhao, Shuyu Pan, Jinjing Gu, Dan Xu 0001 |
Expert Syst. Appl. | 5 |
| 2025 | Multimodal hypergraph network with contrastive learning for sentiment analysis
Zhengpeng Zhao, Qiuxia Yang, Jinjing Gu, Dan Xu 0001 |
Neurocomputing | 6 |
| 2025 | Leveraging Enriched Skeleton Representation With Multi-Relational Metrics for Few-Shot Action RecognitionabstractFew-shot action recognition aims to identify new action classes with limited training samples. Most existing methods overlook the low information content and diversity of skeleton features, failing to exploit useful information in rare samples during meta-training. This leads to poor feature discriminability and recognition accuracy. To address both issues, we propose a novel Enriched Skeleton Representation and Multi-relational Metrics (ESR-MM) method for skeleton-based few-shot action recognition. First, a Frobenius Norm Diversity Loss is introduced to enrich skeleton representation by maximizing the Frobenius norm of the skeleton feature matrix. This mitigates over-smoothing and boosts information content and diversity. Leveraging these enriched features, we propose a multi-relational metrics strategy exploiting cross-sample task-specific information, intra-sample temporal order, and inter-sample distance. Specifically, Support-Adaptive Attention leverages task-specific cues between samples to generate attention-enhanced features. Then, the Bidirectional Temporal Coherent Mean Hausdorff Metric integrates Temporal Coherence Measure into the Bidirectional Mean Hausdorff Metric for class separation by accounting for temporal order. Finally, Prototype-discriminative Contrastive Loss exploits distances from class prototypes to query samples. ESR-MM demonstrates superior performance on two benchmarks. Jingyun Tian, Jinjing Gu, Zhengpeng Zhao |
IEEE Trans. Multim. | 2 |
| 2024 | Dual-path hypernetworks of style and text for one-shot domain adaptation
Zhengpeng Zhao, Qiuxia Yang, Jinjing Gu, Yupan Li, Dan Xu 0001 |
Appl. Intell. | 5 |
| 2024 | Short-term trajectory prediction for individual metro passengers based on multi-level periodicity mining from semantic trajectory
Jinjing Gu, Wei Fan 0006, Wenwen Qin |
Eng. Appl. Artif. Intell. | 1 |
| 2024 | Dynamic hypergraph convolutional network for multimodal sentiment analysis
Dongming Zhou 0001, Jinde Cao, Jinjing Gu, Zhengpeng Zhao, Dan Xu 0001 |
Neurocomputing | 5 |
| 2023 | BIT: Improving Image-text Sentiment Analysis via Learning Bidirectional Image-text InteractionabstractExploring the interaction between image and text has a great strength for image-text sentiment analysis. However, most methods only focus on learning forward interaction in forward image-text features and fail to capture the backward interaction in backward image-text features, which leads to the loss of necessary information embedded in backward interaction. In this paper, Bidirectional Interaction Transformer (BIT) that models both forward and backward image-text interactions is proposed for image-text sentiment analysis. Specifically, we first encode image and text to forward and backward features. Then, these features are fed into Bidirectional Interaction Encoder (BIE) with Forward Interaction and Back Interaction branches to model bidirectional (i.e., forward and backward) image-text interaction. Finally, Two-scale Adaptive Gating Fusion (TAGF) is designed to adaptively fuse the forward and backward interactions learned by BIE. Extensive experiments conducted on two public datasets demonstrate the effectiveness of the proposed model. Xingwang Xiao, Zhengpeng Zhao, Jinjing Gu, Dan Xu 0001 |
IJCNN | 4 |
| 2023 | Collaborative fine-grained interaction learning for image-text sentiment analysis
Xingwang Xiao, Dongming Zhou 0001, Jinde Cao, Jinjing Gu, Zhengpeng Zhao, Dan Xu 0001 |
Knowl. Based Syst. | 5 |
| 2023 | Coherent Visual Storytelling via Parallel Top-Down Visual and Topic AttentionabstractVisual storytelling aims at producing a narrative paragraph for a given photo album automatically. It introduces more new challenges than individual image paragraph descriptions, mainly due to the difficulty in preserving coherent topics and in generating diverse phrases to depict the rich content of a photo album. Existing attention-based models that lack higher-level guiding information always result in a deviation between the generated sentence and the topic expressed by the image. In addition, these widely applied language generation approaches employing standard beam search tend to produce monotonous descriptions. In this work, a coherent visual storytelling (CoVS) framework is designed to address the above-mentioned problems. Specifically, in the encoding phase, an image sequence encoder is designed to efficiently extract visual features of the input photo album. Then, the novel parallel top-down visual and topic attention (PTDVTA) decoder is constructed via a topic-aware neural network, a parallel top-down attention model, and a coherent language generator. Concretely, visual attention focuses on the attributes and the relationships of the objects, while topic attention integrating a topic-aware neural network could improve the coherence of generated sentences. Eventually, a phrase beam search algorithm with$n$-gram hamming diversity is further designed to optimize the expression diversity of the generated story. To justify the proposed CoVS framework, extensive experiments are conducted on the VIST dataset, which shows that CoVS can automatically generate coherent and diverse stories in a more natural way. Moreover, CoVS obtains better performance than state-of-the-art baselines on BLEU-4 and METEOR scores, while maintaining good CIDEr and ROUGH_L scores. The source code of this work can be found inhttps://mic.tongji.edu.cn. Jinjing Gu, Hanli Wang, Ruichao Fan |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Multi-concept Mining for Video Captioning Based on Multiple TasksabstractVideo captioning is a challenging cross-modal task that requires taking full advantage of both vision and language. To identify objects in videos, object detectors are usually employed to extract high-level object-related features, but the fine-grained knowledge from the object detectors are often neglected. Also, there is a fact that not just the task of object detection has the ability to obtain additional knowledge for video understanding. In this paper, multiple tasks are assigned to fully mine multi-concept knowledge in both vision and language, including video-to-video knowledge, video-to-text knowledge and text-to-text knowledge. Moreover, since there is a strong synergy in knowledge, both of global and local word similarities are developed based on the text-to-text knowledge to boost the robustness of the mined semantic knowledge. The mined knowledge can offer the model an extra guidance apart from linguistic prior to generate more semantically appropriate and grammatically correct sentences. The experimental results on the benchmark MSVD and MSR-VTT datasets show that the proposed method makes remarkable improvement on all metrics on MSVD and two out of four metrics on MSR-VTT. Qinyu Zhang 0005, Pengjie Tang, Hanli Wang, Jinjing Gu |
ISCAS | 4 |
| 2022 | Short-term trajectory prediction for individual metro passengers integrating diverse mobility patterns with adaptive location-awareness
Jinjing Gu, Wei Fan 0006 |
Inf. Sci. | 1 |
| 2021 | Visual Storytelling with Hierarchical BERT Semantic GuidanceabstractVisual storytelling, which aims at automatically producing a narrative paragraph for photo album, remains quite challenging due to the complexity and diversity of photo album content. In addition, open-domain photo albums cover a broad range of topics and this results in highly variable vocabularies and expression styles to describe photo albums. In this work, a novel teacher-student visual storytelling framework with hierarchical BERT semantic guidance (HBSG) is proposed to address the above-mentioned challenges. The proposed teacher module consists of two joint tasks, namely, word-level latent topic generation and semantic-guided sentence generation. The first task aims to predict the latent topic of the story. As there is no ground-truth topic information, a pre-trained BERT model based on visual contents and annotated stories is utilized to mine topics. Then the topic vector is distilled to a designed image-topic prediction model. In the semantic-guided sentence generation task, HBSG is introduced for two purposes. The first is to narrow down the language complexity across topics, where the co-attention decoder with vision and semantic is designed to leverage the latent topics to induce topic-related language models. The second is to employ sentence semantic as an online external linguistic knowledge teacher module. Finally, an auxiliary loss is devised to transform linguistic knowledge into the language generation model. Extensive experiments are performed to demonstrate the effectiveness of HBSG framework, which surpasses the state-of-the-art approaches evaluated on the VIST test set. Ruichao Fan, Hanli Wang, Jinjing Gu |
MMAsia | 3 |
| 2021 | Spatio-temporal trajectory estimation based on incomplete Wi-Fi probe data in urban rail transit network
Jinjing Gu, Yanshuo Sun, Shenmeihui Liao |
Knowl. Based Syst. | 1 |
| 2020 | RepeatPadding: Balancing words and sentence length for language comprehension in visual question answering
Yu Long 0003, Pengjie Tang, Zhihua Wei 0001, Jinjing Gu, Hanli Wang |
Inf. Sci. | 4 |