Pengjie Tang

dblp:192/5253 · DBLP profile ↗
← Back
24ranked-venue papers
9as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 4 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 4 since 2021Systems, architecture and hardware · 3 · 3 since 2021Computer networks · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 GLM-EER: Global-Local Memory and Emotion Evaluation Refinement For Emotional Video Description
Chong Ma 0002, Shengbo Chen, Pengjie Tang, Hong Rao, Hanli Wang
Expert Syst. Appl.3
2025 SRVC-LA: Sparse regularization of visual context and latent attention based model for video description
Pengjie Tang, Jiayu Zhang 0002, Hanli Wang, Yunlan Tan, Yun Yi
Neurocomputing1
2025 MGTR-MISS: More Ground Truth Retrieving based Multimodal Interaction and Semantic Supervision for video description
Jiayu Zhang 0002, Pengjie Tang, Yunlan Tan, Hanli Wang
Neural Networks2
2024 Emotion recognition in user-generated videos with long-range correlation-aware network
abstract
Abstract Emotion recognition in user‐generated videos plays an essential role in affective computing. In general, visual information directly affects human emotions, so the visual modality is significant for emotion recognition. Most classic approaches mainly focus on local temporal information of videos, which potentially restricts their capacity to encode the correlation of long‐range context. To address this issue, a novel network is proposed to recognize emotions in videos. To be specific, a spatio‐temporal correlation‐aware block is designed to depict the long‐range correlations between input tokens, where the convolutional layers are used to learn the local correlations and the inter‐image cross‐attention is designed to learn the long‐range and spatio‐temporal correlations between input tokens. To generate diverse and challenging samples, a dual‐augmentation fusion layer is devised, which fuses each frame with its corresponding frame in the temporal domain. To produce rich video clips, a long‐range sampling layer is designed, which generates clips in a wide range of spatial and temporal domains. Extensive experiments are conducted on two challenging video emotion datasets, namely VideoEmotion‐8 and Ekman‐6. The experimental results demonstrate that the proposed method obtains better performance than baseline methods. Moreover, the proposed method achieves state‐of‐the‐art results on the two datasets. The source code of the proposed network is available at: https://github.com/JinChow/LRCANet .
Yun Yi, Hanli Wang, Pengjie Tang, Min Wang 0020
IET Image Process.4
2023 Reversible Circuit Synthesis Method Using Sub-graphs of Shared Functional Decision Diagrams
abstract
Abstract Reversible circuit synthesis methods based on decision diagrams achieve low quantum costs but do not account for quantum bit (qubit) limits for the application of reversible logic in quantum computing. Here, a synthesis method using sub-graphs of shared functional decision diagrams (SFDDs) is proposed for reducing the number of lines when synthesizing reversible circuits. An SFDD is partitioned into sub-graphs by exploiting the longest dominant-active paths, and the sub-graphs are mapped to reversible gate cascades. To further reduce the number of lines, template root matching is presented for reusing circuit lines. Experimental results indicate that the proposed method achieves the known minimum number of lines in many cases and has good scalability. Although the proposed method increases the quantum cost over a prior method based on functional decision diagrams, it significantly reduces the number of lines in most cases. Compared with the one-pass method using quantum multiple-valued decision diagrams, the proposed method reduces the quantum cost without increasing the number of lines in many cases. When compared with the lookup table-based method using a direct mapping flow, the method reduces the number of lines in a few cases. Thus, the method aids in the physical realization of a quantum circuit.
Dengli Bu, Junyi Deng, Pengjie Tang, Shuhong Yang
Comput. J.3
2023 Deep sequential collaborative cognition of vision and language based model for video description
Pengjie Tang, Yunlan Tan, Jiewu Xia
Multim. Tools Appl.1
2022 Spatio-temporal Super-resolution Network: Enhance Visual Representations for Video Captioning
abstract
Video captioning is a sequence-to-sequence task of automatically generating descriptions for given videos. Due to the diversity of video scenes, learning rich representations is critical for video captioning. However, previous works mainly exploited elaborate features but neglected the loss of information caused by frame sampling and image compression. In this paper, we propose a novel spatio-temporal super-resolution (STSR) network which is jointly trained for the video captioning task and the video super-resolution task in an end-to-end fashion. Specifically, a video super-resolution task consists of two subtasks: spatial super-resolution restores high-resolution image features while temporal super-resolution reconstructs missing frame features between two adjacent sampled frames. By sharing multi-modal encoders across both of these two tasks, STSR encourages encoders to capture salient visual contents and learn context-aware representations. Experiments on two benchmark datasets demonstrate that the proposed STSR boosts video captioning performances significantly and outperforms most state-of-the-art approaches.
Quanhui Cao, Pengjie Tang, Hanli Wang
ISCAS2
2022 Multi-concept Mining for Video Captioning Based on Multiple Tasks
abstract
Video captioning is a challenging cross-modal task that requires taking full advantage of both vision and language. To identify objects in videos, object detectors are usually employed to extract high-level object-related features, but the fine-grained knowledge from the object detectors are often neglected. Also, there is a fact that not just the task of object detection has the ability to obtain additional knowledge for video understanding. In this paper, multiple tasks are assigned to fully mine multi-concept knowledge in both vision and language, including video-to-video knowledge, video-to-text knowledge and text-to-text knowledge. Moreover, since there is a strong synergy in knowledge, both of global and local word similarities are developed based on the text-to-text knowledge to boost the robustness of the mined semantic knowledge. The mined knowledge can offer the model an extra guidance apart from linguistic prior to generate more semantically appropriate and grammatically correct sentences. The experimental results on the benchmark MSVD and MSR-VTT datasets show that the proposed method makes remarkable improvement on all metrics on MSVD and two out of four metrics on MSR-VTT.
Qinyu Zhang 0005, Pengjie Tang, Hanli Wang, Jinjing Gu
ISCAS2
2022 Synthesis of Reversible Circuits with Reduced Nearest-Neighbor Cost Using Kronecker Functional Decision Diagrams
Dengli Bu, Pengjie Tang, Haohao Yuan
J. Electron. Test.3
2022 Visual and language semantic hybrid enhancement and complementary for video description
Pengjie Tang, Yunlan Tan, Wenlang Luo
Neural Comput. Appl.1
2022 Emotion Expression With Fact Transfer for Video Description
abstract
Translating a video into natural language is a fundamental but challenging task in visual understanding, since there is a great gap between visual content and linguistic sentence. More attention has been paid to this research field and a number of state-of-the-art results are achieved in recent years. However, the emotions in videos are usually overlooked, leading to the generated description sentences being boring and colorless. In this work, we construct a new dataset for video description with emotion expression, which consists of two parts: a re-annotated subset of the MSVD dataset with emotion embedded and another subset annotated with long sentences and rich emotions based on a video emotion recognition dataset. A fact transfer based framework is designed, which incorporates a fact stream and an emotion stream to generate sentences with emotion expression for video description. In addition, we propose a novel approach for sentence evaluation by balancing facts and emotions. A group of experiments are conducted, and the experimental results demonstrate the effectiveness of the proposed methods, including the idea of dataset construction for video description with emotion expression, model training and testing, and the emotion evaluation metric. The project page (including the code and dataset) can be found inhttps://mic.tongji.edu.cn/ce/70/c9778a183920/page.htm.
Hanli Wang, Pengjie Tang, Qinyu Li
IEEE Trans. Multim.2
2021 CaptionNet: A Tailor-made Recurrent Neural Network for Generating Image Descriptions
abstract
Image captioning is a challenging task of visual understanding and has drawn more attention of researchers. In general, two inputs are required at each time step by the Long Short-Term Memory (LSTM) network used in popular attention based image captioning frameworks, including image features and previous generated words. However, error will be accumulated if the previous words are not accurate and the related semantic is not efficient enough. Facing these challenges, a novel model named CaptionNet is proposed in this work as an improved LSTM specially designed for image captioning. Concretely, only attended image features are allowed to be fed into the memory of CaptionNet through input gates. In this way, the dependency on the previous predicted words can be reduced, forcing model to focus on more visual clues of images at the current time step. Moreover, a memory initialization method called image feature encoding is designed to capture richer semantics of the target image. The evaluation on the benchmark MSCOCO and Flickr30K datasets demonstrates the effectiveness of the proposed CaptionNet model, and extensive ablation studies are performed to verify each of the proposed methods. The project page can be found in https://mic.tongji.edu.cn/3f/9c/c9778a147356/page.htm.
Longyu Yang, Hanli Wang, Pengjie Tang, Qinyu Li
IEEE Trans. Multim.3
2020 RepeatPadding: Balancing words and sentence length for language comprehension in visual question answering
Yu Long 0003, Pengjie Tang, Zhihua Wei 0001, Jinjing Gu, Hanli Wang
Inf. Sci.2
2020 Translating video into language by enhancing visual and language representations
Pengjie Tang, Yunlan Tan
J. Vis. Commun. Image Represent.1
2020 Double-channel language feature mining based model for video description
Pengjie Tang, Jiewu Xia, Yunlan Tan
Multim. Tools Appl.1
2019 Rich Visual and Language Representation with Complementary Semantics for Video Captioning
abstract
It is interesting and challenging to translate a video to natural description sentences based on the video content. In this work, an advanced framework is built to generate sentences with coherence and rich semantic expressions for video captioning. A long short term memory (LSTM) network with an improved factored way is first developed, which takes the inspiration of LSTM with a conventional factored way and a common practice to feed multi-modal features into LSTM at the first time step for visual description. Then, the incorporation of the LSTM network with the proposed improved factored way and un-factored way is exploited, and a voting strategy is utilized to predict candidate words. In addition, for robust and abstract visual and language representation, residuals are employed to enhance the gradient signals that are learned from the residual network (ResNet), and a deeper LSTM network is constructed. Furthermore, three convolutional neural network based features extracted from GoogLeNet, ResNet101, and ResNet152, are fused to catch more comprehensive and complementary visual information. Experiments are conducted on two benchmark datasets, including MSVD and MSR-VTT2016, and competitive performances are obtained by the proposed techniques as compared to other state-of-the-art methods.
Pengjie Tang, Hanli Wang, Qinyu Li
ACM Trans. Multim. Comput. Commun. Appl.1
2018 Image Captioning with Word Level Attention
abstract
Image captioning is an attractive and challenging task to perform automatic image description and a number of works are designed for this task. Most of these researches are based on convolutional neural network (CNN) and recurrent neural network (RNN), where the primary input to language model for word prediction at the current time step is usually the linguistic word generated at the previous time step. In this work, a novel word level attention layer is designed to process image features with two modules for accurate word prediction. The first is a bidirectional spatial embedding module to handle feature maps, then the second module employs attention mechanism to extract word level attention which will be fed into language model. The experimental results on the benchmark MSCOCO dataset demonstrate that the proposed model achieves the state-of-the-art performances with 106.0 on CIDEr and 34.0 on B-4, respectively.
Hanli Wang, Pengjie Tang
ICIP3
2018 Refining Attention: A Sequential Attention Model for Image Captioning
abstract
Visual attention is widely applied to image captioning. Previous works put visual attention and linguistic word into a long short-term memory network together, but neglect the sequential relation of attention at different time steps during word prediction. Moreover, the abstraction degree of visual attention is usually different from that of linguistic word. To address these issues, a sequential attention model is proposed in this work to handle visual attention by considering the corresponding sequential relation, and hence the internal relation among attention at each word prediction step is well utilized to enhance the visual information during sentence decoding. The experimental results on the benchmark MSCOCO and Flickr30K datasets show that the proposed model achieves excellent performances with 108.1 and 34.9 respectively on the evaluation criteria of CIDEr and BLEU-4 for MSCOCO.
Qinyu Li, Hanli Wang, Pengjie Tang
ICME4
2018 Deep sequential fusion LSTM network for image description
Pengjie Tang, Hanli Wang, Sam Kwong
Neurocomputing1
2018 Looking deeper and transferring attention for image captioning
Hanli Wang, Pengjie Tang
Multim. Tools Appl.4
2017 Image captioning with deep LSTM based on sequential residual
abstract
Image captioning is a fundamental task which requires semantic understanding of images and the ability of generating description sentences with proper and correct structure. In consideration of the problem that language models are always shallow in modern image caption frameworks, a deep residual recurrent neural network is proposed in this work with the following two contributions. First, an easy-to-train deep stacked Long Short Term Memory (LSTM) language model is designed to learn the residual function of output distributions by adding identity mappings to multi-layer LSTMs. Second, in order to overcome the over-fitting problem caused by larger-scale parameters in deeper LSTM networks, a novel temporal Dropout method is proposed into LSTM. The experimental results on the benchmark MSCOCO and Flickr30K datasets demonstrate that the proposed model achieves the state-of-the-art performances with 101.1 in CIDEr on MSCOCO and 22.9 in B-4 on Flickr30K, respectively.
Kaisheng Xu, Hanli Wang, Pengjie Tang
ICME3
2017 Richer Semantic Visual and Language Representation for Video Captioning
abstract
Translating and summarizing a video into natural language is an interesting and challenging visual task. In this work, a novel framework is built to generate sentences for videos with more coherence and semantics. A long short term memory (LSTM) network with an improved factored way is first developed, which takes inspiration of the conventional factored way and a common practice of presenting multi-modal features at the first time in LSTM for video captioning. An LSTM network with the combination of improved factored and un-factored ways is exploited, and a voting strategy is employed to predict the words. Then, the residual is used to enhance the gradient signals which is learned from residual network (ResNet), and a deeper LSTM network is constructed. Furthermore, several convolutional neural network (CNN) features from deep models with different architectures are fused to catch more comprehensive and complementary visual information. Experiments are conducted on the MSR-VTT2016 and MSR-VTT2017 grand challenge datasets to demonstrate the effectiveness of each presented techniques as well as the superiority compared to other state-of-the-art methods.
Pengjie Tang, Hanli Wang, Kaisheng Xu
ACM Multimedia1
2017 Photograph aesthetical evaluation and classification with deep convolutional neural networks
Yunlan Tan, Pengjie Tang, Yimin Zhou 0003, Wenlang Luo, Yongping Kang
Neurocomputing2
2017 G-MS2F: GoogLeNet based multi-stage feature fusion of deep CNN for scene recognition
Pengjie Tang, Hanli Wang, Sam Kwong
Neurocomputing1