Zhengpeng Zhao

dblp:00/8387 · DBLP profile ↗
← Back
33ranked-venue papers
0as first author
32since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 10 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Retrieval-based objects and relations prompt for image captioning
Jinjing Gu, Tianbao Qin, Zhengpeng Zhao
Eng. Appl. Artif. Intell.4
2026 Adverse multi-weather image restoration for boosting downstream object detection
Jinjing Gu, Chenggang Yang, Zhengpeng Zhao
Eng. Appl. Artif. Intell.4
2026 Training-free style transfer via frequency domain reorganization of noise in diffusion
Zhengpeng Zhao, Haomin Zhao, Qiuxia Yang, Chengchao Wang 0002
Eng. Appl. Artif. Intell.3
2026 StrCCL: Structure-aware Contrastive Consistency Loss for Artistic Style Transfer
Shuyu Pan, Zhengpeng Zhao, Qiuxia Yang, Jinjing Gu, Dan Xu 0001
Expert Syst. Appl.3
2026 Hybrid prompt learning and multilevel knowledge distillation for multimodal sentiment analysis with missing modalities
Yiqiao Zhai, Qiuxia Yang, Chengchao Wang 0002, Lianmin Zhou, Jue Feng, Fanghong Hu, Zhengpeng Zhao
Expert Syst. Appl.7
2026 Multimodal progressive contrastive learning for sentiment analysis
Lianmin Zhou, Zhengpeng Zhao, Jue Feng, Dan Xu 0001, Jinjing Gu
Neurocomputing3
2026 KidMind: An open framework to develop and benchmark LLMs for empathetic companionship and knowledge reasoning in Chinese child mental health support
Gang Hu 0003, Tian Wei, Jingyao Luo, Zekang Huang, Xinghao Zhao, Shiyuan Chen, Fang Liu 0031, Min Peng 0002, Qianqian Xie, Zhengpeng Zhao
Knowl. Based Syst.11
2026 CoMPLe: Cross-Modal Hybrid Prompt Learning for End-to-End Multimodal Emotion Recognition
abstract
The quality of features directly affects the accuracy of Multimodal Emotion Recognition (MER). A key challenge in this context is the effective extraction of dynamically interactive multimodal features to enrich conversational emotion representations. However, existing approaches are often constrained by non-end-to-end architectures, overlooking the significance of feature extraction in MER. To address the problem of dynamic interaction in emotion feature extraction, this paper introduces an end-to-end network based on Cross-Modal Hybrid Prompt Learning (CoMPLe). The model takes raw video as input and leverages three prompt mechanisms to guide large-scale pre-trained encoders in extracting emotionally salient features with latent correlations. Specifically, we design a cross-modal soft prompt learning strategy to mine complementary information across modalities and dynamically adjust the cross-modal semantic space. To capture stage-dependent characteristics, deep feature prompts are incorporated to progressively learn intra-modal contextual representations. Furthermore, a label prompt mechanism is proposed to construct hard prompt templates from emotion labels. Finally, the highest cosine similarity is computed between each unimodal feature and the label prompt templates to activate factual knowledge relevant to emotion recognition. Experiments on three public datasets show that the end-to-end network proposed in this paper surpasses the existing State-Of-The-Art baselines.
Jue Feng, Zhengpeng Zhao, Lianmin Zhou, Jiale Ye, Dan Xu 0001, Jinjing Gu
IEEE Trans. Affect. Comput.2
2025 Pushing the Limits of BFP on Narrow Precision LLM Inference
abstract
The substantial computational and memory demands of Large Language Models (LLMs) hinder their deployment. Block Floating Point (BFP) has proven effective in accelerating linear operations, a cornerstone of LLM workloads. However, as sequence lengths grow, nonlinear operations, such as Attention, increasingly become performance bottlenecks due to their quadratic computational complexity. These nonlinear operations are predominantly executed using inefficient floating-point formats, which renders the system challenging to optimize software efficiency and hardware overhead. In this paper, we delve into the limitations and potential of applying BFP to nonlinear operations. Given our findings, we introduce a hardware-software co-design framework (DB-Attn), including: (i) DBFP, an advanced BFP version, overcomes nonlinear operation challenges with a pivot-focus strategy for diverse data and an adaptive grouping strategy for flexible exponent sharing. (ii) DH-LUT, a novel lookup table algorithm dedicated to accelerating nonlinear operations with DBFP format. (iii) An RTL-level DBFP-based engine is implemented to support DB-Attn, applicable to FPGA and ASIC. Results show that DB-Attn provides significant performance improvements with negligible accuracy loss, achieving 74% GPU speedup on Softmax of LLaMA and 10x low-overhead performance improvement over SOTA designs.
Hui Wang 0166, Xiaomeng Han, Zhengpeng Zhao, Zhe Jiang 0004
AAAI4
2025 NVR: Vector Runahead on NPUs for Sparse Memory Access
abstract
Deep Neural Networks are increasingly leveraging sparsity to reduce the scaling up of model parameter size. However, reducing wall-clock time through sparsity and pruning remains challenging due to irregular memory access patterns, leading to frequent cache misses. In this paper, we present NPU Vector Runahead (NVR), a prefetching mechanism tailored for NPUs to address cache miss problems in sparse DNN workloads. Rather than optimising memory patterns with high overhead and poor portability, NVR adapts runahead execution to the unique architecture of NPUs. NVR provides a general micro-architectural solution for sparse DNN workloads without requiring compiler or algorithmic support, operating as a decoupled, speculative, lightweight hardware sub-thread alongside the NPU, with minimal hardware overhead (under 5%). NVR achieves an average 90% reduction in cache misses compared to SOTA prefetching in general-purpose processors, delivering 4 x average speedup on sparse workloads versus NPUs without prefetching. Moreover, we investigate the advantages of incorporating a small cache (16 KB) into the NPU combined with NVR. Our evaluation shows that expanding this modest cache delivers 5x higher performance benefits than increasing the $\mathbf{L 2}$ cache size by the same amount.
Hui Wang 0166, Zhengpeng Zhao, Jing Wang 0113, Yushu Du, Chenhao Ma 0006, Xiaomeng Han, Dean You, Jiapeng Guan, Zhe Jiang 0004
DAC2
2025 Layer-wise Parameter Robustness for Continual Test-time Adaptation
abstract
Since inevitable distribution shifts are encountered during test time in practice, test-time adaptation (TTA) presents a promising solution by recalibrating the model online using only an unlabeled test data stream. However, TTA often suffers from issues such as catastrophic forgetting caused by continuously changing environments, as it relies on self-training. Contemporary solutions attempt to mitigate this by anchoring TTA to a static source model, such as stochastic parameter restoration or periodic parameter reset, which restrict model flexibility. Moreover, different layers may exhibit varying sensitivities to distribution shifts, sometimes even showing opposite shift trends, yet prior methods treat all layers homogeneously. Motivated by this, we propose a layer-wise parameter robustness method that autonomously identifies important parameters in different layers for selective weighting by measuring the sharpness of parameter surface. Further in-depth experiments on various benchmarks demonstrate the robustness and effectiveness of our proposed method. Our code is available at https://github.com/ioslide/prda_tta.
Haoyu Xiong, Qiuxia Yang, Tianze Zhong, Zhengpeng Zhao
ICME5
2025 PsyChild: A Child-Centric Psychological Companionship LLM with Fine-Grained Multiturn Dialogue Evaluation Benchmark
Tian Wei, Hongyu Hou, Yating Chen, Zhengpeng Zhao
PRICAI5
2025 Training-free style transfer via content-style image inversion
Songlin Lei, Qiuxia Yang, Zhengpeng Zhao
Comput. Graph.4
2025 AFDFusion: An adaptive frequency decoupling fusion network for multi-modality image
Chengchao Wang 0002, Zhengpeng Zhao, Qiuxia Yang, Rencan Nie, Jinde Cao
Expert Syst. Appl.2
2025 FNContra: Frequency-domain Negative Sample Mining in Contrastive Learning for limited-data image generation
Qiuxia Yang, Zhengpeng Zhao, Shuyu Pan, Jinjing Gu, Dan Xu 0001
Expert Syst. Appl.2
2025 Multimodal hypergraph network with contrastive learning for sentiment analysis
Zhengpeng Zhao, Qiuxia Yang, Jinjing Gu, Dan Xu 0001
Neurocomputing4
2025 Leveraging Enriched Skeleton Representation With Multi-Relational Metrics for Few-Shot Action Recognition
abstract
Few-shot action recognition aims to identify new action classes with limited training samples. Most existing methods overlook the low information content and diversity of skeleton features, failing to exploit useful information in rare samples during meta-training. This leads to poor feature discriminability and recognition accuracy. To address both issues, we propose a novel Enriched Skeleton Representation and Multi-relational Metrics (ESR-MM) method for skeleton-based few-shot action recognition. First, a Frobenius Norm Diversity Loss is introduced to enrich skeleton representation by maximizing the Frobenius norm of the skeleton feature matrix. This mitigates over-smoothing and boosts information content and diversity. Leveraging these enriched features, we propose a multi-relational metrics strategy exploiting cross-sample task-specific information, intra-sample temporal order, and inter-sample distance. Specifically, Support-Adaptive Attention leverages task-specific cues between samples to generate attention-enhanced features. Then, the Bidirectional Temporal Coherent Mean Hausdorff Metric integrates Temporal Coherence Measure into the Bidirectional Mean Hausdorff Metric for class separation by accounting for temporal order. Finally, Prototype-discriminative Contrastive Loss exploits distances from class prototypes to query samples. ESR-MM demonstrates superior performance on two benchmarks.
Jingyun Tian, Jinjing Gu, Zhengpeng Zhao
IEEE Trans. Multim.4
2024 Multimodal Sentiment Analysis Based on 3D Stereoscopic Attention
abstract
In the multimodal (text, audio, and visual) sentiment analysis, the current methods generally consider the bi-modal sentiment interaction, resulting in inadequate mining and fusion of relations between modalities. In this paper, we propose the concept of multimodal 3D (3-Dimensional) stereoscopic attention for the first time, which constructs the tri-modal stereoscopic attention with temporal sequences simultaneously to adequately structure the sentiment interaction. To solve the problems of stereoscopic attention construction such as the increased complexity of algorithms caused by rising dimensions, we propose a progressive construction method with 2D attention as an intermediate process. To implement sentiment relations based on stereoscopic attention to integrating modal information sufficiently, a forward propagation mechanism is proposed, which optimizes the representations of each modality with multimodal modulation. The results on two public datasets confirm the superiority of the proposed method in all metrics to the baselines.
Dongming Zhou 0001, Zhengpeng Zhao, Dan Xu 0001, Jinde Cao
ICASSP5
2024 Dual-path hypernetworks of style and text for one-shot domain adaptation
Zhengpeng Zhao, Qiuxia Yang, Jinjing Gu, Yupan Li, Dan Xu 0001
Appl. Intell.3
2024 Towards diverse image-to-image translation via adaptive normalization layer and contrast learning
Zhengpeng Zhao, Yupan Li, Rencan Nie
Comput. Graph.3
2024 FCLFusion: A frequency-aware and collaborative learning for infrared and visible image fusion
Chengchao Wang 0002, Zhengpeng Zhao, Rencan Nie, Jinde Cao, Dan Xu 0001
Eng. Appl. Artif. Intell.3
2024 PSANet: Automatic colourisation using position-spatial attention for natural images
abstract
Abstract Due to the richness of natural image semantics, natural image colourisation is a challenging problem. Existing methods often suffer from semantic confusion due to insufficient semantic understanding, resulting in unreasonable colour assignments, especially at the edges of objects. This phenomenon is referred to as colour bleeding. The authors have found that using the self‐attention mechanism benefits the model's understanding and recognition of object semantics. However, this leads to another problem in colourisation, namely dull colour. With this in mind, a Position‐Spatial Attention Network(PSANet) is proposed to address the colour bleeding and the dull colour. Firstly, a novel new attention module called position‐spatial attention module (PSAM) is introduced. Through the proposed PSAM module, the model enhances the semantic understanding of images while solving the dull colour problem caused by self‐attention. Then, in order to further prevent colour bleeding on object boundaries, a gradient‐aware loss is proposed. Lastly, the colour bleeding phenomenon is further improved by the combined effect of gradient‐aware loss and edge‐aware loss. Experimental results show that this method can reduce colour bleeding largely while maintaining good perceptual quality.
Peng-Jie Zhu, Qiuxia Yang, Zhengpeng Zhao, Hao Wu 0010, Dan Xu 0001
IET Comput. Vis.5
2024 Dynamic hypergraph convolutional network for multimodal sentiment analysis
Dongming Zhou 0001, Jinde Cao, Jinjing Gu, Zhengpeng Zhao, Dan Xu 0001
Neurocomputing6
2024 Co-space Representation Interaction Network for multimodal sentiment analysis
Zhengpeng Zhao, Dongming Zhou 0001, Dan Xu 0001, Jinde Cao
Knowl. Based Syst.3
2023 BIT: Improving Image-text Sentiment Analysis via Learning Bidirectional Image-text Interaction
abstract
Exploring the interaction between image and text has a great strength for image-text sentiment analysis. However, most methods only focus on learning forward interaction in forward image-text features and fail to capture the backward interaction in backward image-text features, which leads to the loss of necessary information embedded in backward interaction. In this paper, Bidirectional Interaction Transformer (BIT) that models both forward and backward image-text interactions is proposed for image-text sentiment analysis. Specifically, we first encode image and text to forward and backward features. Then, these features are fed into Bidirectional Interaction Encoder (BIE) with Forward Interaction and Back Interaction branches to model bidirectional (i.e., forward and backward) image-text interaction. Finally, Two-scale Adaptive Gating Fusion (TAGF) is designed to adaptively fuse the forward and backward interactions learned by BIE. Extensive experiments conducted on two public datasets demonstrate the effectiveness of the proposed model.
Xingwang Xiao, Zhengpeng Zhao, Jinjing Gu, Dan Xu 0001
IJCNN3
2023 W2GAN: Importance Weight and Wavelet feature guided Image-to-Image translation under limited data
Qiuxia Yang, Zhengpeng Zhao, Dan Xu 0001
Comput. Graph.3
2023 Collaborative fine-grained interaction learning for image-text sentiment analysis
Xingwang Xiao, Dongming Zhou 0001, Jinde Cao, Jinjing Gu, Zhengpeng Zhao, Dan Xu 0001
Knowl. Based Syst.6
2023 Image-Text Sentiment Analysis Via Context Guided Adaptive Fine-Tuning Transformer
Xingwang Xiao, Zhengpeng Zhao, Rencan Nie, Dan Xu 0001, Wenhua Qian, Hao Wu 0010
Neural Process. Lett.3
2023 Unpaired Artistic Portrait Style Transfer via Asymmetric Double-Stream GAN
abstract
With the development of image style transfer technologies, portrait style transfer has attracted growing attention in this research community. In this article, we present an asymmetric double-stream generative adversarial network (ADS-GAN) to solve the problems that caused by cartoonization and other style transfer techniques when they are applied to portrait photos, such as facial deformation, contours missing, and stiff lines. By observing the characteristics between source and target images, we propose an edge contour retention (ECR) regularized loss to constrain the local and global contours of generated portrait images to avoid the portrait deformation. In addition, a content-style feature fusion module is introduced for further learning of the target image style, which uses a style attention mechanism to integrate features and embeds style features into content features of portrait photos according to the attention weights. Finally, a guided filter is introduced in content encoder to smooth the textures and specific details of source image, thereby eliminating its negative impact on style transfer. We conducted overall unified optimization training on all components and got an ADS-GAN for unpaired artistic portrait style transfer. Qualitative comparisons and quantitative analyses demonstrate that the proposed method generates superior results than benchmark work in preserving the overall structure and contours of portrait; ablation and parameter study demonstrate the effectiveness of each component in our framework.
Fanmin Kong, Ivan Lee 0001, Rencan Nie, Zhengpeng Zhao, Dan Xu 0001, Wenhua Qian
IEEE Trans. Neural Networks Learn. Syst.5
2022 Abstract Painting Synthesis via Decremental optimization
abstract
Abstract Existing stroke‐based painting synthesis methods usually fail to achieve good results with limited strokes because these methods use semantically irrelevant metrics to calculate the similarity between the painting and photo domains. Hence, it is hard to see meaningful semantical information from the painting. This paper proposes a painting synthesis method that uses a CLIP (Contrastive‐Language‐Image‐Pretraining) model to build a semantically‐aware metric so that the cross‐domain semantic similarity is explicitly involved. To ensure the convergence of the objective function, we design a new strategy called decremental optimization. Specifically, we define painting as a set of strokes and use a neural renderer to obtain a rasterized painting by optimizing the stroke control parameters through a CLIP‐based loss. The optimization process is initialized with an excessive number of brush strokes, and the number of strokes is then gradually reduced to generate paintings of varying levels of abstraction. Experiments show that our method can obtain vivid paintings, and the results are better than the comparison stroke‐based painting synthesis methods when the number of strokes is limited.
Zhengpeng Zhao, Dan Xu 0001, Qiuxia Yang, Ruxin Wang 0002
Comput. Graph. Forum3
2021 Multi-modal image synthesis combining content-style adaptive normalization and attentive normalization
Rencan Nie, Dan Xu 0001, Zhengpeng Zhao, Wenhua Qian
Comput. Graph.5
2021 Virtual Try-on Network With Attribute Transformation and Local Rendering
abstract
A virtual try-on network has gradually become a popular topic in recent years. It aims to transfer images of in-shop clothes onto the image of a target person. Owing to the diversity of clothing attributes, developing an image-based virtual try-on network is a complicated task for computers to perform and requires significant effort. Existing methods are unsatisfactory as they cannot preserve the characteristics of the clothes or the target person's identity well, thereby affecting the perception of the generated images; therefore, further research is required. To address this problem, we propose a novel try-on method that combines attribute transformation and local rendering. First, we employ pixel-level semantic segmentation to identify the try-on area and provide implementation conditions for local rendering. Second, we construct a learnable attribute transformation module to complete the try-on task for different attributes. Third, we use a learnable clothing warping module to fit the pose and figure of the target person well and establish a novel loss function, called modified style loss (M-SL), to handle clothes with rich details. Finally, we adopt a local rendering strategy, using which only renders the clothing area to ensure that the details of the non-target area are not lost. Extensive experiments are performed to test our method. The results demonstrate that our method outperforms other state-of-the-art methods.
Jun Xu 0028, Rencan Nie, Dan Xu 0001, Zhengpeng Zhao, Wenhua Qian
IEEE Trans. Multim.5
2019 Multi-Feature Fusion for Multimodal Attentive Sentiment Analysis
abstract
Sentiment analysis has been an interesting and challenging task, researchers mostly pay attention to single-modal (image or text) emotion recognition, less attention is paid to joint analysis of multi-modal data. Most existing multi-modal sentiment analysis algorithms combined with attention mechanism focus only on local area of images, ignore the emotional information provided by the global features of the image. Motivated by the research status quo, in this paper, we proposed a novel multi-modal sentiment analysis model, which focuses on local attentive feature also on the global contextual feature from image, then a novel feature fusion mechanism is utilized to fuse features from different modal. In our proposed model, we use a convolutional neural network (CNN) to extract the region maps of images, and use the attention mechanism to acquire attention coefficient, then use a CNN with fewer hidden layers to extract the global feature, a long-short term memory model (LSTM) is utilized to extract textual feature. Finally, a tensor fusion network (TFN) is utilized to fuse all features from different modal. Extensive experiments are conducted on both weakly labeled and manually labeled datasets, and the results demonstrate the superiority of the proposed method.
Man A, Dan Xu 0001, Wenhua Qian, Zhengpeng Zhao, Qiuxia Yang
MMAsia5