Pengjie Wang 0001

dblp:71/7813-1 · also Peng Jie Wang 0001, Peng-Jie Wang 0001 · DBLP profile ↗
← Back
25ranked-venue papers
7as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 4 first-author · 12 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 SigFusion: Unified Signal-Level Self-Supervised Learning Paradigm for Image Fusion
abstract
Image Fusion (IF) aims to integrate complementary features from multiple source images into a single image. However, a key challenge in this field is the lack of large-scale real-world training datasets. Existing models typically rely on either small datasets or synthetic, less realistic datasets. To address this, we propose SigFusion, a unified signal-level self-supervised learning paradigm for various IF tasks.The core idea is to use signal-level Pseudo-Label Generation Networks (PLGN) to automatically synthesize training sets and pseudo labels with real multi-source signal characteristics from vast unlabeled natural images.PLGN includes two critical components: learnable 1D Signal Modulators (SM) and SigFormer. SM learns implicit 1D signal patterns across various source images and embeds them into natural images, reducing the domain gap between synthetic and real datasets. SigFormer integrates Transformer with signal processing methods, establishing an appropriate signal representation space for SM. Its cascaded, multi-level design allows hierarchical feature learning from coarse to fine detail. Moreover, SigFormer can serve as a flexible backbone for IF, as its design adheres to the classic decomposition-reconstruction paradigm. Experimental results demonstrate that SigFusion achieves state-of-the-art performance across multiple IF tasks, including medical image fusion, infrared-visible image fusion, multi-focus image fusion, and multi-exposure image fusion.
Zeyu Wang 0009, Pengjie Wang 0001, Haiyu Song 0002
AAAI4
2026 Infrared and visible image fusion via iterative feature decomposition and deep balanced fusion
Wei Li 0150, Baojia Li 0003, Haiyu Song 0002, Pengjie Wang 0001, Zeyu Wang 0009
Pattern Recognit.4
2026 GeoEdgeFormer: 3D Point Cloud Saliency Detection via Edge-Enhanced Graph-Transformer Network
abstract
The goal of point cloud salient object detection is to identify and segment the most prominent areas or objects within a 3D point cloud. Research on point cloud SOD is still in its early stages, and many existing methods fail to fully utilize the rich geometric information inherent in point clouds. To address this limitation, we introduce GeoEdgeFormer, an effective Edge-Enhanced Transformer Network, designed specif ically for 3D salient object detection. GeoEdgeFormer employs an encoder-decoder architecture featuring two novel components: the Residual Edge Convolution (REC) and the Global Contextual Transformer (GCT). In the encoder, we propose the REC, which is designed to maintain permutation invariance while capturing local geometric information. This component not only improves the model's ability to process complex point cloud data but also enhances its efficiency, making it suitable for dynamic. In the decoder, we introduce the GCT to learn scene-level contextual representations. The GCT integrates global semantics and multi-level features from the encoder into a cohesive global scene context. By effectively combining features from local and global levels, the model achieves a more comprehensive understanding of the scene's semantics, thereby enhancing its generalization ability. Extensive experiments on the PCSOD saliency dataset demonstrate that our proposed GeoEdgeFormer achieves state of-the-art performance.
Zihao Tian, Pengjie Wang 0001, Xuan Qi, Shengfeng He
IEEE Trans. Multim.2
2025 Unsupervised Salient Object Detection with Pseudo-Labels Refinement
Yanfeng Zheng, Pengjie Wang 0001, Xiaosong Yang
CASA2
2025 FaTNET: Feature-alignment transformer network for human pose transfer
Chengzhi Yuan, Lin Gao 0004, Weiwei Xu 0003, Xiaosong Yang, Pengjie Wang 0001
Pattern Recognit.6
2025 Unsupervised Salient Object Detection on Light Field With High-Quality Synthetic Labels
abstract
Most current Light Field Salient Object Detection (LFSOD) methods require full supervision with labor-intensive pixel-level annotations. Unsupervised Light Field Salient Object Detection (ULFSOD) has gained attention due to this limitation. However, existing methods use traditional handcrafted techniques to generate noisy pseudo-labels, which degrades the performance of models trained on them. To mitigate this issue, we present a novel learning-based approach to synthesize labels for ULFSOD. We introduce a prominent focal stack identification module that utilizes light field information (focal stack, depth map, and RGB color image) to generate high-quality pixel-level pseudo-labels, aiding network training. Additionally, we propose a novel model architecture for LFSOD, combining a multi-scale spatial attention module for focal stack information with a cross fusion module for RGB and focal stack integration. Through extensive experiments, we demonstrate that our pseudo-label generation method significantly outperforms existing methods in label quality. Our proposed model, trained with our labels, shows significant improvement on ULFSOD, achieving new state-of-the-art scores across public benchmarks.
Yanfeng Zheng, Zhong Luo, Ying Cao 0001, Xiaosong Yang, Weiwei Xu 0003, Zheng Lin 0005, Pengjie Wang 0001
IEEE Trans. Circuits Syst. Video Technol.8
2025 Taming High-Resolution Auxiliary G-Buffers for Deep Supersampling of Rendered Content
abstract
High-resolution images come with rich color information and texture details. Due to the rapid upgrading of display devices and rendering technologies, high-resolution real-time rendering faces the computational overhead challenge. To address this, the current mainstream solution is to render at a lower resolution and then upsample to the target resolution by supersampling techniques. However, while many prior supersampling approaches have attempted to exploit rich rendered data such as color, depth, motion vectors at low resolution, there is little discussion on how to harness high-frequency information that is readily available in the high-resolution (HR) G-buffers of modern renders. In this article, we seek to investigate how to fully leverage information from HR G-buffers to maximize the visual quality of supersampling results. We propose a neural network for real-time supersampling of rendered content, which is based on several core designs, including gated G-buffers encoder, G-buffers attended encoder and reflection-aware loss. These designs are especially made for the sake of effectively using HR G-buffers, enabling faithful recovery of a variety of high-frequency scene details from low-resolution, highly aliased inputs. Furthermore, a simple occlusion-aware blender is proposed to efficiently rectify dis-occluded features in the warped previous frame, allowing us to better exploit history information to improve temporal stability. The experiments show that our method, equipped with strong ability to harness HR G-buffer information, significantly improves the visual fidelity of high-resolution reconstructions upon previous state-of-the-art methods, even for challenging $4 \times 4$4×4 upsampling, while still being compute-efficient.
Pengjie Wang 0001, Chengzhi Yuan, Jie Guo 0001, Xiaosong Yang, Houjie Li, Ian Stephenson, Jian Chang 0001, Ying Cao 0001
IEEE Trans. Vis. Comput. Graph.1
2024 Personalized facial makeup transfer based on outline correspondence
abstract
Abstract Most existing makeup transfer techniques focus on light makeup styles, which limits the task of makeup transfer to color manipulation issues such as eye shadow and lip gloss. However, the makeup in real life is diverse and personalized, not only the most basic foundation, eye makeup, but also the painted patterns on the face, jewelry decoration and other personalized makeup. Inspired by the painting steps of drawing the outline first and then coloring, we propose a makeup transfer network for personalized makeup, which realizes face makeup transfer by learning outline correspondence. Specifically, we propose the outline feature extraction module and outline loss that can promote outline correspondence. Our network can not only transfer daily light makeup, but also complete transfer for complex facial painting patterns. Experiments show that our method can obtain visually more accurate makeup transfer results. Quantitative and qualitative experimental results show that the method proposed in this paper achieves superior results in extreme makeup transfer compared to the state‐of‐the‐art methods.
Mengying Gao, Pengjie Wang 0001
Comput. Animat. Virtual Worlds2
2024 FrseGAN: Free-style editable facial makeup transfer based on GAN combined with transformer
abstract
Abstract Makeup in real life varies widely and is personalized, presenting a key challenge in makeup transfer. Most previous makeup transfer techniques divide the face into distinct regions for color transfer, frequently neglecting details like eyeshadow and facial contours. Given the successful advancements of Transformers in various visual tasks, we believe that this technology holds large potential in addressing pose, expression, and occlusion differences. To explore this, we propose novel pipeline which combines well‐designed Convolutional Neural Network with Transformer to leverage the advantages of both networks for high‐quality facial makeup transfer. This enables hierarchical extraction of both local and global facial features, facilitating the encoding of facial attributes into pyramid feature maps. Furthermore, a Low‐Frequency Information Fusion Module is proposed to address the problem of large pose and expression variations which exist between the source and reference faces by extracting makeup features from the reference and adapting them to the source. Experiments demonstrate that our method produces makeup faces that are visually more detailed and realistic, yielding superior results.
Pengjie Wang 0001, Xiaosong Yang
Comput. Animat. Virtual Worlds2
2024 Mask-DerainGAN: Learning to remove rain streaks by learning to generate rainy images
Pengjie Wang 0001, Rynson W. H. Lau
Pattern Recognit.1
2024 Few-shot anime pose transfer
abstract
Abstract In this paper, we propose a few-shot method for pose transfer of anime characters—given a source image of an anime character and a target pose, we transfer the pose of the target to the source character. Despite recent advances in pose transfer on real people images, these methods typically require large numbers of training images of different person under different poses to achieve reasonable results. However, anime character images are expensive to obtain they are created with a lot of artistic authoring. To address this, we propose a meta-learning framework for few-shot pose transfer, which can well generalize to an unseen character given just a few examples of the character. Further, we propose fusion residual blocks to align the features of the source and target so that the appearance of the source character can be well transferred to the target pose. Experiments show that our method outperforms leading pose transfer methods, especially when the source characters are not in the training set.
Pengjie Wang 0001, Chengzhi Yuan, Houjie Li, Wen Tang 0004, Xiaosong Yang
Vis. Comput.1
2023 SeTGAN: Semantic-text guided face image generation
abstract
Abstract In this article, we propose a method to jointly control face image generation through semantic segmentation maps and text. Existing semantic segmentation maps lack detailed face attributes such as beards, and it is difficult to explicitly represent the gender of the target person by virtue of the semantic maps. State‐of‐the‐art face image generation methods guided by semantic segmentation maps mostly solved this by introducing the original image for supervision, which cannot accurately control the detailed attributes of the target face. At the same time, the text‐guided image generation method perform poorly in controlling the front and side of the face pose. Therefore, we propose an idea that the semantic segmentation map controls the coarse content of the target image and the text controls the fine details of the target image. Through the well‐designed mapping network and content mixing mechanism, the model in this article flexibly draws on the advantages of the two modes, and can generate high‐resolution images that are diverse, high‐quality, and more faithful to the target in detail attributes than most of previous methods. Extensive experiments demonstrate the superior performance of the proposed method in terms of accuracy and fidelity.
Pengjie Wang 0001
Comput. Animat. Virtual Worlds2
2023 Cycle-attention-derain: unsupervised rain removal with CycleGAN
Dehai Shang, Pengjie Wang 0001
Vis. Comput.4
2022 Salient object detection with image-level binary supervision
Pengjie Wang 0001, Ying Cao 0001, Xin Yang 0011, Huchuan Lu, Rynson W. H. Lau
Pattern Recognit.1
2022 Weakly-Supervised Salient Object Detection on Light Fields
abstract
Most existing salient object detection (SOD) methods are designed for RGB images and do not take advantage of the abundant information provided by light fields. Hence, they may fail to detect salient objects of complex structures and delineate their boundaries. Although some methods have explored multi-view information of light field images for saliency detection, they require tedious pixel-level manual annotations of ground truths. In this paper, we propose a novel weakly-supervised learning framework for salient object detection on light field images based on bounding box annotations. Our method has two major novelties. First, given an input light field image and a bounding-box annotation indicating the salient object, we propose a ground truth label hallucination method to generate a pixel-level pseudo saliency map, to avoid heavy cost of pixel-level annotations. This method generates high quality pseudo ground truth saliency maps to help supervise the training, by exploiting information obtained from the light field (including depths and RGB images). Second, to exploit the multi-view nature of the light field data in learning, we propose a fusion attention module to calibrate the spatial and channel-wise light field representations. It learns to focus on informative features and suppress redundant information from the multi-view inputs. Based on these two novelties, we are able to train a new salient object detector with two branches in a weakly-supervised manner. While the RGB branch focuses on modeling the color contrast in the all-in-focus image for locating the salient objects, the Focal branch exploits the depth and the background spatial redundancy of focal slices for eliminating background distractions. Extensive experiments show that our method outperforms existing weakly-supervised methods and most fully supervised methods.
Pengjie Wang 0001, Ke Xu 0010, Rynson W. H. Lau
IEEE Trans. Image Process.2
2021 Weakly-Supervised Salient Object Detection With Saliency Bounding Boxes
abstract
In this paper, we propose a novel form of weak supervision for salient object detection (SOD) based on saliency bounding boxes, which are minimum rectangular boxes enclosing the salient objects. Based on this idea, we propose a novel weakly-supervised SOD method, by predicting pixel-level pseudo ground truth saliency maps from just saliency bounding boxes. Our method first takes advantage of the unsupervised SOD methods to generate initial saliency maps and addresses the over/under prediction problems, to obtain the initial pseudo ground truth saliency maps. We then iteratively refine the initial pseudo ground truth by learning a multi-task map refinement network with saliency bounding boxes. Finally, the final pseudo saliency maps are used to supervise the training of a salient object detector. Experimental results show that our method outperforms state-of-the-art weakly-supervised methods.
Pengjie Wang 0001, Ying Cao 0001, Rynson W. H. Lau
IEEE Trans. Image Process.2
2020 RGB-D salient object detection via deep fusion of semantics and details
abstract
Abstract In this paper, we address RGB‐D salient object detection task by jointly leveraging semantics and contour details of salient objects. We propose a novel semantics‐and‐details complementary fusion network to adaptively integrate cross‐model and multilevel features. Specifically, we employ two kinds of fusion modules in our model, which are designed for fusing high‐level semantic features and integrating contour detail features of the scene components, respectively. The semantics fusion module aggregates high‐level interdependent semantic relationships by a nonlinear weighted summation of small and medium receptive fields. Meanwhile, the details module integrates multi‐level contour detail features to leverage expressive details of salient objects. We achieve new state‐of‐the‐art salient object detection results on seven RGB‐D datasets, that is, STERE, NJU2000, LFSD, NLPR, SSD, DES, and SIP2019 dataset. Experimental results demonstrate that our method outperforms eleven state‐of‐the‐art salient object detection methods.
Shimin Zhao, Pengjie Wang 0001, Ying Cao 0001, Xin Yang 0011
Comput. Animat. Virtual Worlds3
2019 An Active Learning Framework for Alpha Matting
abstract
Good trimap is essential for high-quality alpha matte. However, making high-quality trimap is hardwork, especially for complex images. In this paper, an active learning framework is proposed to make high quality trimap. There are two active learning methods which are employed: minimization of uncertainty sampling (MUS) and maximization of expected model output change (EMOC). MUS model finds the informative area in image which can decrease the uncertain sampling of alpha matte. EMOC model finds the important areas in image which can give the maximum expected output change of alpha matte. Two methods are combined to define the active map. Active map shows important areas which are informative in image. It can help users to make high quality trimap. The analysis and evaluation of benchmark datasets show that proposed method is effective.
Pengjie Wang 0001, Zhifang Pan, Yanxia Bao
Int. J. Pattern Recognit. Artif. Intell.2
2016 MSKD: multi-split KD-tree design on GPU
Xin Yang 0011, Pengjie Wang 0001, Duanqing Xu
Multim. Tools Appl.3
2015 A Suggestive Interface for Sketch-based Character Posing
abstract
We present a user-friendly suggestive interface for sketch-based character posing. Our interface provides suggestive information on the sketching canvas in succession by combining image retrieval technique with 3D character posing, while the user is drawing. The system highlights the canvas region where the user should draw on and constrains the user's sketches in a reasonable solution space. This is based on an efficient image descriptor, which is used to measure the distance between the user's sketch and 2D views of 3D poses. In order to achieve faster query response, local sensitive hashing is involved in our system. In addition, sampling-based optimization algorithm is adopted to synthesize and optimize the retrieved 3D pose to match the user's sketches the best. Experiments show that our interface can provide smooth suggestive information to improve the reality of sketching poses and shorten the time required for 3D posing.
Pei Lv, Pengjie Wang 0001, Weiwei Xu 0003, Jinxiang Chai
Comput. Graph. Forum2
2014 An Eigen-based motion retrieval method for real-time animation
Pengjie Wang 0001, Rynson W. H. Lau, Jiang Wang 0015, Haiyu Song 0002
Comput. Graph.1
2013 The alpha parallelogram predictor: A lossless compression method for motion capture data
Pengjie Wang 0001, Rynson W. H. Lau, Haiyu Song 0002
Inf. Sci.1
2012 Directional difference chain codes with quasi-lossless compression and run-length encoding
Yongkui Liu 0001, Borut Zalik, Pengjie Wang 0001, David Podgorelec
Signal Process. Image Commun.3
2011 A real-time database architecture for motion capture data
abstract
Due to the popularity of motion capture data in many applications, such as games, movies and virtual environments, huge collections of motion capture data are now available. It is becoming important to store these data in compressed form while being able to retrieve them without much overhead. However, there is little work that addresses both issues together. In this paper, we address these two issues by proposing a novel database architecture. First, we propose a lossless compression algorithm to compress the motion clips, which is based on a novel Alpha Parallelogram Predictor (APP) to estimate the degree of freedom (DOF) of each child joint from its immediate neighbors and parents that have already been processed. Second, we propose to store selected eigenvalues and eigenvectors of each motion clip, which only require a very small amount of memory overheads, for faster filtering of irrelevant motions. Based on this architecture, real-time queries become a three-step process. In the first two steps, we perform a quick filtering to identify relevant motion clips in the database through a two-level indexing structure. In the third step, only a small number of candidate clips are uncompressed and accurately matched with a Dynamic Time Warping algorithm. Our results show that users can efficiently search clips from this losslessly compressed motion database.
Pengjie Wang 0001, Rynson W. H. Lau, Jiang Wang 0015, Haiyu Song 0002
ACM Multimedia1
2007 Compressed vertex chain codes
Yongkui Liu 0001, Pengjie Wang 0001, Borut Zalik
Pattern Recognit.3