Hengyou Wang

dblp:86/8276 · also Heng-You Wang · DBLP profile ↗
← Back
33ranked-venue papers
14as first author
21since 2021 · last 2026
0000-0001-6693-0161ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 7 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Fine-grained weakly supervised video anomaly detection via gated temporal fusion and graph reasoning
Qiang He 0003, Hengyou Wang
Expert Syst. Appl.3
2026 Text image inpainting via Text-Referenced Residual Diffusion Model in latent space
Rongji Ke, Degang Chen 0002, Hengyou Wang
Knowl. Based Syst.3
2026 Deep unfolding low-rank network for image denoising
Hengyou Wang, Qiang He 0003
Multim. Syst.1
2026 A dual modal alignment method of rail image and axle box acceleration based on variable-scale pattern matching
Changlun Zhang, Guocui Zhang, Qiang He 0003, Hengyou Wang, Jinzhao Liu
Pattern Anal. Appl.5
2026 Accurate Absolute Scale With Pseudo Depth Constraint for 3-D Sparse Construction of Monocular RGB Image
abstract
3D Sparse reconstruction and camera pose estimation from monocular sequences are inherently plagued by scale ambiguity, which prevents the recovery of the scene’s absolute scale. Existing solutions are primarily limited in two aspects: (1) traditional geometric methods often depend on additional sensors or restrictive pre-calibration, lacking general applicability; (2) learning-based approaches frequently suffer from scale biases when faced with domain shifts. To overcome these challenges, this paper introduces a unified framework that deeply integrates monocular depth estimation into a traditional reconstruction pipeline. Our method leverages the robustness of modern depth prediction networks to eliminate cumbersome prior assumptions while employing iterative geometric refinement to ensure precise scale accuracy. The proposed framework begins by estimating depth maps via RGB images. First, an absolute scale factor is initialized from the depth priors to calibrate the initial two-view reconstruction, effectively aligning the model into a metric space. Second, during incremental reconstruction, new points and poses are registered via depth-constrained triangulation, with an iterative weighting strategy to enhance stability. Third, we propose a joint bundle adjustment that harmonizes reprojection and geometric errors using an adaptive Mahalanobis distance, which implicitly optimizes the global scale and refines the scene geometry concurrently. Comprehensive evaluations on the ETH3D, TUM-RGBD, and ICL-NUIM benchmarks demonstrate that our framework achieves state-of-the-art performance in absolute scale recovery. The results confirm that our method successfully maintains high geometric fidelity across diverse scenarios, proving its effectiveness and robustness.
Erjie Jiao, Hengyou Wang, Jun Wan 0001, Yihong Wu 0002
IEEE Trans. Circuits Syst. Video Technol.2
2026 SegPainter: User-Controllable Face Inpainting via Mask-Aware Semantic Segmentation-Guided Mamba
abstract
With the growing demand for face editing applications, face inpainting has become an increasingly important subfield within image inpainting research. While many existing methods use semantic segmentation guidance, they apply uniform weighting across all regions of the map. Since the missing areas of the image lack meaningful features, this uniform treatment provides insufficient guidance in missing areas, leading to unrealistic or structurally incoherent results. Moreover, these methods generally lack the capacity to adapt inpainting results to individual user preferences, thereby limiting their effectiveness in personalized face editing. To address these limitations, we propose SegPainter, a Mamba-based architecture for user-controllable face inpainting that enables customized restoration guided by user-defined semantic segmentation maps generated using the Image Segmentation Annotation Tool (ISAT) integrated with Meta's Segment Anything Model (SAM). Specifically, we propose Hard Mask Soft One-Hot Encoding (HMSOE) to adaptively weight regions in the segmentation map based on whether they correspond to known or missing areas of the masked image. This strategy amplifies semantic guidance in missing regions while attenuating it in known regions to avoid over-constraining existing content. We further introduce Semantic-Guided State Space Model (SG-SSM) to dynamically modulate the Mamba layer with semantic features, adapting guidance to the masked image. To enhance the quality of inpainting results, we also propose Tri-Scan Inspection (TSI), a scanning mechanism designed to capture both global and local dependencies while preserving spatial continuity and facial structure. Extensive experiments on the CelebAMask-HQ and FFHQ datasets demonstrate that our framework outperforms state-of-the-art methods, producing sharper and more semantically consistent face inpainting results.
Rongji Ke, Degang Chen 0002, Hengyou Wang, Debby Dan Wang
IEEE Trans. Vis. Comput. Graph.3
2025 Robust graph representation learning with asymmetric debiased contrasts
Wen Li 0015, Wing W. Y. Ng, Hengyou Wang, Jianjun Zhang 0004, Cankun Zhong
Expert Syst. Appl.3
2025 Multiscale contextual joint feature enhancement GAN for semantic image synthesis
Hengyou Wang, Rongxin Ma, Xiang Jiang 0008
Image Vis. Comput.1
2025 FDC-Net: foreground dynamic capture with deep feature enhancement for video anomaly detection
Ruinian Shi, Qiang He 0003, Hengyou Wang, Changlun Zhang
Multim. Syst.3
2025 ERMAV: Efficient and Robust Graph Contrastive Learning via Multiadversarial Views Training
abstract
Graph contrastive learning (GCL) is emerging as a pivotal technique in graph representation learning. However, recent research indicates that GCL is vulnerable to adversarial attacks, while existing robust GCL methods against adversarial attacks are inefficient and lack scalability due to the significant computational expenses of explicit adversarial attacks on the graph structure. To address the shortcomings of existing approaches, we propose an efficient and robust GCL via multiadversarial views training framework, called ERMAV. Specifically, the ERMAV generates two adversarial views by attacking both node attributes and latent representations on randomly sampled subgraphs. The method conducts explicit adversarial attacks on node attributes by attacking node attributes and implicit adversarial attacks on the graph structure by attacking latent representations, which avoids the costly computation of explicit graph structure attacks. Moreover, two efficient attack methods are developed to construct adversarial perturbations, which can dynamically generate different adversarial views to enhance sample diversity in the training phase. Furthermore, to validate the effectiveness and robustness of the proposed framework, extensive experiments of node classification on seven real-world datasets are conducted. Experimental results show that our ERMAV outperforms state-of-the-art GCL methods on the original graphs and is consistently more robust than existing robust GCL methods on a variety of attacked graphs. This demonstrates the strong robustness and great potential of our ERMAV in real-world applications.
Wen Li 0015, Wing W. Y. Ng, Hengyou Wang, Jianjun Zhang 0004, Cankun Zhong, Liang Yang 0002
IEEE Trans. Cybern.3
2024 Global and edge enhanced transformer for semantic segmentation of remote sensing
Hengyou Wang
Appl. Intell.1
2024 ragBERT: Relationship-aligned and grammar-wise BERT model for image captioning
Hengyou Wang, Kani Song, Xiang Jiang 0008, Zhiquan He
Image Vis. Comput.1
2024 Low-rank matrix recovery with total generalized variation for defending adversarial examples
abstract
Low-rank matrix decomposition with first-order total variation (TV) regularization exhibits excellent performance in exploration of image structure. Taking advantage of its excellent performance in image denoising, we apply it to improve the robustness of deep neural networks. However, although TV regularization can improve the robustness of the model, it reduces the accuracy of normal samples due to its over-smoothing. In our work, we develop a new low-rank matrix recovery model, called LRTGV, which incorporates total generalized variation (TGV) regularization into the reweighted low-rank matrix recovery model. In the proposed model, TGV is used to better reconstruct texture information without over-smoothing. The reweighted nuclear norm and L 1 -norm can enhance the global structure information. Thus, the proposed LRTGV can destroy the structure of adversarial noise while re-enhancing the global structure and local texture of the image. To solve the challenging optimal model issue, we propose an algorithm based on the alternating direction method of multipliers. Experimental results show that the proposed algorithm has a certain defense capability against black-box attacks, and outperforms state-of-the-art low-rank matrix recovery methods in image restoration.
Wen Li 0015, Hengyou Wang, Qiang He 0003, Zhiquan He, Wing W. Y. Ng
Frontiers Inf. Technol. Electron. Eng.2
2024 Key Feature Repairing Based on Self-Supervised for Remote Sensing Semantic Segmentation
abstract
As one of the fundamental issues in remote sensing, semantic segmentation has always received widespread attention. However, different from natural images, remote sensing images contain more complex category information, which poses many challenges to researchers, e.g., the lack of large-scale labeled semantic segmentation datasets on remote sensing and the accurate distinguishment of the edge areas between different classes. Recently, self-supervised methods have tried to avoid the issue of greatly dependency on labeled datasets. However, existing self-supervised methods were typically based on randomly masking and repairing images to learn features from unlabeled images. Random masks cannot drive the model focus on the salient information of the image, thus the learned features are not representative. In this paper, to improve the accuracy and generalization ability of the model in remote sensing semantic segmentation, we propose a key feature repairing network based on self-supervised learning, called KFRNet. KFRNet calculates the similarity between each image patch and its surrounding patches and sorts them to find the patches with more prominent feature information for masking and repairing, effective obtaining image context information. Besides, to improve the model’s ability to distinguish different classes of objects, we designed an image comparison branch to obtain the category features of the image by comparing positive and negative samples. The experimental results on the Potsdam and LoveDA datasets show that the proposed method can effectively improve segmentation accuracy. The OA, MIOU, and Fscore indices reached 89.73%, 83.96%, 91.15% (Potsdam) and 70.81%, 53.40%, 68.86% (LoveDA), even surpassing some supervised learning methods.
Hengyou Wang
IEEE Geosci. Remote. Sens. Lett.1
2024 DMFNet: deep matrix factorization network for image compressed sensing
Hengyou Wang, Xiang Jiang 0008
Multim. Syst.1
2023 Robust attention ranking architecture with frequency-domain transform to defend against adversarial samples
Wen Li 0015, Hengyou Wang, Qiang He 0003, Changlun Zhang
Comput. Vis. Image Underst.2
2023 Multiscale object detection based on channel and data enhancement at construction sites
Hengyou Wang, Yanfei Song, Qiang He 0003
Multim. Syst.1
2022 Nonconvex low-rank and sparse tensor representation for multi-view subspace clustering
Shuqin Wang 0001, Yongyong Chen, Yi-Gang Cen, Linna Zhang, Hengyou Wang, Viacheslav V. Voronin
Appl. Intell.5
2022 Structural smoothness low-rank matrix recovery via outlier estimation for image denoising
Hengyou Wang, Wen Li 0015, Lujin Hu, Changlun Zhang, Qiang He 0003
Multim. Syst.1
2021 Structure-Oriented Progressive Low-Rank Image Restoration for Defending Adversarial Attacks
abstract
Deep neural networks recognize objects by analyzing local image details and summarizing their information along the inference layers to derive the final decision. Because of this, they are prone to adversarial attacks. On the other hand, human eyes recognize objects based on their global structures and semantic cues, instead of local image textures. In this work, we propose to develop a structure-oriented progressive low-rank image completion method to remove unneeded texture details from the input images and shift the bias of deep neural networks towards global object structures and semantic cues. We formulate the problem into a low-rank matrix completion problem with progressively smoothed rank functions to avoid local minimums. Our experimental results demonstrate the proposed method is able to successfully remove the insignificant local image details while preserving important global object structures.
Zhiqun Zhao, Hengyou Wang, Hao Sun 0024, Wenming Cao 0001, Zhihai He
ICME2
2021 Removing Adversarial Noise via Low-Rank Completion of High-Sensitivity Points
abstract
Deep neural networks are fragile under adversarial attacks. In this work, we propose to develop a new defense method based on image restoration to remove adversarial attack noise. Using the gradient information back-propagated over the network to the input image, we identify high-sensitivity keypoints which have significant contributions to the image classification performance. We then partition the image pixels into the two groups: high-sensitivity and low-sensitivity points. For low-sensitivity pixels, we use a total variation (TV) norm-based image smoothing method to remove adversarial attack noise. For those high-sensitivity keypoints, we develop a structure-preserving low-rank image completion method. Based on matrix analysis and optimization, we derive an iterative solution for this optimization problem. Our extensive experimental results on the CIFAR-10, SVHN, and Tiny-ImageNet datasets have demonstrated that our method significantly outperforms other defense methods which are based on image de-noising or restoration, especially under powerful adversarial attacks.
Zhiqun Zhao, Hengyou Wang, Hao Sun 0024, Jianhe Yuan, Zhongchao Huang, Zhihai He
IEEE Trans. Image Process.2
2020 L1-norm low-rank linear approximation for accelerating deep neural networks
Zhiqun Zhao, Hengyou Wang, Hao Sun 0024, Zhihai He
Neurocomputing2
2020 Multi-Matrices Low-Rank Decomposition With Structural Smoothness for Image Denoising
abstract
In this paper, we propose a multi-matrices lowrank decomposition method for image denoising. In this new method, the total variation (TV) norm is incorporated into lowrank approximation analysis to achieve structural smoothness and to improve quality of the recovered images. Our proposed mathematical framework for multi-matrices low-rank decomposition combines the nuclear norm, TV norm, and L1norm, which allows us to exploit the low-rank property of natural images, enhance the structural smoothness, and detect and remove large sparse noise. Based on the iterative alternating direction method, we develop an algorithm to solve the proposed challenging optimization problem. We conduct extensive experiments and perform evaluations on multi-images denoising and multi-frames video prediction. Our experimental results demonstrate that the proposed method outperforms the state-of-the-art low-rank matrix recovery methods, particularly for images with large sparse noise.
Hengyou Wang, Yang Li 0091, Yi-Gang Cen, Zhihai He
IEEE Trans. Circuits Syst. Video Technol.1
2019 Kernel-target Alignment Based Multiple Kernel One-class Support Vector Machine
abstract
One-class support vector machine is a hot research topic in the domain of machine learning. It is currently widely used to deal with one-class classification problems or classification problems of class imbalance data, and has good performance in many practical applications. A key problem of one-class support vector machine is the selection of its kernel function and parameters, which has a vital impact on the final performance of the classifier. At present, there is no unified method for how to select the appropriate kernel function and its parameters. In order to solve this problem, the multiple kernel method is introduced into the one-class support vector machine. i.e., a combined kernel is used to replace a single kernel in the one-class support vector machine, where the combined kernel is obtained by weighted summation of several basic kernels, and the kernel weight is calculated by the kernel-target alignment. The experimental results on UCI database show that this method can effectively save training time and solve the selection of kernel parameter and its parameters based on high classification performance.
Qiang He 0003, Qingshuo Zhang, Hengyou Wang
SMC3
2019 Robust discriminant low-rank representation for subspace clustering
Gaoyun An, Yi-Gang Cen, Hengyou Wang, Ruizhen Zhao
Soft Comput.4
2019 A supervised learning to index model for approximate nearest neighbor image retrieval
Shichao Kan, Xinwei Zheng, Yi-Gang Cen, Zhenmin Zhu, Hengyou Wang
Signal Process. Image Commun.6
2018 Reweighted Low-Rank Matrix Analysis With Structural Smoothness for Image Denoising
abstract
In this paper, we develop a new low-rank matrix recovery algorithm for image denoising. We incorporate the total variation (TV) norm and the pixel range constraint into the existing reweighted low-rank matrix analysis to achieve structural smoothness and to significantly improve quality in the recovered image. Our proposed mathematical formulation of the low-rank matrix recovery problem combines the nuclear norm, TV norm, and norm, thereby allowing us to exploit the low-rank property of natural images, enhance the structural smoothness, and detect and remove large sparse noise. Using the iterative alternating direction and fast gradient projection methods, we develop an algorithm to solve the proposed challenging non-convex optimization problem. We conduct extensive performance evaluations on single-image denoising, hyper-spectral image denoising, and video background modeling from corrupted images. Our experimental results demonstrate that the proposed method outperforms the state-of-the-art low-rank matrix recovery methods, particularly for large random noise. For example, when the density of random sparse noise is 30%, for single-image denoising, our proposed method is able to improve the quality of the restored image by up to 4.21 dB over existing methods.
Hengyou Wang, Yi-Gang Cen, Zhiquan He, Zhihai He, Ruizhen Zhao, Fengzhen Zhang
IEEE Trans. Image Process.1
2017 Separable vocabulary and feature fusion for image retrieval based on sparse representation
Yi-Gang Cen, Ruizhen Zhao, Shaohai Hu, Viacheslav V. Voronin, Hengyou Wang
Neurocomputing7
2017 Fast smooth rank function approximation based on matrix tri-factorization
Hengyou Wang, Yi-Gang Cen, Ruizhen Zhao, Viacheslav V. Voronin, Fengzhen Zhang
Neurocomputing1
2017 Analytic separable dictionary learning based on oblique manifold
Fengzhen Zhang, Yi-Gang Cen, Ruizhen Zhao, Hengyou Wang, Lihong Cui, Shaohai Hu
Neurocomputing4
2017 Robust Generalized Low-Rank Decomposition of Multimatrices for Image Recovery
abstract
Low-rank approximation has been successfully used for dimensionality reduction, image noise removal, and image restoration. In existing work, input images are often reshaped to a matrix of vectors before low-rank decomposition. It has been observed that this procedure will destroy the inherent two-dimensional correlation within images. To address this issue, the generalized low-rank approximation of matrices (GLRAM) method has been recently developed, which is able to perform low-rank decomposition of multiple matrices directly without the need for vector reshaping. In this paper, we propose a new robust generalized low-rank matrices decomposition method, which further extends the existing GLRAM method by incorporating rank minimization into the decomposition process. Specifically, our method aims to minimize the sum of nuclear norms and l1-norms. We develop a new optimization method, called alternating direction matrices tri-factorization method, to solve the minimization problem. We mathematically prove the convergence of the proposed algorithm. Our extensive experimental results demonstrate that our method significantly outperforms existing GLRAM methods.
Hengyou Wang, Yi-Gang Cen, Zhihai He, Ruizhen Zhao, Fengzhen Zhang
IEEE Trans. Multim.1
2014 Rank adaptive atomic decomposition for low-rank matrix completion and its application on image recovery
Hengyou Wang, Ruizhen Zhao, Yi-Gang Cen
Neurocomputing1
2010 Performance analysis of cooperative diversity in hierarchical heterogeneous radio access networks
abstract
An outage probability analysis model is presented to evaluate performances and show key factors impacting on performances for the heterogeneous relaying scheme utilized to complete the hierarchical convergence of multiple radio access networks (RANs). The diversity gains related to the multiple access schemes and power constraints are analyzed. Our analytical and simulation results show that the outage probability mainly depends on the number of cooperative relay nodes, multiple access scheme, normalized power, and radio channel characteristics. Meanwhile, the power constraint for relay nodes does not have much impact on the outage performance especially when the broadcast/multicast radio channel condition is good enough.
Hengyou Wang, Mugen Peng, Wenbo Wang 0007, Hequan Wu
IWCMC1