Muwei Jian

dblp:24/922 · DBLP profile ↗
← Back
75ranked-venue papers
31as first author
35since 2021 · last 2026
0000-0002-4249-2264ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 37 · 18 first-author · 15 since 2021Artificial intelligence and machine learning · 20 · 5 first-author · 11 since 2021Databases, data management, data science and information retrieval · 8 · 4 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Computer networks · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 2 since 2021Security and privacy · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author
YearPublicationVenuePosition
2026 Multi-UAV collaborative maritime search via deep reinforcement learning
Hang Tao, Muwei Jian, Hanjiang Luo
Ad Hoc Networks4
2026 A novel feature reconstruction method for bone marrow cell classification
Huixiang Zhi, Muwei Jian, Changqun Nie, Hanjiang Luo
Eng. Appl. Artif. Intell.2
2026 HMamba-3DFT: A hierarchical mamba framework for emotion-driven semantic 3D facial tracking
abstract
• First Mamba-based framework HMamba-3DFT tailored for 3D facial tracking • BSTV-Mamba with BSTS-Scan capture spatiotemporal facial dynamics • Dual optimization integrates dynamic emotion-driven modeling with semantic alignment Monocular video-based 3D face tracking is vital for interactive pattern recognition and human avatars. Most existing image-based methods fail to model temporal dependencies in video, causing jitter and inaccuracies. Furthermore, they also often neglect the continuous multi-modal signals present in facial videos such as expression dynamics and emotional cues that provide essential temporal drivers for facial modeling. To this end, this study first explores the Mamba architecture tailored for 3D facial tracking by proposing a hierarchical Mamba framework, termed HMamba-3DFT. The proposed network can efficiently capture and track variations in 3D facial shapes from a monocular video. To exploit the global spatiotemporal correlations across frames of the dynamic face, we develop a bidirectional spatiotemporal vision Mamba (BSTV-Mamba) module featuring a bidirectional spatiotemporal selective scan (BSTS-Scan) mechanism. To capture temporally evolving multi-modal emotion signals embedded in continuous video sequences, we introduce a dynamic emotion-driven mechanism. Additionally, to mitigate the potential degradation of reconstruction fidelity caused by an over-reliance on emotion-driven cues, we integrate facial semantic alignment with facial emotion driving to enhance the accuracy of emotion-driven facial modeling. This integrated dual-optimization strategy systematically guides the network during training, ensuring that the reconstructed 3D facial mesh not only accurately captures the emotional attributes of the input frames but also benefits from enhanced optimization for more precise reconstruction. Extensive evaluations on benchmark datasets show competitive performance against state-of-the-art methods.
Haodong Jin, Muwei Jian, Derui Ding, Hui Yu 0001
Pattern Recognit.2
2025 CSSANet: A channel shuffle slice-aware network for pulmonary nodule detection
Muwei Jian, Huihui Huang, Rui Wang 0017, Hui Yu 0001
Neurocomputing1
2025 GLMF-NET: global and local multi-scale fusion network for polyp segmentation
Muwei Jian, Yanjie Zhong, Hui Yu 0001
Multim. Syst.1
2025 Self-Supervised Face Deocclusion via 3-D Face Reconstruction With Outlier Segmentation
abstract
Face occlusion poses a challenge for many human–machine systems, such as facial expression perception, social signal analysis, and human identity verification. Accurate face deocclusion is essential for improving the performance of identity recognition, expression recognition, and the robustness of human–machine systems. As a result, this area has attracted significant attention from researchers in recent years. However, most existing methods rely heavily on synthetic occluded face datasets and predefined occlusion masks labels, which limits their applicability in real-world scenarios. To this end, we propose a novel self-supervised generative adversarial networks (GANs)-based framework for face deocclusion in this study, which integrates 3-D facial reconstruction with outlier segmentation guidance. To achieve reliable self-supervised occlusion guidance, we introduce an outlier segmentation module that utilizes statistical priors to generate accurate occlusion masks, facilitating the deocclusion process. Furthermore, we design a GAN-based dual-branch module, which is capable of simultaneously generating the occlusion mask and the deoccluded face. Extensive experiments on the widely used datasets demonstrate the superior performance of our approach on existing methods. Our method achieves 35.71 in peak signal to noise ratio (PSNR) and 0.891 in structural similarity index measure (SSIM) for occluded face restoration, outperforming state-of-the-art techniques.
Haodong Jin, Muwei Jian, Derui Ding, Hui Yu 0001
IEEE Trans. Hum. Mach. Syst.2
2024 An Optimized Scheduling Scheme for UAV-USV Cooperative Search via Multi-Agent Reinforcement Learning Approach
abstract
The collaboration between unmanned aerial vehicles (UAV s) and unmanned surface vehicles (USV s) is critical in mar-itime search scenarios, in which UAV s can leverage high altitude and wide field of view providing real-time target information, while USV s have longer endurance capability of performing precise operations for surface targets. However, UAV s are limited by their battery capacity, which reduces their search duration and range. Furthermore, the traditional fixed charging stations necessitate UAV s to return for recharging, resulting in mission interruptions and decreased search efficiency. To address this problem, we propose an optimized scheduling scheme for UAV-USV cooperative search, in which USV s act as mobile charging stations to provide wireless charging services for UAV s to increase cooperative search mission duration with uninterrupted search execution. Firstly, the USV trajectory optimization problem under the target search constraint is formulated to minimize its energy consumption and maximize the UAV energy utilization. Then, we model the problem as a partially observable Markov decision process (POMDP) and design a scheduling algorithm leveraging the multi-agent deep deterministic policy gradient (MADDPG) method for long endurance UAV-USV collaborative search mission under energy constraints. Numerical simulations confirm the effectiveness of the proposed scheduling scheme.
Pengyan Dong, Hang Tao, Rukhsana Ruby, Muwei Jian, Hanjiang Luo
MSN5
2024 Perceptual loss guided Generative adversarial network for saliency detection
Xiaoxu Cai, Gaige Wang, Jianwen Lou, Muwei Jian, Junyu Dong, Rung Ching Chen, Brett Stevens, Hui Yu 0001
Inf. Sci.4
2024 SwinCT: feature enhancement based low-dose CT images denoising with swin transformer
Muwei Jian, Chengdong Yang
Multim. Syst.1
2024 Video saliency detection via combining temporal difference and pixel gradient
Xiangwei Lu, Muwei Jian, Rui Wang 0017, Peiguang Lin, Hui Yu 0001
Multim. Tools Appl.2
2024 YOLO-AA: an efficient object detection model via strengthening fusion context information
Muwei Jian, Gaige Wang
Multim. Tools Appl.2
2024 Flow-Edge-Net: Video Saliency Detection Based on Optical Flow and Edge-Weighted Balance Loss
abstract
Optical flow networks have been widely utilized for video saliency detection (VSD) due to their effective performance in capturing the motion of objects. However, the use of optical flow blurs the edges of salient objects and leads to the problems of poorly defined object boundaries. To address this issue, we propose an optical flow-based edge-weighted loss function, to train a network called Flow-Edge-Net, which can balance the weights of the foreground and background information at the edges of video frames. It has achieved superior performance in detecting salient boundaries. Specifically, we propose two complementary encoding and decoding networks based on the concept of decoupling. That is, the optical flow network focuses on moving objects, while the edge network, based on the encoder-decoder structure, focuses on edge information. As the two networks output features of the same dimension and are from the same input, our proposed self-designed adaptive weighted feature fusion module can compare and integrate the edge information and location information from the two networks through adaptive weighting. The proposed method has been evaluated on five widely used databases. Experiment results demonstrate the superior performance of the proposed Flow-Edge-Net in locating salient objects, with accurate and refined edges. The proposed method achieves superior performance over the state-of-the-art methods in detecting salient objects in videos.
Muwei Jian, Xiangwei Lu, Yakun Ju, Hui Yu 0001, Kin-Man Lam 0001
IEEE Trans. Comput. Soc. Syst.1
2024 Unsupervised Video Summarization Based on the Diffusion Model of Feature Fusion
abstract
Video summarization (VS) technologies can automatically extract key frames with effective information and thus can help to quickly identify the events or speed up the decision-making process, especially for accidents. With the fast development of deep learning technologies, many generative adversarial network (GAN)- and reinforcement learning (RL)-based unsupervised VS methods have been developed in recent years. However, these methods could suffer from the problems of unstable training and difficulty of reward function formulation, respectively. To this end, we present an unsupervised VS method called diffusion model of feature fusion (DMFF) in this article, which consists of a diffusion module (DM), a feature extraction and compression module (FECM), and a coarse-fine frame selector (CFFS). DM is designed to avoid the training instability problem caused by GAN’s alternate training generator and discriminator. FECM is used to extract and compress video features. CFFS is designed to capture both low-level and high-level features between frames to handle complex and diverse accident videos. Then, high-level local and global features are fused to generate a multigrained final frame score. Experiments on two widely used benchmark datasets, SumMe and TVSum, demonstrate the effectiveness and superiority of the proposed network to the state-of-the-art methods, and the training is more stable.
Qinghao Yu, Hui Yu 0001, Ying Sun 0004, Derui Ding, Muwei Jian
IEEE Trans. Comput. Soc. Syst.5
2024 UniFRD: A Unified Method for Facial Image Restoration Based on Diffusion Probabilistic Model
abstract
This paper presents a Unified Facial image and video Restoration method based on the Diffusion probabilistic model (UniFRD), designed to effectively address both single- and multi-type image degradation. The noise predictor in UniFRD consists of a ViT-based encoder and a novel Separation Fusion Decoding Module (SFDM). The flexible feature optimization strategy allows for decoding complex conditional noise without being limited by degradation patterns. Specifically, SFDM adjusts and refines the channel correlation and expressive power of high-dimensional features step by step, enabling the network to more accurately perceive and enhance the interaction between posterior probabilities and conditional inputs. This process is crucial for improving the visual quality and stability of the restoration results. Extensive experiments demonstrate that even when facial images suffer from both pixel-level and image-level degradation, UniFRD can still guarantee the restoration of rich details and maintain attribute consistency. In summary, compared to existing methods, the solution proposed in this study for facial restoration tasks offers greater generality and adaptability. Moreover, it has high practical value for applications involving faces in complex and unconstrained outdoor scenarios.
Muwei Jian, Rui Wang 0199, Feng Xu 0005, Hui Yu 0001, Kin-Man Lam 0001
IEEE Trans. Circuits Syst. Video Technol.1
2024 Estimating High-Resolution Surface Normals via Low-Resolution Photometric Stereo Images
abstract
Acquiring high-resolution 3D surface structures is a crucial task in computer vision as it provides more detailed surface textures and clearer structures. Photometric stereo can measure per-pixel surface normals of a 3D object using various shading cues. However, obtaining high-resolution images in a linear response photometric stereo imaging system can be challenging. Additionally, photometric stereo, as a per-pixel reconstruction method, requires higher-resolution surface normal maps to accurately depict complex surface structures, particularly in regions that demand more attention and precise reconstruction. Therefore, measuring high-resolution surface normals via low-resolution photometric stereo images is of great importance. Motivated by these, we propose a Super-resolution Photometric Stereo Network, namely SR-PSN. In order to address the issues of measuring the high-resolution surface normals from low-resolution photometric images, we mainly (1) apply a dual-position threshold normalization pre-processing scheme to effectively handle the spatially-varying reflectance of non-Lambertian surfaces, (2) adopt a local affinity feature module to learn the rich structural representation by explicitly revealing the neighbor relationships, (3) employ a parallel multi-scale feature extractor, which preserves high-resolution representations and deep feature extraction, and (4) propose a shared-weight regressor to handle the multi-scale features, to prevent the model collapsing into learning non-important features related to a certain fixed scale. Extensive ablation experiments validate the effectiveness of our proposed modules. Furthermore, quantitative experiments conducted on public benchmarks demonstrate that SR-PSN outperforms state-of-the-art calibrated photometric stereo methods. Notably, SR-PSN achieves superior results while utilizing photometric stereo images with only half the resolution of other methods. It effectively restores the structure of complex surfaces, producing a high-resolution normal map.
Yakun Ju, Muwei Jian, Cong Wang 0018, Junyu Dong, Kin-Man Lam 0001
IEEE Trans. Circuits Syst. Video Technol.2
2023 Periodic-Aware Network for Fine-Grained Action Recognition
Senzi Luo, Jiayin Xiao, Dong Li 0028, Muwei Jian
PRCV (8)4
2023 Inter-slice Correlation Weighted Fusion for Universal Lesion Detection
abstract
Universal lesion detection using computerised tomography (CT) scans is a critical computer-aided diagnosis measure in clinical diagnosis. One of the key issues during the diagnosis is to identify the correlations between sequential slices to improve the feature representation of CT scans. In the process of fusing slice features containing temporal correlations, the correlation between the contextual slices in the channel dimension and the target slices is closely related to the spatial distance in practice. However, convolutional fusion approaches commonly ignore that features of different distances have unequal weights. To tackle this issue, we present a temporal correlation weighted fusion lesion detection network, called TCW-Net. Specifically, for the slices in the channel dimension, we develop a weighted feature fusion module to adjust the more discriminative features using learned weights. Then, we adapt a spatial offset attention mechanism that allows the detection network to pay more attention to the lesion’s slight spatial offset and thus improve the model’s capacity for distinguishing between different lesion features. Extensive experiments carried out on the DeepLesion dataset show that the proposed algorithm has superior performance over the state-of-the-art methods.
Muwei Jian, Rui Wang 0017, Hui Yu 0001
TrustCom1
2023 Unsupervised medical image feature learning by using de-melting reduction auto-encoder
Jinyu Cong, Kuixing Zhang, Muwei Jian, Benzheng Wei
Neurocomputing4
2023 Robust seed selection of foreground and background priors based on directional blocks for saliency-detection system
Muwei Jian, Ruihong Wang, Hui Yu 0001, Junyu Dong, Gongfa Li, Yilong Yin, Kin-Man Lam 0001
Multim. Tools Appl.1
2023 A Comprehensive Survey on Video Saliency Detection With Auditory Information: The Audio-Visual Consistency Perceptual is the Key!
abstract
Video saliency detection (VSD) aims at fast locating the most attractive objects/things/patterns in a given video clip. Existing VSD-related works have mainly relied on the visual system but paid less attention to the audio aspect. In contrast, our audio system is the most vital complementary part of our visual system. Also, audio-visual saliency detection (AVSD), one of the most representative research topics for mimicking human perceptual mechanisms, is currently in its infancy, and none of the existing survey papers have touched on it, especially from the perspective of saliency detection. Thus, the ultimate goal of this paper is to provide an extensive review to bridge the gap between audio-visual fusion and saliency detection. In addition, as another highlight of this review, we have provided a deep insight into key factors that could directly determine AVSD deep models’ performances. We claim that the audio-visual consistency degree (AVC) — a long-overlooked issue, can directly influence the effectiveness of using audio to benefit its visual counterpart when performing saliency detection. Moreover, to make the AVC issue more practical and valuable for future followers, we have newly equipped almost all existing publicly available AVSD datasets with additional frame-wise AVC labels. Based on these upgraded datasets, we have conducted extensive quantitative evaluations to ground our claim on the importance of AVC in the AVSD task. In a word, our ideas and new sets serve as a convenient platform with preliminaries and guidelines, all of which can potentially facilitate future works in further promoting state-of-the-art (SOTA) performance.
Chenglizhao Chen, Mengke Song, Wenfeng Song, Li Guo 0016, Muwei Jian
IEEE Trans. Circuits Syst. Video Technol.5
2022 Face Super-resolution Based on Multi-source References
abstract
This paper proposes a multi-source references (MSR) based face super-resolution (FSR) model. More specifically, to enhance the low-quality large-scale reconstruction of faces without the involvement of face prior knowledge, we propose a multi-source references based FSR framework exploiting a constructed reference library of nonidentity faces and an information mining module for external and internal references. Experimental results show that the proposed model can provide more satisfactory and reliable face super-resolution results than the-state-of-the-art methods.
Rui Wang 0017, Muwei Jian, Paul Smith 0002, Hui Yu 0001
HSI2
2022 Learning conditional photometric stereo with high-resolution features
abstract
Photometric stereo aims to reconstruct 3D geometry by recovering the dense surface orientation of a 3D object from multiple images under differing illumination. Traditional methods normally adopt simplified reflectance models to make the surface orientation computable. However, the real reflectances of surfaces greatly limit applicability of such methods to real-world objects. While deep neural networks have been employed to handle non-Lambertian surfaces, these methods are subject to blurring and errors, especially in high-frequency regions (such as crinkles and edges), caused by spectral bias: neural networks favor low-frequency representations so exhibit a bias towards smooth functions. In this paper, therefore, we propose a self-learning conditional network with multi-scale features for photometric stereo, avoiding blurred reconstruction in such regions. Our explorations include: (i) a multi-scale feature fusion architecture, which keeps high-resolution representations and deep feature extraction, simultaneously, and (ii) an improved gradient-motivated conditionally parameterized convolution (GM-CondConv) in our photometric stereo network, with different combinations of convolution kernels for varying surfaces. Extensive experiments on public benchmark datasets show that our calibrated photometric stereo method outperforms the state-of-the-art.
Yakun Ju, Yuxin Peng 0001, Muwei Jian, Feng Gao 0005, Junyu Dong
Comput. Vis. Media3
2022 NormAttention-PSN: A High-frequency Region Enhanced Photometric Stereo Network with Normalized Attention
Yakun Ju, Boxin Shi, Muwei Jian, Lin Qi 0004, Junyu Dong, Kin-Man Lam 0001
Int. J. Comput. Vis.3
2022 A novel underwater image restoration method based on decomposition network and physical imaging model
abstract
Underwater image restoration is one of the significant research in marine engineering and aquatic robotics. However, due to the propagation characteristics of light and the serious turbidity in underwater, the captured images often have chromatic aberration and scattering blur, which brings great challenges to the restoration of the raw image. In this paper, a revised underwater imaging model is proposed first, which reanalyzes the generation of background light from the atmosphere to the underwater and provides important support for underwater color correction. And then a network framework via the revised model is designed, which can decompose the captured image into different components corresponding to the revised model. The proposed network consists of a decomposition architecture with residual blocks that learns a complete separation of clear image and transmittance features. These two features are used along with the raw image to predict the background light. Finally, combining three constraints of the imaging model, the proposed framework can converge rapidly along the desired direction. By comparison with the performance of the state-of-the-art algorithms, the designed network shows excellent visibility and is capable of removing water on both synthetic and real-world images in different water types.
Yanfang Cui, Yujuan Sun, Muwei Jian, Xiaofeng Zhang 0003, Xin Gao 0010, Yiru Li, Yan Zhang 0175
Int. J. Intell. Syst.3
2022 Face hallucination using multisource references and cross-scale dual residual fusion mechanism
abstract
There is an increasing interest in enhancing the quality of low-resolution (LR) facial images for various social life applications. Existing methods often use domain-specific prior knowledge, which is effective in improving the face super-resolution model's performance. However, it is challenging to obtain rich and accurate prior information from LR inputs in real-world scenarios, which can limit the robustness and generalization ability of the developed face super-resolution model. In this paper, a multisource reference-based face super-resolution Network, namely MSRNet, is proposed. Without considering the prior knowledge of faces, the network can reconstruct a LR face image with a magnitude factor of 8 under the guidance of multiple reference face images of different identities. By constructing an “appearance-alike” reference data set Face_Ref, the designed MSRNet aims to fully exploit the local and spatially similar high frequency information between the distinct references and the current face. More specifically, to effectively combine the information from multiple references, a cross-scale and cross-space feature fusion mechanism is introduced for external and internal references, and then the enhanced local semantics are finally incorporated into the high-resolution face reconstruction. The robustness of face image super-resolution is increased compared to current correlation approaches, since it not only eliminates the need for face prior knowledge but also avoids performing alignment operations on reference faces with multiple expressions and different poses. Experimental results show that the proposed model is able to produce results for face super-resolution that are satisfying and dependable and outperforms the state-of-the-art methods in terms of visual perceptual quality and quantity evaluation.
Rui Wang 0017, Muwei Jian, Hui Yu 0001, Lin Wang 0004, Bo Yang 0001
Int. J. Intell. Syst.2
2022 A deep-shallow and global-local multi-feature fusion network for photometric stereo
Yanru Liu, Yakun Ju, Muwei Jian, Feng Gao 0005, Yuan Rao 0001, Yeqi Hu, Junyu Dong
Image Vis. Comput.3
2022 ConvUNeXt: An efficient convolution neural network for medical image segmentation
Zhimeng Han, Muwei Jian, Gaige Wang
Knowl. Based Syst.2
2022 Visual saliency detection via combining center prior and U-Net
Xiangwei Lu, Muwei Jian, Xing Wang 0002, Hui Yu 0001, Junyu Dong, Kin-Man Lam 0001
Multim. Syst.2
2021 Interactive Attention Sampling Network for Clinical Skin Disease Image Classification
Xulin Chen, Dong Li 0028, Yun Zhang 0001, Muwei Jian
PRCV (3)4
2021 Visual saliency detection by integrating spatial position prior of object with background cues
Muwei Jian, Hui Yu 0001, Guodong Wang 0001, Xianjing Meng, Lu Yang 0005, Junyu Dong, Yilong Yin
Expert Syst. Appl.1
2021 Global context-aware multi-scale features aggregative network for salient object detection
Inam Ullah 0002, Muwei Jian, Sumaira Hussain, Li Lian, Zafar Ali, Imran Qureshi, Jie Guo 0012, Yilong Yin
Neurocomputing2
2021 Integrating object proposal with attention networks for video saliency detection
Muwei Jian, Jiaojin Wang, Hui Yu 0001, Gaige Wang
Inf. Sci.1
2021 DSFMA: deeply supervised fully convolutional neural networks based on multi-level aggregation for saliency detection
Inam Ullah 0002, Muwei Jian, Sumaira Hussain, Jie Guo 0012, Li Lian, Hui Yu 0001, Kashif Shaheed, Yilong Yin
Multim. Tools Appl.2
2021 Underwater image processing and analysis: A review
Muwei Jian, Hanjiang Luo, Xiangwei Lu, Hui Yu 0001, Junyu Dong
Signal Process. Image Commun.1
2021 Deep Multiscale Fusion Hashing for Cross-Modal Retrieval
abstract
Owing to the rapid development of deep learning and the high efficiency of hashing, hashing methods based on deep learning models have been extensively adopted in the area of cross-modal retrieval. In general, in existing deep model-based methods, modality-specific features play an important role during the hash learning. However, most existing methods only use the modality-specific features from the final fully connected layer, ignoring the semantic relevance among modality-specific features with different scales in multiple layers. To address this issue, in this study, we put forward an end-to-end deep hashing method called deep multiscale fusion hashing (DMFH) for cross-modal retrieval. For the proposed DMFH, we first design different network branches for two modalities and then adopt multiscale fusion models for each branch network to fuse the multiscale semantics, which can be used to explore the semantic relevance. Furthermore, the multi-fusion models also embed the multiscale semantics into the final hash codes, making the final hash codes more representative. In addition, the proposed DMFH can learn common hash codes directly without a relaxation, thereby avoiding a loss in accuracy during hash learning. Experimental results on three benchmark datasets prove the relative superiority of the proposed method.
Xiushan Nie, Bowei Wang, Fanchang Hao, Muwei Jian, Yilong Yin
IEEE Trans. Circuits Syst. Video Technol.5
2020 Learning Photometric Stereo via Manifold-based Mapping
abstract
Three-dimensional reconstruction technologies are fundamental problems in computer vision. Photometric stereo recovers the surface normals of a 3D object from varying shading cues, prevailing in its capability for generating fine surface normal. In recent years, deep learning-based photometric stereo methods are capable of improving the surface-normal estimation under general non-Lambertian surfaces, due to its powerful fitting ability on the non-Lambertian surface. These state-of-the-art methods however usually regress the surface normal directly from the high-dimensional features, without exploring the embedded structural information. This results in the underutilization of the information available in the features. Therefore, in this paper, we propose an efficient manifold-based framework for learning-based photometric stereo, which can better map combined high-dimensional feature spaces to low-dimensional manifolds. Extensive experiments show that our method, learning with the low-dimensional manifolds, achieves more accurate surface-normal estimation, outperforming other state-of-the-art methods on the challenging DiLiGenT benchmark dataset.
Yakun Ju, Muwei Jian, Junyu Dong, Kin-Man Lam 0001
VCIP2
2020 A joint guidance-enhanced perceptual encoder and atrous separable pyramid-convolutions for image inpainting
Yongle Zhang 0001, Yingyu Wang, Junyu Dong, Lin Qi 0004, Hao Fan 0004, Xinghui Dong, Muwei Jian, Hui Yu 0001
Neurocomputing7
2020 Enhancing MOEA/D with information feedback models for large-scale many-objective optimization
Gaige Wang, Keqin Li 0001, Wei-Chang Yeh 0001, Muwei Jian, Junyu Dong
Inf. Sci.5
2020 Saliency detection using multiple low-level priors and a propagation mechanism
Muwei Jian, Junyu Dong, Chaoran Cui, Xiushan Nie, Yilong Yin
Multim. Tools Appl.1
2020 A brief survey of visual saliency detection
Inam Ullah 0002, Muwei Jian, Sumaira Hussain, Jie Guo 0012, Hui Yu 0001, Xing Wang 0002, Yilong Yin
Multim. Tools Appl.2
2020 Improving image segmentation based on patch-weighted distance and fuzzy clustering
Xiaofeng Zhang 0003, Muwei Jian, Yujuan Sun, Hua Wang 0012, Caiming Zhang 0001
Multim. Tools Appl.2
2020 Improved face super-resolution generative adversarial networks
Mengxue Wang, Zhenxue Chen, Q. M. Jonathan Wu, Muwei Jian
Mach. Vis. Appl.4
2020 Learning the Traditional Art of Chinese Calligraphy via Three-Dimensional Reconstruction and Assessment
abstract
The traditional art of Chinese calligraphy, reflecting the wisdom of the grass-roots community, is the soul of Chinese culture. Just like many other types of craftsmanship, it is part of the historical heritage and is worth conserving, from generation to generation. Since the movements of an ink brush are in a 3D style when Chinese calligraphy is written, they embody “The Power of Beauty,” comprising various reflectance properties and rough-surface geometry. To truly understand the powerful significance and beauty of the art of Chinese calligraphy, in this paper, a 3D calligraphy reconstruction method, based on Photometric Stereo, is designed to capture the detailed appearance of the calligraphy's 3D surface geometry. For assessment, an Iterative Closest Point (ICP) algorithm is applied for registration of 3D intrinsic shapes between the Chinese calligraphy and the calligraphy fans' handwriting. Through matching these two sets of calligraphy characters, the designed system can give a score to the handwriting of a user. Experiments have been performed on Chinese calligraphy from different historical dynasties to evaluate the effectiveness of the proposed scheme, and experimental results show that the developed system is useful and provides a convenient method of calligraphy appreciation and assessment.
Muwei Jian, Junyu Dong, Maoguo Gong, Hui Yu 0001, Liqiang Nie, Yilong Yin, Kin-Man Lam 0001
IEEE Trans. Multim.1
2020 Social-sensed Image Aesthetics Assessment
abstract
Image aesthetics assessment aims to endow computers with the ability to judge the aesthetic values of images, and its potential has been recognized in a variety of applications. Most previous studies perform aesthetics assessment purely based on image content. However, given the fact that aesthetic perceiving is a human cognitive activity, it is necessary to consider users’ perception of an image when judging its aesthetic quality. In this article, we regard users’ social behavior as the reflection of their perception of images and harness these additional clues to improve image aesthetics assessment. Specifically, we first merge the raw social interactions between users and images into clusters as the social labels of images, so the collective social behavioral information associated with an image can be well represented over a structured and compact space. Then, we develop a novel deep multi-task network to jointly learn social labels in different modalities from social images and apply it to common web images. In this manner, our approach is readily generalized to web images without social behavioral information. Finally, we introduce a high-level fusion sub-network to the aesthetics model, in which the social and visual representations of images are well balanced for aesthetics assessment. Experimental results on two benchmark datasets well verify the effectiveness of our approach and highlight the benefits of different types of social behavioral information for image aesthetics assessment.
Chaoran Cui, Peiguang Lin, Xiushan Nie, Muwei Jian, Yilong Yin
ACM Trans. Multim. Comput. Commun. Appl.4
2019 Visual Saliency Detection Based on Full Convolution Neural Networks and Center Prior
abstract
Video saliency detection aims to mimic the human's visual attention system of perceiving the world via extracting the most attractive regions or objects in the input video. At present, traditional video saliency-detection models have achieved good performance in many applications. However, it is still challenging in exploiting the consistency of spatiotemporal information. In order to tackle this challenge, this paper proposes a video saliency-detection model based on human attention mechanism and full convolution neural networks. First, visual features are extracted from video frames through the fully convolutional networks. The second stage is to spread attention features to the other layer (i. e. the fifth layer) of fully convolutional networks via a weight sharing strategy. Finally, the final result produced by the convolution network is optimized by considering spatial location information with center prior of the salient object. Experimental results show that the performance of the proposed algorithm is superior to other state-of-the-art methods based on the widely used data set for video saliency detection.
Muwei Jian, Jiaojin Wang, Hui Yu 0001
HSI1
2019 Multi-view face hallucination using SVD and a mapping model
Muwei Jian, Chaoran Cui, Xiushan Nie, Huaxiang Zhang 0001, Liqiang Nie, Yilong Yin
Inf. Sci.1
2019 Cascaded one-vs-rest detection network for fine-grained recognition without part annotations
Long Chen 0019, Shengke Wang, Kin-Man Lam 0001, Huiyu Zhou 0001, Muwei Jian, Junyu Dong
Multim. Tools Appl.5
2019 Binary feature representation learning for scene retrieval in micro-video
Jie Guo 0012, Xiushan Nie, Muwei Jian, Yilong Yin
Multim. Tools Appl.3
2019 Assessment of feature fusion strategies in visual attention mechanism for saliency detection
Muwei Jian, Chaoran Cui, Xiushan Nie, Hanjiang Luo, Yilong Yin
Pattern Recognit. Lett.1
2018 Image Aesthetic Distribution Prediction with Fully Convolutional Network
Huidi Fang, Chaoran Cui, Xiang Deng 0002, Xiushan Nie, Muwei Jian, Yilong Yin
MMM (1)5
2018 3D reconstruction of human face from an input image under random lighting condition
abstract
The three-dimensional reconstruction from single input image is quite difficult due to many unknown parameters, such as the light conditions, the surface normal and albedo of the object. However, there are overall similar characteristics for different human faces, such as the shapes and the positions of the eyes, nose; mouth and ears are generally identical. The similar characteristics has been used in this paper to relax the numbers of the input face images, and reconstruct the 3D shape based on a couple statistical model. Moreover, the light condition of the single input image can be different from that of training database. The experiment results show the effectiveness of the proposed method.
Yujuan Sun, Xiaofeng Zhang 0003, Muwei Jian
Int. J. Inf. Comput. Secur.3
2018 Perception-driven procedural texture generation from examples
Jun Liu 0055, Yanhai Gan, Junyu Dong, Lin Qi 0004, Xin Sun 0003, Muwei Jian, Hui Yu 0001
Neurocomputing6
2018 A Visual Secret Sharing Scheme Based on Improved Local Binary Pattern
abstract
A visual secret sharing (VSS) scheme is intended to share secret information in a group to avoid potential treat of interruption and modification. In this paper, we present a novel VSS scheme based on the improved local binary pattern (LBP) operator. It makes full use of local contrast features of LBP for concealing secret image data into different image shares, which can be used to recover the secret easily and exactly. By varying LBP extensions, we can design various kinds of VSS schemes for sharing secret information. Compared to the currently available VSS algorithms, the proposed scheme demonstrates better randomness in shares with less pixel expansion and exact determination in reconstruction with lower computational cost.
Wenyin Zhang, Frank Y. Shih, Shunbo Hu, Muwei Jian
Int. J. Pattern Recognit. Artif. Intell.4
2018 Integrating QDWD with pattern distinctness and local contrast for underwater saliency detection
Muwei Jian, Qiang Qi, Junyu Dong, Yilong Yin, Kin-Man Lam 0001
J. Vis. Commun. Image Represent.1
2018 Saliency detection based on background seeds by object proposals and extended random walk
Muwei Jian, Runxia Zhao, Xin Sun 0003, Hanjiang Luo, Wenyin Zhang, Huaxiang Zhang 0001, Junyu Dong, Yilong Yin, Kin-Man Lam 0001
J. Vis. Commun. Image Represent.1
2018 Saliency detection based on directional patches extraction and principal local color contrast
Muwei Jian, Wenyin Zhang, Hui Yu 0001, Chaoran Cui, Xiushan Nie, Huaxiang Zhang 0001, Yilong Yin
J. Vis. Commun. Image Represent.1
2018 Saliency detection using quaternionic distance based weber local descriptor and level priors
Muwei Jian, Qiang Qi, Junyu Dong, Xin Sun 0003, Yujuan Sun, Kin-Man Lam 0001
Multim. Tools Appl.1
2018 Content-based image retrieval via a hierarchical-local-feature extraction scheme
Muwei Jian, Yilong Yin, Junyu Dong, Kin-Man Lam 0001
Multim. Tools Appl.1
2018 An improved genetic algorithm for three-dimensional reconstruction from a single uniform texture image
Yujuan Sun, Xiaofeng Zhang 0003, Muwei Jian, Shengke Wang, Zeju Wu, Qingtang Su, Beijing Chen
Soft Comput.3
2017 Cascade support vector regression-based facial expression-aware face frontalization
abstract
The main aim of face frontalization is to synthesize the frontal facial appearances from non-frontal facial images. How to estimate the frontal face-shape is a crucial but very challenging problem in the frontalization task. Most existing methods use a single shape template to fit in with frontal facial appearances, which will result in a loss of expression-related information. In this work, we present a novel facial expression-aware face frontalization method which directly learns the pair-wise relations between non-frontal face-shape and its frontal counterpart. The support vector regression is explored to train the pair-wise regression model. Considered the pair-wise relationship is non-linear, an appropriate cascade manner is applied to iteratively adjust and optimize the model. With the estimated frontal shape, facial appearances are synthesized through a texture-fitting process formulated by solving a simple optimization problem. The proposed method has been evaluated on a in-the-wild facial expression database. The experimental results shows an outstanding performance of both visual effects of expression recovery and facial expression recognition.
Yiming Wang 0001, Hui Yu 0001, Junyu Dong, Muwei Jian, Honghai Liu 0001
ICIP4
2017 The OUC-vision large-scale underwater image database
abstract
In this paper, a large-scale underwater image database for underwater salient object detection or saliency detection is presented in detail. This database is called the OUC-VISION underwater image database, which contains 4400 underwater images of 220 individual objects. Each object is captured with four pose variations (the frontal-, the opposite-, the left-, and the right-views of each underwater object) and five spatial locations (the underwater object is located at the top-left corner, the top-right corner, the center, the bottom-left corner, and the bottom-right corner) to obtain 20 images. Meanwhile, this publicly available OUC-VISION database also provides relevant industrial fields, and academic researchers with underwater images under different sources of variations, especially pose, spatial location, illumination, turbidity of water, etc. Ground-truth information is also manually labelled for this database. The OUC-VISION database can not only be widely used to assess and evaluate the performance of the state-of-the-art salient-object detection and saliency-detection algorithms for general images, but also will particularly benefit the development of underwater vision technology in the future.
Muwei Jian, Qiang Qi, Junyu Dong, Yinlong Yin, Wenyin Zhang, Kin-Man Lam 0001
ICME1
2016 Ocean internal waves features extraction by analysis of aerial oblique photography
abstract
Internal waves are a widespread geophysical phenomenon in stratified fluids and studying internal features in the coastal ocean is an important task. Using satellite imagery for studying oceanic internal waves is very popular and studying the low altitude aerial oblique photograph is a new direction. In this paper, we study the images captured from a circling aircraft which track a number of internal wave packets. The captured images are first rectified and photogram metrically mapped to ground coordinates, and then we use canny edge detector to find the internal wave propagation direction. The first several waves' ridges of each ground coordinated images are also exactly labeled and propagation speeds can be achieved. Experiment results show the performance of our algorithms.
Shengke Wang, Long Chen 0019, Jianping Yang, Muwei Jian, Lifang Lin, Junyu Dong
IGARSS5
2016 Reconstruction of normal and albedo of convex Lambertian objects by solving ambiguity matrices using SVD and optimization method
Yujuan Sun, Muwei Jian, Xiaofeng Zhang 0003, Junyu Dong, LinLin Shen, Beijing Chen
Neurocomputing2
2016 Ocean Front Detection From Instant Remote Sensing SST Images
abstract
Identifying fronts manually from satellite images is a tedious and subjective task. Accordingly, edge detection algorithms are introduced for automatic detection of fronts. However, traditional algorithms cannot be applied to cloud-contaminated images, because missing data caused by occasional cloud coverage interferes with front detection. To diminish this risk, this letter proposes a new algorithm for a quick and an accurate detection of fronts from an instant cloud-contaminated sea surface temperature (SST) image, instead of depending on the daily or weekly averaged SST images. This algorithm adopts a data-driven analog interpolation method, which estimates missing values from the historical data of the same region. After reducing the contour between the interpolated data and the original data, an instant front detection algorithm is proposed based on microcanonical multiscale formalism (MMF). The algorithm utilizes MMF to detect singularity exponents (SEs), and then enhances the features detected in a cloud-contaminated region. Finally, a threshold is set to extract fronts from SE. Experimental results on an AVHRR satellite SST image of 12:00 o'clock covering China Coastal waters confirmed the effectiveness of the proposed algorithm.
Yuting Yang 0001, Junyu Dong, Xin Sun 0003, Redouane Lguensat, Muwei Jian
IEEE Geosci. Remote. Sens. Lett.5
2015 Fast 3D face reconstruction based on uncalibrated photometric stereo
Yujuan Sun, Junyu Dong, Muwei Jian, Lin Qi 0004
Multim. Tools Appl.3
2015 Simultaneous Hallucination and Recognition of Low-Resolution Faces Based on Singular Value Decomposition
abstract
In video surveillance, the captured face images are usually of low resolution (LR). Thus, a framework based on singular value decomposition (SVD) for performing both face hallucination and recognition simultaneously is proposed in this paper. Conventionally, LR face recognition is carried out by super-resolving the LR input face first, and then performing face recognition to identify the input face. By considering face hallucination and recognition simultaneously, the accuracy of both the hallucination and the recognition can be improved. In this paper, singular values are first proved to be effective for representing face images, and the singular values of a face image at different resolutions have approximately a linear relation. In our algorithm, each face image is represented using SVD. For each LR input face, the corresponding LR and high-resolution (HR) face-image pairs can then be selected from the face gallery. Based on these selected LR-HR pairs, the mapping functions for interpolating the two matrices in the SVD representation for the reconstruction of HR face images can be learned more accurately. Therefore, the final estimation of the high-frequency details of the HR face images will become more reliable and effective. The experimental results demonstrate that our proposed framework can achieve promising results for both face hallucination and recognition.
Muwei Jian, Kin-Man Lam 0001
IEEE Trans. Circuits Syst. Video Technol.1
2015 Visual-Patch-Attention-Aware Saliency Detection
abstract
The human visual system (HVS) can reliably perceive salient objects in an image, but, it remains a challenge to computationally model the process of detecting salient objects without prior knowledge of the image contents. This paper proposes a visual-attention-aware model to mimic the HVS for salient-object detection. The informative and directional patches can be seen as visual stimuli, and used as neuronal cues for humans to interpret and detect salient objects. In order to simulate this process, two typical patches are extracted individually and in parallel from the intensity channel and the discriminant color channel, respectively, as the primitives. In our algorithm, an improved wavelet-based salient-patch detector is used to extract the visually informative patches. In addition, as humans are sensitive to orientation features, and as directional patches are reliable cues, we also propose a method for extracting directional patches. These two different types of patches are then combined to form the most important patches, which are called preferential patches and are considered as the visual stimuli applied to the HVS for salient-object detection. Compared with the state-of-the-art methods for salient-object detection, experimental results using publicly available datasets show that our produced algorithm is reliable and effective.
Muwei Jian, Kin-Man Lam 0001, Junyu Dong, LinLin Shen
IEEE Trans. Cybern.1
2014 Facial-feature detection and localization based on a hierarchical scheme
Muwei Jian, Kin-Man Lam 0001, Junyu Dong
Inf. Sci.1
2014 Illumination-insensitive texture discrimination based on illumination compensation and enhancement
Muwei Jian, Kin-Man Lam 0001, Junyu Dong
Inf. Sci.1
2014 Face-image retrieval based on singular values and potential-field representation
Muwei Jian, Kin-Man Lam 0001
Signal Process.1
2013 A novel face-hallucination scheme based on singular value decomposition
Muwei Jian, Kin-Man Lam 0001, Junyu Dong
Pattern Recognit.1
2012 Dynamic textures indexing and retrieval based on intrinsic properties
abstract
The indexing and retrieval of dynamic textures remains one of the most challenging tasks in both computer vision and computer graphics. This paper first introduces the concept of intrinsic properties of dynamic textures, including Static Texture Features (STF) and Dynamic Texture Features (DTF), and then proposes an efficient method that utilizes these intrinsic properties for the indexing and retrieval of dynamic textures. Our proposed method is evaluated and compared to existing methods based on a variety of video texture samples. Experimental results show that our scheme can produce promising results.
Muwei Jian, Kin-Man Lam 0001, Junyu Dong
ISCAS1
2011 Capture and fusion of 3d surface texture
Muwei Jian, Junyu Dong
Multim. Tools Appl.1
2009 Automatic correction of non-uniform illumination for 3D surface heightmap reconstruction
abstract
For the three-dimensional surface texture can display the texture information of the object better than the two-dimensional illumination and the view angles; it is widely used in virtual reality and computer games. Photometric Stereo, as one of the effective technologies for capture of three-dimensional surface texture information, has attracted wide attention. Uniform illumination is the essential condition for the capture and reconstruction of three-dimensional surface texture using Photometric Stereo. In practice, non-uniform illumination leads to distorted surface height maps during the capture and reconstruction processes. This paper proposes a simple method to automatic correction of non-uniform illuminate for 3D surface height map reconstruction and to eliminate this kind of distortion and aberration. The experimental result shows the effectiveness of the proposed method.
Muwei Jian, Junyu Dong, Baokang Zhao
ICME1
2007 Wavelet-Based Salient Regions and their Spatial Distribution for Image Retrieval
abstract
In content-based image retrieval, the representation of local properties in an image is one of the most active research issues. This paper proposes a salient region detector based on wavelet transform. The detector can extract the visually meaningful regions on an image and reflect local characteristics. An annular segmentation algorithm based on the distribution of salient regions is designed. It takes not only local image features into account, but also the spatial distribution information of the salient regions. Color moments and Gabor features around the salient regions in every annular region are computed as feature vectors used for indexing the image. We have tested the proposed scheme using a wide range of image samples from the Corel Image Library for content-based image retrieval. The experiments indicate that the method has produced promising results.
Muwei Jian, Junyu Dong
ICME1