VLDB 2026 Research / reviewers in the wild / expert
Jinye Peng 0001
dblp:09/3562-1
· DBLP profile ↗
106ranked-venue papers
5as first author
62since 2021 · last 2026
0000-0003-4286-2576ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 52 · 35 since 2021Graphics, computer vision, multimedia, augmented reality and games · 45 · 2 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 3 first-author · 11 since 2021Databases, data management, data science and information retrieval · 5 · 3 since 2021Computer networks · 3 · 2 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Act-LLM: A whole-process chain for character-centric role-playing with large language models
Xiaoxu Han, Wanqing Zhao, Ziyu Guan, Jinye Peng 0001 |
Expert Syst. Appl. | 4 |
| 2026 | AICE: Three domain conversion network applied to all-in-one image inpainting and color enhancement task
Qiyao Hu, Xianlin Peng, Manli Sun, Shuyi Qu, Jinye Peng 0001 |
Expert Syst. Appl. | 6 |
| 2026 | High-fidelity mural inpainting via progressive reconstruction and damage-aware adaptation
Shuyi Qu, Qingqing Kang, Shenglin Peng, Jun Wang 0078, Qiyao Hu, Xianlin Peng, Jinye Peng 0001 |
Expert Syst. Appl. | 8 |
| 2026 | Dual-branch non-negative matrix factorization guided by information decoupling for multi-view clustering
Mingxia Gong, Chengcai Leng, Zhao Pei, Jinye Peng 0001, Irene Cheng 0001, Anup Basu |
Neurocomputing | 5 |
| 2026 | EmoSkeMPR: ViT-based masked position reconstruction and skeleton feature fusion for multi-task emotion and behavior recognition
Xianlin Peng, Lingjie Kong, Qiyao Hu, Jinye Peng 0001, Gu Fang 0002 |
Inf. Sci. | 4 |
| 2026 | MEMA-ConvLSTM: Spatiotemporal prediction via multi-scale autocorrelation memory and hierarchical fusion
Chengcai Leng, Huaiping Yan, Zhao Pei, Jinye Peng 0001 |
Inf. Sci. | 5 |
| 2026 | MSKICP: multiscale descriptor and KMPE kernel function-improved iterative closest point for point cloud registration
Shengmei Chen, Hao Deng 0018, Jingyi Han, Jinye Peng 0001, Lin Wang 0026 |
Multim. Syst. | 5 |
| 2026 | CalliECD: An Error Correction Diffusion Model for Multi-Style Chinese Calligraphy Generation
Qiyao Hu, Yinyin Luo, Xianlin Peng, Rui Cao 0003, Jinye Peng 0001, Jianping Fan 0001 |
Pattern Recognit. | 5 |
| 2026 | SemiSketch: An ancient mural sketch extraction network based on reference prior and gradient frequency compensation
Jun Wang 0078, Shuyi Qu, Qunxi Zhang, Yirong Ma, Shenglin Peng, Jinye Peng 0001 |
Pattern Recognit. | 7 |
| 2026 | Full-DOF Calibration Method for 3D Point Cloud Acquisition Device Based on BiK Loss FunctionabstractTo accurately estimate full-degree-of-freedom (DOF) model parameters of the 3D point cloud acquisition device, a calibration method by using a calibrator of a simple space ball is proposed. The method can achieve full-DOF (DOF) estimation without additional hardware and step-by-step calculations. Firstly, a measurement model is established according to the rotation characteristics of the 3D point cloud acquisition device. Secondly, with the spherical constraints of the sphere, a nonlinear optimization model of the 3D point cloud acquisition device is established by adopting the bidirectional kernel meanp-power error (BiK) loss function which is robust to the measurement noise and outliers. Finally, the successful history-based differential evolution parameter adaptation (SHADE) algorithm and the Levenberg-Marquardt (LM) algorithm are combined to solve the nonlinear optimization model, and the full-DOF model parameters of the 3D point cloud acquisition device can be estimated. Experimental results demonstrates that the proposed method can accurately estimate the full-DOF model parameters of the 3D point cloud acquisition device. More encouragingly, the effect of measurement noise and outliers can be significantly suppressed. Shengmei Chen, Hao Deng 0018, Jinye Peng 0001, Lin Wang 0026 |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2025 | Iterative Self-Training with Class-Aware Text-to-Image Synthesis for Visual Task LearningabstractGenerative models are widely used to produce synthetic images with annotations, alleviating the burden of image collection and annotation for training deep visual models. However, challenges such as limited image diversity, noisy pseudo labels, and domain gaps between synthetic and real images often undermine their effectiveness in downstream visual tasks. This paper introduces the Iterative Self-Training with Class-Aware Text-to-Image Synthesis (IST-CATS) framework, which addresses these challenges by integrating a class-aware text-to-image synthesis (CATS) component with an iterative self-training (IST) strategy. CATS innovatively introduces a class-aware chain approach to generate detailed descriptions. These descriptions act as prompts for a diffusion model, enabling the creation of a diverse of images accompanied by distinguishable objects against the background. The generated images can be easily pseudo-labeled by an unsupervised instance segmentation method, and then noisy pseudo labels can be effectively purified by a novel feature similarity-based filtering mechanism. The generated images underpin our IST, which progressively enhances vision models and refines pseudo labels through self-training and our proposed label filtering strategy (LabFilt). LabFilt meticulously improves the quality of pseudo labels by employing class-adaptive techniques at both the pixel and object levels, ensuring refined pseudo-label accuracy. IST-CATS demonstrates superior performance in object detection and semantic segmentation compared to traditional synthetic and semi/weakly-supervised methods, effectively addressing data collection and annotation challenges. Xiang Zhang 0018, Wanqing Zhao, Pengyang Li, Hangzai Luo, Sheng Zhong 0006, Jinye Peng 0001, Jianping Fan 0001 |
AAAI | 7 |
| 2025 | ASMCC-Diff: Arbitrary Size Multi-Condition Controllable Chinese Landscape Painting Generation with Diffusion ModelsabstractThanks to the emergence of generative models, Chinese Landscape Painting Generation (CLPG) has garnered increasing attention. However, existing works are primarily limited to relying on text control conditions, lacking more fine-grained control over spatial layout and style. Additionally, they are limited to a fixed size and aspect ratios. But different landscape scenes require different sizes to appear more balanced and natural. Thus, a question arises: Is it possible to design a model that can be controlled by multiple conditions (text, style and sketches) while generating images with various sizes? In this paper, we explore this issue and propose ASMCC-Diff. Specifically, it consists of two modules, i.e. multi-condition controlled image generation module and arbitrary size up-scaling module. The critical insights of multi-condition controlled image generation module are to embed multiple conditions with distinct priorities. The sketch condition serves as the primary flow, guiding the overall structure, while the text and style conditions act as auxiliary components, injected into the diffusion model via a cross-attention module. Additionally, to avoid conflicts between the semantics of style images and text, we use Q-former to separate the semantic and stylistic information of the reference image. For the arbitrary size upscaling module, we first truncate the generation process, up-sample the image to the specified size, and then continue the generation. Furthermore, we introduce a new Chinese landscape painting database that supports multiple conditions, facilitating further research. Experimental results demonstrate the superior performance of our proposed model. The code and dataset will be released. Xiaodan Zhang 0005, Xiteng Zhang, Jianqiang Yan, Qiyao Hu, Jinye Peng 0001, Qiannan Duan |
ECAI | 5 |
| 2025 | Environment-Agnostic Pose: Generating Environment-Independent Object Representations for 6D Pose Estimation
Shaobo Zhang 0006, Wanqing Zhao, Wei Zhao 0019, Ziyu Guan, Jinye Peng 0001 |
ICCV | 6 |
| 2025 | Semantic image segmentation via dynamic curriculum learning
Xiang Zhang 0018, Wanqing Zhao, Chenji Wang, Hangzai Luo, Sheng Zhong 0006, Jinye Peng 0001, Jianping Fan 0001 |
Appl. Intell. | 7 |
| 2025 | An efficient parallel mesh generation method for finite element based analysis of large complex architecture
Wanqing Zhao, Chunnan Li, Tongkun Deng, Jun Wang 0078, Jinye Peng 0001 |
Comput. Aided Des. | 7 |
| 2025 | Block information strategy for multi-modal remote sensing image registration
Yameng Hong, Chengcai Leng, Beihua Liu, Jinye Peng 0001, Irene Cheng 0001, Anup Basu |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | Dual graph-regularized low-rank representation for hyperspectral image denoising
Chengcai Leng, Mingpei Tang, Zhao Pei, Jinye Peng 0001, Anup Basu |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | Orthogonal Diversity Nonnegative Matrix Factorization for multi-view clustering
Xinling Zhang, Chengcai Leng, Jinye Peng 0001, Irene Cheng 0001, Anup Basu |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | MFEL-YOLO for small object detection in UAV aerial images
Ting Hou, Chengcai Leng, Zhao Pei, Jinye Peng 0001, Irene Cheng 0001, Anup Basu |
Expert Syst. Appl. | 5 |
| 2025 | Transformer gate-based interactive U-Net for hyperspectral and multispectral image fusion
Yihao Fu, Lu Liu 0025, Jun Wang 0078, Jinye Peng 0001 |
Expert Syst. Appl. | 5 |
| 2025 | Mamba-GIE: A visual state space models-based generalized image extrapolation method via dual-level adaptive feature fusion
Ruoyi Zhang, Shuyi Qu, Jun Wang 0078, Jinye Peng 0001 |
Expert Syst. Appl. | 5 |
| 2025 | Multi-view data representation via adaptive label propagation nonnegative matrix factorization
Chengcai Leng, Jinye Peng 0001, Zhao Pei, Anup Basu |
Inf. Sci. | 3 |
| 2025 | Learning Cross-View Consistent 3D Keypoints for Object 6D Pose EstimationabstractAccurate 6D object pose estimation from RGB images is crucial for various computer vision applications, such as augmented reality, robotic manipulation and autonomous driving. Existing methods often rely on extensive labeled data, either manually annotated or synthetically generated, which can be laborious and impractical for real-world deployment. To address these challenges, we propose OK-POSE, a keypoint-based 6D object pose estimation method that leverages relative transformations between viewpoints for training. By utilizing pairs of images with object annotations and relative transformation information, OK-POSE automatically learns to detect 3D keypoints of objects, enabling geometrically and visually consistent pose estimation. The simplicity and accessibility of obtaining relative transformation information, which can be acquired from inexpensive binocular cameras or common smartphone devices, significantly reduce labeling costs and mitigate domain gap issues associated with synthetic data. Experimental results demonstrate that OK-POSE achieves competitive performance compared to methods relying on explicit 3D annotations or object 3D models. Moreover, we provide insights into the data collection process and introduce OK-POSE++, an enhanced version with optimized network architecture and loss functions, yielding further improvements in performance. Our approach offers a practical solution for 6D object pose estimation, suitable for real-world applications in scenarios where extensive 3D annotations or object models are unavailable. The code is released athttps://github.com/acmff22/OKPOSE. Shaobo Zhang 0006, Wanqing Zhao, Ziyu Guan, Wei Zhao 0019, Jinye Peng 0001, Jianping Fan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Integrating Recurrent-KAN With SAM Adapter for Blind Hyperspectral UnmixingabstractDue to the limitation of sensors, hyperspectral images contain a large number of mixed pixels. Hyperspectral unmixing techniques decompose these mixed pixels into distinct endmembers and their corresponding abundance values. Traditional methods initialize weights in the decoder and utilize outputs/weights as abundance maps and endmembers——an approach heavily dependent on initial weights that significantly limits performance. This paper proposes a blind hyperspectral unmixing method integrating Recurrent Kolmogorov-Arnold Networks (KAN) with Segment Anything Model (SAM) adapter. The method operates through three sequential stages: feature encoding, endmember extraction, and abundance estimation. Specifically for feature encoding, a HU-SAM adapter is proposed to capture global-local spatial features. For endmember extraction, an iteratively learned Recurrent-KAN module reconstructs endmembers while stabilizing model learning. For abundance estimation, an updated Swin Transformer module is utilized to maintain a lower parameter count. Extensive experiments on real and synthetic datasets demonstrate superior effectiveness of the proposed method over the eight state-of-the-art methods. Yihao Fu, Shenglin Peng, Jun Wang 0078, Jinye Peng 0001, Moncef Gabbouj |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | Incremental semi-supervised graph learning NMF with block-diagonal
Xue Lv, Chengcai Leng, Jinye Peng 0001, Zhao Pei, Irene Cheng 0001, Anup Basu |
Eng. Appl. Artif. Intell. | 3 |
| 2024 | Feature matching based on Gaussian kernel convolution and minimum relative motion
Chengcai Leng, Huaiping Yan, Jinye Peng 0001, Zhao Pei, Anup Basu |
Eng. Appl. Artif. Intell. | 4 |
| 2024 | Bayesian non-negative matrix factorization with Student's t-distribution for outlier removal and data clusteringabstractNon-negative Matrix Factorization (NMF) is an effective way to solve the redundancy of non-negative high-dimensional data. Most of the traditional probability-based NMF methods use Gaussian distribution to model the differences between the matrices before and after decomposition. However, the Gaussian distribution is strongly affected by outliers, and it may not fit all datasets accurately when there are no outliers in the data. In this article, we propose a novel Bayesian NMF with the Student’s t-distribution, i.e., TNMF. specifically, in order to reduce the impact of outliers on the algorithm, we use the Student’s t-distribution to fit the data points instead of the Gaussian distribution. In addition, it is possible to adjust the Degree of Freedom (DF) to make the Student’s t-distribution more flexible than the Gaussian distribution to fit data points when there are no outliers. Next, we combine the Automatic Relevance Determination (ARD) prior in our algorithm to simplify the model and allow for better performance of the algorithm. Finally, the article used 10 datasets to design two kinds of experiments, outlier removal and data clustering. The outlier removal results of this proposed algorithm are significantly better than the other methods, and it performs better in clustering compared to the other methods in the majority of cases. Ruixue Yuan, Chengcai Leng, Jinye Peng 0001, Anup Basu |
Eng. Appl. Artif. Intell. | 4 |
| 2024 | AGD-GAN: Adaptive Gradient-Guided and Depth-supervised generative adversarial networks for ancient mural sketch extraction
Shenglin Peng, Shuyi Qu, Qunxi Zhang, Jun Wang 0078, Jinye Peng 0001 |
Expert Syst. Appl. | 6 |
| 2024 | Confidence-based dynamic cross-modal memory network for image aesthetic assessment
Xiaodan Zhang 0005, Jinye Peng 0001, Xinbo Gao 0001, Bo Hu 0008 |
Pattern Recognit. | 3 |
| 2024 | Coverage Analysis of Single-Swarm mmWave UAV Networks Under Multiple Types of BlockagesabstractMillimeter wave (mmWave)-based unmanned aerial vehicle (UAV) communication is susceptible to blockages, even from humans. Previous studies that primarily focused only on static blockage may not accurately characterize the system performance. This paper investigates the coverage performance of mmWave UAV networks by jointly considering multiple types of blockages under finite homogeneous Poisson point process and Binomial point process, which are commonly employed in finite area scenarios with random and fixed number of UAVs, respectively. Particularly, we derive the average line-of-sight probability and coverage probability under static, dynamic, and self blockages. Simulations verify our theoretical results, demonstrating that: the above system performance predominantly depends on self-blockage if UAVs are at high altitudes. Conversely, at relatively low altitudes, all three types of blockages impact them, with static blockage being the dominant factor. To avoid self-blockage, UAV height should satisfy$h\!\gt \!h_{R}\!+\!\frac {r_{i}}{\tan \varphi _{b}}$, where$h_{R}$is the height of the user equipment (UE),$r_{i}$is the two-dimensional distance of the UAV-UE link,$\varphi _{b}$is the elevation angle between UE and UAV. The required height is proportional to$r_{i}$and increases as distance d between the user and UE decreases, as$\varphi _{b}$is proportional to d. The findings help on designing the network parameters. To our best knowledge, this is the first work to analyze the coverage of mmWave UAV networks under multiple types of blockages. Cunyan Ma, Xiaoya Li 0003, Chen He 0002, Jinye Peng 0001, Kun Yang 0001, Z. Jane Wang 0001 |
IEEE Trans. Commun. | 4 |
| 2024 | Information-Enhanced Network for Noncontact Heart Rate Estimation From Facial VideosabstractRemote photoplethysmography (rPPG) is a vital way of measuring heart rate (HR) to reflect human physical and mental health, which is useful for diagnosing cardiovascular and neurological diseases. Many non-contact HR estimation methods have been proposed gradually in recent years, but the majority of approaches are based on a single-modal HR information source, resulting in ineffective and unsatisfactory estimation results due to noise and insufficient information. This paper proposes a novel information-enhanced network for HR estimation based on multimodal (e.g., RGB and NIR) sources to address these problems. In the network, context and modal difference information are sequentially enhanced from spatiotemporal and modal views for accurately describing HR-aware features, while maximum frequency information is enhanced for inhibiting heartbeat noise. Specifically, a context-enhanced video Swin-Transformer (CET) module is exploited to extract useful rPPG signal features from facial visible-light and near-infrared videos. Then, a novel modal difference enhanced fusion (MDEF) module is designed to acquire a fused rPPG signal, which is taken as the input of the frequency-enhanced estimation (FEE) module to obtain the corresponding HR value. These three modules are integrated and jointly learned in an end-to-end way, and the multimodal combinations can provide highly complementary information for estimating HR value. Experimental and evaluation results on three multimodal datasets show that the proposed model achieves a superior effect compared to the state-of-the-art methods. Zhaoqiang Xia, Xiaobiao Zhang, Jinye Peng 0001, Xiaoyi Feng, Guoying Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Enabling Near-Zero Cost Object Detection in Remote Sensing Imagery via Progressive Self-TrainingabstractDeep learning-based object detection models rely heavily on large-scale and precise annotations for training. However, manually annotating bounding-box annotations for such data is both time-consuming and costly, especially when dealing with high-resolution satellite imagery containing densely packed small-sized objects. To alleviate the burden of manual annotation, we propose a simple yet effective approach, called progressive self-training object detection (PSTDet), to enable accurate object detection in remote sensing imagery without relying on manual annotations. Our PSTDet framework consists of two main components: initial pseudo label generation (IPLG) and progressive self-training with relabeling (PST-R). In IPLG, we leverage unsupervised image clustering, unsupervised instance detection, and geometric constraints to automatically generate high-quality bounding-box annotations for the initial training dataset. This innovative approach significantly reduces the time and expense associated with data annotation, laying a solid foundation for the subsequent progressive self-training stage. The annotations produced by IPLG serve as the training data for PST-R, which enhances the detector and pseudo labels through progressive self-training and our proposed noisy pseudo label filtering strategy (NPLFilter). Our NPLFilter purifies the quality of pseudo labels by integrating geometric constraints, prior knowledge, and category-adaptive thresholds. Experimental results demonstrate that our method achieves significant performance improvement on challenging NWPU VHR-10.v2 and DIOR datasets. Notably, our method far outperforms state-of-the-art weakly supervised methods and compares favorably with fully supervised methods. Xiang Zhang 0018, Xiangteng Jiang, Qiyao Hu, Hangzai Luo, Sheng Zhong 0006, Jinye Peng 0001, Jianping Fan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | Deep Learning-Based Security Analysis of Quantum Random Numbers Generated by Imperfect DevicesabstractThe quantum random number generator (QRNG) is theoretically capable of generating unpredictable random numbers based on the inherent uncertainty of quantum mechanics, which is of paramount importance for information security. However, the security of practical QRNG is susceptible to the influence of unknown classical noise introduced by imperfect measurement devices, leading to potential security threats. In this paper, we propose a deep learning-based prediction model for analyzing the bidirectional security of QRNG where the quantum source is contaminated by classical noise. Firstly, we train a deep learning model capable of evaluating the randomness of mixed entropy source composed of quantum source and classical source, which exhibits excellent performance and effectively avoids the limitations of Statistical tests in evaluating the randomness of the mixture of quantum and classical source results. Secondly, systematically analyzes the impact of non-ideal measurement devices on the practical security of the continuous variable QRNG, which provides an explicit basis for compensating the discrepancy between theory and experiment. Finally, we perform correlation detection on QRNG output sequences with a deep learning model and focus on both the forward and backward security of random numbers. Through bidirectional security detection, random number sequences that may be biased or manipulated can be more accurately and comprehensively evaluated, further preventing potential correlations from opening security holes for eavesdroppers. Lin Wang 0026, Jinye Peng 0001 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2023 | Hyperspectral image classification via deep network with attention mechanism and multigroup strategy
Jun Wang 0078, Jinyue Sun, Erlei Zhang, Jinye Peng 0001 |
Expert Syst. Appl. | 6 |
| 2023 | Hybrid Conv-ViT Network for Hyperspectral Image ClassificationabstractWith the success of ViT (Vision Transformer), Transformer is being increasingly used for hyperspectral image (HSI) classification given its ability to extract global context dependencies. However, existing methods based on transformers tend to classify HSI in the traditional patch-wise manner. Thus, these methods cannot obtain true global features because the inputs of the model are local patches. To solve these problems, a hybrid convolution and ViT network (HCVN) is proposed for HSI classification. HCVN realizes the classification task from the perspective of semantic segmentation, and its input is the entire HSI, which makes it possible to obtain truly meaningful global features. By improving the original ViT, an HCV module is proposed, which enhances the ability of local structure characterization while extracting global features. The HCVN hybrid convolution layer and HCV module realize the extraction and fusion of local and global features. Finally, the dual branch network architecture is used to integrate the spatial and spectral features. Extensive experiments on two datasets verify the effectiveness of the proposed method. Huaiping Yan, Erlei Zhang, Jun Wang 0078, Chengcai Leng, Anup Basu, Jinye Peng 0001 |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2023 | C3N: content-constrained convolutional network for mural image completion
Xianlin Peng, Huayu Zhao, Yongqin Zhang, Qunxi Zhang, Jun Wang 0078, Jinye Peng 0001, Haida Liang |
Neural Comput. Appl. | 8 |
| 2023 | Non-contact PPG signal and heart rate estimation with multi-hierarchical convolutional network
Bin Li 0051, Jinye Peng 0001, Hong Fu |
Pattern Recognit. | 3 |
| 2023 | GGD-GAN: Gradient-Guided dual-Branch adversarial networks for relic sketch generation
Jun Wang 0078, Erlei Zhang, Shan Cui, Qunxi Zhang, Jianping Fan 0001, Jinye Peng 0001 |
Pattern Recognit. | 7 |
| 2023 | Learning to recover lost details from the dark
Maomei Liu, Sheng Zhong 0006, Hangzai Luo, Jinye Peng 0001 |
Pattern Recognit. Lett. | 5 |
| 2023 | Non-Local Color Compensation Network for Intrinsic Image DecompositionabstractSingle image-based intrinsic image decomposition attempts to separate one input image into several intrinsic components, which is inherently an under-constrained problem. Some recent works have been proposed to estimate the intrinsic components using encoder-decoder structures. However, they generally lack exploration of the different component-oriented feature constraints and feature selection processes. In this paper, a non-local color compensation network (NCCNet) is proposed. Firstly, the hue and value channels of HSV color space are used as the complementary information for RGB images for the estimation of albedo and shading, respectively. The color space representation serves as an external constraint, which does not require expensive sensors or complicated computations. Secondly, an integrated non-local attention scheme is proposed to describe the relations of non-adjacent regions with a lower computational complexity compared to traditional methods. Then the non-local and local attention are combined to describe correlations among features and used as feature selectors between the encoder and decoder. Thirdly, the mutual constraint between albedo and shading is also explored in the network to further optimize the process. In order to train the network, a unified mutual exclusion loss function is proposed. Extensive experiments are conducted on several popular datasets, and the proposed NCCNet achieves improved performance with comparable computational cost compared to competing methods. Xiaoyue Jiang, Zhaoqiang Xia, Moncef Gabbouj, Jinye Peng 0001, Xiaoyi Feng |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Toward Blind-Adaptive Remote Sensing Image RestorationabstractWhile deep convolutional neural networks (CNNs) have substantially boosted the performance of low-level vision tasks, they remain largely under-explored in CNN-based remote sensing image restoration. This paper studies the JPEG-LS compressed remote sensing image restoration that faces the following problems. It requires a trade-off in preserving local context information and expanding spatial receptive fields. It needs blind restoration while achieving flexible performance. To this end, we propose a blind-adaptive restoration network, called TBANet, that integrates three modules into an end-to-end network to remedy these problems separately. Specifically, we build a scale-invariant wise-skip ResNet as the baseline to extract more context information. We present a receptive field expansion module by using scale-wise convolution for removing banding artifacts. We design a blind-adaptive controller to provide a deterministic result meanwhile meeting the needs of the user’s preference. In experiments, we compare the restoration accuracy among our model and many different variants of restoration methods on our collected remote sensing image dataset. The proposed network achieves superior performance against state-of-the-art methods in terms of both quantitative metrics and visual quality. Code and models are available at: https://github.com/lmmhh/TBANet. Maomei Liu, Lijia Fan, Sheng Zhong 0006, Hangzai Luo, Jinye Peng 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Movable Object Detection in Remote Sensing Images via Dynamic Automatic LearningabstractThe performance of deep networks for object detection in remote sensing images (RSIs) largely depends on the availability of large-scale training images whose labels are given at the bounding-box level through a labor-intensive manual labeling process. To alleviate the huge burden of providing bounding-box annotations manually for movable objects, we propose a new approach, called dynamic automatic learning (DAL), to progressively learn object detectors. Specifically, a novel initial annotation generation (IAG) strategy is first designed to produce bounding-box annotations for movable objects in multi-temporal remote sensing images. During this process, image-level labels need to be manually labeled for the generated candidates. Next, a detection network learns the detection knowledge from multi-temporal remote sensing images with bounding-box annotations and then transfers the knowledge to generate pseudo boxes for the unlabeled data. Finally, with these pseudo boxes, the object detector can be optimized for generating accurate pseudo boxes iteratively. Furthermore, we introduce a pseudo box filtering (PBF) strategy to purify the quality of pseudo boxes to obtain accurate supervision. Our experiments on the challenging NWPU VHR-10.v2 and DIOR datasets have demonstrated that our DAL approach can achieve competitive results compared to state-of-the-art methods. Xiang Zhang 0018, Hangzai Luo, Sheng Zhong 0006, Jinye Peng 0001, Jianping Fan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Joint beamwidth and resource optimization in ultra-dense MmWave D2D communications
Xiaoya Li 0003, Jinye Peng 0001 |
Wirel. Networks | 3 |
| 2022 | Class Guided Channel Weighting Network for Fine-Grained Semantic SegmentationabstractDeep learning has achieved promising performance on semantic segmentation, but few works focus on semantic segmentation at the fine-grained level. Fine-grained semantic segmentation requires recognizing and distinguishing hundreds of sub-categories. Due to the high similarity of different sub-categories and large variations in poses, scales, rotations, and color of the same sub-category in the fine-grained image set, the performance of traditional semantic segmentation methods will decline sharply. To alleviate these dilemmas, a new approach, named Class Guided Channel Weighting Network (CGCWNet), is developed in this paper to enable fine-grained semantic segmentation. For the large intra-class variations, we propose a Class Guided Weighting (CGW) module, which learns the image-level fine-grained category probabilities by exploiting second-order feature statistics, and use them as global information to guide semantic segmentation. For the high similarity between different sub-categories, we specially build a Channel Relationship Attention (CRA) module to amplify the distinction of features. Furthermore, a Detail Enhanced Guided Filter (DEGF) module is proposed to refine the boundaries of object masks by using an edge contour cue extracted from the enhanced original image. Experimental results on PASCAL VOC 2012 and six fine-grained image sets show that our proposed CGCWNet has achieved state-of-the-art results. Xiang Zhang 0018, Wanqing Zhao, Hangzai Luo, Jinye Peng 0001, Jianping Fan 0001 |
AAAI | 4 |
| 2022 | A classification benchmark for Arabic alphabet phonemes with diacritics in deep neural networks
Eiad Almekhlafi, Moeen Al-Makhlafi, Erlei Zhang, Jun Wang 0078, Jinye Peng 0001 |
Comput. Speech Lang. | 5 |
| 2022 | Automatic learning for object detection
Xiang Zhang 0018, Hangzai Luo, Wanqing Zhao, Sheng Zhong 0006, Jinye Peng 0001, Jianping Fan 0001 |
Neurocomputing | 7 |
| 2022 | Max-Index Based Local Self-Similarity Descriptor for Robust Multi-Modal Image RegistrationabstractIn order to address problems, such as radiation and intensity differences in multi-modal images, this letter proposes a novel idea that integrates maximal indices into the construction of a local self-similarity (LSS) descriptor. The LSS vectors at the same angles but different radial intervals are added to construct the max-index similarity map (MISM) and form the proposed descriptor. This novel descriptor is named max-index-based local self-similarity (MLSS). The MLSS descriptor not only captures the shape similarity between images but is also robust to radiation distortions. Furthermore, a fast and robust algorithm is introduced based on the MLSS descriptor. Comprehensive analysis of accuracy, precision, and computational efficiency shows that the proposed method outperforms five other state-of-the-art methods with stable and better performance on nine pairs of multi-modal test images. Yameng Hong, Chengcai Leng, Xinyue Zhang 0013, Jinye Peng 0001, Licheng Jiao, Anup Basu |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | MTFFN: Multimodal Transfer Feature Fusion Network for Hyperspectral Image ClassificationabstractTransfer learning is an effective way to alleviate the problem of insufficient samples in a hyperspectral image (HSI) classification. However, the present transfer learning-based methods usually transfer knowledge from a single source domain, such as the natural image domain. Therefore, these methods cannot simultaneously transfer spectral and spatial knowledge to the target domain in HSIs. Generally, the natural image has rich spatial structure and texture information, while the HSI has abundant spectral information. To better utilize the knowledge learned from natural image datasets and HSI datasets, we proposed a multimodal transfer feature fusion network (MTFFN) for HSI classification. In MTFFN, a dual-branch network structure is designed to transfer the two-modal knowledge from the natural image domain and the source HSI domain to the target domain in two branches, respectively. A multitask learning strategy is adopted to achieve feature fusion. The fused features are used to generate the final classification result. Moreover, a local attention mechanism is designed to extract more meaningful spectral features. Experiments on two public datasets show that the proposed method is effective (https://github.com/HuaipYan/MTFFN). Huaiping Yan, Erlei Zhang, Jun Wang 0078, Chengcai Leng, Jinye Peng 0001 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2022 | Detail-Aware Multiscale Context Fusion Network for Cloud DetectionabstractIn recent years, a large number of convolutional neural networks-based cloud detection algorithms have been proposed for remote sensing image preprocessing and most of them have an encoder-decoder structure. However, downsampling and upsampling operations, as the basic components of these methods, inevitably lead to the loss of detailed information in high-level features, which affects cloud detection performance. At the same time, the physical characteristics of the cloud, such as the variable size and irregular structure, also put forward requirements for the multi-scale feature representation ability of the network. To this end, we propose a novel cloud detection network named DMNet, which contains a Dense Feature Enhancement Module (DFEM) and a Multi-scale Context Fusion Spatial Attention Module (MCFSAM). DFEM aims to achieve information complementarity by exploiting the different properties of the features at different levels of the encoder, so as to strengthen the detailed information of high-level features and make the low-level features have more semantics. MCFSAM introduces a Multi-scale Context Fusion Block (MCFB) in spatial attention, which enables the network to densely capture contextual information at different scales and further emphasize useful features in the spatial dimension. Extensive experiments on GF-1 wide field-of-view satellite imagery (GF-1 WFV) dataset and Moderate-Resolution Imaging Spectroradiometer (MODIS) dataset demonstrate that our method outperforms other state-of-the-art cloud detection algorithms. Xiang Zhang 0018, Hangzai Luo, Sheng Zhong 0006, Jinye Peng 0001, Jianping Fan 0001 |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2022 | Pain fingerprinting using multimodal sensing: pilot studyabstractAbstract Pain is a complex phenomenon, the experience of which varies widely across individuals. At worst, chronic pain can lead to anxiety and depression. Cost-effective strategies are urgently needed to improve the treatment of pain, and thus we propose a novel home-based pain measurement system for the longitudinal monitoring of pain experience and variation in different patients with chronic low back pain. The autonomous nervous system and audio-visual features are analyzed from heart rate signals, voice characteristics and facial expressions using a unique measurement protocol. Self-reporting is utilized for the follow-up of changes in pain intensity, induced by well-designed physical maneuvers, and for studying the consecutive trends in pain. We describe the study protocol, including hospital measurements and questionnaires and the implementation of the home measurement devices. We also present different methods for analyzing the multimodal data: electroencephalography, audio, video and heart rate. Our intention is to provide new insights using technical methodologies that will be beneficial in the future not only for patients with low back pain but also patients suffering from any chronic pain. Anja Keskinarkaus, Ruijing Yang, Angelos Fylakis, Md. Surat-E.-Mostafa, Arto J. Hautala, Yong Hu 0003, Jinye Peng 0001, Guoying Zhao 0001, Tapio Seppänen, Jaro Karppinen |
Multim. Tools Appl. | 7 |
| 2022 | Contour-enhanced CycleGAN framework for style transfer from scenery photos to Chinese landscape paintings
Xianlin Peng, Shenglin Peng, Qiyao Hu, Jinye Peng 0001, Jianping Fan 0001 |
Neural Comput. Appl. | 4 |
| 2022 | Guided Filter Network for Semantic Image SegmentationabstractThe existing publicly available datasets with pixel-level labels contain limited categories, and it is difficult to generalize to the real world containing thousands of categories. In this paper, we propose an approach to generate object masks with detailed pixel-level structures/boundaries automatically to enable semantic image segmentation of thousands of targets in the real world without manually labelling. A Guided Filter Network (GFN) is first developed to learn the segmentation knowledge from an existed dataset, and such GFN then transfers the learned segmentation knowledge to generate initial coarse object masks for the target images. These coarse object masks are treated as pseudo labels to self-optimize the GFN iteratively in the target images. Our experiments on six image sets have demonstrated that our proposed approach can generate object masks with detailed pixel-level structures/boundaries, whose quality is comparable to the manually-labelled ones. Our proposed approach also achieves better performance on semantic image segmentation than most existing weakly-supervised, semi-supervised, and domain adaptation approaches under the same experimental conditions. Xiang Zhang 0018, Wanqing Zhao, Wei Zhang 0016, Jinye Peng 0001, Jianping Fan 0001 |
IEEE Trans. Image Process. | 4 |
| 2022 | Spatio-Temporal Pain Estimation Network With Measuring Pseudo Heart Rate GainabstractPain is a significant indicator that shows people are suffering from an unwell experience and its automatic estimation has attracted much interest in recent years. Of late, most estimation methods are designed to capture the dynamic pain information from visual signals while a few physiological-signal based methods can provide extra potential cues to analyze the pain more accurately. However, it is still challenging to capture the physiological data from patients as it requires contact devices and patients’ cooperation. In this paper, we propose to leverage the pseudo physiological information by generating new modal data from the original visual videos and jointly estimating the pain by an end-to-end network. To extract the representations from bi-modal data, we design a spatio-temporal pain estimation network, which employs a dual-branch framework for extracting pain-aware visual and pseudo physiological features separately and fuses the features in a probabilistic way. The inherent vital sign, i.e., heart rate gain (HRG), from pseudo physiological information can be utilized as an auxiliary signal and integrated with the visual pain estimation framework. Moreover, specially-designed 3D convolution filters and attention structures are employed to extract spatio-temporal features for both branches. To use the HRG as an auxiliary way for pain estimation, we propose a probabilistic inference model by jointly considering the visual branch and physiological branch, which makes our model estimate the pain comprehensively. Experiments on two publicly-available datasets show the effectiveness of introducing the pseudo modality, and the proposed method can outperform the state-of-the-art methods. Dong Huang 0003, Xiaoyi Feng, Haixi Zhang, Zitong Yu, Jinye Peng 0001, Guoying Zhao 0001, Zhaoqiang Xia |
IEEE Trans. Multim. | 5 |
| 2021 | Keypoint-Graph-Driven Learning Framework for Object Pose EstimationabstractMany recent 6D pose estimation methods exploited object 3D models to generate synthetic images for training because labels come for free. However, due to the domain shift of data distributions between real images and synthetic images, the network trained only on synthetic images fails to capture robust features in real images for 6D pose estimation. We propose to solve this problem by making the network insensitive to different domains, rather than taking the more difficult route of forcing synthetic images to be similar to real images. Inspired by domain adaption methods, a Domain Adaptive Keypoints Detection Network (DAKDN) including a domain adaption layer is used to minimize the discrepancy of deep features between synthetic and real images. A unique challenge here is the lack of ground truth labels (i.e., keypoints) for real images. Fortunately, the geometry relations between keypoints are invariant under real/synthetic domains. Hence, we propose to use the domain-invariant geometry structure among keypoints as a "bridge" constraint to optimize DAKDN for 6D pose estimation across domains. Specifically, DAKDN employs a Graph Convolutional Network (GCN) block to learn the geometry structure from synthetic images and uses the GCN to guide the training for real images. The 6D poses of objects are calculated using Perspective-n-Point (PnP) algorithm based on the predicted keypoints. Experiments show that our method outperforms state-of-the-art approaches without manual poses labels and competes with approaches using manual poses labels. Shaobo Zhang 0006, Wanqing Zhao, Ziyu Guan, Xianlin Peng, Jinye Peng 0001 |
CVPR | 5 |
| 2021 | Non-contact Pain Recognition from Video Sequences with Remote Physiological Measurements PredictionabstractAutomatic pain recognition is paramount for medical diagnosis and treatment. The existing works fall into three categories: assessing facial appearance changes, exploiting physiological cues, or fusing them in a multi-modal manner. However, (1) appearance changes are easily affected by subjective factors which impedes objective pain recognition. Besides, the appearance-based approaches ignore long-range spatial-temporal dependencies that are important for modeling expressions over time; (2) the physiological cues are obtained by attaching sensors on human body, which is inconvenient and uncomfortable. In this paper, we present a novel multi-task learning framework which encodes both appearance changes and physiological cues in a non-contact manner for pain recognition. The framework is able to capture both local and long-range dependencies via the proposed attention mechanism for the learned appearance representations, which are further enriched by temporally attended physiological cues (remote photoplethysmography, rPPG) that are recovered from videos in the auxiliary task. This framework is dubbed rPPG-enriched Spatio-Temporal Attention Network (rSTAN) and allows us to establish the state-of-the-art performance of non-contact pain recognition on publicly available pain databases. It demonstrates that rPPG predictions can be used as an auxiliary task to facilitate non-contact automatic pain recognition. Ruijing Yang, Ziyu Guan, Zitong Yu, Xiaoyi Feng, Jinye Peng 0001, Guoying Zhao 0001 |
IJCAI | 5 |
| 2021 | Hierarchical bilinear convolutional neural network for image classificationabstractAbstract Image classification is one of the mainstream tasks of computer vision. However, the most existing methods use labels of the same granularity level for training. This leads to ignoring the hierarchy that may help to differentiate different visual objects better. Embedding hierarchical information into the convolutional neural networks (CNNs) can effectively regulate the semantic space and thus reduce the ambiguity of prediction. To this end, a multi‐task learning framework, named as Hierarchical Bilinear Convolutional Neural Network (HB‐CNN), is developed by seamlessly integrating CNNs with multi‐task learning over the hierarchical visual concept structures. Specifically, the labels with a tree structure are used as the supervision to hierarchically train multiple branch networks. In this way, the model can not only learn additional information (e.g. context information) as the coarse‐level category features, but also focus the learned fine‐level category features on the object properties. To smoothly pass hierarchical conceptual information and encourage feature reuse, a connectivity pattern is proposed to connect features at different levels. Furthermore, a bilinear module is embedded to generalise various orderless texture feature descriptors so that our model can capture more discriminative features. The proposed method is extensively evaluated on the CIFAR‐10, CIFAR‐100, and ‘Orchid’ Plant image sets. The experimental results show the effectiveness and superiority of our method. Xiang Zhang 0018, Hangzai Luo, Sheng Zhong 0006, Ziyu Guan, Long Chen 0007, Jinye Peng 0001, Jianping Fan 0001 |
IET Comput. Vis. | 8 |
| 2021 | Learning noise-decoupled affine models for extreme low-light image enhancement
Maomei Liu, Sheng Zhong 0006, Hangzai Luo, Jinye Peng 0001 |
Neurocomputing | 5 |
| 2021 | RBPR: A hybrid model for the new user cold start problem in recommender systems
Junmei Feng, Zhaoqiang Xia, Xiaoyi Feng, Jinye Peng 0001 |
Knowl. Based Syst. | 4 |
| 2021 | RMCNet: Random Multiscale Convolutional Network for Hyperspectral Image ClassificationabstractTo address the limitation of the high-dimensionality features and single spatial scale in the spectral–spatial classification of hyperspectral image (HSI), we propose a random multiscale convolutional network (RMCNet) that combines a multiscale dimensionality reduction module (MDRM) and the RMCNet for improving classification accuracy. The MDRM is based on multiscale superpixel segmentations, which implements dimensionality reduction leading to relieve the Hughes problem and reduce the computation burden in deep learning. Then, the multiscale spectral–spatial features are extracted by the RMCNet to adaptive various complex scenes in HSI. Finally, the multiscale spectral–spatial features act as inputs of support vector machine for classification. In the experiments, three benchmark HSIs are used to evaluate the performance of the proposed method. The experimental results demonstrate that the RMCNet can yield a competitive performance compared with the state-of-the-art methods. Jun Wang 0078, Erlei Zhang, Yongqin Zhang, Jinye Peng 0001 |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2021 | A relic sketch extraction framework based on detail-aware hierarchical deep network
Jinye Peng 0001, Jun Wang 0078, Erlei Zhang, Qunxi Zhang, Yongqin Zhang, Xianlin Peng |
Signal Process. | 1 |
| 2021 | PSMD-Net: A Novel Pan-Sharpening Method Based on a Multiscale Dense NetworkabstractPan sharpening is used to fuse a low-resolution multispectral (MS) image and a high-resolution panchromatic (PAN) image to obtain a high-resolution MS image. This article proposes PSMD-Net, an end-to-end pan-sharpening method based on a multi-scale dense network. A shallow feature extraction layer (SFEL) extracts the shallow features from the original images, and these are used as an input to a global dense feature fusion (GDFF) network to learn the global features for image reconstruction. A multiscale dense block (MDB) is designed to fully extract the spatial and spectral information from the shallow features in the GDFF network. In the proposed network, multiple MDBs are stacked to extract rich, multi-scale dense hierarchical features, and a global dense connection (GDC) is designed to allow direct connections from the state of the current MDB to all subsequent MDBs to extract more advanced features. The extracted hierarchical features are sent to the global feature fusion layer (GFFL) to adaptively learn the global features for image reconstruction. Finally, global residual learning (GRL) is adopted to force the network to pay more attention to the changing part of the image. We perform experiments on simulated and real data from WorldView-2 and WorldView-3 satellites. Visual and quantitative assessment results demonstrate that PSMD-Net yields higher-resolution fusion images than the state-of-the-art methods. Jinye Peng 0001, Lu Liu 0025, Jun Wang 0078, Erlei Zhang, Xuan Zhu 0003, Yongqin Zhang, Jie Feng 0003, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | Deep Multiple Instance Hashing for Fast Multi-Object Image SearchabstractMulti-keyword query is widely supported in text search engines. However, an analogue in image retrieval systems, multi-object query, is rarely studied. Meanwhile, traditional object-based image retrieval methods often involve multiple steps separately. In this work, we propose a weakly-supervised Deep Multiple Instance Hashing (DMIH) approach for multi-object image retrieval. Our DMIH approach, which leverages a popular CNN model to build the end-to-end relation between a raw image and the binary hash codes of its multiple objects, can support multi-object queries effectively and integrate object detection with hashing learning seamlessly. We treat object detection as a binary multiple instance learning (MIL) problem and such instances are automatically extracted from multi-scale convolutional feature maps. We also design a conditional random field (CRF) module to capture both the semantic and spatial relations among different class labels. For hashing training, we sample image pairs to learn their semantic relationships in terms of hash codes of the most probable proposals for owned labels as guided by object predictors. The two objectives benefit each other in a multi-task learning scheme. Finally, a two-level inverted index method is proposed to further speed up the retrieval of multi-object queries. Our DMIH approach outperforms state-of-the-arts on public benchmarks for object-based image retrieval and achieves promising results for multi-object queries. Wanqing Zhao, Ziyu Guan, Hangzai Luo, Jinye Peng 0001, Jianping Fan 0001 |
IEEE Trans. Image Process. | 4 |
| 2020 | Learning Deep Network for Detecting 3D Object Keypoints and 6D PosesabstractThe state-of-art 6D object pose detection methods use convolutional neural networks to estimate objects' 6D poses from RGB images. However, they require huge numbers of images with explicit 3D annotations such as 6D poses, 3D bounding boxes and 3D keypoints, either obtained by manual labeling or inferred from synthetic images generated by 3D CAD models. Manual labeling for a large number of images is a laborious task, and we usually do not have the corresponding 3D CAD models of objects in real environment. In this paper, we develop a keypoint-based 6D object pose detection method (and its deep network) called Object Keypoint based POSe Estimation (OK-POSE). OK-POSE employs relative transformation between viewpoints for training. Specifically, we use pairs of images with object annotation and relative transformation information between their viewpoints to automatically discover objects' 3D keypoints which are geometrically and visually consistent. Then, the 6D object pose can be estimated using a keypoint-based geometric reasoning method with a reference viewpoint. The relative transformation information can be easily obtained from any cheap binocular cameras or most smartphone devices, thus greatly lowering the labeling cost. Experiments have demonstrated that OK-POSE achieves acceptable performance compared to methods relying on the object's 3D CAD model or a great deal of 3D labeling. These results show that our method can be used as a suitable alternative when there are no 3D CAD models or a large number of 3D annotations. Wanqing Zhao, Shaobo Zhang 0006, Ziyu Guan, Wei Zhao 0019, Jinye Peng 0001, Jianping Fan 0001 |
CVPR | 5 |
| 2020 | Aggregating diverse deep attention networks for large-scale plant species identification
Haixi Zhang, Zhenzhong Kuang, Xianlin Peng, Guiqing He, Jinye Peng 0001, Jianping Fan 0001 |
Neurocomputing | 5 |
| 2020 | 6D object pose estimation via viewpoint relation reasoning
Wanqing Zhao, Shaobo Zhang 0006, Ziyu Guan, Hangzai Luo, Jinye Peng 0001, Jianping Fan 0001 |
Neurocomputing | 6 |
| 2020 | Image denoising via structure-constrained low-rank approximation
Yongqin Zhang, Ruiwen Kang, Xianlin Peng, Jun Wang 0078, Jihua Zhu, Jinye Peng 0001, Hangfan Liu |
Neural Comput. Appl. | 6 |
| 2019 | Multilayer feature descriptors fusion CNN models for fine-grained visual recognitionabstractAbstract Fine‐grained image classification is a challenging topic in the field of computer vision. General models based on first‐order local features cannot achieve acceptable performance because the features are not so efficient in capturing fine‐grained difference. A bilinear convolutional neural network (CNN) model exhibits that a second‐order statistical feature is more efficient in capturing fine‐grained difference than a first‐order local feature. However, this framework only considers the extraction of a second‐order feature descriptor, using a single convolutional layer. The potential effective classification features of other convolutional layers are ignored, resulting in loss of recognition accuracy. In this paper, a multilayer feature descriptors fusion CNN model is proposed. It fully considers the second‐order feature descriptors and the first‐order local feature descriptor generated by different layers. Experimental verification was carried out on fine‐grained classification benchmark data sets, CUB‐200‐2011, Stanford Cars, and FGVC‐aircraft. Compared with the bilinear CNN model, the proposed method has improved accuracy by 0.8%, 1.1%, and 5.5%. Compared with the compact bilinear pooling model, there is an accuracy increase of 0.64%, 1.63%, and 1.45%, respectively. In addition, the proposed model effectively uses multiple 1×1 convolution kernels to reduce dimension. The experimental results show that the multilayer low‐dimensional second‐order feature descriptors fusion model has comparable recognition accuracy of the original model. Yong Hou, Hangzai Luo, Wanqing Zhao, Xiang Zhang 0018, Jun Wang 0078, Jinye Peng 0001 |
Comput. Animat. Virtual Worlds | 6 |
| 2019 | Plant recognition via leaf shape and margin features
Xiang Zhang 0018, Wanqing Zhao, Hangzai Luo, Long Chen 0007, Jinye Peng 0001, Jianping Fan 0001 |
Multim. Tools Appl. | 5 |
| 2018 | Incorporating high-level and low-level cues for pain intensity estimationabstractPain is a transient physical reaction that exhibits on human faces. Automatic pain intensity estimation is of great importance in clinical and health-care applications. Pain expression is identified by a set of deformations of facial features. Hence, features are essential for pain estimation. In this paper, we propose a novel method that encodes low-level descriptors and powerful high-level deep features by a weighting process, to form an efficient representation of facial images. To obtain a powerful and compact low-level representation, we explore the way of using second-order pooling over the local descriptors. Instead of direct concatenation, we develop an efficient fusion approach that unites the low-level local descriptors and the high-level deep features. To the best of our knowledge, this is the first approach that incorporates the low-level local statistics together with the high-level deep features in pain intensity estimation. Experiments are evaluated on the benchmark databases of pain. The results demonstrate that the proposed low-to-high-level representation outperforms other methods and achieves promising results. Ruijing Yang, Xiaopeng Hong, Jinye Peng 0001, Xiaoyi Feng, Guoying Zhao 0001 |
ICPR | 3 |
| 2018 | Tag-based Weakly-supervised Hashing for Image RetrievalabstractWe are concerned with using user-tagged images to learn proper hashing functions for image retrieval. The benefits are two-fold: (1) we could obtain abundant training data for deep hashing models; (2) tagging data possesses richer semantic information which could help better characterize similarity relationships between images. However, tagging data suffers from noises, vagueness and incompleteness. Different from previous unsupervised or supervised hashing learning, we propose a novel weakly-supervised deep hashing framework which consists of two stages: weakly-supervised pre-training and supervised fine-tuning. The second stage is as usual. In the first stage, rather than performing supervision on tags, the framework introduces a semantic embedding vector (sem-vector) for each image and performs learning of hashing and sem-vectors jointly. By carefully designing the optimization problem, it can well leverage tagging information and image content for hashing learning. The framework is general and does not depend on specific deep hashing methods. Empirical results on real world datasets show that when it is integrated with state-of-art deep hashing methods, the performance increases by 8-10%. Ziyu Guan, Fei Xie 0007, Wanqing Zhao, Long Chen 0007, Wei Zhao 0019, Jinye Peng 0001 |
IJCAI | 7 |
| 2018 | Dual regularized multi-view non-negative matrix factorization for clustering
Peng Luo 0007, Jinye Peng 0001, Ziyu Guan, Jianping Fan 0001 |
Neurocomputing | 2 |
| 2018 | Locally linear spatial pyramid hash for large-scale image search
Wanqing Zhao, Hangzai Luo, Jinye Peng 0001, Jianping Fan 0001 |
Multim. Tools Appl. | 3 |
| 2017 | Single Image Super-Resolution by Learned Double Sparsity Dictionaries Combining Bootstrapping Method
Na Ai, Jinye Peng 0001, Jun Wang 0078, Lin Wang 0026 |
ICANN (2) | 2 |
| 2017 | Similar Trademark Image Retrieval Integrating LBP and Convolutional Neural Network
Xiaoyi Feng, Zhaoqiang Xia, Shijie Pan, Jinye Peng 0001 |
ICIG (3) | 5 |
| 2017 | Intrinsic Image Decomposition: A Comprehensive Review
Yupeng Ma, Xiaoyi Feng, Xiaoyue Jiang, Zhaoqiang Xia, Jinye Peng 0001 |
ICIG (1) | 5 |
| 2017 | Multi-orientation Scene Text Detection Leveraging Background Suppression
Xihan Wang, Xiaoyi Feng, Zhaoqiang Xia, Jinye Peng 0001, Eric Granger |
ICIG (1) | 4 |
| 2017 | Deep Multiple Instance Hashing for Object-based Image RetrievalabstractMulti-keyword query is widely supported in text search engines. However, an analogue in image retrieval systems, multi-object query, is rarely studied. Meanwhile, traditional object-based image retrieval methods often involve multiple steps separately and need expensive location labeling for detecting objects. In this work, we propose a weakly-supervised Deep Multiple Instance Hashing (DMIH) framework for object-based image retrieval. DMIH integrates object detection and hashing learning on the basis of a popular CNN model to build the end-to-end relation between a raw image and the binary hashing codes of multiple objects in it. Specifically, we cast the object detection of each object class as a binary multiple instance learning problem where instances are object proposals extracted from multi-scale convolutional feature maps. For hashing training, we sample image pairs to learn their semantic relationships in terms of hash codes of the most probable proposals for owned labels as guided by object predictors. The two objectives benefit each other in learning. DMIH outperforms state-of-the-arts on public benchmarks for object-based image retrieval and achieves promising results for multi-object queries. Wanqing Zhao, Ziyu Guan, Hangzai Luo, Jinye Peng 0001, Jianping Fan 0001 |
IJCAI | 4 |
| 2017 | Spatial pyramid deep hashing for large-scale image retrieval
Wanqing Zhao, Hangzai Luo, Jinye Peng 0001, Jianping Fan 0001 |
Neurocomputing | 3 |
| 2017 | An automatic image-text alignment method for large-scale web image retrieval
Baopeng Zhang, Yanyun Qu, Jinye Peng 0001, Jianping Fan 0001 |
Multim. Tools Appl. | 3 |
| 2017 | MapReduce-based clustering for near-duplicate image identification
Wanqing Zhao, Hangzai Luo, Jinye Peng 0001, Jianping Fan 0001 |
Multim. Tools Appl. | 3 |
| 2017 | HD-MTL: Hierarchical Deep Multi-Task Learning for Large-Scale Visual RecognitionabstractIn this paper, a hierarchical deep multi-task learning (HD-MTL) algorithm is developed to support large-scale visual recognition (e.g., recognizing thousands or even tens of thousands of atomic object classes automatically). First, multiple sets of multi-level deep features are extracted from different layers of deep convolutional neural networks (deep CNNs), and they are used to achieve more effective accomplishment of the coarseto- fine tasks for hierarchical visual recognition. A visual tree is then learned by assigning the visually-similar atomic object classes with similar learning complexities into the same group, which can provide a good environment for determining the interrelated learning tasks automatically. By leveraging the inter-task relatedness (inter-class similarities) to learn more discriminative group-specific deep representations, our deep multi-task learning algorithm can train more discriminative node classifiers for distinguishing the visually-similar atomic object classes effectively. Our hierarchical deep multi-task learning (HD-MTL) algorithm can integrate two discriminative regularization terms to control the inter-level error propagation effectively, and it can provide an end-to-end approach for jointly learning more representative deep CNNs (for image representation) and more discriminative tree classifier (for large-scale visual recognition) and updating them simultaneously. Our incremental deep learning algorithms can effectively adapt both the deep CNNs and the tree classifier to the new training images and the new object classes. Our experimental results have demonstrated that our HD-MTL algorithm can achieve very competitive results on improving the accuracy rates for large-scale visual recognition. Jianping Fan 0001, Zhenzhong Kuang, Yu Zheng 0006, Ji Zhang 0005, Jun Yu 0002, Jinye Peng 0001 |
IEEE Trans. Image Process. | 7 |
| 2016 | Spontaneous micro-expression spotting via geometric deformation modeling
Zhaoqiang Xia, Xiaoyi Feng, Jinye Peng 0001, Xianlin Peng, Guoying Zhao 0001 |
Comput. Vis. Image Underst. | 3 |
| 2016 | Single image super-resolution by combining self-learning and example-based learning methods
Na Ai, Jinye Peng 0001, Xuan Zhu 0003, Xiaoyi Feng |
Multim. Tools Appl. | 2 |
| 2015 | Multi-view Semantic Learning for Data Representation
Peng Luo 0007, Jinye Peng 0001, Ziyu Guan, Jianping Fan 0001 |
ECML/PKDD (1) | 2 |
| 2015 | A regularized optimization framework for tag completion and image retrieval
Zhaoqiang Xia, Xiaoyi Feng, Jinye Peng 0001, Jun Wu 0022, Jianping Fan 0001 |
Neurocomputing | 3 |
| 2015 | SISR via trained double sparsity dictionaries
Na Ai, Jinye Peng 0001, Xuan Zhu 0003, Xiaoyi Feng |
Multim. Tools Appl. | 2 |
| 2015 | Automatic tag-to-region assignment via multiple instance learning
Zhaoqiang Xia, Yi Shen 0005, Xiaoyi Feng, Jinye Peng 0001, Jianping Fan 0001 |
Multim. Tools Appl. | 4 |
| 2015 | Cost-sensitive learning of hierarchical tree classifiers for large-scale image classification and novel category detection
Jianping Fan 0001, Ji Zhang 0005, Jinye Peng 0001 |
Pattern Recognit. | 4 |
| 2015 | Hierarchical Learning of Tree Classifiers for Large-Scale Plant Species IdentificationabstractIn this paper, a hierarchical multi-task structural learning algorithm is developed to support large-scale plant species identification, where a visual tree is constructed for organizing large numbers of plant species in a coarse-to-fine fashion and determining the inter-related learning tasks automatically. For a given parent node on the visual tree, it contains a set of sibling coarse-grained categories of plant species or sibling fine-grained plant species, and a multi-task structural learning algorithm is developed to train their inter-related classifiers jointly for enhancing their discrimination power. The inter-level relationship constraint, e.g., a plant image must first be assigned to a parent node (high-level non-leaf node) correctly if it can further be assigned to the most relevant child node (low-level non-leaf node or leaf node) on the visual tree, is formally defined and leveraged to learn more discriminative tree classifiers over the visual tree. Our experimental results have demonstrated the effectiveness of our hierarchical multi-task structural learning algorithm on training more discriminative tree classifiers for large-scale plant species identification. Jianping Fan 0001, Jinye Peng 0001 |
IEEE Trans. Image Process. | 3 |
| 2015 | Hierarchical Classification of Large-Scale Patient Records for Automatic Treatment StratificationabstractIn this paper, a hierarchical learning algorithm is developed for classifying large-scale patient records, e.g., categorizing large-scale patient records into large numbers of known patient categories (i.e., thousands of known patient categories) for automatic treatment stratification. Our hierarchical learning algorithm can leverage tree structure to train more discriminative max-margin classifiers for high-level nodes and control interlevel error propagation effectively. By ruling out unlikely groups of patient categories (i.e., irrelevant high-level nodes) at an early stage, our hierarchical approach can achieve log-linear computational complexity, which is very attractive for big data applications. Our experiments on one specific medical domain have demonstrated that our hierarchical approach can achieve very competitive results on both classification accuracy and computational efficiency as compared with other state-of-the-art techniques. Jinye Peng 0001, Naiquan (Nigel) Zheng, Jianping Fan 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2015 | Multi-View Concept Learning for Data RepresentationabstractReal-world datasets often involve multiple views of data items, e.g., a Web page can be described by both its content and anchor texts of hyperlinks leading to it; photos in Flickr could be characterized by visual features, as well as user contributed tags. Different views provide information complementary to each other. Synthesizing multi-view features can lead to a comprehensive description of the data items, which could benefit many data analytic applications. Unfortunately, the simple idea of concatenating different feature vectors ignores statistical properties of each view and usually incurs the “curse of dimensionality” problem. We propose Multi-view Concept Learning (MCL), a novel nonnegative latent representation learning algorithm for capturing conceptual factors from multi-view data. MCL exploits both multi-view information and label information. The key idea is to learn a common latent space across different views which (1) captures the semantic relationships between data items through graph embedding regularization on labeled items, and (2) allows each latent factor to be associated with a subset of views via sparseness constraints. In this way, MCL could capture flexible conceptual patterns hidden in multi-view features. Experiments on a toy problem and three real-world datasets show that MCL performs well and outperforms baseline methods. Ziyu Guan, Lijun Zhang 0005, Jinye Peng 0001, Jianping Fan 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2014 | Cross-modality based celebrity face naming for news image collections
Xueping Su, Jinye Peng 0001, Xiaoyi Feng, Jun Wu 0022, Jianping Fan 0001 |
Multim. Tools Appl. | 2 |
| 2014 | Integrating bilingual search results for automatic junk image filtering
Chunlei Yang, Jinye Peng 0001, Xiaoyi Feng, Jianping Fan 0001 |
Multim. Tools Appl. | 2 |
| 2013 | Multiple Instance Learning for Automatic Image Annotation
Zhaoqiang Xia, Jinye Peng 0001, Xiaoyi Feng, Jianping Fan 0001 |
MMM (2) | 2 |
| 2013 | Cross-modal social image clustering and tag cleansing
Jinye Peng 0001, Yi Shen 0005, Jianping Fan 0001 |
J. Vis. Commun. Image Represent. | 1 |
| 2013 | Multi-label multi-instance learning with missing object tags
Yi Shen 0005, Jinye Peng 0001, Xiaoyi Feng, Jianping Fan 0001 |
Multim. Syst. | 2 |
| 2013 | Image collection summarization via dictionary learning for sparse representation
Chunlei Yang, Jialie Shen 0001, Jinye Peng 0001, Jianping Fan 0001 |
Pattern Recognit. | 3 |
| 2013 | Leveraging Social Supports for Improving Personal Expertise on ACL Reconstruction and RehabilitationabstractIn this paper, a social health support system is developed to assist both ACL (anterior cruciate ligament) patients and clinicians on making better decisions and choices for ACL reconstruction and rehabilitation. By providing a good platform to enable more effective sharing of personal expertise and ACL treatments, our social health support system can allow: (1) ACL patients to identify the best-matching social groups and locate the most suitable expertise for personal health management; and (2) clinicians to easily locate the best-matching ACL patients and learn from well-done treatments, so that they can make better decisions for new ACL patients (who have similar ACL injuries and close social principles with those best-matching ACL patients) and prescribe safer and more effective knee rehabilitation treatments. Jinye Peng 0001, Naiquan (Nigel) Zheng, Jianping Fan 0001 |
IEEE J. Biomed. Health Informatics | 1 |
| 2012 | Image collection summarization via dictionary learning for sparse representationabstractIn this paper, a novel framework is developed to achieve effective summarization of large-scale image collection by treating the problem of automatic image summarization as the problem of dictionary learning for sparse representation, e.g., the summarization task can be treated as a dictionary learning task (i.e., the given image set can be reconstructed sparsely with this dictionary). For image set of a specific category or a mixture of multiple categories, we have built a sparsity model to reconstruct all its images by using a subset of most representative images (i.e., image summary); and we adopted the simulated annealing algorithm to learn such sparse dictionary by minimizing an explicit optimization function. By investigating their reconstruction ability under sparsity constrain and diversity constrain, we have quantitatively measure the performance of various summarization algorithms. Our experimental results have shown that our dictionary learning for sparse representation algorithm can obtain more accurate summary as compared with other baseline algorithms. Chunlei Yang, Jinye Peng 0001, Jianping Fan 0001 |
CVPR | 2 |
| 2012 | Learning inter-related visual dictionary for object recognitionabstractObject recognition is challenging especially when the objects from different categories are visually similar to each other. In this paper, we present a novel joint dictionary learning (JDL) algorithm to exploit the visual correlation within a group of visually similar object categories for dictionary learning where a commonly shared dictionary and multiple category-specific dictionaries are accordingly modeled. To enhance the discrimination of the dictionaries, the dictionary learning problem is formulated as a joint optimization by adding a discriminative term on the principle of the Fisher discrimination criterion. As well as presenting the JDL model, a classification scheme is developed to better take advantage of the multiple dictionaries that have been trained. The effectiveness of the proposed algorithm has been evaluated on popular visual benchmarks. Yi Shen 0005, Jinye Peng 0001, Jianping Fan 0001 |
CVPR | 3 |
| 2012 | Quantitative Characterization of Semantic Gaps for Learning Complexity Estimation and Inference Model SelectionabstractIn this paper, a novel data-driven algorithm is developed for achieving quantitative characterization of the semantic gaps directly in the visual feature space, where the visual feature space is the common space for concept classifier training and automatic concept detection. By supporting quantitative characterization of the semantic gaps, more effective inference models can automatically be selected for concept classifier training by: (1) identifying the image concepts with small semantic gaps (i.e., the isolated image concepts with high inner-concept visual consistency) and training their one-against-all SVM concept classifiers independently; (2) determining the image concepts with large semantic gaps (i.e., the visually-related image concepts with low inner-concept visual consistency) and training their inter-related SVM concept classifiers jointly; and (3) using more image instances to achieve more reliable training of the concept classifiers for the image concepts with large semantic gaps. Our experimental results on NUS-WIDE and ImageNet image sets have obtained very promising results. Jianping Fan 0001, Xiaofei He 0001, Jinye Peng 0001, Ramesh Jain 0001 |
IEEE Trans. Multim. | 4 |
| 2011 | Efficient Large-Scale Image Data Set Exploration: Visual Concept Network and Image Summarization
Chunlei Yang, Xiaoyi Feng, Jinye Peng 0001, Jianping Fan 0001 |
MMM (2) | 3 |
| 2011 | Towards More Precise Social Image-Tag Alignment
Jinye Peng 0001, Xiaoyi Feng, Jianping Fan 0001 |
MMM (2) | 2 |
| 2010 | Constructing distributed hippocratic video databases for privacy-preserving online patient training and counselingabstractDigital video now plays an important role in supporting more profitable online patient training and counseling, and integration of patient training videos from multiple competitive organizations in the health care network will result in better offerings for patients. However, privacy concerns often prevent multiple competitive organizations from sharing and integrating their patient training videos. In addition, patients with infectious or chronic diseases may not want the online patient training organizations to identify who they are or even which video clips they are interested in. Thus, there is an urgent need to develop more effective techniques to protect both video content privacy and access privacy . In this paper, we have developed a new approach to construct a distributed Hippocratic video database system for supporting more profitable online patient training and counseling. First, a new database modeling approach is developed to support concept-oriented video database organization and assign a degree of privacy of the video content for each database level automatically. Second, a new algorithm is developed to protect the video content privacy at the level of individual video clip by filtering out the privacy-sensitive human objects automatically. In order to integrate the patient training videos from multiple competitive organizations for constructing a centralized video database indexing structure, a privacy-preserving video sharing scheme is developed to support privacy-preserving distributed classifier training and prevent the statistical inferences from the videos that are shared for cross-validation of video classifiers. Our experiments on large-scale video databases have also provided very convincing results. Jinye Peng 0001, Noboru Babaguchi, Hangzai Luo, Yuli Gao, Jianping Fan 0001 |
IEEE Trans. Inf. Technol. Biomed. | 1 |
| 2009 | An Interactive Approach for Filtering Out Junk Images From Keyword-Based Google Search ResultsabstractThe keyword-based Google images search engine is now becoming very popular for online image search. Unfortunately, only the text terms that are explicitly or implicitly linked with the images are used for image indexing but the associated text terms may not have exact correspondence with the underlying image semantics, thus the keyword-based Google images search engine may return large amounts of junk images which are irrelevant to the given keyword-based queries. Based on this observation, we have developed an interactive approach to filter out the junk images from keyword-based Google images search results and our approach consists of the following major components. a) A kernel-based image clustering technique is developed to partition the returned images into multiple clusters and outliers. b) Hyperbolic visualization is incorporated to display large amounts of returned images according to their nonlinear visual similarity contexts, so that users can assess the relevance between the returned images and their real query intentions interactively and select one or multiple images to express their query intentions and personal preferences precisely. c) An incremental kernel learning algorithm is developed to translate the users' query intentions and personal preferences for updating the mixture-of-kernels and generating better hypotheses to achieve more accurate clustering of the returned images and filter out the junk images more effectively. Experiments on diverse keyword-based queries from Google images search engine have obtained very positive results. Our junk image filtering system is released for public evaluation at: http://www.cs.uncc.edu/~jfan/google-demo/. Yuli Gao, Jinye Peng 0001, Hangzai Luo, Daniel A. Keim, Jianping Fan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2008 | Incorporating feature hierarchy and boosting to achieve more effective classifier training and concept-oriented video summarization and skimmingabstractFor online medical education purposes, we have developed a novel scheme to incorporate the results of semantic video classification to select the most representative video shots for generating concept-oriented summarization and skimming ofsurgery education videos. First, salient objects are used as the video patterns for feature extraction to achieve a good representation of the intermediate video semantics. The salient objects are defined as the salient video compounds that can be used to characterize the most significant perceptual properties of the corresponding real world physical objects in a video, and thus the appearances of such salient objects can be used to predict the appearances of the relevant semantic video concepts in a specific video domain. Second, a novelmulti-modal boostingalgorithm is developed to achieve more reliable video classifier training by incorporating feature hierarchy and boosting to dramatically reduce both the training cost and the size of training samples, thus it can significantly speed up SVM (support vector machine) classifier training. In addition, the unlabeled samples are integrated to reduce the human efforts on labeling large amount of training samples. Finally, the results of semantic video classification are incorporated to enable concept-oriented video summarization and skimming. Experimental results in a specific domain ofsurgery education videosare provided. Hangzai Luo, Yuli Gao, Xiangyang Xue 0001, Jinye Peng 0001, Jianping Fan 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |