EDBT 2026 Demo / reviewers in the wild / expert
Jinxing Li 0003
dblp:119/7206-3
· DBLP profile ↗
75ranked-venue papers
16as first author
55since 2021 · last 2026
0000-0001-5156-0305ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 35 · 7 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 35 · 7 first-author · 27 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Computer networks · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Domain Information Removal With Decision Region Enlargement for Unseen Conditions Fault DiagnosisabstractMost existing domain generalization methods for fault diagnosis focus on extracting domain-invariant features from multiple source domains. However, the lack of target data severely restricts these domain-invariant features to only the source domain distribution, leading to inadequate generalization performance under unseen working conditions. To address this critical limitation, we propose a novel method named domain information removal with decision region enlargement. Specifically, for domain information elimination, we design dual encoders and dual classifiers to separately extract and classify fault-related features and domain-related features. A distribution discriminator is then introduced to minimize the mutual information between fault-features and domain-features, thereby yielding purer fault-discriminative representations. To further enhance the discriminability of fault features, learnable comparative anchors are employed to strengthen intra-class compactness and inter-class separability. This process clarifies the decision region for each fault category, enabling more robust adaptation to accurate identification of diverse faults in the target domain. Extensive experiments demonstrate the superiority of our proposed method. Yu Gao 0024, Zhanpei Zhang, Shilong Sun 0001, Jinxing Li 0003, Guangming Lu 0002 |
IEEE Internet Things J. | 4 |
| 2026 | Multi -directional decision fusion for black-box source-free anomaly detection
Yu Gao 0024, Shilong Sun 0001, Zhanpei Zhang, Jinxing Li 0003, Guangming Lu 0002 |
Pattern Recognit. | 4 |
| 2026 | Hierarchical Multi-Criteria Representation Fusion for Robust Incomplete Multimodal Sentiment Analysis
Yijing Dai, Yingjian Li 0001, Jinxing Li 0003, Guangming Lu 0002 |
IEEE Trans. Affect. Comput. | 3 |
| 2026 | Pseudocylindrical Convolutions for Learned Omnidirectional Image CompressionabstractEquirectangular projection (ERP) is a convenient form to store omnidirectional images, but it is neither equal-area nor conformal, creating challenges for subsequent visual communication. When used for image compression, ERP amplifies sampling density and deforms objects near the poles, hindering perceptually optimal bit allocation. Here, we present one of the earliest endeavors to apply deep neural networks to omnidirectional image compression. We first propose parametric pseudocylindrical representations that generalize common pseudocylindrical map projections. A tractable greedy algorithm is introduced to identify (sub-)optimal representation configurations, guided by a proxy objective for rate-distortion performance. We then develop pseudocylindrical convolutions, which can be efficiently implemented by standard convolutions with “pseudocylindrical padding.” To demonstrate the utility of the proposed pseudocylindrical representations and convolutions, we implement an end-to-end omnidirectional image compression method, consisting of an analysis transform, a uniform quantizer, a synthesis transform, and an entropy model. Experiments show that our optimized method achieves consistently better rate-distortion performance compared to the state-of-the-art. Mu Li 0005, Kede Ma, Jinxing Li 0003, David Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | IMPRESS: Incomplete Human Motion Prediction via Motion Recovery and Structural-Semantic FusionabstractHuman motion prediction is a key task in computer vision and human-robot interaction, which has received much attention in recent years. However, existing approaches suffer from two issues: 1) They typically rely only on complete data and overlook real-world challenges such as missing observations. 2) Recent works fail to capture the diverse relations among body parts in different action categories, which limits their prediction performance. To address the above problems, we propose a novel Incomplete human Motion Prediction method through motion Re covery and Structure-Semantic fusion (IMPRESS). Specifically, for motion recovery, we introduce a wavelet-based self-attention module. It captures motion details from high-frequency features and extracts global trends from low-frequency components. To enhance the relations among different body parts, we design a structure-semantic fusion graph convolutional network. Moreover, we employ a dual-channel sliding window attention mechanism to capture motion periodicity, enabling smoother predictions. Extensive experiments on two benchmark datasets (Human3.6M, CMU-MoCap) demonstrate that IMPRESS achieves state-of-the-art average prediction performance under both complete and incomplete observations. Jinxing Li 0003, Jie Wen 0001, Yong Xu 0001 |
IEEE Trans. Image Process. | 3 |
| 2026 | Soft Supervision-Guided Spatial-Temporal Refinement Network for Video-Based Visible-Infrared Person Re-IdentificationabstractThanks to automatic switch between visible and infrared modes, person re-identification (Re-ID) in 24-hour has been possible through cross-modal retrieval. Instead of exploiting still images, video-based cross-modal person Re-ID is studied in this paper. Specifically, a large-scale dataset 'HITSZ-PVCM' is first collected, consisting of as many as 1,681 identities and 839,632 frames. Generally, videos contain much richer pedestrian appearances. However, most existing works only generate temporal representations by whole frames, inevitably losing fine-grained details. Furthermore, training a network by metric losses (e.g., center loss) is a common strategy, while such point-to-point constraints are too strong and limit model generalization due to existing diversity among intra-class samples. Here, we propose a Soft Supervision guided Spatial-Temporal Refinement (S3TR) network to tackle these problems. Specifically, S3TR refines each frame guided by a coarse temporal feature, so that more discriminative features are extracted and transformed to a sequential representation. Followed by a global-local mutual learning module, the modality gap is then erased without losing fine-grained details. Furthermore, we propose a novel soft-clustering center loss to measure intra-/inter-class similarity/dissimilarity in a group-to-group way, efficiently improving model generalization. To the best of our knowledge, HITSZ-PVCM is the largest dataset and S3TR achieves superior performances compared with state-of-the-arts. Jinxing Li 0003, Chuhao Zhou, Huafeng Li 0001, Guangming Lu 0002, Yong Xu 0001, David Zhang 0001 |
IEEE Trans. Image Process. | 1 |
| 2025 | UniFuse: A Unified All-In-One Framework for Multi-Modal Medical Image Fusion Under Diverse Degradations and Misalignments
Dayong Su, Huafeng Li 0001, Jinxing Li 0003, Yu Liu 0023 |
ICCV | 4 |
| 2025 | Latitude-oriented hierarchical enhancement network for omnidirectional image super-resolution
Xin Wang 0160, Jinxing Li 0003, Shiqi Wang 0001, Yong Xu 0001 |
Inf. Process. Manag. | 3 |
| 2025 | Breaking the Paired Sample Barrier in Person Re-Identification: Leveraging Unpaired Samples for Domain GeneralizationabstractDomain generalization (DG) for person re-identification (Re-ID) aims to train models on labeled source domains that generalize well to unseen target domains. However, DG for Re-ID faces a major challenge: existing methods rely solely on labeled paired samples to train DG models and are unable to effectively leverage unpaired samples across cameras. In many cases, cross-camera paired samples are extremely scarce and difficult to annotate. To overcome this limitation, we introduce a novel method specifically tailored for Re-ID. This method leverages cross-camera unpaired samples in model training, thereby reducing the dependence on cross-camera paired samples. We refer to this technique as Unpaired-driven DG (U-DG) person Re-ID. The proposed method leverages a robust image encoder to extract identity-consistent features across various camera views. This capability is further enhanced by integrating a multi-camera person identity classifier, which boosts the encoder’s ability to capture consistent identities, even when viewed from different camera perspectives. To address the scarcity of cross-camera paired samples, we devise a unique model training strategy in our method. Specifically, we use the feature vector from the person identity classifier as a single identity prototype. This prototype serves as a reference for generating identity-related prompts across cameras, effectively compensating for the scarcity of cross-camera paired samples during model training. Additionally, we employ a learnable perturbation prompt to mimic appearance variations exhibited by the same individual across different cameras. Our U-DG offers numerous advantages: it can effectively leverage a large number of unpaired samples for model training, compensating for the scarcity of cross-camera paired samples. Moreover, it does not rely solely on cross-camera paired samples, thereby facilitating the construction of training samples. Experimental results on multiple challenging datasets demonstrate that our approach achieves performance comparable to typical DG person Re-ID, highlighting its feasibility and effectiveness. The source code of our method is available athttps://github.com/lhf12278/DGPS. Huafeng Li 0001, Yaoxin Liu, Jinxing Li 0003, Zhengtao Yu 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | Disentangling Inter- and Intra-Video Relations for Multi-Event Video-Text Retrieval and GroundingabstractVideo-text retrieval aims to precisely search for videos most relevant to text queries within a video corpus. However, existing methods are largely limited to single-text (single-event) queries and are not effective at handling multi-text (multi-event) queries. Furthermore, these methods typically focus solely on retrieval and do not attempt to locate multiple events within the retrieved videos. To address these limitations, our paper proposes a novel method named Disentangling Inter- and Intra-Video Relations, which jointly addresses multi-event video-text retrieval and grounding. This method leverages both inter-video and intra-video event relationships to enhance retrieval and grounding performance. At the retrieval level, we devise a Relational Event-Centric Video-Text Retrieval module based on the principle that comprehensive textual information leads to precise correspondence between text and video. It incorporates event relationship features at different hierarchical levels and exploits the hierarchical structure of video relationships to achieve multi-level contrastive learning between events and videos. This approach enhances the richness, accuracy, and comprehensiveness of event descriptions, improving alignment precision between text and video and enabling effective differentiation among videos. For event grounding, we propose Event Contrast-Driven Video Grounding, which accounts for positional differences among events on the 2D temporal score map and achieves precise grounding of multiple events through divergence learning for their locations. Our solution not only provides efficient text-to-video retrieval but also accurately grounds events within the retrieved videos, addressing the shortcomings of existing methods. Extensive experimental results on the ActivityNet Captions and Charades-STA benchmark datasets demonstrate the superior performance of our method, validating its effectiveness. The innovation of this research lies in introducing a new joint framework for video-text retrieval and multi-event grounding while offering new ideas for further research and applications in related fields. The code is available at https://github.com/X7J92/MVT-RG. Mengzhao Wang 0002, Huafeng Li 0001, Jinxing Li 0003, Dapeng Tao, Zhengtao Yu 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | Caption Assisted Multimodal Large Language Model for Video Moment RetrievalabstractMultimodal Large Language Models (MLLMs) have demonstrated significant potential across various multimodal tasks, including retrieval, summarization, and reasoning. However, it remains a substantial challenge for MLLMs to understand and precisely retrieve specific moments from a video, which require fine-grained spatial and temporal understanding of a video. To overcome this, we propose the Caption Assisted MLLM from Coarse to finE (CALCE), a novel two-stage framework designed for enhanced moment retrieval. Our pipeline begins with a first stage where captions extracted from the audio are utilized to assist the MLLM to provide a robust foundation for precise moment retrieval. To efficiently manage memory consumption from this additional data, a clustering algorithm is applied to the sparsely sampled video frames, categorizing them into key frames and non-key frames. The second stage focuses on recalling missed moments and achieving more fine-grained moment boundaries by adopting a higher sampling rate. In this process, predictions from the first stage cast votes for their correlated densely sampled frames, thereby filtering out less relevant frames. By repeating the process of the first stage with these selected frames, CALCE progressively retrieves video moments from coarse to precise. Experiments on QVHighlights and Charades-STA demonstrate the effectiveness of CALCE, which outperforms existing state-of-the-art methods. The code is available at https://github.com/tjhd1475/CALCE. Peiyu Xie, Jinxing Li 0003, Guangming Lu 0002, Yong Xu 0001, David Zhang 0001 |
IEEE Trans. Image Process. | 2 |
| 2025 | Dual-Task Mutual Reinforcing Embedded Joint Video Paragraph Retrieval and GroundingabstractVideo Paragraph Grounding (VPG) aims to precisely locate the most appropriate moments within a video that are relevant to a given textual paragraph query. However, existing methods typically rely on large-scale annotated temporal labels and assume that the correspondence between videos and paragraphs is known. This is impractical in real-world applications, as constructing temporal labels requires significant labor costs, and the correspondence is often unknown. To address this issue, we propose a Dual-task Mutual Reinforcing Embedded Joint Video Paragraph Retrieval and Grounding method (DMR-JRG). In this method, retrieval and grounding tasks are mutually reinforced rather than being treated as separate issues. DMR-JRG mainly consists of two branches: a retrieval branch and a grounding branch. The retrieval branch uses inter-video contrastive learning to roughly align the global features of paragraphs and videos, reducing modality differences and constructing a coarse-grained feature space to break free from the need for correspondence between paragraphs and videos. Additionally, this coarse-grained feature space further facilitates the grounding branch in extracting fine-grained contextual representations. In the grounding branch, we achieve precise cross-modal matching and grounding by exploring the consistency between local, global, and temporal dimensions of video segments and textual paragraphs. By synergizing these dimensions, we construct a fine-grained feature space for video and textual features, greatly reducing the need for large-scale annotated temporal labels. Meanwhile, we design a grounding reinforcement retrieval module (GRRM) that brings the coarse-grained feature space of the retrieval branch closer to the fine-grained feature space of the grounding branch, thereby reinforcing retrieval branch through grounding branch, and finally achieving mutual reinforcement between tasks. Extensive experiments on three challenging datasets demonstrate the effectiveness of our proposed method. The code is available athttps://github.com/X7J92/DMR-JRG. Mengzhao Wang 0002, Huafeng Li 0001, Jinxing Li 0003, Minghong Xie, Dapeng Tao |
IEEE Trans. Multim. | 4 |
| 2025 | Focus Affinity Perception and Super-Resolution Embedding for Multifocus Image FusionabstractDespite the fact that there is a remarkable achievement on multifocus image fusion, most of the existing methods only generate a low-resolution image if the given source images suffer from low resolution. Obviously, a naive strategy is to independently conduct image fusion and image super-resolution. However, this two-step approach would inevitably introduce and enlarge artifacts in the final result if the result from the first step meets artifacts. To address this problem, in this article, we propose a novel method to simultaneously achieve image fusion and super-resolution in one framework, avoiding step-by-step processing of fusion and super-resolution. Since a small receptive field can discriminate the focusing characteristics of pixels in detailed regions, while a large receptive field is more robust to pixels in smooth regions, a subnetwork is first proposed to compute the affinity of features under different types of receptive fields, efficiently increasing the discriminability of focused pixels. Simultaneously, in order to prevent from distortion, a gradient embedding-based super-resolution subnetwork is also proposed, in which the features from the shallow layer, the deep layer, and the gradient map are jointly taken into account, allowing us to get an upsampled image with high resolution. Compared with the existing methods, which implemented fusion and super-resolution independently, our proposed method directly achieves these two tasks in a parallel way, avoiding artifacts caused by the inferior output of image fusion or super-resolution. Experiments conducted on the real-world dataset substantiate the superiority of our proposed method compared with state of the arts. Huafeng Li 0001, Jinxing Li 0003, Yu Liu 0023, Guangming Lu 0002, Yong Xu 0001, Zhengtao Yu 0001, David Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Prototype-Guided Dual-Transformer Reasoning for Video Individual CountingabstractVideo Individual Counting (VIC), which focuses on accurately tallying the total number of individuals in a video without duplication, is crucial for urban public space management and densely-populated areas planning. Existing methods suffer from limitations in terms of expensive manual annotation, and the efficiency of location or detection algorithms. In this work, we contribute a novel Prototype-guided Dual-Transformer Reasoning framework, termed PDTR, which takes both similarity and difference of adjacent frames into account to achieve accurate counting in an end-to-end regression manner. Specifically, we first design a multi-receptive field feature fusion module to acquire initial comprehensive representations. Subsequently, the dynamic prototype generation module memorizes consistent representations of similar information to generate prototypes. Additionally, to further dig out the shared and private features from different frames, a prototype cross-guided decoder and a privacy-decoupling module are designed. Extensive experiments conducted on two existing VIC datasets, consistently demonstrate the superiority of PDTR over state-of-the-art baselines. Yishu Liu 0001, Huafeng Li 0001, Jinxing Li 0003, Guangming Lu 0002 |
ACM Multimedia | 4 |
| 2024 | Optimal Transport-based Labor-free Text Prompt Modeling for Sketch Re-identificationabstractSketch Re-identification (Sketch Re-ID),
which aims to retrieve target person from an image gallery based on a sketch query, is crucial for criminal investigation, law enforcement, and missing person searches.
Existing methods aim to alleviate the modality gap by employing semantic metrics constraints or auxiliary modal guidance. However, they incur expensive labor costs and inevitably omit fine-grained modality-consistent information due to the abstraction of sketches.
To address this issue, this paper proposes a novel $\textit{Optimal Transport-based Labor-free Text Prompt Modeling}$ (OLTM) network, which hierarchically extracts coarse- and fine-grained similarity representations guided by textual semantic information without any additional annotations.
Specifically, multiple target attributes are flexibly obtained by a pre-trained visual question answering (VQA) model. Subsequently, a text prompt reasoning module employs learnable prompt strategy and optimal transport algorithm to extract discriminative global and local text representations, which serve as a bridge for hierarchical and multi-granularity modal alignment between sketch and image modalities.
Additionally, instead of measuring the similarity of two samples by only computing their distance, a novel triplet assignment loss is further proposed, in which the whole data distribution also contributes to optimizing the inter/intra-class distances. Extensive experiments conducted on two public benchmarks consistently demonstrate the robustness and superiority of our OLTM over state-of-the-art methods. Tingting Ren, Jie Wen 0001, Jinxing Li 0003 |
NeurIPS | 4 |
| 2024 | A coarse-to-fine registration network based on affine transformation and multi-scale pyramid
Dongming Li 0002, Yingjian Li 0001, Jinxing Li 0003, Guangming Lu 0002 |
Expert Syst. Appl. | 3 |
| 2024 | Multi-modal graph context extraction and consensus-aware learning for emotion recognition in conversationabstractMulti-modal emotion recognition in conversation is challenging because of the difficulty to jointly leverage the information from heterogeneous text, acoustic, and visual modalities . Recent context-aware methods usually design a graph structure to model dependencies of utterances and speakers, or integrate Multi-modal information. However, they typically lack a sufficient extraction of unimodal context, and rarely explore the emotion consensus prototypes among different samples with the same label. For solving these problems, in this paper, we propose a Graph Context extraction and Consensus-aware Learning (GCCL) framework to excavate context-sensitive fusion features and simulate the emotion evocation process during the emotion consensus learning. Specifically, GCCL contains a well-designed graph-based module to capture speaker, temporal and modality dependencies and integrate information from different modalities. Then, we design an emotion consensus learning unit to mine the most typical feature of each category in each modality. A speaker-guided contrastive learning loss is further proposed to guarantee the diversity between different individuals and the semantic consistency between distinct modalities. Moreover, we construct a consensus-aware unit with an attention-based memory mechanism to preserve semantic correlations among different samples on the category-level. Extensive experimental results on two conversational datasets demonstrate that the proposed GCCL outperforms the state-of-art methods. Code is available at https://github.com/gityider/GCCL . Yijing Dai, Jinxing Li 0003, Yingjian Li 0001, Guangming Lu 0002 |
Knowl. Based Syst. | 2 |
| 2024 | Omnidirectional image super-resolution via position attention network
Xin Wang 0160, Shiqi Wang 0001, Jinxing Li 0003, Mu Li 0005, Yong Xu 0001 |
Neural Networks | 3 |
| 2024 | Multimodal Decoupled Distillation Graph Neural Network for Emotion Recognition in ConversationabstractGraph Neural Networks (GNNs) have attracted increasing attentions for multimodal Emotion Recognition in Conversation (ERC) due to their good performance in contextual understanding. However, most existing GNN-based methods suffer from two challenges: 1) How to explore and propagate appropriate information in a conversational graph. Typical GNNs in ERC neglect to mine the emotion commonality and discrepancy in the local neighborhood, leading to learn similar embbedings for connected nodes. However, the embeddings of these connected nodes are supposed to be distinguishable as they belong to different speakers with different emotions. 2) Most existing works apply simple concatenation or co-occurrence prior for modality combination, failing to fully capture the emotional information of multiple modalities in relationship modeling. In this paper, we propose a multimodal Decoupled Distillation Graph Neural Network (D2GNN) to address the above challenges. Specifically, D2GNN decouples the input features into emotion-aware and emotion-agnostic ones on the emotion category-level, aiming to capture emotion commonality and implicit emotion information, respectively. Moreover, we design a new message passing mechanism to separately propagate emotion-aware and -agnostic knowledge between nodes according to speaker dependency in two GNN-based modules, exploring the correlations of utterances and alleviating the similarities of embeddings. Furthermore, a multimodal distillation unit is performed to obtain the distinguishable embeddings by aggregating unimodal decoupled features. Experimental results on two ERC benchmarks demonstrate the superiority of the proposed model. Code is available at https://github.com/gityider/D2GNN. Yijing Dai, Yingjian Li 0001, Dongpeng Chen, Jinxing Li 0003, Guangming Lu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | U²-Former: Nested U-Shaped Transformer for Image Restoration via Multi-View Contrastive LearningabstractWhile Transformer has achieved remarkable performance in various high-level vision tasks, it is still challenging to exploit the full potential of Transformer in image restoration. The crux lies in the limited depth of applying Transformer in the typical encoder-decoder framework for image restoration, resulting from heavy self-attention computation load and inefficient communications across different depth (scales) of layers. In this paper, we present a deep and effective Transformer-based network for image restoration, termed as U2-Former, which is able to employ self-attention of Transformer as the core operation for feature learning to perform image restoration in a deep encoding and decoding space. Specifically, it leverages the nested U-shaped structure to facilitate the interactions across different layers with different scales of feature maps. Furthermore, we optimize the computational efficiency for the basic Transformer block by introducing a simple yet effective feature-filtering mechanism to compress the token representation. Apart from the typical supervision ways for image restoration, our U2-Former also performs multi-view contrastive learning, which constructs positive pairs in various aspects, to learn noise-sensitive but content-irrelevant features and further decouple the noise component from the background image. Extensive experiments on various image restoration tasks, including reflection removal, rain streak removal and dehazing respectively, demonstrate the effectiveness of the proposed U2-Former. Xin Feng 0005, Haobo Ji, Wenjie Pei, Jinxing Li 0003, Guangming Lu 0002, David Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Text Position-Aware Pixel Aggregation Network With Adaptive Gaussian Threshold: Detecting Text in the WildabstractOver recent years, deep learning has significantly boosted scene text detection performance, and current segmentation-based scene text detectors can achieve compact bounding boxes for irregular texts. However, it is also challenging to tackle crowded or overlapping texts for these existing methods due to conglutination between adjacent text instances in segmentation results. To address these issues, we propose a more accurate scene text detector, Text Position-Aware Pixel Aggregation Network, termed TPPAN. Specifically, a Gaussian threshold representation is adaptively learned instead of a constant setting in Adaptively Text Kernel Thresholding (ATKT) module to obtain more accurate text kernels. Then Text Position-Aware Region Pixel Aggregation (TPAR-PA) module predicts the text regions in relative positions and generates more accurate text contours. Adequate experiments have demonstrated that the resulting detector has achieved state-of-the-art performance on multi-oriented and curved scene text benchmarks. Jiayu Xu 0002, Ailiang Lin, Jinxing Li 0003, Guangming Lu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Disentanglement Learning With Adaptive Centroid Alignment for Multiple Target Domains Fault DiagnosisabstractMost current domain adaptation methods for fault diagnosis focus on single target domain. However, test data often comes from multiple target domains, as machines work under different operating conditions, subsequently generating a more complex and extensive distribution of target data. Unfortunately, single target domain adaptation methods are not adaptive for multiple target domains adaptation (MTDA), which results in transfer performance degradation. To this end, a novel disentanglement learning with adaptive centroid alignment is proposed for MTDA. Specifically for disentanglement learning, two encoders and two classifiers are constructed independently for fault-related and domain-related feature extractions and classifications. Followed by the dual-adversarial strategy, only fault-related but domain-irrelevant features are extracted. Furthermore, to achieve the category alignment, we also propose an adaptive centroid alignment strategy, so that the feature centroids of the same fault category in different domains are enforced to be close to each other. Extensive experiments demonstrate the superiority of our proposed method compared with other popular approaches. Yu Gao 0024, Xutao Zheng, Jinxing Li 0003, Lijun Zong, Hongpeng Yin, Huafeng Li 0001, Guangming Lu 0002 |
IEEE Trans. Ind. Informatics | 3 |
| 2024 | Learning Content-Weighted Pseudocylindrical Representation for 360° Image CompressionabstractLearned 360° image compression methods using equirectangular projection (ERP) often confront a non-uniform sampling issue, inherent to sphere-to-rectangle projection. While uniformly or nearly uniformly sampling representations, along with their corresponding convolution operations, have been proposed to mitigate this issue, these methods often concentrate solely on uniform sampling rates, thus neglecting the content of the image. In this paper, we urge that different contents within 360° images have varying significance and advocate for the adoption of a content-adaptive parametric representation in 360° image compression, which takes into account both the content and sampling rate. We first introduce the parametric pseudocylindrical representation and corresponding convolution operation, upon which we build a learned 360° image codec. Then, we model the hyperparameter of the representation as the output of a network, derived from the image's content and its spherical coordinates. We treat the optimization of hyperparameters for different 360° images as distinct compression tasks and propose a meta-learning algorithm to jointly optimize the codec and the metaknowledge, i.e., the hyperparameter estimation network. A significant challenge is the lack of a direct derivative from the compression loss to the hyperparameter network. To address this, we present a novel method to relax the rate-distortion loss as a function of the hyperparameters, enabling gradient-based optimization of the metaknowledge. Experimental results on omnidirectional images demonstrate that our method achieves state-of-the-art performance and superior visual quality. Mu Li 0005, Youneng Bao, Xiaohang Sui, Jinxing Li 0003, Guangming Lu 0002, Yong Xu 0001 |
IEEE Trans. Image Process. | 4 |
| 2024 | Discrepancy and Structure-Based Contrast for Test-Time Adaptive RetrievalabstractDomain adaptive hashing has received increasing attention since it is capable of enhancing the performance of retrieval if the target domain for testing meets domain shift. However, owing to data security and transmission constraints nowadays, abundant source data is often not available. Towards this end, this paper investigates a novel yet practical problem named test-time adaptive hashing, which aims to enhance the performance of hashing models without access to the source domain data when tested on the target domain with domain shift. This problem is challenging due to both fugacious domain shift and label scarcity on the target domain. In this paper, we propose a novel hashing approach namedDiscrepancy andStructure-basedContrast (DISC) for effective test-time adaptive retrieval. In particular, DISC first trains the hashing model using the source domain data and stores the distribution of each class in the hidden space. During test-time adaptation, we generate simulated source features based on stored distributions and compare class-specific distributions across domains using maximum mean discrepancy (MMD) to overcome potential domain shift. Furthermore, to tackle the label scarcity, we estimate the graph structure using deep features on the target domain, which guides effective hashing contrastive learning for generating discriminative and domain-invariant hash codes. Extensive experiments on various benchmark datasets validate the superiority of our proposed DISC compared with a range of competing baselines. Zeyu Ma 0001, Yizhi Luo, Xiao Luo 0001, Jinxing Li 0003, Chong Chen 0002, Xian-Sheng Hua 0001, Guangming Lu 0002 |
IEEE Trans. Multim. | 5 |
| 2024 | Logical Relation Inference and Multiview Information Interaction for Domain Adaptation Person Re-IdentificationabstractDomain adaptation person re-identification (Re-ID) is a challenging task, which aims to transfer the knowledge learned from the labeled source domain to the unlabeled target domain. Recently, some clustering-based domain adaptation Re-ID methods have achieved great success. However, these methods ignore the inferior influence on pseudo-label prediction due to the different camera styles. The reliability of the pseudo-label plays a key role in domain adaptation Re-ID, while the different camera styles bring great challenges for pseudo-label prediction. To this end, a novel method is proposed, which bridges the gap of different cameras and extracts more discriminative features from an image. Specifically, an intra-to-intermechanism is introduced, in which samples from their own cameras are first grouped and then aligned at the class level across different cameras followed by our logical relation inference (LRI). Thanks to these strategies, the logical relationship between simple classes and hard classes is justified, preventing sample loss caused by discarding the hard samples. Furthermore, we also present a multiview information interaction (MvII) module that takes features of different images from the same pedestrian as patch tokens, obtaining the global consistency of a pedestrian that contributes to the discriminative feature extraction. Unlike the existing clustering-based methods, our method employs a two-stage framework that generates reliable pseudo-labels from the views of the intracamera and intercamera, respectively, to differentiate the camera styles, subsequently increasing its robustness. Extensive experiments on several benchmark datasets show that the proposed method outperforms a wide range of state-of-the-art methods. The source code has been released at https://github.com/lhf12278/LRIMV. Fan Li 0006, Jinxing Li 0003, Huafeng Li 0001, Bob Zhang 0001, Dapeng Tao, Xinbo Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | HARR: Learning Discriminative and High-Quality Hash Codes for Image RetrievalabstractThis article studies deep unsupervised hashing, which has attracted increasing attention in large-scale image retrieval. The majority of recent approaches usually reconstruct semantic similarity information, which then guides the hash code learning. However, they still fail to achieve satisfactory performance in reality for two reasons. On the one hand, without accurate supervised information, these methods usually fail to produce independent and robust hash codes with semantics information well preserved, which may hinder effective image retrieval. On the other hand, due to discrete constraints, how to effectively optimize the hashing network in an end-to-end manner with small quantization errors remains a problem. To address these difficulties, we propose a novel unsupervised hashing method called HARR to learn discriminative and high-quality hash codes. To comprehensively explore semantic similarity structure, HARR adopts the Winner-Take-All hash to model the similarity structure. Then similarity-preserving hash codes are learned under the reliable guidance of the reconstructed similarity structure. Additionally, we improve the quality of hash codes by a bit correlation reduction module, which forces the cross-correlation matrix between a batch of hash codes under different augmentations to approach the identity matrix. In this way, the generated hash bits are expected to be invariant to disturbances with minimal redundancy, which can be further interpreted as an instantiation of the information bottleneck principle. Finally, for effective hashing network training, we minimize the cosine distances between real-value network outputs and their binary codes for small quantization errors. Extensive experiments demonstrate the effectiveness of our proposed HARR. Zeyu Ma 0001, Siwei Wang 0010, Xiao Luo 0001, Zhonghui Gu, Chong Chen 0002, Jinxing Li 0003, Xian-Sheng Hua 0001, Guangming Lu 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2023 | Differential Enhanced Siamese Segmentation Network for Printed Label Defect DetectionabstractMany vision-based methods have been widely used to detect defects in industrial printed labels. However, most of them still face challenges of detecting unseen defects, low-contrast defects, and false detections caused by artifacts. To address these problems, we propose a differential enhanced Siamese segmentation network (DESS-Net) for defect detection. This method is based on Siamese similarity comparison which has a better generalization ability for unseen defects. Moreover, we introduce the differential feature enhancement (DFE) modules into the Siamese network to focus on multiple differential feature information which contributes to identifying defects and reducing false detections caused by artifacts. Additionally, a multi-scale feature fusion (MFF) module is further designed to fuse multiple low-level differential features, which is conducive to recovering fine boundaries of low-contrast defects. Experimental results show that our DESS-Net outperforms other compared methods. Dongming Li 0002, Yingjian Li 0001, Jinxing Li 0003, Guangming Lu 0002 |
ICIP | 3 |
| 2023 | Video-based Visible-Infrared Person Re-Identification via Style Disturbance Defense and Dual InteractionabstractVideo-based visible-infrared person re-identification (VVI-ReID) aims to retrieve video sequences of the same pedestrian from different modalities. The key of VVI-ReID is to learn discriminative sequence-level representations that are invariant to both intra- and inter-modal discrepancies. However, most works only focus on the elimination of modality-gap while ignore the distractors within the modality. Moreover, existing sequence-level representation learning approaches are limited to a single video, failing to mine the correlations among multiple videos of the same pedestrian. In this paper, we propose a Style Augmentation, Attack and Defense network with Graph-based dual interaction (SAADG) to guarantee the semantic consistency against both intra-modal discrepancies and inter-modal gap. Specifically, we first generate diverse styles for video frames by random style variation in image spaces. Followed by the style attack and defense, the intra- and inter-modal discrepancies are modeled as different types of style disturbance (attack), and our model achieves to keep the id-related content invariant under such attack. Besides, a graph-based dual interaction module is further introduced to fully explore the cross-view and cross-modal correlations among various videos of the same identity, which are then transferred to the sequence-level representations. Extensive experiments on the public SYSU-MM01 and HITSZ-VCM datasets show that our approach achieves the remarkable performance compared with state-of-the-arts. The code is available at https://github.com/ChuhaoZhou99/SAADG_VVIReID. Chuhao Zhou, Jinxing Li 0003, Huafeng Li 0001, Guangming Lu 0002, Yong Xu 0001, Min Zhang 0005 |
ACM Multimedia | 2 |
| 2023 | Deep adaptive hiding network for image hiding using attentive frequency extraction and gradual depth extraction
Le Zhang 0016, Yao Lu 0008, Jinxing Li 0003, Fanglin Chen 0001, Guangming Lu 0002, David Zhang 0001 |
Neural Comput. Appl. | 3 |
| 2023 | Facial Expression Recognition in the Wild Using Multi-Level Features and Attention MechanismsabstractLearning discriminative features is of vital importance for automatic facial expression recognition (FER) in the wild. In this article, we propose a novel Slide-Patch and Whole-Face Attention model with SE blocks (SPWFA-SE), which jointly perceives the discriminative locality characteristics and informative global features of the face for effective FER. Specifically, the well-designed slide patches are proposed to extract local features. Different from the existing methods, our slide patches not only can maintain the information at the edge area of patches, but also do not need to detect facial landmarks. Moreover, to make the model adaptively focus on the distinguishable regions, an attention module is proposed in the patch level to learn the weight of each patch. Furthermore, squeeze-and-excitation blocks are explored in the channel level to learn the weight of each channel. As such, the proposed multi-level feature extraction and attention mechanisms can enhance the representative ability of the learned features. Extensive experiments on five challenging datasets demonstrate that our method can achieve state-of-the-art performance. Cross database experiments on another three databases show the superior generalization performance of our model. Furthermore, complexity analysis results show that our model contains fewer parameters with fast training advantages than other competing models. Yingjian Li 0001, Guangming Lu 0002, Jinxing Li 0003, Zheng Zhang 0006, David Zhang 0001 |
IEEE Trans. Affect. Comput. | 3 |
| 2023 | HIPA: Hierarchical Patch Transformer for Single Image Super ResolutionabstractTransformer-based architectures start to emerge in single image super resolution (SISR) and have achieved promising performance. However, most existing vision Transformer-based SISR methods still have two shortcomings: (1) they divide images into the same number of patches with a fixed size, which may not be optimal for restoring patches with different levels of texture richness; and (2) their position encodings treat all input tokens equally and hence, neglect the dependencies among them. This paper presents a HIPA, which stands for a novel Transformer architecture that progressively recovers the high resolution image using a hierarchical patch partition. Specifically, we build a cascaded model that processes an input image in multiple stages, where we start with tokens with small patch sizes and gradually merge them to form the full resolution. Such a hierarchical patch mechanism not only explicitly enables feature aggregation at multiple resolutions but also adaptively learns patch-aware features for different image regions, e.g., using a smaller patch for areas with fine details and a larger patch for textureless regions. Meanwhile, a new attention-based position encoding scheme for Transformer is proposed to let the network focus on which tokens should be paid more attention by assigning different weights to different tokens, which is the first time to our best knowledge. Furthermore, we also propose a multi-receptive field attention module to enlarge the convolution receptive field from different branches. The experimental results on several public datasets demonstrate the superior performance of the proposed HIPA over previous methods quantitatively and qualitatively. We will share our code and models when the paper is accepted. Yiming Qian, Jinxing Li 0003, Yee-Hong Yang, Feng Wu 0001, David Zhang 0001 |
IEEE Trans. Image Process. | 3 |
| 2023 | From Global to Local: Multi-Patch and Multi-Scale Contrastive Similarity Learning for Unsupervised Defocus Blur DetectionabstractDefocus blur detection (DBD), which aims to detect out-of-focus or in-focus pixels from a single image, has been widely applied to many vision tasks. To remove the limitation on the abundant pixel-level manual annotations, unsupervised DBD has attracted much attention in recent years. In this paper, a novel deep network named Multi-patch and Multi-scale Contrastive Similarity (M2CS) learning is proposed for unsupervised DBD. Specifically, the predicted DBD mask from a generator is first exploited to re-generate two composite images by transporting the estimated clear and unclear areas from the source image to realistic full-clear and full-blurred images, respectively. To encourage these two composite images to be completely in-focus or out-of-focus, a global similarity discriminator is exploited to measure the similarity of each pair in a contrastive way, through which each two positive samples (two clear images or two blurred images) are enforced to be close while each two negative samples (a clear image and a blurred image) are inversely far. Since the global similarity discriminator only focuses on the blur-level of a whole image and there do exist some fail-detected pixels which only cover a small part of areas, a set of local similarity discriminators are further designed to measure the similarity of image patches in multiple scales. Thanks to this joint global and local strategy, as well as the contrastive similarity learning, the two composite images are more efficiently moved to be all-clear or all-blurred. Experimental results on real-world datasets substantiate the superiority of our proposed method both in quantification and visualization. The source code is released at: https://github.com/jerysaw/M2CS. Jinxing Li 0003, Beicheng Liang, Xiangwei Lu, Mu Li 0005, Guangming Lu 0002, Yong Xu 0001 |
IEEE Trans. Image Process. | 1 |
| 2023 | Learning Context-Based Nonlocal Entropy Modeling for Image CompressionabstractThe entropy of the codes usually serves as the rate loss in the recent learned lossy image compression methods. Precise estimation of the probabilistic distribution of the codes plays a vital role in reducing the entropy and boosting the joint rate-distortion performance. However, existing deep learning based entropy models generally assume the latent codes are statistically independent or depend on some side information or local context, which fails to take the global similarity within the context into account and thus hinders the accurate entropy estimation. To address this issue, we propose a special nonlocal operation for context modeling by employing the global similarity within the context. Specifically, due to the constraint of context, nonlocal operation is incalculable in context modeling. We exploit the relationship between the code maps produced by deep neural networks and introduce the proxy similarity functions as a workaround. Then, we combine the local and the global context via a nonlocal attention block and employ it in masked convolutional networks for entropy modeling. Taking the consideration that the width of the transforms is essential in training low distortion models, we finally produce a U-net block in the transforms to increase the width with manageable memory consumption and time complexity. Experiments on Kodak and Tecnick datasets demonstrate the priority of the proposed context-based nonlocal attention block in entropy modeling and the U-net block in low distortion situations. On the whole, our model performs favorably against the existing image compression standards and recent deep image compression models. Mu Li 0005, Kai Zhang 0008, Jinxing Li 0003, Wangmeng Zuo, Radu Timofte, David Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | A Survey on Incomplete Multiview ClusteringabstractConventional multiview clustering seeks to partition data into respective groups based on the assumption that all views are fully observed. However, in practical applications, such as disease diagnosis, multimedia analysis, and recommendation system, it is common to observe that not all views of samples are available in many cases, which leads to the failure of the conventional multiview clustering methods. Clustering on such incomplete multiview data is referred to as incomplete multiview clustering (IMC). In view of the promising application prospects, the research of IMC has noticeable advances in recent years. However, there is no survey to summarize the current progresses and point out the future research directions. To this end, we review the recent studies of IMC. Importantly, we provide some frameworks to unify the corresponding IMC methods and make an in-depth comparative analysis for some representative methods from theoretical and experimental perspectives. Finally, some open problems in the IMC field are offered for researchers. The related codes are released athttps://github.com/DarrenZZhang/Survey_IMC. Jie Wen 0001, Zheng Zhang 0006, Lunke Fei, Bob Zhang 0001, Yong Xu 0001, Zhao Zhang 0001, Jinxing Li 0003 |
IEEE Trans. Syst. Man Cybern. Syst. | 7 |
| 2022 | PPR-Net: Patch-Based Multi-scale Pyramid Registration Network for Defect Detection of Printed Label
Dongming Li 0002, Yingjian Li 0001, Jinxing Li 0003, Guangming Lu 0002 |
ACCV (2) | 3 |
| 2022 | Learning Modal-Invariant and Temporal-Memory for Video-based Visible-Infrared Person Re-IdentificationabstractThanks for the cross-modal retrieval techniques, visible-infrared (RGB-IR) person re-identification (Re-ID) is achieved by projecting them into a common space, allowing person Re-ID in 24-hour surveillance systems. However, with respect to the probe-to- gallery, almost all existing RGB-IR based cross-modal person Re-ID methods focus on image-to-image matching, while the video-to-video matching which contains much richer spatial- and temporal-information remains under-explored. In this paper, we primarily study the video-based cross-modal per-son Re-ID method. To achieve this task, a video-based RGB-IR dataset is constructed, in which 927 valid identities with 463,259 frames and 21,863 tracklets captured by 12 RGB/IR cameras are collected. Based on our constructed dataset, we prove that with the increase of frames in a tracklet, the performance does meet more enhancement, demonstrating the significance of video-to-video matching in RGB-IR person Re-ID. Additionally, a novel method is further proposed, which not only projects two modalities to a modal-invariant subspace, but also extracts the temporal-memory for motion-invariant. Thanks to these two strategies, much better results are achieved on our video-based cross-modal person Re-ID. The code and dataset are released at: https://github.com/VCM-project233/MITML. Jinxing Li 0003, Zeyu Ma 0001, Huafeng Li 0001, Kaixiong Xu, Guangming Lu 0002, David Zhang 0001 |
CVPR | 2 |
| 2022 | Improved Deep Unsupervised Hashing with Fine-grained Semantic Similarity Mining for Multi-Label Image RetrievalabstractIn this paper, we study deep unsupervised hashing, a critical problem for approximate nearest neighbor research. Most recent methods solve this problem by semantic similarity reconstruction for guiding hashing network learning or contrastive learning of hash codes. However, in multi-label scenarios, these methods usually either generate an inaccurate similarity matrix without reflection of similarity ranking or suffer from the violation of the underlying assumption in contrastive learning, resulting in limited retrieval performance. To tackle this issue, we propose a novel method termed HAMAN, which explores semantics from a fine-grained view to enhance the ability of multi-label image retrieval. In particular, we reconstruct the pairwise similarity structure by matching fine-grained patch features generated by the pre-trained neural network, serving as reliable guidance for similarity preserving of hash codes. Moreover, a novel conditional contrastive learning on hash codes is proposed to adopt self-supervised learning in multi-label scenarios. According to extensive experiments on three multi-label datasets, the proposed method outperforms a broad range of state-of-the-art methods. Zeyu Ma 0001, Xiao Luo 0001, Yingjie Chen 0002, Mi-Xiao Hou, Jinxing Li 0003, Minghua Deng, Guangming Lu 0002 |
IJCAI | 5 |
| 2022 | ConTrans: Improving Transformer with Convolutional Attention for Medical Image Segmentation
Ailiang Lin, Jiayu Xu 0002, Jinxing Li 0003, Guangming Lu 0002 |
MICCAI (5) | 3 |
| 2022 | Printed label defect detection using twice gradient matching based on improved cosine similarity measure
Dongming Li 0002, Jinxing Li 0003, Yuanyi Fan, Guangming Lu 0002, Jie Ge |
Expert Syst. Appl. | 2 |
| 2022 | Dual-stream Reciprocal Disentanglement Learning for domain adaptation person re-identification
Huafeng Li 0001, Kaixiong Xu, Jinxing Li 0003, Zhengtao Yu 0001 |
Knowl. Based Syst. | 3 |
| 2022 | Touchless palmprint recognition based on 3D Gabor template and block feature refinement
Zhaoqun Li, Jinxing Li 0003, Wei Jia 0001, David Zhang 0001 |
Knowl. Based Syst. | 4 |
| 2022 | Learning Informative and Discriminative Features for Facial Expression Recognition in the WildabstractThe informativeness and discriminativeness of features collaboratively ensure high-accuracy Facial Expression Recognition (FER) in the wild. Most of existing methods use the single-path deep convolutional neural network with softmax loss for basic FER, while they cannot deal with the challenging situations of the compound FER in the wild, because they fail to learn informative and discriminative features in a targeted manner. To this end, we present an Informative and Discriminative Feature Learning (IDFL) framework that consists of two key components: the Multi-Path Attention Convolutional Neural Network (MPACNN) and Balanced Separate loss (BS loss), for both basic and compound high-accuracy FER in the wild. Specifically, MPACNN leverages different paths to learn diverse features. These features are then adaptively fused into informative ones via an attention module, such that the model can adequately capture detailed information for both basic and compound FER. The BS loss maximizes the inter-class distance of features and minimizes the intra-class one. In this way, the features are discriminative enough for high-accuracy FER in the wild. Particularly, the BS loss is invoked as the objective function of MPACNN, so the model can learn informative and discriminative features at the same time, yielding better performance. Seven databases are utilized to evaluate the proposed method, and the results demonstrate that our method achieves state-of-the-art performance on both basic and compound expressions with good generalization ability. Moreover, our model contains fewer parameters and can be trained faster than other related models. Yingjian Li 0001, Yao Lu 0008, Bingzhi Chen, Zheng Zhang 0006, Jinxing Li 0003, Guangming Lu 0002, David Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Multiscale Conditional Regularization for Convolutional Neural NetworksabstractWith the increased model size of convolutional neural networks (CNNs), overfitting has become the main bottleneck to further improve the performance of networks. Currently, the weighting regularization methods have been proposed to address the overfitting problem and they perform satisfactorily. Since these regularization methods cannot be used in all the networks and they are usually not flexible enough in different phases of the training and test processes, this article proposes a multiscale conditional (MSC) regularization method. MSC divides the intermediate features into different scales and then generates new data for each scale features, respectively. In addition, the new data are generated by employing the information from two conditions: 1) each sample feature and 2) each layer pattern. Finally, a self-identity structure is proposed to supplement the features with the generated data. Therefore, MSC can adaptively and efficiently generate much finer and individualized data to make the entire regularization more flexible. Furthermore, MSC is more general and can be applied to all kinds of networks through the proposed self-identity structure. The experimental results on all the benchmark datasets showed that the proposed MSC regularization method achieves the best performances in all the networks. Yao Lu 0008, Guangming Lu 0002, Jinxing Li 0003, Yuanrong Xu, Zheng Zhang 0006, David Zhang 0001 |
IEEE Trans. Cybern. | 3 |
| 2022 | Addi-Reg: A Better Generalization-Optimization Tradeoff Regularization Method for Convolutional Neural NetworksabstractIn convolutional neural networks (CNNs), generating noise for the intermediate feature is a hot research topic in improving generalization. The existing methods usually regularize the CNNs by producing multiplicative noise (regularization weights), called multiplicative regularization (Multi-Reg). However, Multi-Reg methods usually focus on improving generalization but fail to jointly consider optimization, leading to unstable learning with slow convergence. Moreover, Multi-Reg methods are not flexible enough since the regularization weights are generated from a definite manual-design distribution. Besides, most popular methods are not universal enough, because these methods are only designed for the residual networks. In this article, we, for the first time, experimentally and theoretically explore the nature of generating noise in the intermediate features for popular CNNs. We demonstrate that injecting noise in the feature space can be transformed to generating noise in the input space, and these methods regularize the networks in a Mini-batch in Mini-batch (MiM) sampling manner. Based on these observations, this article further discovers that generating multiplicative noise can easily degenerate the optimization due to its high dependence on the intermediate feature. Based on these studies, we propose a novel additional regularization (Addi-Reg) method, which can adaptively produce additional noise with low dependence on intermediate feature in CNNs by employing a series of mechanisms. Particularly, these well-designed mechanisms can stabilize the learning process in training, and our Addi-Reg method can pertinently learn the noise distributions for every layer in CNNs. Extensive experiments demonstrate that the proposed Addi-Reg method is more flexible and universal, and meanwhile achieves better generalization performance with faster convergence against the state-of-the-art Multi-Reg methods. Yao Lu 0008, Zheng Zhang 0006, Guangming Lu 0002, Yicong Zhou, Jinxing Li 0003, David Zhang 0001 |
IEEE Trans. Cybern. | 5 |
| 2022 | TDPN: Texture and Detail-Preserving Network for Single Image Super-ResolutionabstractSingle image super-resolution (SISR) using deep convolutional neural networks (CNNs) achieves the state-of-the-art performance. Most existing SISR models mainly focus on pursuing high peak signal-to-noise ratio (PSNR) and neglect textures and details. As a result, the recovered images are often perceptually unpleasant. To address this issue, in this paper, we propose a texture and detail-preserving network (TDPN), which focuses not only on local region feature recovery but also on preserving textures and details. Specifically, the high-resolution image is recovered from its corresponding low-resolution input in two branches. First, a multi-reception field based branch is designed to let the network fully learn local region features by adaptively selecting local region features in different reception fields. Then, a texture and detail-learning branch supervised by the textures and details decomposed from the ground-truth high resolution image is proposed to provide additional textures and details for the super-resolution process to improve the perceptual quality. Finally, we introduce a gradient loss into the SISR field and define a novel hybrid loss to strengthen boundary information recovery and to avoid overly smooth boundary in the final recovered high-resolution image caused by using only the MAE loss. More importantly, the proposed method is model-agnostic, which can be applied to most off-the-shelf SISR networks. The experimental results on public datasets demonstrate the superiority of our TDPN on most state-of-the-art SISR methods in PSNR, SSIM and perceptual quality. We will share our code on https://github.com/tocaiqing/TDPN. Jinxing Li 0003, Huafeng Li 0001, Yee-Hong Yang, Feng Wu 0001, David Zhang 0001 |
IEEE Trans. Image Process. | 2 |
| 2022 | AVLSM: Adaptive Variational Level Set Model for Image Segmentation in the Presence of Severe Intensity Inhomogeneity and High NoiseabstractIntensity inhomogeneity and noise are two common issues in images but inevitably lead to significant challenges for image segmentation and is particularly pronounced when the two issues simultaneously appear in one image. As a result, most existing level set models yield poor performance when applied to this images. To this end, this paper proposes a novel hybrid level set model, named adaptive variational level set model (AVLSM) by integrating an adaptive scale bias field correction term and a denoising term into one level set framework, which can simultaneously correct the severe inhomogeneous intensity and denoise in segmentation. Specifically, an adaptive scale bias field correction term is first defined to correct the severe inhomogeneous intensity by adaptively adjusting the scale according to the degree of intensity inhomogeneity while segmentation. More importantly, the proposed adaptive scale truncation function in the term is model-agnostic, which can be applied to most off-the-shelf models and improves their performance for image segmentation with severe intensity inhomogeneity. Then, a denoising energy term is constructed based on the variational model, which can remove not only common additive noise but also multiplicative noise often occurred in medical image during segmentation. Finally, by integrating the two proposed energy terms into a variational level set framework, the AVLSM is proposed. The experimental results on synthetic and real images demonstrate the superiority of AVLSM over most state-of-the-art level set models in terms of accuracy, robustness and running time. Yiming Qian, Sanping Zhou, Jinxing Li 0003, Yee-Hong Yang, Feng Wu 0001, David Zhang 0001 |
IEEE Trans. Image Process. | 4 |
| 2022 | End-to-End Optimized 360° Image CompressionabstractThe 360° image that offers a 360-degree scenario of the world is widely used in virtual reality and has drawn increasing attention. In 360° image compression, the spherical image is first transformed into a planar image with a projection such as equirectangular projection (ERP) and then saved with the existing codecs. The ERP images that represent different circles of latitude with the same number of pixels suffer from the unbalance sampling problem, resulting in inefficiency using planar compression methods, especially for the deep neural network (DNN) based codecs. To tackle this problem, we introduce a latitude adaptive coding scheme for DNNs by allocating variant numbers of codes for different regions according to the latitude on the sphere. Specifically, taking both the number of allocated codes for each region and their entropy into consideration, we introduce a flexible regional adaptive rate loss for region-wise rate controlling. Latitude adaptive constraints are then introduced to prevent spending too many codes on the over-sampling regions. Furthermore, we introduce viewport-based distortion loss by calculating the average distortion on a set of viewports. We optimize and test our model on a large 360° dataset containing 19,790 images collected from the Internet. The experiment results demonstrate the superiority of the proposed latitude adaptive coding scheme. On the whole, our model outperforms the existing image compression standards, including JPEG, JPEG2000, HEVC Intra Coding, and VVC Intra Coding, and helps to save around 15% bits compared to the baseline learned image compression model for planar images. Mu Li 0005, Jinxing Li 0003, Shuhang Gu, Feng Wu 0001, David Zhang 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | BPFNet: A Unified Framework for Bimodal Palmprint Alignment and Fusion
Zhaoqun Li, Jinxing Li 0003, David Zhang 0001 |
ICONIP (6) | 4 |
| 2021 | Highly shared Convolutional Neural Networks
Yao Lu 0008, Guangming Lu 0002, Yicong Zhou, Jinxing Li 0003, Yuanrong Xu, David Zhang 0001 |
Expert Syst. Appl. | 4 |
| 2021 | Fully shared convolutional neural networks
Yao Lu 0008, Guangming Lu 0002, Jinxing Li 0003, Zheng Zhang 0006, Yuanrong Xu |
Neural Comput. Appl. | 3 |
| 2021 | Consensus guided incomplete multi-view spectral clustering
Jie Wen 0001, Huijie Sun, Lunke Fei, Jinxing Li 0003, Zheng Zhang 0006, Bob Zhang 0001 |
Neural Networks | 4 |
| 2021 | Shared Linear Encoder-Based Multikernel Gaussian Process Latent Variable Model for Visual ClassificationabstractMultiview learning has been widely studied in various fields and achieved outstanding performances in comparison to many single-view-based approaches. In this paper, a novel multiview learning method based on the Gaussian process latent variable model (GPLVM) is proposed. In contrast to existing GPLVM methods which only assume that there are transformations from the latent variable to the multiple observed inputs, our proposed method simultaneously takes a back constraint into account, encoding multiple observations to the latent variable by enjoying the Gaussian process (GP) prior. Particularly, to overcome the difficulty of the covariance matrix calculation in the encoder, a linear projection is designed to map different observations to a consistent subspace first. The obtained variable in this subspace is then projected to the latent variable in the manifold space with the GP prior. Furthermore, different from most GPLVM methods which strongly assume that the covariance matrices follow a certain kernel function, for example, radial basis function (RBF), we introduce a multikernel strategy to design the covariance matrix, being more reasonable and adaptive for the data representation. In order to apply the presented approach to the classification, a discriminative prior is also embedded to the learned latent variables to encourage samples belonging to the same category to be close and those belonging to different categories to be far. Experimental results on three real-world databases substantiate the effectiveness and superiority of the proposed method compared with state-of-the-art approaches. Jinxing Li 0003, Guangming Lu 0002, Bob Zhang 0001, Jane You, David Zhang 0001 |
IEEE Trans. Cybern. | 1 |
| 2021 | Layer-Output Guided Complementary Attention Learning for Image Defocus Blur DetectionabstractDefocus blur detection (DBD), which has been widely applied to various fields, aims to detect the out-of-focus or in-focus pixels from a single image. Despite the fact that the deep learning based methods applied to DBD have outperformed the hand-crafted feature based methods, the performance cannot still meet our requirement. In this paper, a novel network is established for DBD. Unlike existing methods which only learn the projection from the in-focus part to the ground-truth, both in-focus and out-of-focus pixels, which are completely and symmetrically complementary, are taken into account. Specifically, two symmetric branches are designed to jointly estimate the probability of focus and defocus pixels, respectively. Due to their complementary constraint, each layer in a branch is affected by an attention obtained from another branch, effectively learning the detailed information which may be ignored in one branch. The feature maps from these two branches are then passed through a unique fusion block to simultaneously get the two-channel output measured by a complementary loss. Additionally, instead of estimating only one binary map from a specific layer, each layer is encouraged to estimate the ground truth to guide the binary map estimation in its linked shallower layer followed by a top-to-bottom combination strategy, gradually exploiting the global and local information. Experimental results on released datasets demonstrate that our proposed method remarkably outperforms state-of-the-art algorithms. Jinxing Li 0003, Lingxiao Yang, Shuhang Gu, Guangming Lu 0002, Yong Xu 0001, David Zhang 0001 |
IEEE Trans. Image Process. | 1 |
| 2021 | Harmonization Shared Autoencoder Gaussian Process Latent Variable Model With Relaxed Hamming DistanceabstractMultiview learning has shown its superiority in visual classification compared with the single-view-based methods. Especially, due to the powerful representation capacity, the Gaussian process latent variable model (GPLVM)-based multiview approaches have achieved outstanding performances. However, most of them only follow the assumption that the shared latent variables can be generated from or projected to the multiple observations but fail to exploit the harmonization in the back constraint and adaptively learn a classifier according to these learned variables, which would result in performance degradation. To tackle these two issues, in this article, we propose a novel harmonization shared autoencoder GPLVM with a relaxed Hamming distance (HSAGP-RHD). Particularly, an autoencoder structure with the Gaussian process (GP) prior is first constructed to learn the shared latent variable for multiple views. To enforce the agreement among various views in the encoder, a harmonization constraint is embedded into the model by making consistency for the view-specific similarity. Furthermore, we also propose a novel discriminative prior, which is directly imposed on the latent variable to simultaneously learn the fused features and adaptive classifier in a unit model. In detail, the centroid matrix corresponding to the centroids of different categories is first obtained. A relaxed Hamming distance (RHD)-based measurement is subsequently presented to measure the similarity and dissimilarity between the latent variable and centroids, not only allowing us to get the closed-form solutions but also encouraging the points belonging to the same class to be close, while those belonging to different classes to be far. Due to this novel prior, the category of the out-of-sample is also allowed to be simply assigned in the testing phase. Experimental results conducted on three real-world data sets demonstrate the effectiveness of the proposed method compared with state-of-the-art approaches. Jinxing Li 0003, Bob Zhang 0001, Guangming Lu 0002, Yong Xu 0001, Feng Wu 0001, David Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2021 | Fast Pore Comparison for High Resolution Fingerprint Images Based on Multiple Co-Occurrence Descriptors and Local Topology SimilaritiesabstractPore-based fingerprint recognition has been researched for decades. Many algorithms have been proposed to improve the recognition accuracy of the system. However, the accuracies are always improved at the cost of speed. This article proposes a novel method to compare the pores in high-resolution fingerprint images using the popular coarse-to-fine strategy. A multiple spatial pairwise local co-occurrence descriptor is proposed to improve the calculation of the similarities between pores. It calculates multiple local co-occurrence statistics for each pore using its neighbors. The proposed method can establish correspondences between pores more accurately. The refinement of the correspondences is then achieved by using a local topology-preserving matching algorithm. The algorithm uses rotational invariant local structures and pore pair local topology similarities to calculate the cost of each correspondence. It can remove the mismatches more accurately and efficiently. The experimental results on two high-resolution fingerprint image databases show that the proposed algorithm perform well in both accuracy and speed comparing to the existing algorithms. Yuanrong Xu, Yao Lu 0008, Guangming Lu 0002, Jinxing Li 0003, David Zhang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2020 | Online Multi-view Subspace Learning with Mixed NoiseabstractMulti-view learning reveals the latent correlation between different input modalities and has achieved outstanding performances in many fields. Recent approaches aim to find a low-dimensional subspace to reconstruct each view, in which the gross residual or noise follows either Gaussian or Laplacian distribution. However, the noise distribution is often more complex in practical applications, and a deterministic distribution assumption is incapable of modeling it. Additionally, referring to time-changed data, e.g., videos, the noise is temporal smooth, preventing us from processing the data with the whole input, as have generally been done in many existing multi-view learning methods. To tackle these problems, a novel online multi-view subspace learning is proposed in this paper. Particularly, our proposed method not only estimates a transformation for each view to extract the correlation among various views, but also introduces a Mixture of Gausssians (MoG) model into the multi-view data, successfully exploiting numbers of Gaussian Distributions to adaptively fit a wider range of the complex noise. Furthermore, we further design a novel online Expectation Maximization (EM) algorithm, being capable of efficiently processing the dynamic data. Experimental results substantiate the effectiveness and superiority of our approach. Jinxing Li 0003, Hongwei Yong, Feng Wu 0001, Mu Li 0005 |
ACM Multimedia | 1 |
| 2020 | Similarity and diversity induced paired projection for cross-modal retrieval
Jinxing Li 0003, Mu Li 0005, Guangming Lu 0002, Bob Zhang 0001, Hongpeng Yin, David Zhang 0001 |
Inf. Sci. | 1 |
| 2020 | High-parameter-efficiency convolutional neural networks
Yao Lu 0008, Guangming Lu 0002, Jinxing Li 0003, Yuanrong Xu, David Zhang 0001 |
Neural Comput. Appl. | 3 |
| 2020 | A supervised non-negative matrix factorization model for speech emotion recognition
Mi-Xiao Hou, Jinxing Li 0003, Guangming Lu 0002 |
Speech Commun. | 2 |
| 2020 | DRPL: Deep Regression Pair Learning for Multi-Focus Image FusionabstractIn this paper, a novel deep network is proposed for multi-focus image fusion, named Deep Regression Pair Learning (DRPL). In contrast to existing deep fusion methods which divide the input image into small patches and apply a classifier to judge whether the patch is in focus or not, DRPL directly converts the whole image into a binary mask without any patch operation, subsequently tackling the difficulty of the blur level estimation around the focused/defocused boundary. Simultaneously, a pair learning strategy, which takes a pair of complementary source images as inputs and generates two corresponding binary masks, is introduced into the model, greatly imposing the complementary constraint on each pair and making a large contribution to the performance improvement. Furthermore, as the edge or gradient does exist in the focus part while there is no similar property for the defocus part, we also embed a gradient loss to ensure the generated image to be all-in-focus. Then the structural similarity index (SSIM) is utilized to make a trade-off between the reference and fused images. Experimental results conducted on the synthetic and real-world datasets substantiate the effectiveness and superiority of DRPL compared with other state-of-the-art approaches. The testing code can be found in https://github.com/sasky1/DPRL. Jinxing Li 0003, Xiaobao Guo, Guangming Lu 0002, Bob Zhang 0001, Yong Xu 0001, Feng Wu 0001, David Zhang 0001 |
IEEE Trans. Image Process. | 1 |
| 2020 | Lesion Location Attention Guided Network for Multi-Label Thoracic Disease Classification in Chest X-RaysabstractTraditional clinical experiences have shown the benefit of lesion location attention for improving clinical diagnosis tasks. Inspired by this point of interest, in this paper we propose a novel lesion location attention guided network named LLAGnet to focus on the discriminative features from lesion locations for multi-label thoracic disease classification in chest X-rays (CXRs). By revealing the equivalence of the region-level attention (RLA) and channel-level attention (CLA), we find that the RLA is available as priors for object localization while the CLA implicitly provides high weights to the attractive channels, which both enable lesion location attention excitation. To integrate the advantages from both mechanisms, the proposed LLAGnet is structured with two corresponding attention modules, i.e., the RLA and CLA modules. Specifically, the RLA module consists of the global and local branches. And the weakly supervised attention mechanism embedded in the global branch can obtain visual regions of lesion locations by back-propagating gradients. Then the optimal attention region is amplified and applied to the local branch to provide more fine-grained features for the image classification. Finally, the CLA module adaptively enhances the weights of channel-wise features from the lesion locations by modeling interdependencies among channels. Extensive experiments on the ChestX-ray14 dataset clearly substantiate the effectiveness of LLAGnet as compared with the state-of-the-art baselines. Bingzhi Chen, Jinxing Li 0003, Guangming Lu 0002, David Zhang 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2020 | Label Co-Occurrence Learning With Graph Convolutional Networks for Multi-Label Chest X-Ray Image ClassificationabstractExisting multi-label medical image learning tasks generally contain rich relationship information among pathologies such as label co-occurrence and interdependency, which is of great importance for assisting in clinical diagnosis and can be represented as the graph-structured data. However, most state-of-the-art works only focus on regression from the input to the binary labels, failing to make full use of such valuable graph-structured information due to the complexity of graph data. In this paper, we propose a novel label co-occurrence learning framework based on Graph Convolution Networks (GCNs) to explicitly explore the dependencies between pathologies for the multi-label chest X-ray (CXR) image classification task, which we term the "CheXGCN". Specifically, the proposed CheXGCN consists of two modules, i.e., the image feature embedding (IFE) module and label co-occurrence learning (LCL) module. Thanks to the LCL model, the relationship between pathologies is generalized into a set of classifier scores by introducing the word embedding of pathologies and multi-layer graph information propagation. During end-to-end training, it can be flexibly integrated into the IFE module and then adaptively recalibrate multi-label outputs with these scores. Extensive experiments on the ChestX-Ray14 and CheXpert datasets have demonstrated the effectiveness of CheXGCN as compared with the state-of-the-art baselines. Bingzhi Chen, Jinxing Li 0003, Guangming Lu 0002, Hongbing Yu, David Zhang 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2020 | Relaxed Asymmetric Deep Hashing Learning: Point-to-Angle MatchingabstractDue to the powerful capability of the data representation, deep learning has achieved a remarkable performance in supervised hash function learning. However, most of the existing hashing methods focus on point-to-point matching that is too strict and unnecessary. In this article, we propose a novel deep supervised hashing method by relaxing the matching between each pair of instances to a point-to-angle way. Specifically, an inner product is introduced to asymmetrically measure the similarity and dissimilarity between the real-valued output and the binary code. Different from existing methods that strictly enforce each element in the real-valued output to be either +1 or -1, we only encourage the output to be close to its corresponding semantic-related binary code under the cross-angle. This asymmetric product not only projects both the real-valued output and the binary code into the same Hamming space but also relaxes the output with wider choices. To further exploit the semantic affinity, we propose a novel Hamming-distance-based triplet loss, efficiently making a ranking for the positive and negative pairs. An algorithm is then designed to alternatively achieve optimal deep features and binary codes. Experiments on four real-world data sets demonstrate the effectiveness and superiority of our approach to the state of the art. Jinxing Li 0003, Bob Zhang 0001, Guangming Lu 0002, Jane You, Yong Xu 0001, Feng Wu 0001, David Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2020 | SRGC-Nets: Sparse Repeated Group Convolutional Neural NetworksabstractGroup convolution is widely used in many mobile networks to remove the filter's redundancy from the channel extent. In order to further reduce the redundancy of group convolution, this article proposes a novel repeated group convolutional (RGC) kernel, which has M primary groups, and each primary group includes N tiny groups. In every primary group, the same convolutional kernel is repeated in all the tiny groups. The RGC filter is the first kernel to remove the redundancy from group extent. Based on RGC, a sparse RGC (SRGC) kernel is also introduced in this article, and its corresponding network is called SRGC neural networks (SRGC-Net). The SRGC kernel is the summation of RGC kernel and pointwise group convolutional (PGC) kernel. The number of PGC's groups is M . Accordingly, in each primary group, besides the center locations in all channels, the values of parameters located in other N-1 tiny groups are all zero. Therefore, SRGC can significantly reduce the parameters. Moreover, it can also effectively retrieve spatial and channel-difference features by utilizing RGC and PGC to preserve the richness of produced features. Comparative experiments were performed on the benchmark classification data sets. Compared with the traditional popular networks, SRGC-Nets can perform better with timely reducing the model size and computational complexity. Furthermore, it can also achieve better performances than other latest state-of-the-art mobile networks on most of the databases and effectively decrease the test and training runtime. Yao Lu 0008, Guangming Lu 0002, Jinxing Li 0003, David Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2019 | Super Sparse Convolutional Neural NetworksabstractTo construct small mobile networks without performance loss and address the over-fitting issues caused by the less abundant training datasets, this paper proposes a novel super sparse convolutional (SSC) kernel, and its corresponding network is called SSC-Net. In a SSC kernel, every spatial kernel has only one non-zero parameter and these non-zero spatial positions are all different. The SSC kernel can effectively select the pixels from the feature maps according to its non-zero positions and perform on them. Therefore, SSC can preserve the general characteristics of the geometric and the channels’ differences, resulting in preserving the quality of the retrieved features and meeting the general accuracy requirements. Furthermore, SSC can be entirely implemented by the “shift” and “group point-wise” convolutional operations without any spatial kernels (e.g., “3×3”). Therefore, SSC is the first method to remove the parameters’ redundancy from the both spatial extent and the channel extent, leading to largely decreasing the parameters and Flops as well as further reducing the img2col and col2img operations implemented by the low leveled libraries. Meanwhile, SSC-Net can improve the sparsity and overcome the over-fitting more effectively than the other mobile networks. Comparative experiments were performed on the less abundant CIFAR and low resolution ImageNet datasets. The results showed that the SSC-Nets can significantly decrease the parameters and the computational Flops without any performance losses. Additionally, it can also improve the ability of addressing the over-fitting problem on the more challenging less abundant datasets. Yao Lu 0008, Guangming Lu 0002, Bob Zhang 0001, Yuanrong Xu, Jinxing Li 0003 |
AAAI | 5 |
| 2019 | Separate Loss for Basic and Compound Facial Expression Recognition in the WildabstractIn the past few years, facial expression recognition has made great progress because of the development of convolutional neural networks. However, the features learned only using the softmax loss are not discriminative enough for highly accurate facial expression recognition in the wild, especially for the compound facial expression recognition. To enhance the discriminative power of the learned features, we propose the separate loss for both basic and compound facial expression recognition in the wild in this paper. Such loss maximizes intra-class similarity while minimizing the similarity between different classes. The qualitative and quantitative analysis shows that the features learned using such loss function are characterized by intra-class compactness and inter-class separation. Experiments are performed on two databases in the wild and the proposed method achieves state-of-the-art results on both basic and compound expressions. Furthermore, another two databases are used to perform cross database experiments to show the generalization ability of our method. Yingjian Li 0001, Yao Lu 0008, Jinxing Li 0003, Guangming Lu 0002 |
ACML | 3 |
| 2019 | Stable Pore Detection for High-Resolution Fingerprint based on a CNN DetectorabstractHigh-resolution fingerprint images contain three levels of features. Pores, as one of the level 3 features, have wide attention due to its significant contribution to the recognition accuracy. An accurate and stable pore detection algorithm plays a key role on the pore-based fingerprint recognition system. This paper proposes a pore detection method for high-resolution fingerprint images. The method uses fully convolutional network combined with the focal loss and shortcut structure to detect pores. The proposed algorithm is tested on the high-resolution fingerprint database. Experimental results show that our method outperforms the existing algorithms in accuracy, stability and matching performance. Zuolin Shen, Yuanrong Xu, Jinxing Li 0003, Guangming Lu 0002 |
ICIP | 3 |
| 2019 | Mask-Most Net: Mask Approximation Based Multi-oriented Scene Text Detection NetworkabstractIn this paper, a novel multi-task cascade framework, which jointly takes the detection and the segmentation into account, is presented for the scene text detection. To address the issue of multi-oriented scene text detection, we propose an instance-level mask approximation method through the auxiliary regression task on center and corner points. Specifically, the text instance in the image is first coarsely detected, followed by a contextual module which can capture more accurate instances. To cope with the scale variation existing in these detected instances, a combination of high-level semantic and low-level features is further exploited, achieving more robust and better performance. A series of experiments conducted on different benchmark datasets demonstrate the effectiveness of the proposed method. Xiaobao Guo, Jinxing Li 0003, Bingzhi Chen, Guangming Lu 0002 |
ICME | 2 |
| 2019 | Body surface feature-based multi-modal Learning for Diabetes Mellitus detection
Jinxing Li 0003, Bob Zhang 0001, Guangming Lu 0002, Jane You, David Zhang 0001 |
Inf. Sci. | 1 |
| 2019 | Visual Classification With Multikernel Shared Gaussian Process Latent Variable ModelabstractMultiview learning methods often achieve improvement compared with single-view-based approaches in many applications. Due to the powerful nonlinear ability and probabilistic perspective of Gaussian process (GP), some GP-based multiview efforts were presented. However, most of these methods make a strong assumption on the kernel function (e.g., radial basis function), which limits the capacity of the real data modeling. In order to address this issue, in this paper, we propose a novel multiview approach by combining a multikernel and GP latent variable model. Instead of designing a deterministic kernel function, multiple kernel functions are established to automatically adapt various types of data. Considering a simple way of obtaining latent variables at the testing stage, a projection from the observed space to the latent space as a back constraint has also been simultaneously introduced into the proposed method. Additionally, different from some existing methods which apply the classifiers off-line, a hinge loss is embedded into the model to jointly learn the classification hyperplane, encouraging the latent variables belonging to the different classes to be separated. An efficient algorithm based on the gradient decent technique is constructed to optimize our method. Finally, we apply the proposed approach to three real-world datasets and the associated results demonstrate the effectiveness and superiority of our model compared with other state-of-the-art methods. Jinxing Li 0003, Bob Zhang 0001, Guangming Lu 0002, Hu Ren, David Zhang 0001 |
IEEE Trans. Cybern. | 1 |
| 2018 | A Probabilistic Hierarchical Model for Multi-View and Multi-Feature ClassificationabstractSome recent works in classification show that the data obtained from various views with different sensors for an object contributes to achieving a remarkable performance. Actually, in many real-world applications, each view often contains multiple features, which means that this type of data has a hierarchical structure, while most of existing works do not take these features with multi-layer structure into consideration simultaneously. In this paper, a probabilistic hierarchical model is proposed to address this issue and applied for classification. In our model, a latent variable is first learned to fuse the multiple features obtained from a same view, sensor or modality. Particularly, mapping matrices corresponding to a certain view are estimated to project the latent variable from a shared space to the multiple observations. Since this method is designed for the supervised purpose, we assume that the latent variables associated with different views are influenced by their ground-truth label. In order to effectively solve the proposed method, the Expectation-Maximization (EM) algorithm is applied to estimate the parameters and latent variables. Experimental results on the extensive synthetic and two real-world datasets substantiate the effectiveness and superiority of our approach as compared with state-of-the-art. Jinxing Li 0003, Hongwei Yong, Bob Zhang 0001, Mu Li 0005, Lei Zhang 0006, David Zhang 0001 |
AAAI | 1 |
| 2018 | Shared Linear Encoder-based Gaussian Process Latent Variable Model for Visual ClassificationabstractMulti-view learning has shown its powerful potential in many applications and achieved outstanding performances compared with the single-view based methods. In this paper, we propose a novel multi-view learning model based on the Gaussian Process Latent Variable Model (GPLVM) to learn a shared latent variable in the manifold space with a linear and gaussian process prior based back projection. Different from existing GPLVM methods which only consider a mapping from the latent space to the observed space, the proposed method simultaneously takes a back projection from the observation to the latent variable into account. Concretely, due to the various dimensions of different views, a projection for each view is first learned to linearly map its observation to a subspace. The gaussian process prior is then imposed on another transformation to non-linearly and efficiently map the learned subspace to a shared manifold space. In order to apply the proposed approach to the classification, a discriminative regularization is also embedded to exploit the label information. Experimental results on three real-world databases substantiate the effectiveness and superiority of the proposed approach as compared with several state-of-the-art approaches. Jinxing Li 0003, Bob Zhang 0001, Guangming Lu 0002, David Zhang 0001 |
ACM Multimedia | 1 |
| 2018 | Shared Autoencoder Gaussian Process Latent Variable Model for Visual ClassificationabstractMultiview learning reveals the latent correlation among different modalities and utilizes the complementary information to achieve a better performance in many applications. In this paper, we propose a novel multiview learning model based on the Gaussian process latent variable model (GPLVM) to learn a set of nonlinear and nonparametric mapping functions and obtain a shared latent variable in the manifold space. Different from the previous work on the GPLVM, the proposed shared autoencoder Gaussian process (SAGP) latent variable model assumes that there is an additional mapping from the observed data to the shared manifold space. Due to the introduction of the autoencoder framework, both nonlinear projections from and to the observation are considered simultaneously. Additionally, instead of fully connecting used in the conventional autoencoder, the SAGP achieves the mappings utilizing the GP, which remarkably reduces the number of estimated parameters and avoids the phenomenon of overfitting. To make the proposed method adaptive for classification, a discriminative regularization is embedded into the proposed method. In the optimization process, an efficient algorithm based on the alternating direction method and gradient decent techniques is designed to solve the encoder and decoder parts alternatively. Experimental results on three real-world data sets substantiate the effectiveness and superiority of the proposed approach as compared with the state of the art. Jinxing Li 0003, Bob Zhang 0001, David Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2017 | Joint discriminative and collaborative representation for fatty liver disease diagnosis
Jinxing Li 0003, Bob Zhang 0001, David Zhang 0001 |
Expert Syst. Appl. | 1 |
| 2017 | Joint similar and specific learning for diabetes mellitus and impaired glucose regulation detection
Jinxing Li 0003, David Zhang 0001, Bob Zhang 0001 |
Inf. Sci. | 1 |