VLDB 2026 Research / reviewers in the wild / expert
Jun Li 0033
dblp:116/1011-33
· DBLP profile ↗
75ranked-venue papers
7as first author
52since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 44 · 3 first-author · 30 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 4 first-author · 19 since 2021Databases, data management, data science and information retrieval · 11 · 1 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Label Selection Algorithm with Boolean Interpolative Matrix Decomposition for Multi-Label Learning
Tianqi Ji, Jun Li 0033 |
DEXA (2) | 2 |
| 2026 | Dual feature masking stage-wise knowledge distillation for object detection
Zhourui Zhang, Jun Li 0033, Zhijian Wu, Jifeng Shen |
Eng. Appl. Artif. Intell. | 2 |
| 2026 | IRDFusion: Iterative relation-map difference guided feature fusion for multispectral object detection
Jifeng Shen, Haibo Zhan, Heng Fan 0001, Xiaohui Yuan 0001, Jun Li 0033, Wankou Yang |
Pattern Recognit. | 6 |
| 2026 | AE-IPP: Adversarial Example Enabled Identity Privacy Preserving With mmWave SignalsabstractDespite convenience and reliability of mmWave-based action recognition, it still raises privacy concerns on identity leakage threat since human behaviors could meanwhile expose massive user information in real-world applications. Existing solutions attempt to send anonymized features extracted from mmWave signals; however, features not only reduce the application flexibility but also increase the privacy disclosure risk due to original data reconstruction. Instead, in this paper we propose a de-identification system, AE-IPP, which customizes learned noises into the raw data to generate adversarial examples for identity privacy and action utility balance. In other words, the noises are sample-specific perturbations that are automatically learned for each sample through our presented network. To achieve the performance balance and ensure robustness to other models, we are faced with two challenges, including the decoupling of action and identity information and the transferability of models. To this end, AE-IPP focuses on respective attention areas by leveraging task-specific gradients and designs a dynamic attention mechanism to update the attention weights according to the final optimization objective. Moreover, we present a multidirectional perturbation strategy to improve the model generalization capabilities, enabling robust de-identification. Extensive experiments on mmWave datasets demonstrate the superiority of our method over state-of-the-art approaches. Biyun Sheng, Wangquan Qin, Jun Li 0033, Li Lu 0008, Tie Qiu 0001, Fu Xiao 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2025 | Recover and Match: Open-Vocabulary Multi-Label Recognition through Knowledge-Constrained Optimal TransportabstractIdentifying multiple novel classes in an image, known as open-vocabulary multi-label recognition, is a challenging task in computer vision. Recent studies explore the transfer of powerful vision-language models such as CLIP. However, these approaches face two critical challenges: (1) The local semantics of CLIP are disrupted due to its global pre-training objectives, resulting in unreliable regional predictions. (2) The matching property between image regions and candidate labels has been neglected, relying instead on naive feature aggregation such as average pooling, which leads to spurious predictions from irrelevant regions. In this paper, we present RAM (Recover And Match), a novel framework that effectively addresses the above issues. To tackle the first problem, we propose Ladder Local Adapter (LLA) to enforce refocusing on local regions, recovering local semantics in a memory-friendly way. For the second issue, we propose Knowledge-Constrained Optimal Transport (KCOT) to suppress meaningless matching to non-GT labels by formulating the task as an optimal transport problem. As a result, RAM achieves state-of-the-art performance on various datasets from three distinct domains, and shows great potential to boost the existing methods. Code: https://github.com/EricTan7/RAM. Zichang Tan, Jun Li 0033, Ajian Liu 0001, Jun Wan 0001, Zhen Lei 0001 |
CVPR | 3 |
| 2025 | A Parameter-free Quadratic Programming Feature Selection Algorithm Based on Minimizing Redundancy and Fixing RelevanceabstractFeature selection for classification aims to find an optimal subset of original features through minimizing featurefeature redundancy and maximizing feature-class relevance. After estimating redundancy matrix and relevance vector, constructing a constrained optimization problem has become one of the main-stream approaches recently. However, existing methods either involve some tunable parameters or utilize a complicated optimization procedure. In this paper, we propose a parameterfree quadratic programming feature selection model which minimizes redundancy and fixes relevance, where the former is formulated as a minimum of quadratic term associated with redundancy matrix and the latter as a linear equality constraint with relevance vector, plus a non-negative limit. Further, we apply a first-order Frank-Wolfe optimization method with a sub-linear convergence rate to solve our model, where at each iteration there exist an analytical solution for linear programming problem, an analytical step size and three recursive formulae for objective function, solution vector and gradient vector, which dramatically speeds up our optimization procedure. Our solution technique has an empirical time complexity of$O\left(D^{2}\right)$, where$D$is the number of features. When Hilbert-Schmidt independence criterion (HSIC) is to estimate redundancy matrix and relevance vector, our comprehensive experiments on four extremely highdimensional biological data sets show that the proposed method is superior to seven existing main-stream techniques, including mRMR, InfFS, FoANOCCO, FoHSIC, QPFS, HSICLasso and GRM. Jun Li 0033 |
ICTAI | 2 |
| 2025 | SAMKD: Spatial-aware Adaptive Masking Knowledge Distillation for Object DetectionabstractMost of recent attention-guided feature masking distillation methods perform knowledge transfer via global teacher attention maps without delving into fine-grained clues. Instead, performing distillation at finer granularity is conducive to uncovering local details supplementary to global knowledge transfer and reconstructing comprehensive student features. In this study, we propose a Spatial-aware Adaptive Masking Knowledge Distillation (SAMKD) framework for accurate object detection. Different from previous feature distillation methods which mainly perform single-scale feature masking, we develop spatially hierarchical feature masking distillation scheme, such that the object-aware locality is encoded during coarse-to-fine distillation process for improved feature reconstruction. In addition, our spatial-aware feature distillation strategy is combined with a masking logit distillation scheme in which region-specific feature difference between teacher and student networks is utilized to adaptively guide the distillation process. Thus, it can help the student model to better learn from the teacher counterpart with improved knowledge transfer and reduced gap. Extensive experiments for detection task demonstrate the superiority of our method. For example, when FCOS is used as teacher detector with ResNet101 backbone, our method improves the student network from 35.3% to 38.8% mAP, outperforming state-of-the-art distillation methods including MGD, FreeKD and DMKD. Zhourui Zhang, Jun Li 0033, Jiayan Li |
SMC | 2 |
| 2025 | Anchor-Based Adaptive Similarity Graph Learning for Semi-Supervised ClassificationabstractABSTRACT Recently, graph semi‐supervised classification for different datasets is faced with some problems, such as low classification accuracy, and error labels in labeled data. To alleviate these problems, we propose a model called anchor‐based adaptive similarity graph learning for semi‐supervised classification (AAGSSL). This model leverages anchors to construct weight matrix associated with anchor and data, and obtains the sparse initial affinity graph by special matrix factorization. It adaptively learns a new similarity graph close to the initial affinity graph, which reduces the model's reliance of classification accuracy on the initial affinity graph. The model enhances the tolerance of error labels in the labeled data and accelerates the process of obtaining predictive labels by adjusting corresponding matrix internal parameters which introduced in model and employing label propagation, respectively. The feasibility and effectiveness of the model were verified in experiments on artificial datasets. Focus on image datasets classification, the comparative experiments on real benchmark image datasets verified the advantages of the proposed model in handling error labels and finding new class. And we additionally evaluate the impact of several parameters on classification performance and choose the best hyperparameters. Yali Peng 0004, Shigang Liu, Jun Li 0033 |
Concurr. Comput. Pract. Exp. | 5 |
| 2025 | FrePrompter: Frequency self-prompt for all-in-one image restoration
Zhijian Wu, Jun Li 0033, Dingjiang Huang |
Pattern Recognit. | 4 |
| 2025 | UBG: An Unreal BattleGround Benchmark With Object-Aware Hierarchical Proximal Policy OptimizationabstractThe deep reinforcement learning (DRL) has made significant progress in various simulation environments. However, applying DRL methods to real-world scenarios poses certain challenges due to limitations in visual fidelity, scene complexity, and task diversity within existing environments. To address limitations and explore the potential ability of DRL, we developed a 3-D open-world first-person shooter (FPS) game called Unreal BattleGround (UBG) using the unreal engine (UE). UBG provides a realistic 3-D environment with variable complexity, random scenes, diverse tasks, and multiple scene interaction methods. This benchmark involves far more complex state-action spaces than classic pseudo-3-D FPS games (e.g., ViZDoom), making it challenging for DRL to learn human-level decision sequences. Then, we propose the object-aware hierarchically proximal policy optimization (OaH-PPO) method in the UBG. It involves a two-level hierarchy, where the high-level controller is tasked with learning option control, and the low-level workers focus on mastering subtasks. To boost the learning of subtasks, we propose three modules: an object-aware module for extracting depth detection information from the environment, potential-based intrinsic reward shaping for efficient exploration, and annealing imitation learning (IL) to guide the initialization. Experimental results have demonstrated the broad applicability of the UBG and the effectiveness of the OaH-PPO. We will release the code of the UBG and OaH-PPO after publication. Longyu Niu, Baihui Li, Xingjian Fan, Jun Li 0033, Junliang Xing, Jun Wan 0001, Zhen Lei 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | A Survey on Composed Image RetrievalabstractComposed Image Retrieval (CIR) processes a query consisting of a reference image and a modification text, aiming to retrieve target images that not only resemble the reference image visually but also reflect the modification described in the caption. Unlike traditional image retrieval methods that rely on a single modality, CIR integrates visual and textual information, enabling more nuanced and constraint-based query representations. This unique capability has garnered growing interest from researchers. Despite its potential, the field lacks a systematic review that comprehensively examines its advancements and trends. This article seeks to fill this gap by providing a detailed review of CIR research developments over the past 5 years. It categorizes existing methods into supervised approaches, which leverage triplet-labeled data for model training, and zero-shot approaches, which utilize unlabeled data to address CIR challenges. Additionally, the widely used benchmark datasets and evaluation indicators are comprehensively introduced. A comparative analysis of state-of-the-art methods across five datasets is also conducted, providing insights into their strengths and limitations. Ultimately, this article also outlines potential research directions for the future. Longye Du, Shuaiyu Deng, Ying Li 0016, Jun Li 0033, Qi Tian 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2024 | Compound Text-Guided Prompt Tuning via Image-Adaptive CuesabstractVision-Language Models (VLMs) such as CLIP have demonstrated remarkable generalization capabilities to downstream tasks. However, existing prompt tuning based frameworks need to parallelize learnable textual inputs for all categories, suffering from massive GPU memory consumption when there is a large number of categories in the target dataset. Moreover, previous works require to include category names within prompts, exhibiting subpar performance when dealing with ambiguous category names. To address these shortcomings, we propose Compound Text-Guided Prompt Tuning (TGP-T) that significantly reduces resource demand while achieving superior performance. We introduce text supervision to the optimization of prompts, which enables two benefits: 1) releasing the model reliance on the pre-defined category names during inference, thereby enabling more flexible prompt generation; 2) reducing the number of inputs to the text encoder, which decreases GPU memory consumption significantly. Specifically, we found that compound text supervisions, i.e., category-wise and content-wise, is highly effective, since they provide inter-class separability and capture intra-class variations, respectively. Moreover, we condition the prompt generation on visual features through a module called Bonder, which facilitates the alignment between prompts and visual features. Extensive experiments on few-shot recognition and domain generalization demonstrate that TGP-T achieves superior performance with consistently lower training costs. It reduces GPU memory usage by 93% and attains a 2.5% performance gain on 16-shot ImageNet. The code is available at https://github.com/EricTan7/TGP-T. Jun Li 0033, Yizhuang Zhou, Jun Wan 0001, Zhen Lei 0001, Xiangyu Zhang 0005 |
AAAI | 2 |
| 2024 | Factorized Learning Assisted with Large Language Model for Gloss-free Sign Language TranslationabstractPrevious Sign Language Translation (SLT) methods achieve superior performance by relying on gloss annotations. However, labeling high-quality glosses is a labor-intensive task, which limits the further development of SLT. Although some approaches work towards gloss-free SLT through jointly training the visual encoder and translation network, these efforts still suffer from poor performance and inefficient use of the powerful Large Language Model (LLM). Most seriously, we find that directly introducing LLM into SLT will lead to insufficient learning of visual representations as LLM dominates the learning curve. To address these problems, we propose Factorized Learning assisted with Large Language Model (FLa-LLM) for gloss-free SLT. Concretely, we factorize the training process into two stages. In the visual initialing stage, we employ a lightweight translation model after the visual encoder to pre-train the visual encoder. In the LLM fine-tuning stage, we freeze the acquired knowledge in the visual encoder and integrate it with a pre-trained LLM to inspire the LLM’s translation potential. This factorized training strategy proves to be highly effective as evidenced by significant improvements achieved across three SLT datasets which are all conducted under the gloss-free setting. Benjia Zhou, Jun Li 0033, Jun Wan 0001, Zhen Lei 0001 |
LREC/COLING | 3 |
| 2024 | A Label Embedding Algorithm Based on Maximizing Normalized Cross-Covariance Operator
Yulin Xue, Jun Li 0033 |
DEXA (1) | 4 |
| 2024 | DMKD: Improving Feature-Based Knowledge Distillation for Object Detection Via Dual Masking AugmentationabstractRecent mainstream masked distillation methods function by reconstructing selectively masked areas of a student network from the feature map of its teacher counterpart. In these methods, the masked regions need to be properly selected, such that reconstructed features encode sufficient discrimination and representation capability like the teacher feature. However, previous masked distillation methods only focus on spatial masking, making the resulting masked areas biased towards spatial importance without encoding informative channel clues. In this study, we devise a Dual Masked Knowledge Distillation (DMKD) framework which can capture both spatially important and channel-wise informative clues for comprehensive masked feature reconstruction. More specifically, we employ dual attention mechanism for guiding the respective masking branches, leading to reconstructed feature encoding dual significance. Furthermore, fusing the reconstructed features is achieved by self-adjustable weighting strategy for effective feature distillation. Our experiments on object detection task demonstrate that the student networks achieve performance gains of 4.1% and 4.3% with the help of our method when RetinaNet and Cascade Mask R-CNN are respectively used as the teacher networks, while outperforming the other state-of-the-art distillation methods. Zhijian Wu, Jun Li 0033, Xili Wan |
ICASSP | 4 |
| 2024 | HURDNet: Heterogeneous UNet Structure With Range-Null Space Decomposition for Hyperspectral Image ReconstructionabstractHyperspectral image (HSI) reconstruction aims to recover a 3D HSI from its degraded 2D measurement. Despite achieving significant progress, most methods do not consider the data-consistency condition, leading to biased reconstructed results. Besides, current UNet structures mainly utilize homogeneous modules for feature reconstruction without fully exploring the heterogeneous structures. To overcome the aforementioned drawbacks, we propose a heterogeneous UNet structure with range-null space decomposition for HSI reconstruction termed HURDNet. Specifically, we leverage range-null space decomposition for data-consistency calibration, whilst design a precise average calibration pseudo-inverse operator (ACPIO) for finely modeling the inverse HSI degradation process, such that 3D HSIs can be precisely reconstructed. In addition, we explore heterogeneous structure design in UNet architecture, and develop a pixel-wise long-term non-local Transformer structure (PLNT) particularly for feature reconstruction. Numerous experimental results demonstrate that our HURDNet achieves state-of-the-art performance on both simulation and real HSI datasets. Yeheng Zhu, Zhijian Wu, Jun Li 0033 |
ICME | 3 |
| 2024 | Two-dimensional Seismic Velocity Inversion via Enhanced Multi-view Convolutional Neural Networks for RegressionabstractSeismic exploration is a mainstream technique to find subsurface raw oil and gas in energy industry, which stimulates seismic wave on the Earth’s surface, receives seismic waves from the subsurface layers, processes these seismic data and lastly infers oil and gas existence. All these steps are associated with a physical parameter: velocity. Nowadays, this parameter is estimated via a so-called seismic waveform inversion method, which is to derive a subsurface velocity model of the Earth by inverting seismic data observed at the surface. Besides this traditional physics-guided full waveform inversion, recently, more attention has paid on data-driven inversion techniques, mainly deep learning based pre-stack inversion, which builds a nonlinear mapping from a multi-shot record to a velocity profile for two-dimensional situation. Usually, a multi-shot record is widely simulated as a multiple channel color image. However, the multiple shots are located at the distinct surface positions, which is different from the fixed camera position of images, and whose shot gathers include less complementary and much redundant information. In order to characterize this situation, we regard each shot gather as a view in machine learning and propose an enhanced multi-view convolutional neural network for regression (MVCNNR) to velocity inversion, in this paper. Our MVCNNR generalizes multi-view convolutional neural network for 3D shape recognition into a pixel-level regression network that further is added both addition connections in residual network and skipping concatenation connections in U-net, to improve its performance. Experiments on four types of velocity datasets (Layered, Faulted, SaltBody, and SaltDome) show that our MVCNNR is superior to two representative data-driven techniques (FCNVMB and VelocityGAN), via quantitative evaluation and visualization analysis. Chuang Pan, Qingzhen Wang, Jun Li 0033 |
IJCNN | 4 |
| 2024 | Compacter: A Lightweight Transformer for Image RestorationabstractAlthough deep learning-based methods have made significant advances in the field of image restoration (IR), they often suffer from excessive model parameters. To tackle this problem, this work proposes a compact Transformer (Compacter) for lightweight image restoration by making several key designs. We employ the concepts of projection sharing, adaptive interaction, and heterogeneous aggregation to develop a novel Compact Adaptive Self-Attention (CASA). Specifically, CASA utilizes shared projection to generate Query, Key, and Value to simultaneously model spatial and channel-wise self-attention. The adaptive interaction process is then used to propagate and integrate global information from two different dimensions, thus enabling omnidirectional relational interaction. Finally, a depth-wise convolution is incorporated on Value to complement heterogeneous local information, enabling global-local coupling. Moreover, we propose a Dual Selective Gated Module (DSGM) to dynamically encapsulate the globality into each pixel for context-adaptive aggregation. Extensive experiments demonstrate that our Compacter achieves state-of-the-art performance for a variety of lightweight IR tasks with approximately 400K parameters. Zhijian Wu, Jun Li 0033, Dingjiang Huang |
ACM Multimedia | 2 |
| 2024 | Multi-head Hashing with Orthogonal Decomposition for Cross-modal Retrieval
Jun Li 0033, Zhijian Wu, Bo Yang 0019 |
MMM (2) | 2 |
| 2024 | NCL++: Nested Collaborative Learning for long-tailed visual recognition
Zichang Tan, Jun Li 0033, Jinhao Du, Jun Wan 0001, Zhen Lei 0001, Guodong Guo |
Pattern Recognit. | 2 |
| 2024 | RUN: Rethinking the UNet Architecture for Efficient Image RestorationabstractRecent advanced image restoration (IR) methods typically stack homogeneous operators hierarchically in the UNet architecture. To achieve higher accuracy, these models are now going deeper and more complex, making them resource-intensive. After comprehensively reviewing different operators within modern networks, we provide an in-depth analysis of their individual favorable properties and invent a novel efficient IR network by redesigning the UNet architecture (RUN) with heterogeneous operators. Specifically, we propose three heterogeneous operators for different relational interactions concerning the specificity of different hierarchical features of the UNet architecture. First, the spatial self-attention block (SSA Block) processes high-resolution top-level features by modeling pixel interactions from the spatial dimension. Second, the channel self-attention block (CSA Block) performs channel recalibration and information transmission for the bottom-level features with rich channels. Finally, a simple and efficient convolution block (Conv Block) is used to facilitate middle-order information propagation, which complements the self-attention mechanism to achieve local-global coupling. Based on these designs, our RUN enables more comprehensive information dissemination and interaction regardless of topological distance, thus achieving superior performance while maintaining desirable computational budgets. Extensive experiments show that our RUN achieves state-of-the-art results for a variety of IR tasks, including image deblurring, image denoising, image deraining, and low-light image enhancement. Zhijian Wu, Jun Li 0033, Chang Xu 0002, Dingjiang Huang, Steven C. H. Hoi |
IEEE Trans. Multim. | 2 |
| 2023 | A Label Embedding Method via Conditional Covariance Maximization for Multi-label Classification
Yunqian Li, Jun Li 0033 |
DEXA (2) | 3 |
| 2023 | LogoNet: A Fine-Grained Network for Instance-Level Logo Sketch Retrieval
Binbin Feng, Jun Li 0033 |
ICIG (4) | 2 |
| 2023 | Reversible Column Networks
Yizhuang Zhou, Jianjian Sun, Xiangwen Kong, Jun Li 0033, Xiangyu Zhang 0005 |
ICLR | 6 |
| 2023 | Retinex Meets Transformer: Bridging Illumination and Reflectance Maps for Low-Light Image Enhancement
Yilong Cheng, Zhijian Wu, Jun Li 0033 |
ICONIP (12) | 3 |
| 2023 | Label Selection Algorithm Based on Ant Colony Optimization and Reinforcement Learning for Multi-label Classification
Yulin Xue, Jun Li 0033 |
ICONIP (5) | 3 |
| 2023 | A Novel Label Selection Algorithm Based on Principal Component Analysis and Sparse Approximation Solution for Multi-label ClassificationabstractIn multi-label classification, an instance may be associated with multiple labels simultaneously and thus the class labels are correlated one another. As various applications emerge, besides large instance size and high feature dimensionality, the dimensionality of label space also grows quickly, which would increase computational costs and even deteriorate classification performance. To this end, dimensionality reduction strategy is applied to label space via exploiting label correlation information, which covers label embedding and label selection techniques. Recently a lot of label embedding work has been conducted, but less attention has been paid to label selection techniques due to its difficulty. In this case, it is still an open problem how to design more effective label selection techniques for multi-label classification. Column subset selection problem (CSSP) originally is a mathematical issue in matrix theory to select a small portion of columns from a large-scale matrix for more interpretable data summarization. Therefore, such a CSSP naturally becomes an attractive mathematical representation for label selection, which is NP-hard and generally is solved via greedy strategy. In this paper, we build a two-stage label selection algorithm. At first, we apply principal component analysis (PCA) to reduce the dimensionality of the label matrix to obtain a low dimensional real matrix as the right side term in linear systems. Then, we use sparse approximation (SA) solution for linear systems to choose several informative columns from the label matrix as approximations of the low dimensional real matrix, which ultimately obtains the sub-optimal label subset of original label matrix. This new label selection method based on PCA and SA is referred to as PCASA simply. Our proposed method is validated experimentally to work well on six benchmark data sets with more than 100 labels. Yulin Xue, Jun Li 0033 |
ICTAI | 3 |
| 2023 | AMD: Adaptive Masked Distillation for Object DetectionabstractAs a general model compression paradigm, feature-based knowledge distillation allows the student model to learn expressive features from the teacher counterpart. In this paper, we mainly focus on designing an effective feature-distillation framework and propose a spatial-channel adaptive masked distillation (AMD) network for object detection. More specifically, in order to accurately reconstruct important feature regions, we first perform attention-guided feature masking on the feature map of the student network, such that we can identify the important features via spatially adaptive feature masking instead of random masking in the previous methods. In addition, we employ a simple and efficient module to allow the student network channel to be adaptive, improving its model capability in object perception and detection. In contrast to the previous methods, more crucial object-aware features can be reconstructed and learned from the proposed network, which is conducive to accurate object detection. The empirical experiments demonstrate the superiority of our method: with the help of our proposed distillation method, the student networks report 41.3%, 42.4%, and 42.7% mAP scores when RetinaNet, Cascade Mask-RCNN and RepPoints are respectively used as the teacher framework for object detection, which outperforms the previous state-of-the-art distillation methods including FGD and MGD. Jun Li 0033, Xili Wan |
IJCNN | 3 |
| 2023 | Separable Modulation Network for Efficient Image Super-ResolutionabstractDeep learning-based models have demonstrated unprecedented success in image super-resolution (SR) tasks. However, more attention has been paid to lightweight SR models lately, due to the increasing demand for on-device inference. In this paper, we propose a novel Separable Modulation Network (SMN) for efficient image SR. The key parts of the SMN are the Separable Modulation Unit (SMU) and the Locality Self-enhanced Network (LSN). SMU enables global relational interactions but significantly eases the process by separating spatial modulation from channel aggregation, hence making the long-range interaction efficient. Specifically, spatial modulation extracts global contexts from spatial, and channel aggregation condenses all global context features into the channel modulator, ultimately the aggregated contexts are fused into the final features. In addition, LSN allows guiding the network to focus on more refined image attributes by encoding local contextual information. By coupling two complementary components, SMN can capture both short- and long-range contexts for accurate image reconstruction. Extensive experimental results demonstrate that our SMN achieves state-of-the-art performance among the existing efficient SR methods with less complexity. Zhijian Wu, Jun Li 0033, Dingjiang Huang |
ACM Multimedia | 2 |
| 2023 | Online Change-Point Detection in High-Dimensional Covariance Structure with Application to Dynamic NetworksabstractIn this paper, we develop an online change-point detection procedure in the covariance structure of high-dimensional data. A new stopping rule is proposed to terminate the process as early as possible when a change in covariance structure occurs. The stopping rule allows spatial and temporal dependence and can be applied to non-Gaussian data. An explicit expression for the average run length is derived, so that the level of threshold in the stopping rule can be easily obtained with no need to run time-consuming Monte Carlo simulations. We also establish an upper bound for the expected detection delay, the expression of which demonstrates the impact of data dependence and magnitude of change in the covariance structure. Simulation studies are provided to confirm accuracy of the theoretical results. The practical usefulness of the proposed procedure is illustrated by detecting the change of brain’s covariance network in a resting-state fMRI data set. The implementation of the methodology is provided in the R package OnlineCOV. Lingjun Li, Jun Li 0033 |
J. Mach. Learn. Res. | 2 |
| 2023 | EfficientFace: an efficient deep network with feature enhancement for accurate face detection
Guangtao Wang, Jun Li 0033, Zhijian Wu, Jifeng Shen, Wankou Yang |
Multim. Syst. | 2 |
| 2023 | A Multi-scale Dilated Residual Convolution Network for Image Denoising
Xinlei Jia, Yali Peng 0004, Bao Ge, Jun Li 0033, Shigang Liu, Wenan Wang |
Neural Process. Lett. | 4 |
| 2023 | A two-phase projective dictionary pair learning-based classification scheme for positive and unlabeled learning
Yali Peng 0004, Shigang Liu, Bao Ge, Jun Li 0033 |
Pattern Anal. Appl. | 5 |
| 2023 | SFHN: Spatial-Frequency Domain Hybrid Network for Image Super-ResolutionabstractDeep convolutional neural networks (CNNs) have demonstrated tremendous success in image super-resolution (SR). According to the frequency principle, the vanilla CNNs fit the target function from low to high frequencies during the training process. It implies an implicit bias that CNNs tend to fit the training data by a low-frequency function. This is detrimental to SR task which essentially recovers the missing high-frequency cues from degraded images. To address this issue, a novel spatial-frequency domain hybrid network (SFHN) is proposed for image SR in this paper. More specifically, it contains multiple spatial-frequency domain hybrid convolution blocks (SFBlocks) for extracting both spatial and frequency components. In particular, frequency information is complementary to spatial one, mitigating the disadvantage of the spatial CNNs in capturing high-frequency content resulting from the inherent bias of the networks. In contrast to convolution blocks in the previous SR methods, the proposed SFBlock is configured with additional frequency-domain convolution branch driven by the Fourier transform, which allows to process frequency content. In addition, we introduce a spectral loss based on Fast Fourier Transform (FFT) to fully leverage the performance of our SFHN, as it prevents the loss of important frequency content during training. Experimental results on public benchmarks demonstrate that our SFHN achieves promising performance superior to the state-of-the-art methods. Zhijian Wu, Jun Li 0033, Chang Xu 0002, Dingjiang Huang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | General vs. Long-Tailed Age Estimation: An Approach to Kill Two Birds With One StoneabstractFacial age estimation has received a lot of attention for its diverse application scenarios. Most existing studies treat each sample equally and aim to reduce the average estimation error for the entire dataset, which can be summarized as General Age Estimation. However, due to the long-tailed distribution prevalent in the dataset, treating all samples equally will inevitably bias the model toward the head classes (usually the adult with a majority of samples). Driven by this, some works suggest that each class should be treated equally to improve performance in tail classes (with a minority of samples), which can be summarized as Long-tailed Age Estimation. However, Long-tailed Age Estimation usually faces a performance trade-off, i.e., achieving improvement in tail classes by sacrificing the head classes. In this paper, our goal is to design a unified framework to perform well on both tasks, killing two birds with one stone. To this end, we propose a simple, effective, and flexible training paradigm named GLAE, which is two-fold. First, we propose Feature Rearrangement (FR) and Pixel-level Auxiliary learning (PA) for better feature utilization to improve the overall age estimation performance. Second, we propose Adaptive Routing (AR) for selecting the appropriate classifier to improve performance in the tail classes while maintaining the head classes. Moreover, we introduce a new metric, named Class-wise Mean Absolute Error (CMAE), to equally evaluate the performance of all classes. Our GLAE provides a surprising improvement on Morph II, reaching the lowest MAE and CMAE of 1.14 and 1.27 years, respectively. Compared to the previous best method, MAE dropped by up to 34%, which is an unprecedented improvement, and for the first time, MAE is close to 1 year old. Extensive experiments on other age benchmark datasets, including CACD, MIVIA, and Chalearn LAP 2015, also indicate that GLAE outperforms the state-of-the-art approaches significantly. Zenghao Bao, Zichang Tan, Jun Li 0033, Jun Wan 0001, Xibo Ma, Zhen Lei 0001 |
IEEE Trans. Image Process. | 3 |
| 2022 | Nested Collaborative Learning for Long-Tailed Visual RecognitionabstractThe networks trained on the long-tailed dataset vary remarkably, despite the same training settings, which shows the great uncertainty in long-tailed learning. To alleviate the uncertainty, we propose a Nested Collaborative Learning (NCL), which tackles the problem by collaboratively learning multiple experts together. NCL consists of two core components, namely Nested Individual Learning (NIL) and Nested Balanced Online Distillation (NBOD), which focus on the individual supervised learning for each single expert and the knowledge transferring among multiple experts, respectively. To learn representations more thoroughly, both NIL and NBOD are formulated in a nested way, in which the learning is conducted on not just all categories from a full perspective but some hard categories from a partial perspective. Regarding the learning in the partial perspective, we specifically select the negative categories with high predicted scores as the hard categories by using a proposed Hard Category Mining (HCM). In the NCL, the learning from two perspectives is nested, highly related and complementary, and helps the network to capture not only global and robust features but also meticulous distinguishing ability. Moreover, self-supervision is further utilized for feature enhancement. Extensive experiments manifest the superiority of our method with outperforming the state-of-the-art whether by using a single model or an ensemble. Code is available at https://github.com/Bazinga699/NCL Jun Li 0033, Zichang Tan, Jun Wan 0001, Zhen Lei 0001, Guodong Guo |
CVPR | 1 |
| 2022 | Multi-label Online Streaming Feature Selection Algorithms via Extending Alpha-Investing Strategy
Tianqi Ji, Xizhi Guo, Yunqian Li, Jun Li 0033 |
DaWaK | 5 |
| 2022 | Label Selection Algorithm Based on Iteration Column Subset Selection for Multi-label Classification
Jun Li 0033 |
DEXA (1) | 2 |
| 2022 | Infinite Label Selection Method for Mutil-label Classification
Jun Li 0033 |
ICONIP (4) | 2 |
| 2022 | Subspace-based self-weighted multiview fusion for instance retrieval
Zhijian Wu, Jun Li 0033, Wankou Yang |
Inf. Sci. | 2 |
| 2022 | Dependency maximization forward feature selection algorithms based on normalized cross-covariance operator and its approximated form for high-dimensional data
Wenkai Lu, Jun Li 0033, Hongli Yuan |
Inf. Sci. | 3 |
| 2022 | Dual-Complementary Convolution Network for Remote-Sensing Image DenoisingabstractRemote-sensing images serve as key data sources which play a crucial role in recording the target information of ground features. Due to the limitations of the existing imaging equipment, environments, and transmission conditions, the obtained remote-sensing images are usually contaminated by noise in real-world scenarios. To address this problem, we propose a dual-complementary convolution network (DCCNet), including structural and detailed subnetwork, for repairing the structure and details of noisy remote-sensing images. More specifically, they generate multiresolution inputs via discrete wavelet transform and shuffling operation, respectively. Since the convolution operation is imposed on low-resolution inputs, the network parameters are considerably reduced. Experimental evaluations demonstrate that our proposed network exhibits superior performance to other competing methods in remote-sensing public datasets. The code of the DCCNet is available athttps://github.com/20155104009/DCCNet. Xinlei Jia, Yali Peng 0004, Jun Li 0033, Bao Ge, Yunhong Xin, Shigang Liu |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Regularized label relaxation with negative technique for image classification
Yali Peng 0004, Shigang Liu, Jun Li 0033 |
Multim. Tools Appl. | 4 |
| 2021 | Multi-label Feature Selection Algorithm via Maximizing Label Correlation-Aware Relevance and Minimizing Redundance with Mutation Binary Particle Swarm Optimization
Yuanyuan Tao, Jun Li 0033 |
DaWaK | 3 |
| 2021 | A Globally Optimal Label Selection Method via Genetic Algorithm for Multi-label Classification
Tianqi Ji, Jun Li 0033 |
DEXA (2) | 2 |
| 2021 | Label Selection Algorithm Based on Boolean Interpolative Decomposition with Sequential Backward Selection for Multi-label Classification
Tianqi Ji, Jun Li 0033 |
ICDAR (2) | 2 |
| 2021 | Binary Multi-view Image Re-ranking
Zhijian Wu, Jun Li 0033 |
ICIG (1) | 2 |
| 2021 | UVIP: Robust UWB aided Visual-Inertial Positioning System for Complex Indoor EnvironmentsabstractIndoor positioning without GPS is a challenge task, especially, in complex scenes or when sensors fail. In this paper, we develop an ultra-wideband aided visual-inertial positioning system (UVIP) which aims to achieve accurate and robust positioning results in complex indoor environments. To this end, a point-line-based stereo visual-inertial odometry (PL-sVIO) is firstly designed to improve the positioning accuracy in structured or low-textured scenarios by making use of line features. Secondly, a loop closure method is proposed to suppress the drift of PL-sVIO based on image patch features described by a CNN for handing the situation of a large environment and viewpoint variation. Thirdly, an accurate relocalization approach is presented for the case when the visual sensor fails. In this scheme, a top-to-down matching strategy from image to point and line features is presented to improve relocalization performance. Finally, the UWB sensor is combined with the visual-inertial system to further improve the accuracy and robustness of the positioning system and provide the results in a fixed reference frame. Thus, desirable real-time positioning results are derived for complex indoor scenes. Evaluations on challenging public datasets and real-world experiments are conducted to demonstrate that the proposed UVIP can provide more accurate and robust positioning results in complex indoor environments, even in the case when the visual sensor fails or in the absence of UWB anchors. Bo Yang 0019, Jun Li 0033, Hong Zhang 0013 |
ICRA | 2 |
| 2021 | An Interconnected Feature Pyramid Networks for object detection
Qiang Wang 0023, Lukuan Zhou, Yuncong Yao, Yong Wang 0032, Jun Li 0033, Wankou Yang |
J. Vis. Commun. Image Represent. | 5 |
| 2021 | Beyond ITQ: Efficient binary multi-view subspace learning for instance retrieval
Zhijian Wu, Jun Li 0033, Wankou Yang |
J. Vis. Commun. Image Represent. | 2 |
| 2021 | A Two-Step Classification Method Based on Collaborative Representation for Positive and Unlabeled Learning
Yali Peng 0004, Shigang Liu, Jun Li 0033 |
Neural Process. Lett. | 5 |
| 2021 | Subspace-based multi-view fusion for instance-level image retrieval
Jun Li 0033, Bo Yang 0019, Wankou Yang, Changyin Sun 0001 |
Vis. Comput. | 1 |
| 2020 | Efficient Binary Multi-view Subspace Learning for Instance-Level Image Retrieval
Zhijian Wu, Jun Li 0033 |
ICONIP (4) | 2 |
| 2020 | Multi-label Feature Selection Method via Maximizing Correlation-based Criterion with Mutation Binary Bat AlgorithmabstractMulti-label feature selection is a vital pre-processing step to reduce computational complexity, improve classification performance and enhance model interpretability, via selecting a discriminative subset of features from original high-dimensional features. Correlation-based feature selection (CFS) criterion measures the relevance between features and labels, and the redundancy among features, which has been combined with hill climbing and genetic algorithm to execute multi-label feature selection task. However, it is an open problem to search for more effective optimization tools for CFS. In this paper, through adding a mutation operation, we modify existing binary bat algorithm to build its mutation version (MBBA), to adjust the number of "1" components to be a fix size. Then a new multi-label feature selection approach is proposed via maximizing CFS criterion using MBBA, to select a fixed number of discriminative features. Our experiments on four data sets show that our proposed method is superior to three state-of-the-art approaches, according to four sample-based performance evaluation metrics for multi-label classification. Yuanyuan Tao, Jun Li 0033 |
IJCNN | 2 |
| 2020 | DenseUNet: densely connected UNet for electron microscopy image segmentationabstractElectron microscopy (EM) image segmentation plays an important role in computer‐aided diagnosis of specific pathogens or disease. However, EM image segmentation is a laborious task and needs to impose experts knowledge, which can take up valuable time from research. Convolutional neural network (CNN)‐based methods have been proposed for EM image segmentation and achieved considerable progress. Among those CNN‐based methods, UNet is regarded as the state‐of‐the‐art method. However, the UNet usually has millions of parameters to increase training difficulty and is limited by the issue of vanishing gradients. To address those problems, the authors present a novel highly parameter efficient method called DenseUNet, which is inspired by the approach that takes particular advantage of recent advances in both UNet and DenseNet. In addition, they successfully apply the weighted loss, which enables us to boost the performance of segmentation. They conduct several comparative experiments on the ISBI 2012 EM dataset. The experimental results show that their method can achieve state‐of‐the‐art results on EM image segmentation without any further post‐processing module or pre‐training. Moreover, due to smart design of the model, their approach has much less parameters than currently published encoder–decoder architecture variants for this dataset. Yue Cao 0009, Shigang Liu, Yali Peng 0004, Jun Li 0033 |
IET Image Process. | 4 |
| 2020 | ☆ - Discriminative dictionary learning algorithm based on sample diversity and locality of atoms for face recognition
Shigang Liu, Xiaosheng Wu, Jun Li 0033, Tao Lei 0003 |
J. Vis. Commun. Image Represent. | 4 |
| 2020 | Sparsity adaptive matching pursuit for face recognition
Yali Peng 0004, Shigang Liu, Jun Li 0033 |
J. Vis. Commun. Image Represent. | 4 |
| 2020 | Regularized Negative Label Relaxation Least Squares Regression for Face Recognition
Yali Peng 0004, Shigang Liu, Jun Li 0033 |
Neural Process. Lett. | 4 |
| 2020 | Weighted constraint based dictionary learning for image classification
Yali Peng 0004, Lingjun Li, Shigang Liu, Jun Li 0033 |
Pattern Recognit. Lett. | 5 |
| 2020 | Discriminative Multi-View Subspace Feature Learning for Action RecognitionabstractAlthough deep features have achieved the state-of-the-art performance in action recognition recently, the hand-crafted shallow features still play a critical role in characterizing human actions for taking advantage of visual contents in an intuitive way such as edge features. Therefore, the shallow features can serve as auxiliary visual cues supplementary to deep representations. In this paper, we propose a discriminative subspace learning model (DSLM) to explore the complementary properties between the hand-crafted shallow feature representations and the deep features. As for the RGB action recognition, this is the first work attempting to mine multi-level feature complementaries by the multi-view subspace learning scheme. To sufficiently capture the complementary information among heterogeneous features, we construct the DSLM by integrating the multi-view reconstruction error and classification error into an unified objective function. To be specific, we first use Fisher Vector to encode improved dense trajectories (iDT+FV) for shallow representations and two-stream convolutional neural network models (T-CNN) for generating deep features. Moreover, the presented DSLM algorithm projects multi-level features onto a shared discriminative subspace with the complementary information and discriminating capacity simultaneously incorporated. Finally, the action types of test samples are identified by the margins from the learned compact representations to the decision boundary. The experimental results on three datasets demonstrate the effectiveness of the proposed method. Biyun Sheng, Jun Li 0033, Fu Xiao 0001, Qun Li 0002, Wankou Yang, Junwei Han 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Discriminative Multi-View Privileged Information Learning for Image Re-RankingabstractConventional multi-view re-ranking methods usually perform asymmetrical matching between the region of interest (ROI) in the query image and the whole target image for similarity computation. Due to the inconsistency in the visual appearance, this practice tends to degrade the retrieval accuracy particularly when the image ROI, which is usually interpreted as the image objectness, accounts for a smaller region in the image. Since Privileged Information (PI), which can be viewed as the image prior, is able to characterize well the image objectness, we are aiming at leveraging PI for further improving the performance of multi-view re-ranking in this paper. Towards this end, we propose a discriminative multi-view re-ranking approach in which both the original global image visual contents and the local auxiliary PI features are simultaneously integrated into a unified training framework for generating the latent subspaces with sufficient discriminating power. For the on-the-fly re-ranking, since the multi-view PI features are unavailable, we only project the original multi-view image representations onto the latent subspace, and thus the re-ranking can be achieved by computing and sorting the distances from the multi-view embeddings to the separating hyperplane. Extensive experimental evaluations on the two public benchmarks, Oxford5k and Paris6k, reveal that our approach provides further performance boost for accurate image re-ranking, whilst the comparative study demonstrates the advantage of our method against other multi-view re-ranking methods. Jun Li 0033, Chang Xu 0002, Wankou Yang, Changyin Sun 0001, Hong Zhang 0013 |
IEEE Trans. Image Process. | 1 |
| 2019 | A Label Embedding Method for Multi-label Classification via Exploiting Local Label Correlations
Xidong Wang, Jun Li 0033 |
ICONIP (5) | 2 |
| 2019 | Virtual samples and sparse representation-based classification algorithm for face recognitionabstractDue to the environment and equipment are not controllable, the process of face image acquisition is inevitable to be interfered by external factors, and there are usually only a small number of available face images. Insufficient samples are not conducive to face recognition. Therefore, it is a popular scheme to produce virtual samples based on the available training samples. In this study, the authors first take the symmetry of human face into account, and propose a novel method to generate virtual samples. Then a representation‐based classification method and the score fusion strategy are applied to both original face images and virtual images to perform face recognition. Several sparse representation‐based classification algorithms are compared on ORL, FERET and GT databases. Experimental results show that the authors’ method is effective for improving the face recognition. Yali Peng 0004, Lingjun Li, Shigang Liu, Jun Li 0033 |
IET Comput. Vis. | 4 |
| 2019 | Pedestrian Proposal and Refining Based on the Shared Pixel Differential FeatureabstractWe design a pedestrian proposal and refining system tailored for fast pedestrian detection. The pedestrian proposal is based on pixel differential feature (PDF), which is a light weighted feature with a high recall rate. For the pedestrian refining, we propose an aggregated region feature (ARF) to distill the co-existing dominant pixel differential patterns in a local region to reject hard false positives. Albeit discriminative, ARF largely relies on the size of anchored regions and the scale of the PDF, which hinders its performance in real-world applications. Although multi-scale PDF with spatial pyramid somewhat alleviates this problem, it is computationally expensive and thus infeasible in practice. To address this issue, we further propose a directional radius pooling method to extract discriminative information in each orientation of PDF while reducing the feature dimensionality with a more compact size. The pedestrian proposal and refining framework is built on the shared pixel differential feature map which is very computationally efficient. More specifically, a set of pedestrian proposals generated from the single-scale PDF are first obtained in images. Second, multi-scale ARF in spatial pyramid is used to fuse information from different scales and spatial resolutions for anchor regions. Third, the directional radius pooling method is proposed to extract dominant information of each orientation in the anchor regions. The pedestrian proposal and refining are finally integrated for accurate pedestrian detection. The extensive experimental evaluations on five public benchmarks show that our method achieves state-of-the-art results while running at 18 fps for $480\times640$ images. Jifeng Shen, Lei Zhu 0010, Jun Li 0033, Wankou Yang, Haibin Ling |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2019 | ROMIR: Robust Multi-View Image Re-RankingabstractIn multi-view re-ranking, multiple heterogeneous visual features are usually projected onto a low-dimensional subspace, and thus the resulting latent representation can be used for the subsequent similarity-based ranking. Albeit effective, this standard mechanism underplays the intrinsic structure underlying the latent subspace and does not take into account the substantial noise in the original spaces. In this paper, we propose a robust multi-view image re-ranking strategy. Due to the dramatic variability in image visual appearance, it is necessary to uncover the shared components underlying those query-related instances that are visually unlike for improving the re-ranking accuracy. Consequently, it is reasonable to assume the latent subspace enjoys the low-rank property and thus the subspace recovery can be achieved via the low-rank modeling accordingly. In addition, since the real-world data are usually partially contaminated, we employ `2;1-norm based sparsity constraint to appropriately model the sample-specific mapping noise for enhancing the model robustness. In order to produce discriminative representations, we encode a similarity preserving term in our multi-view embedding framework. As a result, the sample separability is maximally maintained in the latent subspace with sufficient discriminative power. The extensive evaluations on public landmark benchmarks demonstrate the efficacy and superiority of the proposed method. Jun Li 0033, Chang Xu 0002, Wankou Yang, Changyin Sun 0001, Kotagiri Ramamohanarao, Dacheng Tao |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2018 | Extended sparse representation-based classification method for face recognition
Yali Peng 0004, Lingjun Li, Shigang Liu, Jun Li 0033 |
Mach. Vis. Appl. | 4 |
| 2018 | l2, 1-norm minimization based negative label relaxation linear regression for feature selection
Yali Peng 0004, Paramjit S. Sehdev, Shigang Liu, Jun Li 0033 |
Pattern Recognit. Lett. | 4 |
| 2018 | Crowd Counting via Weighted VLAD on a Dense Attribute Feature MapabstractCrowd counting is an important task in computer vision, which has many applications in video surveillance. Although the regression-based framework has achieved great improvements for crowd counting, how to improve the discriminative power of image representation is still an open problem. Conventional holistic features used in crowd counting often fail to capture semantic attributes and spatial cues of the image. In this paper, we propose integrating semantic information into learning locality-aware feature (LAF) sets for accurate crowd counting. First, with the help of a convolutional neural network, the original pixel space is mapped onto a dense attribute feature map, where each dimension of the pixelwise feature indicates the probabilistic strength of a certain semantic class. Then, LAF built on the idea of spatial pyramids on neighboring patches is proposed to explore more spatial context and local information. Finally, the traditional vector of locally aggregated descriptor (VLAD) encoding method is extended to a more generalized form weighted-VLAD (W-VLAD) in which diverse coefficient weights are taken into consideration. Experimental results validate the effectiveness of our presented method. Biyun Sheng, Chunhua Shen, Guosheng Lin, Jun Li 0033, Wankou Yang, Changyin Sun 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2017 | Robust optimal control for time-delay systems with dynamic uncertainties via ADPabstractThis paper considers a robust optimal control design for a class of nonlinear discrete-time systems with unknown time-varying delays and dynamic uncertainties. An iterative control strategy based on adaptive dynamic programming (ADP) has been proposed. Neural networks are applied to realize the state prediction, the control input estimation and the performance index function approximation. The estimated control input and performance index function are updated iteratively. Furthermore, it has been proven that the approximated performance index function can converge to the optimal solution of the Hamilton-Jacobia-Bellman (HJB) equation. Finally, the proposed algorithm has been conducted to a numerical simulation. The simulation results demonstrate the effectiveness of the new design. Lu Dong 0002, Jun Li 0033, Wankou Yang, Changyin Sun 0001 |
IJCNN | 2 |
| 2017 | SPA: Spatially Pooled Attributes for image retrieval
Jun Li 0033, Chang Xu 0002, Wankou Yang, Changyin Sun 0001 |
Neurocomputing | 1 |
| 2017 | Filtered shallow-deep feature channels for pedestrian detection
Biyun Sheng, Qichang Hu, Jun Li 0033, Wankou Yang, Baochang Zhang 0001, Changyin Sun 0001 |
Neurocomputing | 3 |
| 2017 | A novel pixel neighborhood differential statistic feature for pedestrian and face detection
Jifeng Shen, Jun Li 0033, Wankou Yang, Haibin Ling |
Pattern Recognit. | 3 |
| 2017 | Discriminative Multi-View Interactive Image Re-RankingabstractGiven an unreliable visual patterns and insufficient query information, content-based image retrieval is often suboptimal and requires image re-ranking using auxiliary information. In this paper, we propose a discriminative multi-view interactive image re-ranking (DMINTIR), which integrates user relevance feedback capturing users' intentions and multiple features that sufficiently describe the images. In DMINTIR, heterogeneous property features are incorporated in the multi-view learning scheme to exploit their complementarities. In addition, a discriminatively learned weight vector is obtained to reassign updated scores and target images for re-ranking. Compared with other multi-view learning techniques, our scheme not only generates a compact representation in the latent space from the redundant multi-view features but also maximally preserves the discriminative information in feature encoding by the large-margin principle. Furthermore, the generalization error bound of the proposed algorithm is theoretically analyzed and shown to be improved by the interactions between the latent space and discriminant function learning. Experimental results on two benchmark data sets demonstrate that our approach boosts baseline retrieval quality and is competitive with the other state-of-the-art re-ranking strategies. Jun Li 0033, Chang Xu 0002, Wankou Yang, Changyin Sun 0001, Dacheng Tao |
IEEE Trans. Image Process. | 1 |
| 2016 | SERVE: Soft and Equalized Residual VEctors for image retrieval
Jun Li 0033, Chang Xu 0002, Mingming Gong, Junliang Xing, Wankou Yang, Changyin Sun 0001 |
Neurocomputing | 1 |
| 2014 | Kernel inverse Fisher discriminant analysis for face recognition
Zhongxi Sun, Jun Li 0033, Changyin Sun 0001 |
Neurocomputing | 2 |