EDBT 2026 Demo / reviewers in the wild / expert
Zhong Ji
dblp:36/6466
· DBLP profile ↗
133ranked-venue papers
68as first author
84since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 74 · 30 first-author · 46 since 2021Graphics, computer vision, multimedia, augmented reality and games · 54 · 31 first-author · 34 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 5 first-author · 7 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-authorComputer networks · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | iEBAKER: Improved remote sensing image-text retrieval framework via eliminate before align and keyword explicit reasoning
Yan Zhang 0135, Zhong Ji, Changxu Meng, Yanwei Pang |
Expert Syst. Appl. | 2 |
| 2026 | A causality based multi-task framework for enhancing image and text interactions
Zhaomeng Cheng, Zhong Ji, Yan Zhang 0135 |
Knowl. Based Syst. | 2 |
| 2026 | SD2-SNN: Self-distillation and structural decomposition framework for SNNs in continual learning
Zhenhao Xie, Xia Xiao 0001, Yanwei Pang, Zhong Ji |
Neural Networks | 5 |
| 2026 | Parameter-Efficient Fine-Tuning for Continual Learning: A Neural Tangent Kernel PerspectiveabstractParameter-efficient fine-tuning for continual learning (PEFT-CL) has shown promise in adapting pre-trained models to sequential tasks while mitigating catastrophic forgetting problem. However, understanding the mechanisms that dictate continual performance in this paradigm remains elusive. To unravel this mystery, we undertake a rigorous analysis of PEFT-CL dynamics to derive relevant metrics for continual scenarios using Neural Tangent Kernel (NTK) theory. With the aid of NTK as a mathematical analysis tool, we recast the challenge of test-time forgetting into the quantifiable generalization gaps during training, identifying three key factors that influence these gaps and the performance of PEFT-CL: training sample size, task-level feature orthogonality, and regularization. To address these challenges, we introduce NTK-CL, a novel framework that eliminates task-specific parameter storage while adaptively generating task-relevant features. Aligning with theoretical guidance, NTK-CL triples the feature representation of each sample, theoretically and empirically reducing the magnitude of both task-interplay and task-specific generalization gaps. Grounded in NTK analysis, our framework imposes an adaptive exponential moving average mechanism and constraints on task-level feature orthogonality, maintaining intra-task NTK forms while attenuating inter-task NTK forms. Ultimately, by fine-tuning optimizable parameters with appropriate regularization, NTK-CL achieves state-of-the-art performance on established PEFT-CL benchmarks. This work provides a theoretical foundation for understanding and improving PEFT-CL models, offering insights into the interplay between feature representation, task orthogonality, and generalization, contributing to the development of more efficient continual learning systems. Jingren Liu, Zhong Ji, Yunlong Yu 0001, Jiale Cao, Yanwei Pang, Jungong Han, Xuelong Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | A Fresh Look at Generalized Category Discovery Through Non-Negative Matrix FactorizationabstractGeneralized Category Discovery (GCD) aims to classify both base and novel images using labeled base data. However, current approaches inadequately address the intrinsic optimization of the co-occurrence matrix A¯ based on cosine similarity, failing to achieve zero base-novel regions and adequate sparsity in base and novel domains. To address these deficiencies, we propose a Non-Negative Generalized Category Discovery (NN-GCD) framework. By establishing within the Symmetric Non-negative Matrix Factorization (SNMF) framework: (i) the equivalence between ideal k-means clustering and ideal SNMF, and (ii) the equivalence between SNMF solvers and Non-negative Contrastive Learning (NCL) optimization, we reformulate both the optimization of A¯ and k-means clustering as an NCL optimization problem. Moreover, to satisfy the non-negative constraints and make a GCD model converge to a near-ideal region, we propose a GELU activation function and an NMF NCE loss. To transition A¯ from a near-ideal state to the desired A¯∗, we introduce a hybrid sparse regularization approach to impose sparsity constraints. Experimental results show NN-GCD outperforms state-of-the-art methods on GCD benchmarks, achieving an average accuracy of 66.9% on the Semantic Shift Benchmark, surpassing prior counterparts by 2.5%. Zhong Ji, Jingren Liu, Yanwei Pang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | Cyclic Pseudo-Label Generation and Refinement for Weakly Supervised Referring Expression GroundingabstractWeakly supervised Referring Expression Grounding (WREG) aims at grounding the target region based on a given expression, where the mapping between regions and expressions is unknown during training. Recent WREG methods leverage the strategy of generating pseudo-labels utilizing Vision-Language Pre-training (VLP) to avoid the cross-modal heterogeneous gaps arising from the two-stage reconstruction strategy. However, mainstream VLPs are trained with image-text alignment data, which makes the generated labels inapplicable to REG task. Furthermore, due to the constraints of WREG data, it is challenging to ensure the quality of the pseudo-labels. To this end, we propose a Cyclic Pseudo-label Generation and Refinement (CPGR) method to alleviate the above limitations. Specifically, we cycle through the process of Generation-Refinement-Grounding to alleviate the impact of missing region annotations. We perform REG task-adaptive fine-tuning on BLIP-2 to generate REG-style descriptions with Region-Centrality. Then, we design a Pseudo-label Refinement module by utilizing cross-modal token attention to enhance the reliability of pseudo-labels and ensure their Reference-Discrimination. Experiments on five benchmark datasets demonstrate that our proposed method outperforms the current state-of-the-art weakly supervised methods. Our code and models will be released at https://github.com/5jiahe/CPGR. Jiahe Wu, Zhong Ji, Yanwei Pang, Jungong Han |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Underlying Semantic Diffusion for Effective and Efficient In-Context LearningabstractDiffusion models have emerged as a powerful framework for tasks like image controllable generation and dense prediction. However, existing models often struggle to capture underlying semantics (e.g., edges, textures, shapes) and effectively utilize in-context learning, limiting their contextual understanding and image generation quality. Furthermore, high computational costs and slow inference speeds hinder their real-time applications. To address these challenges, we propose Underlying Semantic Diffusion (US-Diffusion), an enhanced diffusion model that improves underlying semantics learning, computational efficiency, and in-context learning capabilities on multi-task scenarios. We introduce Separate & Gather Adapter (SGA), which decouples input conditions for different tasks while sharing the architecture, enabling better in-context learning and generalization across diverse visual domains. We also present a Feedback-Aided Learning (FAL) framework, which leverages feedback signals to guide the model in capturing semantic details and dynamically adapting to task-specific contextual cues. Furthermore, we propose a plug-and-play Efficient Sampling Strategy (ESS) for dense sampling at time steps with high-noise levels, which aims at optimizing training and inference efficiency while maintaining strong in-context learning performance. Experimental results demonstrate that US-Diffusion outperforms the state-of-the-art method, achieving an average reduction of 7.47 in FID on Map2Image tasks and an average reduction of 0.026 in RMSE on Image2Map tasks, while achieving approximately $9.45\times $ faster inference speed. Our method also demonstrates superior training efficiency and in-context learning capabilities, excelling in new datasets and tasks, highlighting its robustness and adaptability across diverse visual domains. The source code will be released at https://github.com/dragon-cao/US-Diffusion. Zhong Ji, Weilong Cao, Yan Zhang 0135, Yanwei Pang, Jungong Han |
IEEE Trans. Image Process. | 1 |
| 2026 | Interpretable Few-Shot Image Classification via Prototypical Concept-Guided Mixture of LoRA ExpertsabstractSelf-Explainable Models (SEMs) rely on Prototypical Concept Learning (PCL) to enable their visual recognition processes more interpretable, but they often struggle in data-scarce settings where insufficient training samples lead to suboptimal performance. To address this limitation, we propose a Few-Shot Prototypical Concept Classification (FSPCC) framework that systematically mitigates two key challenges under low-data regimes: parametric imbalance and representation misalignment. Specifically, our approach leverages a Mixture of LoRA Experts (MoLE) for parameter-efficient adaptation, ensuring a balanced allocation of trainable parameters between the backbone and the PCL module. Meanwhile, cross-module concept guidance enforces tight alignment between the backbone's feature representations and the prototypical concept activation patterns. In addition, we incorporate a multi-level feature preservation strategy that fuses spatial and semantic cues across various layers, thereby enriching the learned representations and mitigating the challenges posed by limited data availability. Finally, to enhance interpretability and minimize concept overlap, we introduce a geometry-aware concept discrimination loss that enforces orthogonality among concepts, encouraging more disentangled and transparent decision boundaries. Experimental results on six popular benchmarks (CUB-200-2011, mini-ImageNet, CIFAR-FS, Stanford Cars, FGVC-Aircraft, and DTD) demonstrate that our approach consistently outperforms existing SEMs by a notable margin, with 4.2%-8.7% relative gains in 5-way 5-shot classification. These findings highlight the efficacy of coupling concept learning with few-shot adaptation to achieve both higher accuracy and clearer model interpretability, paving the way for more transparent visual recognition systems. Zhong Ji, Rongshuai Wei, Jingren Liu, Yanwei Pang, Jungong Han |
IEEE Trans. Image Process. | 1 |
| 2026 | Multi-Stage Knowledge Integration of Vision-Language Models for Continual LearningabstractVision Language Models (VLMs), pre-trained on large-scale image-text datasets, enable zero-shot predictions for unseen data but may underperform on specific unseen tasks. Continual learning (CL) can help VLMs effectively adapt to new data distributions without joint training, but faces challenges of catastrophic forgetting and generalization forgetting. Although significant progress has been achieved by distillation-based methods, they exhibit two severe limitations. One is the popularly adopted single-teacher paradigm fails to impart comprehensive knowledge, The other is the existing methods inadequately leverage the multimodal information in the original training dataset, instead they rely on additional data for distillation, which increases computational and storage overhead. To mitigate both limitations, by drawing on Knowledge Integration Theory (KIT), we propose a Multi-Stage Knowledge Integration network (MulKI) to emulate the human learning process in distillation methods. MulKI achieves this through four stages, including Eliciting Ideas, Adding New Ideas, Distinguishing Ideas, and Making Connections. During the four stages, we first leverage prototypes to align across modalities, eliciting cross-modal knowledge, then adding new knowledge by constructing fine-grained intra- and inter-modality relationships with prototypes. After that, knowledge from two teacher models is adaptively distinguished and re-weighted. Finally, we connect between models from intra- and inter-task, integrating preceding and new knowledge. Our method demonstrates significant improvements in maintaining zero-shot capabilities while supporting continual learning across diverse downstream tasks, showcasing its potential in adapting VLMs to evolving data distributions. Zhong Ji, Jingren Liu, Yanwei Pang, Jungong Han |
IEEE Trans. Image Process. | 2 |
| 2025 | SGD: Street View Synthesis with Gaussian Splatting and Diffusion PriorabstractNovel View Synthesis (NVS) for street scenes plays a critical role in the autonomous driving simulation. Current mainstream methods, such as Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS), struggle to maintain rendering quality at the viewpoint that deviates significantly from the training viewpoints. This issue stems from the sparse training views captured by a fixed camera on a moving vehicle. To tackle this problem, we propose a novel approach that enhances the capacity of 3DGS by leveraging prior from a Diffusion Model along with complementary multi-modal data. Specifically, we first fine-tune a Diffusion Model by adding images from adjacent frames as condition, meanwhile exploiting depth data from LiDAR point clouds to supply additional spatial information. Then we apply the fine-tuned Diffusion Model to regularize the 3DGS at unseen views during training. Experimental results validate the effectiveness of our method compared with current state-of-the-art models, and demonstrate its advance in rendering images from broader views. Zhongrui Yu, Haoran Wang 0004, Jinze Yang, Jiale Cao, Zhong Ji, Mingming Sun 0001 |
WACV | 6 |
| 2025 | Concept agent network for zero-base generalized few-shot learning
Xuan Wang 0016, Zhong Ji, Xiyao Liu 0002, Yanwei Pang, Xuelong Li 0001 |
Appl. Intell. | 2 |
| 2025 | Mitigating forgetting in the adaptation of CLIP for few-shot classification
Jiale Cao, Yuanheng Liu, Zhong Ji, Jingren Liu, Ai-Ping Yang, Yanwei Pang |
Comput. Vis. Image Underst. | 3 |
| 2025 | A lightweight cross-axis semantic interaction network with receptive-field-based attention for industrial surface defect detection
Xuening Li, Yan Xu 0016, Zhong Ji, Shouxiang Wang, Qianyu Zhao |
Expert Syst. Appl. | 3 |
| 2025 | Video Wire Inpainting via Hierarchical Feature Mixture
Zhong Ji, Yimu Su, Yan Zhang 0135, Shuangming Yang, Yanwei Pang |
Image Vis. Comput. | 1 |
| 2025 | Hierarchical and complementary experts transformer with momentum invariance for image-text retrieval
Yan Zhang 0135, Zhong Ji, Yanwei Pang, Jungong Han |
Knowl. Based Syst. | 2 |
| 2025 | Radlora: a smart low-rank adaptive approach for radiological image classification
Yuze Gao, Jingren Liu, Zhong Ji |
Multim. Syst. | 5 |
| 2025 | Token-aware and step-aware acceleration for Stable Diffusion
Ting Zhen, Jiale Cao, Xuebin Sun, Zhong Ji, Yanwei Pang |
Pattern Recognit. | 5 |
| 2025 | Bidirectional Error-Aware Fusion Network for Video InpaintingabstractExisting video inpainting approaches tend to adopt vision transformers with rare customized designs, which poses two limitations. Firstly, the conventional self-attention mechanism treats tokens from invalid and valid regions equally and mingles them, which may incur blurriness. Secondly, these approaches merely employ forward frames as references, while ignoring the past inpainted frames, which are also valuable in enhancing temporal consistency and offering more available information. In this paper, we propose a new video inpainting network, called Bidirectional Error-Aware Fusion Network (BEAF-Net). Concretely, on one hand, we propose a tailored Error-Aware Transformer (EAT) that discerns different tokens by assigning dynamic weights to bridle the use of erroneous tokens. Meanwhile, each EAT is equipped with a Spatial Feature Enhancement (SFE) layer to synthesize features with multi-scales. On the other hand, we apply a pair of EATs to utilize forward reference frames and past inpainted frames simultaneously, and a proposed Bidirectional Fusion (BiF) layer is exerted to blend the aggregation results adaptively. By coupling these novel designs, our proposed BEAF-Net completely leverages the location priors, multi-scale perception, and past predictions to produce more faithful and consistent inpainting results. We corroborate our BEAF-Net on two commonly-used video inpainting datasets: DAVIS and Youtube-VOS, where the experimental results demonstrate BEAF-Net compares favorably with state-of-the-art solutions. Video examples can be found athttps://github.com/JCATCV/BEAF-Net. Zhong Ji, Feng Zheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Raformer: Redundancy-Aware Transformer for Video Wire InpaintingabstractVideo Wire Inpainting (VWI) is a prominent application in video inpainting, aimed at flawlessly removing wires in films or TV series, offering significant time and labor savings compared to manual frame-by-frame removal. However, wire removal poses greater challenges due to the wires being longer and slimmer than objects typically targeted in general video inpainting tasks, and often intersecting with people and background objects irregularly, which adds complexity to the inpainting process. Recognizing the limitations posed by existing video wire datasets, which are characterized by their small size, poor quality, and limited variety of scenes, we introduce a new VWI dataset with a novel mask generation strategy, namely Wire Removal Video Dataset 2 (WRV2) and Pseudo Wire-Shaped (PWS) Masks. WRV2 dataset comprises over 4,000 videos with an average length of 80 frames, designed to facilitate the development and efficacy of inpainting models. Building upon this, our research proposes the Redundancy-Aware Transformer (Raformer) method that addresses the unique challenges of wire removal in video inpainting. Unlike conventional approaches that indiscriminately process all frame patches, Raformer employs a novel strategy to selectively bypass redundant parts, such as static background segments devoid of valuable information for inpainting. At the core of Raformer is the Redundancy-Aware Attention (RAA) module, which isolates and accentuates essential content through a coarse-grained, window-based attention mechanism. This is complemented by a Soft Feature Alignment (SFA) module, which refines these features and achieves end-to-end feature alignment. Extensive experiments on both the traditional video inpainting datasets and our proposed WRV2 dataset demonstrate that Raformer outperforms other state-of-the-art methods. Our codes and the WRV2 dataset will be made available at: https://github.com/Suyimu/WRV2. Zhong Ji, Yimu Su, Yan Zhang 0135, Yanwei Pang, Jungong Han |
IEEE Trans. Image Process. | 1 |
| 2025 | Frequency-Spatial Complementation: Unified Channel-Specific Style Attack for Cross-Domain Few-Shot LearningabstractCross-Domain Few-Shot Learning (CD-FSL) addresses the challenges of recognizing targets with out-of-domain data when only a few instances are available. Many current CD-FSL approaches primarily focus on enhancing the generalization capabilities of models in spatial domain, which neglects the role of the frequency domain in domain generalization. To take advantage of frequency domain in processing global information, we propose a Frequency-Spatial Complementation (FSC) model, which combines frequency domain information with spatial domain information to learn domain-invariant information from attacked data style. Specifically, we design a Frequency and Spatial Fusion (FusionFS) module to enhance the ability of the model to capture style-related information. Besides, we propose two attack strategies, i.e., the Gradient-guided Unified Style Attack (GUSA) strategy and the Channel-specific Attack Intensity Calculation (CAIC) strategy, which conduct targeted attacks on different channels to provide more diversified style data during the training phase, especially in single-source domain scenarios where the source domain data style is homogeneous. Extensive experiments across eight target domains demonstrate that our method significantly improves the model's performance under various styles. Zhong Ji, Zhilong Wang 0001, Xiyao Liu 0002, Yunlong Yu 0001, Yanwei Pang, Jungong Han |
IEEE Trans. Image Process. | 1 |
| 2025 | Balancing Feature Alignment and Uniformity for Few-Shot ClassificationabstractIn Few-Shot Learning (FSL), the objective is to correctly recognize new samples from novel classes with only a few available samples per class. Existing methods in FSL primarily focus on learning transferable knowledge from base classes by maximizing the information between feature representations and their corresponding labels. However, this approach may suffer from the "supervision collapse" issue, which arises due to a bias towards the base classes. In this paper, we propose a solution to address this issue by preserving the intrinsic structure of the data and enabling the learning of a generalized model for the novel classes. Following the InfoMax principle, our approach maximizes two types of mutual information (MI): between the samples and their feature representations, and between the feature representations and their class labels. This allows us to strike a balance between discrimination (capturing class-specific information) and generalization (capturing common characteristics across different classes) in the feature representations. To achieve this, we adopt a unified framework that perturbs the feature embedding space using two low-bias estimators. The first estimator maximizes the MI between a pair of intra-class samples, while the second estimator maximizes the MI between a sample and its augmented views. This framework effectively combines knowledge distillation between class-wise pairs and enlarges the diversity in feature representations. By conducting extensive experiments on popular FSL benchmarks, our proposed approach achieves comparable performances with state-of-the-art competitors. For example, we achieved an accuracy of 69.53% on the miniImageNet dataset and 77.06% on the CIFAR-FS dataset for the 5-way 1-shot task. Yunlong Yu 0001, Dingyi Zhang, Zhong Ji, Xi Li 0001, Jungong Han, Zhongfei Zhang |
IEEE Trans. Image Process. | 3 |
| 2025 | Visual Semantic Contextualization Network for Multi-Query Image RetrievalabstractMulti-Query Image Retrieval (MQIR) aims to establish connections between vision and language by exploring fine-grained region-query alignments. It is still a challenging task owing to its intrinsical ambiguity, where a query matches with multiple semantically similar regions and introduces misleading noises. Although researchers have made great efforts to alleviate the ambiguity in many retrieval-related tasks, there are few attempts considering this bottleneck in MQIR, which greatly limits present performance. To this end, we propose a novel Visual Semantic Contextualization Network (VSCN) to mitigate ambiguity by capturing the contextual knowledge within each image-text pair. Specifically, we first develop a Context Semantic Perception (CSP) module to capture the dual-level context, where a visual context transformer explores the intra-context within regions, and a cross-modal context transformer mines the inter-context among concatenated visual-linguistic embeddings. Then, to yield superior contextual understanding, we strengthen the connotations in context via a Context Semantic Interaction (CSI) module. Particularly, knowledge distillation is first employed to transfer the CLIP-guided semantic into the regional intra-context to complement the potential background information. Then, the intra-context & inter-context interaction is conducted via the self-attention mechanism to link the dual-level context and obtain the interacted contextual knowledge. Our method is evaluated on the Visual Genome dataset and substantially outperforms the state-of-the-art methods (30.3% improvements on Recall@1 in the first round). Our source codes will be released athttps://github.com/zhli-cs/VSCN. Zhong Ji, Zhihao Li 0006, Yan Zhang 0135, Yanwei Pang, Xuelong Li 0001 |
IEEE Trans. Multim. | 1 |
| 2025 | Imbalance Mitigation for Continual Learning via Knowledge Decoupling and Dual Enhanced Contrastive LearningabstractContinual learning (CL) aims at studying how to learn new knowledge continuously from data streams without catastrophically forgetting the previous knowledge. One of the key problems is catastrophic forgetting, that is, the performance of the model on previous tasks declines significantly after learning the subsequent task. Several studies addressed it by replaying samples stored in the buffer when training new tasks. However, the data imbalance between old and new task samples results in two serious problems: information suppression and weak feature discriminability. The former refers to the information in the sufficient new task samples suppressing that in the old task samples, which is harmful to maintaining the knowledge since the biased output worsens the consistency of the same sample's output at different moments. The latter refers to the feature representation being biased to the new task, which lacks discrimination to distinguish both old and new tasks. To this end, we build an imbalance mitigation for CL (IMCL) framework that incorporates a decoupled knowledge distillation (DKD) approach and a dual enhanced contrastive learning (DECL) approach to tackle both problems. Specifically, the DKD approach alleviates the suppression of the new task on the old tasks by decoupling the model output probability during the replay stage, which better maintains the knowledge of old tasks. The DECL approach enhances both low- and high-level features and fuses the enhanced features to construct contrastive loss to effectively distinguish different tasks. Extensive experiments on three popular datasets show that our method achieves promising performance under task incremental learning (Task-IL), class incremental learning (Class-IL), and domain incremental learning (Domain-IL) settings. Zhong Ji, Zhanyu Jiao, Qiang Wang 0056, Yanwei Pang, Jungong Han |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | Conformal Loss-Controlling PredictionabstractConformal prediction (CP) is a learning framework controlling prediction coverage of prediction sets, which can be built on any learning algorithm for point prediction. This work proposes a learning framework named conformal loss-controlling prediction, which extends CP to the situation where the value of a loss function needs to be controlled. Different from existing works about risk-controlling prediction sets and conformal risk control with the purpose of controlling the expected values of loss functions, the proposed approach in this article focuses on the loss for any test object, which is an extension of CP from miscoverage loss to some general loss. The controlling guarantee is proved under the assumption of exchangeability of data in finite-sample cases and the framework is tested empirically for classification with a class-varying loss and statistical postprocessing of numerical weather forecasting applications, which are introduced as point-wise classification and point-wise regression problems. All theoretical analysis and experimental results confirm the effectiveness of our loss-controlling approach. Di Wang 0026, Ping Wang 0015, Zhong Ji, Hong-Yue Li |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | On the Approximation Risk of Few-Shot Class-Incremental Learning
Xuan Wang 0016, Zhong Ji, Xiyao Liu 0002, Yanwei Pang, Jungong Han |
ECCV (51) | 2 |
| 2024 | Eliminate Before Align: A Remote Sensing Image-Text Retrieval Framework with Keyword Explicit ReasoningabstractMountains of researches center around the Remote Sensing Image-Text Retrieval (RSITR), aiming at retrieving the corresponding targets based on the given query. Among them, the transfer of Foundation Models (FMs), such as CLIP, to remote sensing domain shows promising results. However, existing FM-based approaches neglect the negative impact of weakly correlated sample pairs and the key distinctions among remote sensing texts, leading to biased and superficial exploration of sample pairs. To address these challenges, we propose a novel Eliminate Before Align strategy with Keyword Explicit Reasoning framework (EBAKER) for RSITR. Specifically, we devise an innovative Eliminate Before Align (EBA) strategy to filter out the weakly correlated sample pairs to mitigate their deviations from optimal embedding space during alignment. Moreover, we introduce a Keyword Explicit Reasoning (KER) module to facilitate the positive role of subtle key concept differences. Without bells and whistles, our method achieves a one-step transformation from FM to RSITR task, obviating the necessity for extra pretraining on remote sensing data. Extensive experiments on three popular benchmark datasets validate that our proposed EBAKER method outperform the state-of-the-art methods with fewer training data. Our source code will be released soon. Zhong Ji, Changxu Meng, Yan Zhang 0135, Haoran Wang 0004, Yanwei Pang, Jungong Han |
ACM Multimedia | 1 |
| 2024 | Uncertainty-aware enhanced dark experience replay for continual learning
Qiang Wang 0056, Zhong Ji, Yanwei Pang, Zhongfei Zhang |
Appl. Intell. | 2 |
| 2024 | Modality-experts coordinated adaptation for large multimodal models
Yan Zhang 0135, Zhong Ji, Yanwei Pang, Jungong Han, Xuelong Li 0001 |
Sci. China Inf. Sci. | 2 |
| 2024 | A cognition-driven framework for few-shot class-incremental learning
Xuan Wang 0016, Zhong Ji, Yanwei Pang, Yunlong Yu 0001 |
Neurocomputing | 2 |
| 2024 | Unified feature learning network for few-shot fault diagnosis
Yan Xu 0016, Xinyao Ma, Xuan Wang 0016, Jinjia Wang, Zhong Ji |
Neurocomputing | 6 |
| 2024 | Towards Unsupervised Referring Expression Comprehension with Visual Semantic Parsing
Zhong Ji, Di Wang 0026, Yanwei Pang, Xuelong Li 0001 |
Knowl. Based Syst. | 2 |
| 2024 | Diversity-Infused Network for Unsupervised Few-Shot Remote Sensing Scene ClassificationabstractFew-shot Remote Sensing Scene Classification (RSSC) confronts challenges due to its dependence on extensive labeled datasets. Addressing this, we propose the Diversity-Infused Network (DIN), an unsupervised paradigm for few-shot RSSC, utilizing unlabeled data in training and adapting to novel classes with limited labeled samples. Within an augmentation-based framework, DIN includes a Random Augmentation Sampling (RAS) strategy for task diversity in the meta-training stage, and a Channel-Driven Metric Learning (CDML) module to decode complex channel interactions, enhancing information diversity. Additionally, DIN presents a Multi-Mutual Information (Multi-MI) objective function to balance the architecture and reduce the unreliability and potential biases from pseudo-training. Experimental results demonstrate that DIN surpasses other unsupervised approaches by over 11% in 1-shot and nearly 9% in 5-shot settings on WHU-RS19, and closely approaches supervised methods, with less than 1% difference in both settings. Liyuan Hou 0002, Zhong Ji, Xuan Wang 0016, Yunlong Yu 0001, Yanwei Pang |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2024 | Hierarchical matching and reasoning for multi-query image retrieval
Zhong Ji, Zhihao Li 0006, Yan Zhang 0135, Haoran Wang 0004, Yanwei Pang, Xuelong Li 0001 |
Neural Networks | 1 |
| 2024 | Tolerant Self-Distillation for image classification
Mushui Liu, Yunlong Yu 0001, Zhong Ji, Jungong Han, Zhongfei Zhang |
Neural Networks | 3 |
| 2024 | Multi-task hierarchical convolutional network for visual-semantic cross-modal retrieval
Zhong Ji, Zhigang Lin, Haoran Wang 0004, Yanwei Pang, Xuelong Li 0001 |
Pattern Recognit. | 1 |
| 2024 | Progressive Semantic Reconstruction Network for Weakly Supervised Referring Expression GroundingabstractWeakly supervised Referring Expression Grounding (REG) aims to localize the target entity in an image based on a given expression, where the mapping between image regions and expressions is unknown during training. It faces two primary challenges. Firstly, conventional methods involve selecting regions to generate reconstructed texts for computing the backpropagation loss between regions and expressions. However, semantic deviations in text reconstruction may result in significant cross-modal bias, leading to substantial losses even in cases of correctly matched regions. Secondly, the absence of region-level ground truth in weakly supervised REG results in a lack of stable and reliable supervision during training. To tackle these challenges, we propose a Progressive Semantic Reconstruction Network (PSRN), which utilizes a two-level matching-reconstruction process based on the key triad and adaptive phrases, respectively. We leverage progressive semantic reconstruction with a three-staged training strategy to mitigate the deviations in the reconstructed texts. Additionally, we introduce a Constrained Interactions operation and an Attention Coordination mechanism to facilitate additional bidirectional supervision between the two matching processes. Experiments on three benchmark datasets of RefCOCO, RefCOCO+ and RefCOCOg demonstrate that the proposed PSRN has the competing results. Our source code will be released athttps://github.com/5jiahe/psrn. Zhong Ji, Jiahe Wu, Ai-Ping Yang, Jungong Han |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Matching Multi-Scale Feature Sets in Vision Transformer for Few-Shot ClassificationabstractRecently, Transformer-based few-shot classification methods are widely exploited. However, they only leverage feature information at a single scale, resulting in weak feature representations, which cannot fully capture the rich information contained in a limited number of images regarding diverse objects with different scales, even those belonging to the same category. To mitigate this issue, we propose a multi-scale feature sets matching scheme in vision Transformer for few-shot classification, and name it FSViT, which can sufficiently extract discriminative features from the few number of labeled support examples. Concretely, we establish a patch-based multi-scale feature representation based on the feature extractors of FSViT, where we introduce an attention-aware grid pooling operation to merge adjacent patches with various scales to obtain multi-scale feature sets. Moreover, we devise a multi-scale patch matching metric to aggregate the measurement of similarity over the multi-scale feature sets for few-shot classification. Extensive experiments demonstrate the effectiveness of the proposed FSViT in both 1-shot and 5-shot scenarios on standard single-domain and cross-domain few-shot classification, especially improving the state-of-the-art recognition accuracy by 1.27% and 1.33% on average on the Mini-ImageNet and CFAIR-FS datasets, respectively. The code of FSViT is available athttps://github.com/codeshop715/FSViT. Mingchen Song, Fengqin Yao, Guoqiang Zhong 0001, Zhong Ji, Xiaowei Zhang 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | MCD-Net: Toward RGB-D Video Inpainting in Real-World ScenesabstractVideo inpainting gains an increasing amount of attention ascribed to its wide applications in intelligent video editing. However, despite tremendous progress made in RGB video inpainting, the existing RGB-D video inpainting models are still incompetent to inpaint real-world RGB-D videos, as they simply fuse color and depth via explicit feature concatenation, neglecting the natural modality gap. Moreover, current RGB-D video inpainting datasets are synthesized with homogeneous and delusive RGB-D data, which is far from real-world application and cannot provide comprehensive evaluation. To alleviate these problems and achieve real-world RGB-D video inpainting, on one hand, we propose a Mutually-guided Color and Depth Inpainting Network (MCD-Net), where color and depth are reciprocally leveraged to inpaint each other implicitly, mitigating the modality gap and fully exploiting cross-modal association for inpainting. On the other hand, we build a Video Inpainting with Depth (VID) dataset to supply diverse and authentic RGB-D video data with various object annotation masks to enable comprehensive evaluation for RGB-D video inpainting under real-world scenes. Experimental results on the DynaFill benchmark and our collected VID dataset demonstrate our MCD-Net not only yields the state-of-the-art quantitative performance but successfully achieves high-quality RGB-D video inpainting under real-world scenes. All resources are available at https://github.com/JCATCV/MCD-Net. Zhong Ji, Chengjie Wang 0001, Feng Zheng 0001 |
IEEE Trans. Image Process. | 2 |
| 2024 | NTK-Guided Few-Shot Class Incremental LearningabstractThe proliferation of Few-Shot Class Incremental Learning (FSCIL) methodologies has highlighted the critical challenge of maintaining robust anti-amnesia capabilities in FSCIL learners. In this paper, we present a novel conceptualization of anti-amnesia in terms of mathematical generalization, leveraging the Neural Tangent Kernel (NTK) perspective. Our method focuses on two key aspects: ensuring optimal NTK convergence and minimizing NTK-related generalization loss, which serve as the theoretical foundation for cross-task generalization. To achieve global NTK convergence, we introduce a principled meta-learning mechanism that guides optimization within an expanded network architecture. Concurrently, to reduce the NTK-related generalization loss, we systematically optimize its constituent factors. Specifically, we initiate self-supervised pre-training on the base session to enhance NTK-related generalization potential. These self-supervised weights are then carefully refined through curricular alignment, followed by the application of dual NTK regularization tailored specifically for both convolutional and linear layers. Through the combined effects of these measures, our network acquires robust NTK properties, ensuring optimal convergence and stability of the NTK matrix and minimizing the NTK-related generalization loss, significantly enhancing its theoretical generalization. On popular FSCIL benchmark datasets, our NTK-FSCIL surpasses contemporary state-of-the-art approaches, elevating end-session accuracy by 2.9% to 9.3%. Jingren Liu, Zhong Ji, Yanwei Pang, Yunlong Yu 0001 |
IEEE Trans. Image Process. | 2 |
| 2024 | Model Attention Expansion for Few-Shot Class-Incremental LearningabstractFew-Shot Class-Incremental Learning (FSCIL) aims at incrementally learning new knowledge from limited training examples without forgetting previous knowledge. However, we observe that existing methods face a challenge known as supervision collapse, where the model disproportionately emphasizes class-specific features of base classes at the detriment of novel class representations, leading to restricted cognitive capabilities. To alleviate this issue, we propose a new framework, Model aTtention Expansion for Few-Shot Class-Incremental Learning (MTE-FSCIL), aimed at expanding the model attention fields to improve transferability without compromising the discriminative capability for base classes. Specifically, the framework adopts a dual-stage training strategy, comprising pre-training and meta-training stages. In the pre-training stage, we present a new regularization technique, named the Reserver (RS) loss, to expand the global perception and reduce over-reliance on class-specific features by amplifying feature map activations. During the meta-training stage, we introduce the Repeller (RP) loss, a novel pair-based loss that promotes variation in representations and improves the model's recognition of sample uniqueness by scattering intra-class samples within the embedding space. Furthermore, we propose a Transformational Adaptation (TA) strategy to enable continuous incorporation of new knowledge from downstream tasks, thus facilitating cross-task knowledge transfer. Extensive experimental results on mini-ImageNet, CIFAR100, and CUB200 datasets demonstrate that our proposed framework consistently outperforms the state-of-the-art methods. Xuan Wang 0016, Zhong Ji, Yunlong Yu 0001, Yanwei Pang, Jungong Han |
IEEE Trans. Image Process. | 2 |
| 2024 | USER: Unified Semantic Enhancement With Momentum Contrast for Image-Text RetrievalabstractAs a fundamental and challenging task in bridging language and vision domains, Image-Text Retrieval (ITR) aims at searching for the target instances that are semantically relevant to the given query from the other modality, and its key challenge is to measure the semantic similarity across different modalities. Although significant progress has been achieved, existing approaches typically suffer from two major limitations: (1) It hurts the accuracy of the representation by directly exploiting the bottom-up attention based region-level features where each region is equally treated. (2) It limits the scale of negative sample pairs by employing the mini-batch based end-to-end training mechanism. To address these limitations, we propose a Unified Semantic Enhancement Momentum Contrastive Learning (USER) method for ITR. Specifically, we delicately design two simple but effective Global representation based Semantic Enhancement (GSE) modules. One learns the global representation via the self-attention algorithm, noted as Self-Guided Enhancement (SGE) module. The other module benefits from the pre-trained CLIP module, which provides a novel scheme to exploit and transfer the knowledge from an off-the-shelf model, noted as CLIP-Guided Enhancement (CGE) module. Moreover, we incorporate the training mechanism of MoCo into ITR, in which two dynamic queues are employed to enrich and enlarge the scale of negative sample pairs. Meanwhile, a Unified Training Objective (UTO) is developed to learn from mini-batch based and dynamic queue based samples. Extensive experiments on the benchmark MSCOCO and Flickr30K datasets demonstrate the superiority of both retrieval accuracy and inference efficiency. For instance, compared with the existing best method NAAF, the metric R@1 of our USER on the MSCOCO 5K Testing set is improved by 5% and 2.4% on caption retrieval and image retrieval without any external knowledge or pre-trained model while enjoying over 60 times faster inference speed. Our source code will be released at https://github.com/zhangy0822/USER. Yan Zhang 0135, Zhong Ji, Di Wang 0026, Yanwei Pang, Xuelong Li 0001 |
IEEE Trans. Image Process. | 2 |
| 2024 | Semantic-Aware Dynamic Generation Networks for Few-Shot Human-Object Interaction RecognitionabstractRecognizing human-object interaction (HOI) aims at inferring various relationships between actions and objects. Although great progress in HOI has been made, the long-tail problem and combinatorial explosion problem are still practical challenges. To this end, we formulate HOI as a few-shot task to tackle both challenges and design a novel dynamic generation method to address this task. The proposed approach is called semantic-aware dynamic generation networks (SADG-Nets). Specifically, SADG-Net first assigns semantic-aware task representations for different batches of data, which further generates dynamic parameters. It obtains the features that highlight intercategory discriminability and intracategory commonality adaptively. In addition, we also design a dual semantic-aware encoder module (DSAE-Module), that is, verb-aware and noun-aware branches, to yield both action and object prototypes of HOI for each task space, which generalizes to novel combinations by transferring similarities among interactions. Extensive experimental results on two benchmark datasets, that is, humans interacting with common objects (HICO)-FS and trento universal HOI (TUHOI)-FS, illustrate that our SADG-Net achieves superior performance over state-of-the-art approaches, which proves its impressive effectiveness on few-shot HOI recognition. Zhong Ji, Xiyao Liu 0002, Changxin Gao, Yanwei Pang, Ling Shao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Multi-feature self-attention super-resolution network
Ai-Ping Yang, Zihao Wei, Jinbin Wang, Jiale Cao, Zhong Ji, Yanwei Pang |
Vis. Comput. | 5 |
| 2023 | Lightweight MIMO-WNet for single image deblurring
Mushui Liu, Yingming Li, Zhong Ji |
Neurocomputing | 4 |
| 2023 | Mutual mentor: Online contrastive distillation network for general continual learning
Qiang Wang 0056, Zhong Ji, Jin Li 0054, Yanwei Pang |
Neurocomputing | 2 |
| 2023 | Self-taught cross-domain few-shot learning with weakly supervised object localization and task-decomposition
Xiyao Liu 0002, Zhong Ji, Yanwei Pang, Zhi Han |
Knowl. Based Syst. | 2 |
| 2023 | Spatial attention-guided deformable fusion network for salient object detection
Ai-Ping Yang, Simeng Cheng, Jiale Cao, Zhong Ji, Yanwei Pang |
Multim. Syst. | 5 |
| 2023 | Zero-shot classification with unseen prototype learning
Zhong Ji, Biying Cui, Yunlong Yu 0001, Yanwei Pang, Zhongfei Zhang |
Neural Comput. Appl. | 1 |
| 2023 | Dual Distillation Discriminator Networks for Domain Adaptive Few-Shot Learning
Xiyao Liu 0002, Zhong Ji, Yanwei Pang, Zhi Han |
Neural Networks | 2 |
| 2023 | Visual-quality-driven unsupervised image dehazing
Ai-Ping Yang, Jinbin Wang, Jiale Cao, Zhong Ji, Yanwei Pang |
Neural Networks | 6 |
| 2023 | COREN: Multi-Modal Co-Occurrence Transformer Reasoning Network for Image-Text Retrieval
Zhong Ji, Yanwei Pang, Zhongfei Zhang |
Neural Process. Lett. | 2 |
| 2023 | G2LP-Net: Global to Local Progressive Video Inpainting NetworkabstractThe self-attention based video inpainting methods have achieved promising progress by establishing long-range correlation over the whole video. However, existing methods generally relied on the global self-attention that directly searches missing contents among all reference frames but lacks accurate matching and effective organization on contents, which often blurs the result owing to the loss of local textures. In this paper, we propose a Global-to-Local Progressive Inpainting Network (G2LP-Net) consisting of the following innovative ideas. First, we present a global to local self-attention mechanism by incorporating local self-attention into global self-attention to improve searching efficiency and accuracy, where the self-attention is implemented in multi-scale regions to fully exploit local redundancy for the texture recovery. Second, we propose a progressive video inpainting (PVI) method to organize the generated contents, which completes the target video frames from periphery to core to ensure reliable contents serve first. Last, we develop a window-sliding method for sampling reference frames to obtain rich available information for inpainting. In addition, we release a wire-removal video (WRV) dataset that consists of 150 video clips masked by wires to evaluate the video inpainting on irregularly slender regions. Both quantitative and qualitative experiments on benchmark datasets, DAVIS, YouTube-VOS and our WRV dataset have demonstrated the superiority of our proposed G2LP-Net method. Zhong Ji, Yimu Su, Yanwei Pang, Xuelong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Consensus Knowledge Exploitation for Partial Query Based Image RetrievalabstractPartial Query based Image Retrieval (PQIR) enables a search engine to perform an interactive retrieval given by only an initial query and actively provide alternative feedbacks for a user to refine a set of retrieval results. It alleviates the deficiency in practice interactive image retrieval that requires the user to laboriously provide detailed feedbacks, and enables the retrieval on-the-fly with the incomplete initial query. Although significant progress has been made, existing works remain have challenge in actively providing more discriminative feedbacks. To address this challenge, we propose a novel Attributes&Objects-based Consensus Extraction and Representation (AoCer) framework. Specifically, we formulate a simple but effective Attribute&Object Feedback (AOF) paradigm, which employs both attributes and objects as intermediate feedbacks to carry out multiple rounds of interaction. To mine the intrinsic associations among concepts and enhance their feature representations, we further propose an Interventional Consensus Representation Learning (ICRL) module, which mainly constructs an interventional concept graph to yield the Interventional Consensus Representation (ICR). In addition, a Dual-Head Feedback Sampler (DHFS) is developed to sample objects and attributes for conducting the next round retrieval. Extensive experiments demonstrate the superiority of the proposed framework. Our source code will be released athttps://github.com/zhangy0822/AoCer. Yan Zhang 0135, Zhong Ji, Yanwei Pang, Xuelong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Dual Contrastive Network for Few-Shot Remote Sensing Image Scene ClassificationabstractFew-shot remote sensing image scene classification (FS-RSISC) aims at classifying remote sensing images with only a few labeled samples. The main challenges lie in small inter-class variances and large intra-class variances, which are the inherent property of remote sensing images. To address these challenges, we propose a transfer-based Dual Contrastive Network (DCN), which incorporates two auxiliary supervised contrastive learning branches during the training process. Specifically, one is a Context-guided Contrastive Learning (CCL) branch and the other is a Detail-guided Contrastive Learning (DCL) branch, which focus on inter-class discriminability and intra-class invariance, respectively. In the CCL branch, we first devise a Condenser Network to capture context features, and then leverage a supervised contrastive learning on top of the obtained context features to facilitate the model to learn more discriminative features. In the DCL branch, a Smelter Network is designed to highlight the significant local detail information. And then we construct a supervised contrastive learning based on the detail feature maps to fully exploit the spatial information in each map, enabling the model to concentrate on invariant detail features. Extensive experiments on four public benchmark remote sensing datasets demonstrate the competitive performance of our proposed DCN. Zhong Ji, Liyuan Hou 0002, Xuan Wang 0016, Gang Wang 0060, Yanwei Pang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Knowledge-Aided Momentum Contrastive Learning for Remote-Sensing Image Text RetrievalabstractRemote sensing image-text retrieval (RSITR) has attracted widespread attention due to its great potential for rapid information mining ability on remote sensing images. Although significant progress has been achieved, existing methods typically overlook the challenge posed by the extremely analogous descriptions, where the subtle differences remain largely unexploited or, in some cases, are entirely disregarded. To address the limitation, we propose a Knowledge Aided Momentum Contrastive Learning (KAMCL) method for RSITR. Specifically, we propose a novel Knowledge Aided Learning framework, including knowledge initialization, construction, filtration, and alignment operations, which aims at providing valuable concepts and learning discriminative representations. On this basis, we integrate Momentum Contrastive Learning to promote the capture of key concepts within the representation via expanding the scale of negative sample pairs. Moreover, we design a hierarchical aggregator module to better capture the multi-level information from remote sensing images. Finally, we introduce an innovative two-step training strategy designed to effectively harness the synergy among concepts and leverage their respective functionalities. Extensive experiments conducted on the three public datasets showcase the remarkable performance of our approach in terms of retrieval accuracy and computational efficiency. For instance, compared with the existing state-of-the-art method, our method exhibits notable performance improvements of 2.65% on the RSICD dataset, simultaneously achieving improvements in inference efficiency by 48%. Our source code will be released at https://github.com/mcx-mcx/KAMCL. Zhong Ji, Changxu Meng, Yan Zhang 0135, Yanwei Pang, Xuelong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Memorizing Complementation Network for Few-Shot Class-Incremental LearningabstractFew-shot Class-Incremental Learning (FSCIL) aims at learning new concepts continually with only a few samples, which is prone to suffer the catastrophic forgetting and overfitting problems. The inaccessibility of old classes and the scarcity of the novel samples make it formidable to realize the trade-off between retaining old knowledge and learning novel concepts. Inspired by that different models memorize different knowledge when learning novel concepts, we propose a Memorizing Complementation Network (MCNet) to ensemble multiple models that complements the different memorized knowledge with each other in novel tasks. Additionally, to update the model with few novel samples, we develop a Prototype Smoothing Hard-mining Triplet (PSHT) loss to push the novel samples away from not only each other in current task but also the old distribution. Extensive experiments on three benchmark datasets, e.g., CIFAR100, miniImageNet and CUB200, have demonstrated the superiority of our proposed method. Zhong Ji, Zhishen Hou, Xiyao Liu 0002, Yanwei Pang, Xuelong Li 0001 |
IEEE Trans. Image Process. | 1 |
| 2023 | Complementary Calibration: Boosting General Continual Learning With Collaborative Distillation and Self-SupervisionabstractGeneral Continual Learning (GCL) aims at learning from non independent and identically distributed stream data without catastrophic forgetting of the old tasks that don't rely on task boundaries during both training and testing stages. We reveal that the relation and feature deviations are crucial problems for catastrophic forgetting, in which relation deviation refers to the deficiency of the relationship among all classes in knowledge distillation, and feature deviation refers to indiscriminative feature representations. To this end, we propose a Complementary Calibration (CoCa) framework by mining the complementary model's outputs and features to alleviate the two deviations in the process of GCL. Specifically, we propose a new collaborative distillation approach for addressing the relation deviation. It distills model's outputs by utilizing ensemble dark knowledge of new model's outputs and reserved outputs, which maintains the performance of old tasks as well as balancing the relationship among all classes. Furthermore, we explore a collaborative self-supervision idea to leverage pretext tasks and supervised contrastive learning for addressing the feature deviation problem by learning complete and discriminative features for all classes. Extensive experiments on six popular datasets show that our CoCa framework achieves superior performance against state-of-the-art methods. Code is available at https://github.com/lijincm/CoCa. Zhong Ji, Jin Li 0054, Qiang Wang 0056, Zhongfei Zhang |
IEEE Trans. Image Process. | 1 |
| 2023 | Asymmetric Cross-Scale Alignment for Text-Based Person SearchabstractText-based person search (TBPS) is of significant importance in intelligent surveillance, which aims to retrieve pedestrian images with high semantic relevance to a given text description. This retrieval task is characterized with both modal heterogeneity and fine-grained matching. To implement this task, one needs to extract multi-scale features from both image and text domains, and then perform the cross-modal alignment. However, most existing approaches only consider the alignment confined at their individual scales, e.g., an image-sentence or a region-phrase scale. Such a strategy adopts the presumable alignment in feature extraction, while overlooking the cross-scale alignment, e.g., image-phrase. In this paper, we present a transformer-based model to extract multi-scale representations, and perform Asymmetric Cross-Scale Alignment (ACSA) to precisely align the two modalities. Specifically, ACSA consists of a global-level alignment module and an asymmetric cross-attention module, where the former aligns an image and texts on a global scale, and the latter applies the cross-attention mechanism to dynamically align the cross-modal entities in region/image-phrase scales. Extensive experiments on two benchmark datasets CUHK-PEDES and RSTPReid demonstrate the effectiveness of our approach. Zhong Ji, Junhua Hu, Deyin Liu, Lin Wu 0001, Ye Zhao 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | Knowledge Distillation Classifier Generation Network for Zero-Shot LearningabstractIn this article, we present a conceptually simple but effective framework called knowledge distillation classifier generation network (KDCGN) for zero-shot learning (ZSL), where the learning agent requires recognizing unseen classes that have no visual data for training. Different from the existing generative approaches that synthesize visual features for unseen classifiers' learning, the proposed framework directly generates classifiers for unseen classes conditioned on the corresponding class-level semantics. To ensure the generated classifiers to be discriminative to the visual features, we borrow the knowledge distillation idea to both supervise the classifier generation and distill the knowledge with, respectively, the visual classifiers and soft targets trained from a traditional classification network. Under this framework, we develop two, respectively, strategies, i.e., class augmentation and semantics guidance, to facilitate the supervision process from the perspectives of improving visual classifiers. Specifically, the class augmentation strategy incorporates some additional categories to train the visual classifiers, which regularizes the visual classifier weights to be compact, under supervision of which the generated classifiers will be more discriminative. The semantics-guidance strategy encodes the class semantics into the visual classifiers, which would facilitate the supervision process by minimizing the differences between the generated and the real-visual classifiers. To evaluate the effectiveness of the proposed framework, we have conducted extensive experiments on five datasets in image classification, i.e., AwA1, AwA2, CUB, FLO, and APY. Experimental results show that the proposed approach performs best in the traditional ZSL task and achieves a significant performance improvement on four out of the five datasets in the generalized ZSL task. Yunlong Yu 0001, Bin Li 0038, Zhong Ji, Jungong Han, Zhongfei Zhang |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | CODER: Coupled Diversity-Sensitive Momentum Contrastive Learning for Image-Text Retrieval
Haoran Wang 0004, Dongliang He, Boyang Xia, Fu Li 0003, Zhong Ji, Errui Ding, Jingdong Wang 0001 |
ECCV (36) | 8 |
| 2022 | Learning from Students: Online Contrastive Distillation Network for General Continual LearningabstractThe goal of General Continual Learning (GCL) is to preserve learned knowledge and learn new knowledge with constant memory from an infinite data stream where task boundaries are blurry. Distilling the model's response of reserved samples between the old and the new models is an effective way to achieve promise performance on GCL. However, it accumulates the inherent old model's response bias and is not robust to model changes. To this end, we propose an Online Contrastive Distillation Network (OCD-Net) to tackle these problems, which explores the merit of the student model in each time step to guide the training process of the student model. Concretely, the teacher model is devised to help the student model to consolidate the learned knowledge, which is trained online via integrating the model weights of the student model to accumulate the new knowledge. Moreover, our OCD-Net incorporates both relation and adaptive response to help the student model alleviate the catastrophic forgetting, which is also beneficial for the teacher model preserves the learned knowledge. Extensive experiments on six benchmark datasets demonstrate that our proposed OCD-Net significantly outperforms state-of-the-art approaches in 3.26%~8.71% with various buffer sizes. Our code is available at https://github.com/lijincm/OCD-Net. Jin Li 0054, Zhong Ji, Gang Wang 0060, Qiang Wang 0056, Feng Gao 0005 |
IJCAI | 2 |
| 2022 | Masked Feature Generation Network for Few-Shot LearningabstractIn this paper, we present a feature-augmentation approach called Masked Feature Generation Network (MFGN) for Few-Shot Learning (FSL), a challenging task that attempts to recognize the novel classes with a few visual instances for each class. Most of the feature-augmentation approaches tackle FSL tasks via modeling the intra-class distributions. We extend this idea further to explicitly capture the intra-class variations in a one-to-many manner. Specifically, MFGN consists of an encoder-decoder architecture, with an encoder that performs as a feature extractor and extracts the feature embeddings of the available visual instances (the unavailable instances are seen to be masked), along with a decoder that performs as a feature generator and reconstructs the feature embeddings of the unavailable visual instances from both the available feature embeddings and the masked tokens. Equipped with this generative architecture, MFGN produces nontrivial visual features for the novel classes with limited visual instances. In extensive experiments on four FSL benchmarks, MFGN performs competitively and outperforms the state-of-the-art competitors on most of the few-shot classification tasks. Dingyi Zhang, Zhong Ji |
IJCAI | 3 |
| 2022 | Boosting Video-Text Retrieval with Explicit High-Level SemanticsabstractVideo-text retrieval (VTR) is an attractive yet challenging task for multi-modal understanding, which aims to search for relevant video (text) given a query (video). Existing methods typically employ completely heterogeneous visual-textual information to align video and text, whilst lacking the awareness of homogeneous high-level semantic information residing in both modalities. To fill this gap, in this work, we propose a novel visual-linguistic aligning model named HiSE for VTR, which improves the cross-modal representation by incorporating explicit high-level semantics. First, we explore the hierarchical property of explicit high-level semantics, and further decompose it into two levels, i.e. discrete semantics and holistic semantics. Specifically, for visual branch, we exploit an off-the-shelf semantic entity predictor to generate discrete high-level semantics. In parallel, a trained video captioning model is employed to output holistic high-level semantics. As for the textual modality, we parse the text into three parts including occurrence, action and entity. In particular, the occurrence corresponds to the holistic high-level semantics, meanwhile both action and entity represent the discrete ones. Then, different graph reasoning techniques are utilized to promote the interaction between holistic and discrete high-level semantics. Extensive experiments demonstrate that, with the aid of explicit high-level semantics, our method achieves the superior performance over state-of-the-art methods on three benchmark datasets, including MSR-VTT, MSVD and DiDeMo. Haoran Wang 0004, Dongliang He, Fu Li 0003, Zhong Ji, Jungong Han, Errui Ding |
ACM Multimedia | 5 |
| 2022 | Heterogeneous memory enhanced graph reasoning network for cross-modal retrieval
Zhong Ji, Yanwei Pang, Xuelong Li 0001 |
Sci. China Inf. Sci. | 1 |
| 2022 | Teachers cooperation: team-knowledge distillation for multiple cross-domain few-shot learning
Zhong Ji, Jingwei Ni, Xiyao Liu 0002, Yanwei Pang |
Frontiers Comput. Sci. | 1 |
| 2022 | Meta hyperbolic networks for zero-shot learning
Yan Xu 0016, Lifu Mu, Zhong Ji, Xiyao Liu 0002, Jungong Han |
Neurocomputing | 3 |
| 2022 | Saliency detection network with two-stream encoder and interactive decoder
Ai-Ping Yang, Simeng Cheng, Shangyang Song, Jinbin Wang, Zhong Ji, Yanwei Pang, Jiale Cao |
Neurocomputing | 5 |
| 2022 | Local spatial alignment network for few-shot learning
Yunlong Yu 0001, Dingyi Zhang, Sidi Wang, Zhong Ji, Zhongfei Zhang |
Neurocomputing | 4 |
| 2022 | Reinforced pedestrian attribute recognition with group optimization reward
Zhong Ji, Zhenfei Hu, Yanwei Pang |
Image Vis. Comput. | 1 |
| 2022 | Hierarchical Correlations Replay for Continual Learning
Qiang Wang 0056, Zhong Ji, Yanwei Pang, Zhongfei Zhang |
Knowl. Based Syst. | 3 |
| 2022 | Non-linear perceptual multi-scale network for single image super-resolution
Ai-Ping Yang, Jinbin Wang, Zhong Ji, Yanwei Pang, Jiale Cao, Zihao Wei |
Neural Networks | 4 |
| 2022 | SMAN: Stacked Multimodal Attention Network for Cross-Modal Image-Text RetrievalabstractThis article focuses on tackling the task of the cross-modal image-text retrieval which has been an interdisciplinary topic in both computer vision and natural language processing communities. Existing global representation alignment-based methods fail to pinpoint the semantically meaningful portion of images and texts, while the local representation alignment schemes suffer from the huge computational burden for aggregating the similarity of visual fragments and textual words exhaustively. In this article, we propose a stacked multimodal attention network (SMAN) that makes use of the stacked multimodal attention mechanism to exploit the fine-grained interdependencies between image and text, thereby mapping the aggregation of attentive fragments into a common space for measuring cross-modal similarity. Specifically, we sequentially employ intramodal information and multimodal information as guidance to perform multiple-step attention reasoning so that the fine-grained correlation between image and text can be modeled. As a consequence, we are capable of discovering the semantically meaningful visual regions or words in a sentence which contributes to measuring the cross-modal similarity in a more precise manner. Moreover, we present a novel bidirectional ranking loss that enforces the distance among pairwise multimodal instances to be closer. Doing so allows us to make full use of pairwise supervised information to preserve the manifold structure of heterogeneous pairwise data. Extensive experiments on two benchmark datasets demonstrate that our SMAN consistently yields competitive performance compared to state-of-the-art methods. Zhong Ji, Haoran Wang 0004, Jungong Han, Yanwei Pang |
IEEE Trans. Cybern. | 1 |
| 2022 | Semantic-Guided Class-Imbalance Learning Model for Zero-Shot Image ClassificationabstractIn this article, we focus on the task of zero-shot image classification (ZSIC) that equips a learning system with the ability to recognize visual images from unseen classes. In contrast to the traditional image classification, ZSIC more easily suffers from the class-imbalance issue since it is more concerned with the class-level knowledge transferring capability. In the real world, the sample numbers of different categories generally follow a long-tailed distribution, and the discriminative information in the sample-scarce seen classes is hard to transfer to the related unseen classes in the traditional batch-based training manner, which degrades the overall generalization ability a lot. To alleviate the class-imbalance issue in ZSIC, we propose a sample-balanced training process to encourage all training classes to contribute equally to the learned model. Specifically, we randomly select the same number of images from each class across all training classes to form a training batch to ensure that the sample-scarce classes contribute equally as those classes with sufficient samples during each iteration. Considering that the instances from the same class differ in class representativeness, we further develop an efficient semantic-guided feature fusion model to obtain the discriminative class visual prototype for the following visual-semantic interaction process via distributing different weights to the selected samples based on their class representativeness. Extensive experiments on three imbalanced ZSIC benchmark datasets for both traditional ZSIC and generalized ZSIC tasks demonstrate that our approach achieves promising results, especially for the unseen categories that are closely related to the sample-scarce seen categories. Besides, the experimental results on two class-balanced datasets show that the proposed approach also improves the classification performance against the baseline model. Zhong Ji, Xuejie Yu, Yunlong Yu 0001, Yanwei Pang, Zhongfei Zhang |
IEEE Trans. Cybern. | 1 |
| 2022 | DGIG-Net: Dynamic Graph-in-Graph Networks for Few-Shot Human-Object InteractionabstractFew-shot learning (FSL) for human-object interaction (HOI) aims at recognizing various relationships between human actions and surrounding objects only from a few samples. It is a challenging vision task, in which the diversity and interactivity of human actions result in great difficulty to learn an adaptive classifier to catch ambiguous interclass information. Therefore, traditional FSL methods usually perform unsatisfactorily in complex HOI scenes. To this end, we propose dynamic graph-in-graph networks (DGIG-Net), a novel graph prototypes framework to learn a dynamic metric space by embedding a visual subgraph to a task-oriented cross-modal graph for few-shot HOI. Specifically, we first build a knowledge reconstruction graph to learn latent representations for HOI categories by reconstructing the relationship among visual features, which generates visual representations under the category distribution of every task. Then, a dynamic relation graph integrates both reconstructible visual nodes and dynamic task-oriented semantic information to explore a graph metric space for HOI class prototypes, which applies the discriminative information from the similarities among actions or objects. We validate DGIG-Net on multiple benchmark datasets, on which it largely outperforms existing FSL approaches and achieves state-of-the-art results. Xiyao Liu 0002, Zhong Ji, Yanwei Pang, Jungong Han, Xuelong Li 0001 |
IEEE Trans. Cybern. | 2 |
| 2022 | Information Symmetry Matters: A Modal-Alternating Propagation Network for Few-Shot LearningabstractSemantic information provides intra-class consistency and inter-class discriminability beyond visual concepts, which has been employed in Few-Shot Learning (FSL) to achieve further gains. However, semantic information is only available for labeled samples but absent for unlabeled samples, in which the embeddings are rectified unilaterally by guiding the few labeled samples with semantics. Therefore, it is inevitable to bring a cross-modal bias between semantic-guided samples and nonsemantic-guided samples, which results in an information asymmetry problem. To address this problem, we propose a Modal-Alternating Propagation Network (MAP-Net) to supplement the absent semantic information of unlabeled samples, which builds information symmetry among all samples in both visual and semantic modalities. Specifically, the MAP-Net transfers the neighbor information by the graph propagation to generate the pseudo-semantics for unlabeled samples guided by the completed visual relationships and rectify the feature embeddings. In addition, due to the large discrepancy between visual and semantic modalities, we design a Relation Guidance (RG) strategy to guide the visual relation vectors via semantics so that the propagated information is more beneficial. Extensive experimental results on three semantic-labeled datasets, i.e., Caltech-UCSD-Birds 200-2011, SUN Attribute Database and Oxford 102 Flower, have demonstrated that our proposed method achieves promising performance and outperforms the state-of-the-art approaches, which indicates the necessity of information symmetry. Zhong Ji, Zhishen Hou, Xiyao Liu 0002, Yanwei Pang, Jungong Han |
IEEE Trans. Image Process. | 1 |
| 2022 | Task-Oriented High-Order Context Graph Networks for Few-Shot Human-Object Interaction RecognitionabstractFew-shot human-object interaction (FS-HOI) recognition aims at inferring new interactions between human actions and surrounding objects merely with a few available instances. It is beneficial to alleviate the long-tail and combinatorial explosion problems in human-object interaction (HOI). Nevertheless, the existing FS-HOI methods only focus on modeling the relationships between labeled samples and unlabeled samples in the Euclidean domain, which neglects the rich relational structures of the visual information among labeled samples and between human actions and objects. Accordingly, we tackle the few-shot HOI task in the non-Euclidean domain and present a graph-based model, namely, task-oriented high-order context graph network (THCG-Net). It contains a task attention module (TA-Module) and a high-order context graph module (HG-Module). In TA-Module, an attention mechanism is designed by utilizing task information to build a task-oriented space, in which the discriminative information for the current task (episode) is captured by embedding the visual features into the task-oriented space. The HG-Module is proposed to construct a task-level graph and takes the context information as high-order knowledge, which provides discriminative guidance for propagating visual information. It captures the discriminability among different categories while highlights the commonality of related categories adaptively, which effectively transfers knowledge to related categories. Extensive experimental results on two benchmark datasets, HICO-FS and TUHOI-FS, are provided. It demonstrates that our THCG-Net significantly outperforms the state-of-the-art approaches, which proves its impressive effectiveness in recognizing various human actions and surrounding objects in few-shot scenarios. Zhong Ji, Xiyao Liu 0002, Yanwei Pang, Ling Shao 0001, Zhongfei Zhang |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2021 | Step-Wise Hierarchical Alignment Network for Image-Text MatchingabstractImage-text matching plays a central role in bridging the semantic gap between vision and language. The key point to achieve precise visual-semantic alignment lies in capturing the fine-grained cross-modal correspondence between image and text. Most previous methods rely on single-step reasoning to discover the visual-semantic interactions, which lacks the ability of exploiting the multi-level information to locate the hierarchical fine-grained relevance. Different from them, in this work, we propose a step-wise hierarchical alignment network (SHAN) that decomposes image-text matching into multi-step cross-modal reasoning process. Specifically, we first achieve local-to-local alignment at fragment level, following by performing global-to-local and global-to-global alignment at context level sequentially. This progressive alignment strategy supplies our model with more complementary and sufficient semantic clues to understand the hierarchical correlations between image and text. The experimental results on two benchmark datasets demonstrate the superiority of our proposed method. Zhong Ji, Haoran Wang 0004 |
IJCAI | 1 |
| 2021 | Triple discriminator generative adversarial network for zero-shot image classification
Zhong Ji, Jiangtao Yan, Qiang Wang 0056, Yanwei Pang, Xuelong Li 0001 |
Sci. China Inf. Sci. | 1 |
| 2021 | Reweighting and information-guidance networks for Few-Shot Learning
Zhong Ji, Xingliang Chai, Yunlong Yu 0001, Zhongfei Zhang |
Neurocomputing | 1 |
| 2021 | TT-SVD: An Efficient Sparse Decision-Making Model With Two-Way Trust Recommendation in the AI-Enabled IoT SystemsabstractThe convergence of AI and IoT enables data to be quickly explored and turned into vital decisions, and however, there are still some challenging issues to be further addressed. For example, lacking of enough data in AI-based decision making [so-called sparse decision making (SDM)] will decrease the efficiency dramatically, or even disable the intelligent IoT networks. Taking the intelligent IoT networks as the network infrastructure, the recommendation systems have been facing such SDM problems. A naive solution is to introduce trust information. However, trust information may also face the difficulty of sparse trust evidence (also known as sparse trust problem). In our work, an accurate SDM model with two-way trust recommendation in the AI-enabled IoT systems is proposed, named TT-SVD. Our model incorporates both trust information and rating information more thoroughly, which can efficiently alleviate the above-mentioned sparse trust problem and therefore be able to solve the cold start and data sparsity problems. Specifically, we first consider the twofold trust influences from both trustees and trusters, which can be represented by a factor named trust propensity. To this end, we propose a dual model, including a truster model (TrusterSVD) and a trustee model (TrusteeSVD) based on an existing rating-only recommendation model called SVD++, which are integrated by the weighted average and yield the final model, TT-SVD. The experimental results show that our model outperforms the state-of-the-art, including SVD and TrustSVD in both the “all users” and “cold start users” cases, and the accuracy improvement can reach a maximum of 29%. Complexity analysis shows that our model is equally suitable for the case of large sparse data sets. In summary, our model can effectively solve the sparse decision problem by introducing the two-way trust recommendation, and hence improve the efficiency of the intelligent recommendation systems. Guangquan Xu, Litao Jiao, Meiqi Feng, Zhong Ji, Emmanouil A. Panaousis, Si Chen 0009, James Xi Zheng |
IEEE Internet Things J. | 5 |
| 2021 | Coordinating Experience Replay: A Harmonious Experience Retention approach for Continual Learning
Zhong Ji, Qiang Wang 0056, Zhongfei Zhang |
Knowl. Based Syst. | 1 |
| 2021 | A semi-supervised zero-shot image classification method based on soft-target
Zhong Ji, Qiang Wang 0056, Biying Cui, Yanwei Pang, Xianbin Cao 0001, Xuelong Li 0001 |
Neural Networks | 1 |
| 2021 | Few-Shot Human-Object Interaction Recognition With Semantic-Guided Attentive Prototypes NetworkabstractExtreme instance imbalance among categories and combinatorial explosion make the recognition of Human-Object Interaction (HOI) a challenging task. Few studies have addressed both challenges directly. Motivated by the success of few-shot learning that learns a robust model from a few instances, we formulate HOI as a few-shot task in a meta-learning framework to alleviate the above challenges. Due to the fact that the intrinsical characteristic of HOI is diverse and interactive, we propose a Semantic-guided Attentive Prototypes Network (SAPNet) framework to learn a semantic-guided metric space where HOI recognition can be performed by computing distances to attentive prototypes of each class. Specifically, the model generates attentive prototypes guided by the category names of actions and objects, which highlight the commonalities of images from the same class in HOI. In addition, we design two alternative prototypes calculation methods, i.e., Prototypes Shift (PS) approach and Hallucinatory Graph Prototypes (HGP) approach, which explore to learn a suitable category prototypes representations in HOI. Finally, in order to realize the task of few-shot HOI, we reorganize 2 HOI benchmark datasets with 2 split strategies, i.e., HICO-NN, TUHOI-NN, HICO-NF, and TUHOI-NF. Extensive experimental results on these datasets have demonstrated the effectiveness of our proposed SAPNet approach. Zhong Ji, Xiyao Liu 0002, Yanwei Pang, Wanli Ouyang, Xuelong Li 0001 |
IEEE Trans. Image Process. | 1 |
| 2021 | Deep Attentive Video Summarization With Distribution Consistency LearningabstractThis article studies supervised video summarization by formulating it into a sequence-to-sequence learning framework, in which the input and output are sequences of original video frames and their predicted importance scores, respectively. Two critical issues are addressed in this article: short-term contextual attention insufficiency and distribution inconsistency. The former lies in the insufficiency of capturing the short-term contextual attention information within the video sequence itself since the existing approaches focus a lot on the long-term encoder-decoder attention. The latter refers to the distributions of predicted importance score sequence and the ground-truth sequence is inconsistent, which may lead to a suboptimal solution. To better mitigate the first issue, we incorporate a self-attention mechanism in the encoder to highlight the important keyframes in a short-term context. The proposed approach alongside the encoder-decoder attention constitutes our deep attentive models for video summarization. For the second one, we propose a distribution consistency learning method by employing a simple yet effective regularization loss term, which seeks a consistent distribution for the two sequences. Our final approach is dubbed as Attentive and Distribution consistent video Summarization (ADSum). Extensive experiments on benchmark data sets demonstrate the superiority of the proposed ADSum approach against state-of-the-art approaches. Zhong Ji, Yanwei Pang, Xi Li 0001, Jungong Han |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2020 | SGAP-Net: Semantic-Guided Attentive Prototypes Network for Few-Shot Human-Object Interaction RecognitionabstractExtreme instance imbalance among categories and combinatorial explosion make the recognition of Human-Object Interaction (HOI) a challenging task. Few studies have addressed both challenges directly. Motivated by the success of few-shot learning that learns a robust model from a few instances, we formulate HOI as a few-shot task in a meta-learning framework to alleviate the above challenges. Due to the fact that the intrinsic characteristic of HOI is diverse and interactive, we propose a Semantic-Guided Attentive Prototypes Network (SGAP-Net) to learn a semantic-guided metric space where HOI recognition can be performed by computing distances to attentive prototypes of each class. Specifically, the model generates attentive prototypes guided by the category names of actions and objects, which highlight the commonalities of images from the same class in HOI. In addition, we design a novel decision method to alleviate the biases produced by different patterns of the same action in HOI. Finally, in order to realize the task of few-shot HOI, we reorganize two HOI benchmark datasets, i.e., HICO-FS and TUHOI-FS, to realize the task of few-shot HOI. Extensive experimental results on both datasets have demonstrated the effectiveness of our proposed SGAP-Net approach. Zhong Ji, Xiyao Liu 0002, Yanwei Pang, Xuelong Li 0001 |
AAAI | 1 |
| 2020 | GTNet: Generative Transfer Network for Zero-Shot Object DetectionabstractWe propose a Generative Transfer Network (GTNet) for zero-shot object detection (ZSD). GTNet consists of an Object Detection Module and a Knowledge Transfer Module. The Object Detection Module can learn large-scale seen domain knowledge. The Knowledge Transfer Module leverages a feature synthesizer to generate unseen class features, which are applied to train a new classification layer for the Object Detection Module. In order to synthesize features for each unseen class with both the intra-class variance and the IoU variance, we design an IoU-Aware Generative Adversarial Network (IoUGAN) as the feature synthesizer, which can be easily integrated into GTNet. Specifically, IoUGAN consists of three unit models: Class Feature Generating Unit (CFU), Foreground Feature Generating Unit (FFU), and Background Feature Generating Unit (BFU). CFU generates unseen features with the intra-class variance conditioned on the class semantic embeddings. FFU and BFU add the IoU variance to the results of CFU, yielding class-specific foreground and background features, respectively. We evaluate our method on three public datasets and the results demonstrate that our method performs favorably against the state-of-the-art ZSD approaches. Shizhen Zhao, Changxin Gao, Yuanjie Shao, Lerenhan Li, Changqian Yu, Zhong Ji, Nong Sang |
AAAI | 6 |
| 2020 | Episode-Based Prototype Generating Network for Zero-Shot LearningabstractWe introduce a simple yet effective episode-based training framework for zero-shot learning (ZSL), where the learning system requires to recognize unseen classes given only the corresponding class semantics. During training, the model is trained within a collection of episodes, each of which is designed to simulate a zero-shot classification task. Through training multiple episodes, the model progressively accumulates ensemble experiences on predicting the mimetic unseen classes, which will generalize well on the real unseen classes. Based on this training framework, we propose a novel generative model that synthesizes visual prototypes conditioned on the class semantic prototypes. The proposed model aligns the visual-semantic interactions by formulating both the visual prototype generation and the class semantic inference into an adversarial framework paired with a parameter-economic Multi-modal Cross-Entropy Loss to capture the discriminative information. Extensive experiments on four datasets under both traditional ZSL and generalized ZSL tasks show that our model outperforms the state-of-the-art approaches by large margins. Yunlong Yu 0001, Zhong Ji, Jungong Han, Zhongfei Zhang |
CVPR | 2 |
| 2020 | Consensus-Aware Visual-Semantic Embedding for Image-Text Matching
Haoran Wang 0004, Zhong Ji, Yanwei Pang |
ECCV (24) | 3 |
| 2020 | Deep attentive and semantic preserving video summarization
Zhong Ji, Fang Jiao, Yanwei Pang, Ling Shao 0001 |
Neurocomputing | 1 |
| 2020 | Dual triplet network for image zero-shot learning
Zhong Ji, Yanwei Pang, Ling Shao 0001 |
Neurocomputing | 1 |
| 2020 | Multimodal Alignment and Attention-Based Person Search via Natural Language DescriptionabstractVisual Internet of Things (VIoT) has been widely deployed in the field of social security. However, how to enable it to be intelligent is an urgent yet challenging task. In this article, we address the task of searching persons with natural language description query in a public safety surveillance system, which is a practical and demanding technique in VIoT. It is a fine-grained many-to-many cross-modal problem and more challenging than those with the image and the attribute as queries. The existing attempts are still weak in bridging the semantic gap between visual modality from different camera sensors and text modality from natural language descriptions. We propose a deep person search approach with a natural language description query by employing the attention mechanism (AM) and multimodal alignment (MA) method to supervise the cross-modal mapping. Particularly, the AM consists of two self-attention modules and one cross-attention module, where the former aims at learning discriminative representations and the latter supervises each other with their own information to offer accurate guidance to a common space. The MA approach contains three alignment processes with a novel cross-ranking loss function to make different matching pairs separable in a common space. Extensive experiments on large-scale CUHK-PEDES demonstrate the superiority of the proposed approach. Zhong Ji, Shengjia Li |
IEEE Internet Things J. | 1 |
| 2020 | Lightweight group convolutional network for single image super-resolution
Ai-Ping Yang, Bingwang Yang, Zhong Ji, Yanwei Pang, Ling Shao 0001 |
Inf. Sci. | 3 |
| 2020 | Multi-modal generative adversarial network for zero-shot learning
Zhong Ji, Junyue Wang, Yunlong Yu 0001, Zhongfei Zhang |
Knowl. Based Syst. | 1 |
| 2020 | Stacked squeeze-and-excitation recurrent residual network for visual-semantic matching
Haoran Wang 0004, Zhong Ji, Zhigang Lin, Yanwei Pang, Xuelong Li 0001 |
Pattern Recognit. | 2 |
| 2020 | Improved prototypical networks for few-Shot learning
Zhong Ji, Xingliang Chai, Yunlong Yu 0001, Yanwei Pang, Zhongfei Zhang |
Pattern Recognit. Lett. | 1 |
| 2020 | Pedestrian attribute recognition based on multiple time steps attention
Zhong Ji, Zhenfei Hu, Erlu He, Jungong Han, Yanwei Pang |
Pattern Recognit. Lett. | 1 |
| 2020 | Cross-modal guidance based auto-encoder for multi-video summarization
Zhong Ji, Yanwei Pang, Xuelong Li 0001 |
Pattern Recognit. Lett. | 1 |
| 2020 | Video Summarization With Attention-Based Encoder-Decoder NetworksabstractThis paper addresses the problem of supervised video summarization by formulating it as a sequence-to-sequence learning problem, where the input is a sequence of original video frames, and the output is a keyshot sequence. Our key idea is to learn a deep summarization network with attention mechanism to mimic the way of selecting the keyshots of human. To this end, we propose a novel video summarization framework named attentive encoder-decoder networks for video summarization (AVS), in which the encoder uses a bidirectional long short-term memory (BiLSTM) to encode the contextual information among the input video frames. As for the decoder, two attention-based LSTM networks are explored by using additive and multiplicative objective functions, respectively. Extensive experiments are conducted on two video summarization benchmark datasets, i.e., SumMe and TVSum. The results demonstrate the superiority of the proposed AVS-based approaches against the state-of-the-art approaches, with remarkable improvements on both datasets. Zhong Ji, Kailin Xiong, Yanwei Pang, Xuelong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | Deep Ranking for Image Zero-Shot Multi-Label ClassificationabstractDuring the past decade, both multi-label learning and zero-shot learning have attracted huge research attention, and significant progress has been made. Multi-label learning algorithms aim to predict multiple labels given one instance, while most existing zero-shot learning approaches target at predicting a single testing label for each unseen class via transferring knowledge from auxiliary seen classes to target unseen classes. However, relatively less effort has been made on predicting multiple labels in the zero-shot setting, which is nevertheless a quite challenging task. In this work, we investigate and formalize a flexible framework consisting of two components, i.e., visual-semantic embedding and zero-shot multi-label prediction. First, we present a deep regression model to project the visual features into the semantic space, which explicitly exploits the correlations in the intermediate semantic layer of word vectors and makes label prediction possible. Then, we formulate the label prediction problem as a pairwise one and employ Ranking SVM to seek the unique multi-label correlations in the embedding space. Furthermore, we provide a transductive multi-label zeroshot prediction approach that exploits the testing data manifold structure. We demonstrate the effectiveness of the proposed approach on three popular multi-label datasets with state-of-theart performance obtained on both conventional and generalized ZSL settings. Zhong Ji, Biying Cui, Yu-Gang Jiang 0001, Tao Xiang 0002, Timothy M. Hospedales, Yanwei Fu 0001 |
IEEE Trans. Image Process. | 1 |
| 2020 | Attribute-Guided Network for Cross-Modal Zero-Shot HashingabstractZero-shot hashing (ZSH) aims at learning a hashing model that is trained only by instances from seen categories but can generate well to those of unseen categories. Typically, it is achieved by utilizing a semantic embedding space to transfer knowledge from seen domain to unseen domain. Existing efforts mainly focus on single-modal retrieval task, especially image-based image retrieval (IBIR). However, as a highlighted research topic in the field of hashing, cross-modal retrieval is more common in real-world applications. To address the cross-modal ZSH (CMZSH) retrieval task, we propose a novel attribute-guided network (AgNet), which can perform not only IBIR but also text-based image retrieval (TBIR). In particular, AgNet aligns different modal data into a semantically rich attribute space, which bridges the gap caused by modality heterogeneity and zero-shot setting. We also design an effective strategy that exploits the attribute to guide the generation of hash codes for image and text within the same network. Extensive experimental results on three benchmark data sets (AwA, SUN, and ImageNet) demonstrate the superiority of AgNet on both cross-modal and single-modal zero-shot image retrieval tasks. Zhong Ji, Yunlong Yu 0001, Yanwei Pang, Jungong Han |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2019 | Saliency-Guided Attention Network for Image-Sentence MatchingabstractThis paper studies the task of matching image and sentence, where learning appropriate representations to bridge the semantic gap between image contents and language appears to be the main challenge. Unlike previous approaches that predominantly deploy symmetrical architecture to represent both modalities, we introduce a Saliency-guided Attention Network (SAN) that is characterized by building an asymmetrical link between vision and language to efficiently learn a fine-grained cross-modal correlation. The proposed SAN mainly includes three components: saliency detector, Saliency-weighted Visual Attention (SVA) module, and Saliency-guided Textual Attention (STA) module. Concretely, the saliency detector provides the visual saliency information to drive both two attention modules. Taking advantage of the saliency information, SVA is able to learn more discriminative visual features. By fusing the visual information from SVA and intra-modal information as a multi-modal guidance, STA affords us powerful textual representations that are synchronized with visual clues. Extensive experiments demonstrate SAN can improve the state-of-the-art results on the benchmark Flickr30K and MSCOCO datasets by a large margin. Zhong Ji, Haoran Wang 0004, Jungong Han, Yanwei Pang |
ICCV | 1 |
| 2019 | Dual-Path in Dual-Path Network for Single Image DehazingabstractRecently, deep learning-based single image dehazing method has been a popular approach to tackle dehazing. However, the existing dehazing approaches are performed directly on the original hazy image, which easily results in image blurring and noise amplifying. To address this issue, the paper proposes a DPDP-Net (Dual-Path in Dual-Path network) framework by employing a hierarchical dual path network. Specifically, the first-level dual-path network consists of a Dehazing Network and a Denoising Network, where the Dehazing Network is responsible for haze removal in the structural layer, and the Denoising Network deals with noise in the textural layer, respectively. And the second-level dual-path network lies in the Dehazing Network, which has an AL-Net (Atmospheric Light Network) and a TM-Net (Transmission Map Network), respectively. Concretely, the AL-Net aims to train the non-uniform atmospheric light, while the TM-Net aims to train the transmission map that reflects the visibility of the image. The final dehazing image is obtained by nonlinearly fusing the output of the Denoising Network and the Dehazing Network. Extensive experiments demonstrate that our proposed DPDP-Net achieves competitive performance against the state-of-the-art methods on both synthetic and real-world images. Ai-Ping Yang, Zhong Ji, Yanwei Pang, Ling Shao 0001 |
IJCAI | 3 |
| 2019 | Small and Dense Commodity Object Detection with Multi-Scale Receptive Field AttentionabstractSmall and dense commodity object detection is highly valued to the applications in practical scenario. Unlike existing approaches mostly focus on detecting generic objects, this paper studies the problem of specific commodity detection, which is characterized by searching for small and dense instances with similar appearances. Since there is no available dataset or benchmark specialized for exploring this issue, we release a Small and Dense Object Dataset of Milk Tea (SDOD-MT) for promoting the research. Besides, our main solutions for mitigating the detection performance drop caused by the existence of small and dense objects can be concluded as two items. First, for the sake of highlighting the information of positive objects in the feature map, we propose a Multi-Scale Receptive Field (MSRF) attention to generate an attention map to weight the importance on each location of the image feature. Second, for eliminating the negative impact for detection performance brought by the issue of sample imbalance, we present a new loss function named ω-focal loss, which significantly improves the detection accuracy of the categories with few objects. Incorporating these two components into an end-to-end deep architecture, we propose a one-stage detecting framework, dubbed CommodityNet. Extensive experimental results on SDODMT demonstrate that the proposed approach achieves a superior performance on small dense object detection. Zhong Ji, Qiankun Kong, Haoran Wang 0004, Yanwei Pang |
ACM Multimedia | 1 |
| 2019 | Class-specific synthesized dictionary model for Zero-Shot Learning
Zhong Ji, Junyue Wang, Yunlong Yu 0001, Yanwei Pang, Jungong Han |
Neurocomputing | 1 |
| 2019 | Multi-video summarization with query-dependent weighted archetypal analysis
Zhong Ji, Yanwei Pang, Xuelong Li 0001 |
Neurocomputing | 1 |
| 2019 | Query-aware sparse coding for web multi-video summarization
Zhong Ji, Yaru Ma, Yanwei Pang, Xuelong Li 0001 |
Inf. Sci. | 1 |
| 2019 | Image-attribute reciprocally guided attention network for pedestrian attribute recognition
Zhong Ji, Erlu He, Haoran Wang 0004, Ai-Ping Yang |
Pattern Recognit. Lett. | 1 |
| 2019 | Zero-Shot Learning via Latent Space EncodingabstractZero-shot learning (ZSL) is typically achieved by resorting to a class semantic embedding space to transfer the knowledge from the seen classes to unseen ones. Capturing the common semantic characteristics between the visual modality and the class semantic modality (e.g., attributes or word vector) is a key to the success of ZSL. In this paper, we propose a novel encoder-decoder approach, namely latent space encoding (LSE), to connect the semantic relations of different modalities. Instead of requiring a projection function to transfer information across different modalities like most previous work, LSE performs the interactions of different modalities via a feature aware latent space, which is learned in an implicit way. Specifically, different modalities are modeled separately but optimized jointly. For each modality, an encoder-decoder framework is performed to learn a feature aware latent space via jointly maximizing the recoverability of the original space from the latent space and the predictability of the latent space from the original space. To relate different modalities together, their features referring to the same concept are enforced to share the same latent codings. In this way, the common semantic characteristics of different modalities are generalized with the latent representations. Another property of the proposed approach is that it is easily extended to more modalities. Extensive experimental results on four benchmark datasets [animal with attribute, Caltech UCSD birds, aPY, and ImageNet] clearly demonstrate the superiority of the proposed approach on several ZSL tasks, including traditional ZSL, generalized ZSL, and zero-shot retrieval. Yunlong Yu 0001, Zhong Ji, Jichang Guo, Zhongfei Zhang |
IEEE Trans. Cybern. | 2 |
| 2019 | Nonlinear Thermoacoustic Imaging Based on Temperature-Dependent Thermoelastic ResponseabstractIn this paper, a novel nonlinear thermoacoustic imaging (NTAI) system is developed based on the temperature-dependent thermoelastic response under microwave irradiation. Specifically, we consider the high-pulse repetition frequency (HPRF) microwave regime, where the tissue temperature increases after microwave irradiation. In this circumstance, the temperature-dependent thermodynamic parameters of the tissue change; thus, the influence of absorbed microwave energy on the thermoacoustic (TA) amplitude is reflected in the change of thermodynamic parameters. Hence, this temperature dependence yields a nonlinear relationship between the TA amplitude and the absorbed microwave energy in the HPRF regime. In this paper, we obtain the nonlinear coefficient as a new parameter to distinguish biological tissues by performing differential measurements in the NTAI system. We further experimentally extract the nonlinear coefficient in phantom and ex vivo mouse, revealing the feasibility of applications in biological tissues identification with high contrast. Zhong Ji, Sihua Yang, Da Xing |
IEEE Trans. Medical Imaging | 2 |
| 2018 | Stacked Semantics-Guided Attention Model for Fine-Grained Zero-Shot LearningabstractZero-Shot Learning (ZSL) is generally achieved via aligning the semantic relationships between the visual features and the corresponding class semantic descriptions. However, using the global features to represent fine-grained images may lead to sub-optimal results since they neglect the discriminative differences of local regions. Besides, different regions contain distinct discriminative information. The important regions should contribute more to the prediction. To this end, we propose a novel stacked semantics-guided attention (S2GA) model to obtain semantic relevant features by using individual class semantic features to progressively guide the visual features to generate an attention map for weighting the importance of different local regions. Feeding both the integrated visual features and the class semantic features into a multi-class classification architecture, the proposed framework can be trained end-to-end. Extensive experimental results on CUB and NABird datasets show that the proposed approach has a consistent improvement on both fine-grained zero-shot classification and retrieval tasks. Yunlong Yu 0001, Zhong Ji, Yanwei Fu 0001, Jichang Guo, Yanwei Pang, Zhongfei Zhang |
NeurIPS | 2 |
| 2018 | Semantic softmax loss for zero-shot learning
Zhong Ji, Yunlong Yu 0001, Jichang Guo, Yanwei Pang |
Neurocomputing | 1 |
| 2018 | Fusion-Attention Network for person search with free-form natural language
Zhong Ji, Shengjia Li, Yanwei Pang |
Pattern Recognit. Lett. | 1 |
| 2018 | Hypergraph dominant set based multi-video summarization
Zhong Ji, Yanwei Pang, Xuelong Li 0001 |
Signal Process. | 1 |
| 2018 | Transductive Zero-Shot Learning With a Self-Training Dictionary ApproachabstractAs an important and challenging problem in computer vision, zero-shot learning (ZSL) aims at automatically recognizing the instances from unseen object classes without training data. To address this problem, ZSL is usually carried out in the following two aspects: 1) capturing the domain distribution connections between seen classes data and unseen classes data and 2) modeling the semantic interactions between the image feature space and the label embedding space. Motivated by these observations, we propose a bidirectional mapping-based semantic relationship modeling scheme that seeks for cross-modal knowledge transfer by simultaneously projecting the image features and label embeddings into a common latent space. Namely, we have a bidirectional connection relationship that takes place from the image feature space to the latent space as well as from the label embedding space to the latent space. To deal with the domain shift problem, we further present a transductive learning approach that formulates the class prediction problem in an iterative refining process, where the object classification capacity is progressively reinforced through bootstrapping-based model updating over highly reliable instances. Experimental results on four benchmark datasets (animal with attribute, Caltech-UCSD Bird2011, aPascal-aYahoo, and SUN) demonstrate the effectiveness of the proposed approach against the state-of-the-art approaches. Yunlong Yu 0001, Zhong Ji, Xi Li 0001, Jichang Guo, Zhongfei Zhang, Haibin Ling, Fei Wu 0001 |
IEEE Trans. Cybern. | 2 |
| 2018 | Transductive Zero-Shot Learning With Adaptive Structural EmbeddingabstractZero-shot learning (ZSL) endows the computer vision system with the inferential capability to recognize new categories that have never seen before. Two fundamental challenges in it are visual-semantic embedding and domain adaptation in cross-modality learning and unseen class prediction steps, respectively. This paper presents two corresponding methods named Adaptive STructural Embedding (ASTE) and Self-PAced Selective Strategy (SPASS) for both challenges. Specifically, ASTE formulates the visual-semantic interactions in a latent structural support vector machine framework by adaptively adjusting the slack variables to embody different reliablenesses among training instances. To alleviate the domain shift problem in ZSL, SPASS borrows the idea from self-paced learning by iteratively selecting the unseen instances from reliable to less reliable to gradually adapt the knowledge from the seen domain to the unseen domain. Consequently, by combining SPASS and ASTE, we present a self-paced Transductive ASTE (TASTE) method to progressively reinforce the classification capacity. Extensive experiments on three benchmark data sets (i.e., AwA, CUB, and aPY) demonstrate the superiorities of ASTE and TASTE. Furthermore, we also propose a fast training (FT) strategy to improve the efficiency of most existing ZSL methods. The FT strategy is surprisingly simple and general enough, which speeds up the training time of most existing ZSL methods by 4~300 times while holding the previous performance. Yunlong Yu 0001, Zhong Ji, Jichang Guo, Yanwei Pang |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2017 | Deep pedestrian attribute recognition based on LSTMabstractAutomatically recognizing attributes such as gender, age, footwear and clothing style from pedestrian images at far distance is an important task in surveillance scenarios. However, the appearance diversity and ambiguity in these images make it a challenging task. This paper presents an end-to-end Neural Pedestrian Attribute Recognition (Neural PAR) model to address these challenges. Rather than taking it as a recognition problem like previous methods, Neural PAR formulates it as an end-to-end image to attribute description problem. To this end, the training images and their corresponding attributes are used as inputs. Specifically, the attributes are concatenated into different attribute descriptions to well contextualize the potential relationships among them. Then, a neural network model is trained based on CNN and LSTM to learn the complex relations between visual features and their corresponding attributes. Extensive experiments show that the proposed Neural PAR significantly outperforms the state-of-the-art methods on the benchmark PETA dataset. Zhong Ji, Weixiong Zheng, Yanwei Pang |
ICIP | 1 |
| 2017 | Zero-shot learning with regularized cross-modality ranking
Yunlong Yu 0001, Zhong Ji, Jichang Guo, Yanwei Pang |
Neurocomputing | 2 |
| 2017 | Manifold regularized cross-modal embedding for zero-shot learning
Zhong Ji, Yunlong Yu 0001, Yanwei Pang, Jichang Guo, Zhongfei Zhang |
Inf. Sci. | 1 |
| 2017 | Zero-shot learning with Multi-Battery Factor Analysis
Zhong Ji, Yunlong Yu 0001, Yanwei Pang, Zhongfei Zhang |
Signal Process. | 1 |
| 2016 | Visual search reranking with RElevant Local Discriminant Analysis
Peiguang Jing, Zhong Ji, Yunlong Yu 0001, Zhongfei Zhang |
Neurocomputing | 2 |
| 2016 | Relevance and irrelevance graph based marginal Fisher analysis for image search reranking
Zhong Ji, Yanwei Pang, Yuan Yuan 0001 |
Signal Process. | 1 |
| 2016 | Ultrashort Microwave-Pumped Real-Time Thermoacoustic Breast Tumor Imaging SystemabstractWe report the design of a real-time thermoacoustic (TA) scanner dedicated to imaging deep breast tumors and investigate its imaging performance. The TA imaging system is composed of an ultrashort microwave pulse generator and a ring transducer array with 384 elements. By vertically scanning the transducer array that encircles the breast phantom, we achieve real-time, 3D thermoacoustic imaging (TAI) with an imaging speed of 16.7 frames per second. The stability of the microwave energy and its distribution in the cling-skin acoustic coupling cup are measured. The results indicate that there is a nearly uniform electromagnetic field in each XY-imaging plane. Three plastic tubes filled with salt water are imaged dynamically to evaluate the real-time performance of our system, followed by 3D imaging of an excised breast tumor embedded in a breast phantom. Finally, to demonstrate the potential for clinical applications, the excised breast of a ewe embedded with an ex vivo human breast tumor is imaged clearly with a contrast of about 1:2.8. The high imaging speed, large field of view, and 3D imaging performance of our dedicated TAI system provide the potential for clinical routine breast screening. Fanghao Ye, Zhong Ji, Wenzheng Ding, Cunguang Lou, Sihua Yang, Da Xing |
IEEE Trans. Medical Imaging | 2 |
| 2015 | Semi-supervised LPP algorithms for learning-to-rank-based visual search reranking
Zhong Ji, Yanwei Pang, Huanfen Zhang |
Inf. Sci. | 1 |
| 2015 | Relevance Preserving Projection and Ranking for Web Image Search RerankingabstractAn image search reranking (ISR) technique aims at refining text-based search results by mining images' visual content. Feature extraction and ranking function design are two key steps in ISR. Inspired by the idea of hypersphere in one-class classification, this paper proposes a feature extraction algorithm named hypersphere-based relevance preserving projection (HRPP) and a ranking function called hypersphere-based rank (H-Rank). Specifically, an HRPP is a spectral embedding algorithm to transform an original high-dimensional feature space into an intrinsically low-dimensional hypersphere space by preserving the manifold structure and a relevance relationship among the images. An H-Rank is a simple but effective ranking algorithm to sort the images by their distances to the hypersphere center. Moreover, to capture the user's intent with minimum human interaction, a reversed k-nearest neighbor (KNN) algorithm is proposed, which harvests enough pseudorelevant images by requiring that the user gives only one click on the initially searched images. The HRPP method with reversed KNN is named one-click-based HRPP (OC-HRPP). Finally, an OC-HRPP algorithm and the H-Rank algorithm form a new ISR method, H-reranking. Extensive experimental results on three large real-world data sets show that the proposed algorithms are effective. Moreover, the fact that only one relevant image is required to be labeled makes it has a strong practical significance. Zhong Ji, Yanwei Pang, Xuelong Li 0001 |
IEEE Trans. Image Process. | 1 |
| 2013 | Image Search Reranking with Semi-supervised LPP and Ranking SVM
Zhong Ji, Yanru Yu, Yuting Su 0001, Yanwei Pang |
MMM (1) | 1 |
| 2013 | Ranking Fisher discriminant analysis
Zhong Ji, Peiguang Jing, Tianshi Yu, Yuting Su 0001, Changshu Liu |
Neurocomputing | 1 |
| 2013 | Balance between object and background: Object-enhanced features for scene image classification
Zhong Ji, Yuting Su 0001, Zhanjie Song, Shikai Xing |
Neurocomputing | 1 |
| 2013 | Rank canonical correlation analysis and its application in visual search reranking
Zhong Ji, Peiguang Jing, Yuting Su 0001, Yanwei Pang |
Signal Process. | 1 |
| 2013 | Ranking Graph Embedding for Learning to RerankabstractDimensionality reduction is a key step to improving the generalization ability of reranking in image search. However, existing dimensionality reduction methods are typically designed for classification, clustering, and visualization, rather than for the task of learning to rank. Without using of ranking information such as relevance degree labels, direct utilization of conventional dimensionality reduction methods in ranking tasks generally cannot achieve the best performance. In this paper, we show that introducing ranking information into dimensionality reduction significantly increases the performance of image search reranking. The proposed method transforms graph embedding, a general framework of dimensionality reduction, into ranking graph embedding (RANGE) by modeling the global structure and the local relationships in and between different relevance degree sets, respectively. The proposed method also defines three types of edge weight assignment between two nodes: binary, reconstruction, and global. In addition, a novel principal components analysis based similarity calculation method is presented in the stage of global graph construction. Extensive experimental results on the MSRA-MM database demonstrate the effectiveness and superiority of the proposed RANGE method and the image search reranking framework. Yanwei Pang, Zhong Ji, Peiguang Jing, Xuelong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2011 | Diversifying the Image Relevance Reranking with Absorbing Random WalksabstractImage visual reranking holds the simple search mechanism preferred by typical users, and exploits the visual information and image analysis methods in another way. Therefore, it integrates characteristics of real-time and accuracy, and has great importance to establish practical image search system. A novel reranking method named DIRRA is proposed in this paper, in which absorbing random walks is utilized to enhance the diversity as well as relevance of the initial search results. Four kinds of image visual features are extracted firstly, and then a graph is built, where nodes are images and edges are the similarities between images. Next, the first item is decided by teleporting random walks on the graph, and the other items are decided by absorbing random walks on the graph at last. Experiments are performed on a web image database including 10 queries, which prove the reranking results are both diverse and relevant, and practical to improve user's satisfaction in web search. Zhong Ji, Yuting Su 0001, Yanwei Pang, Xiaojie Qu |
ICIG | 1 |
| 2007 | Anchorperson Shot Detection in MPEG DomainabstractIn this paper, a refined ASD algorithm in MPEG compressed domain is proposed. The new method is expected to outperform the existing strategies based on the following two improvements. One is that an effective face detection method is introduced. It aims to further remove the false alarms in the candidates generated from the previous modules, employing the chrominance DC coefficients. The other is the utilization of a new robust metric, which represents the dissimilarity of the video frames in the unsupervised clustering module. The proposed algorithm has been evaluated by six different TV channels. Compared with the state-of-the-art methods, the new algorithm is effective and computationally efficient. Zhong Ji, Chuntian Zhang, Yuting Su 0001 |
ICME | 1 |
| 2003 | Detection of EEG basic rhythm feature by using band relative intensity ratio (BRIR)abstractIn the clinical analysis and processing for EEG, because of the difference of ages and pathology, it is possible for abnormal waves to appear, related with pathology. Also, the restraint of normal rhythms could be abnormal. But at present doctors estimate if a certain rhythm is restrained only by eye or by some simple analysis methods in clinical EEG detection, which will inevitably lead to some errors and are not observable. By "the virtual EEG record and analysis instrument" introduced in this paper, all kinds of characteristic waveforms (e.g. epileptic wave and spikes wave etc.) can be detected and analyzed in time-frequency domain. From the view of clinical application, the concept of band relative intensity ratio (BRIR) is introduced with time-frequency domain analysis, by the use of which we can obtain the relative intensity of all basic rhythms in a certain time period, and this is believed to provide a good assisting analysis method. Zhong Ji |
ICASSP (6) | 1 |
| 2003 | Detection of EEG basic rhythm feature by using band relative intensity ratio(BRIR)abstractIn the clinical analysis and processing for EEG, because of the difference of ages and pathology, it is possible to appear abnormal waves related with pathology. Also, the restrain of normal rhythms could be abnormal. But at present doctors estimate if certain rhythm is restrained only by using the method of eyeballing or by some simply analysis methods in clinical EEG detection, which will inevitably lead to some errors and not observable. By "the virtual EEG record and analysis instrument" introduced in this paper, all kinds of characteristic waveforms (e.g. epileptic wave and spikes wave etc.) can be detected and analyzed in time-frequency domain. From the view of clinical application, the concept of band relative intensity ratio (BRIR) is introduced with time-frequency domain analysis, by the use of which, we can obtain the relative intensity of all basic rhythms in a certain time period, and this is believed to provide a good assistant analysis method. Zhong Ji |
ICME | 1 |