Mengmeng Jing

dblp:219/2183 · DBLP profile ↗
← Back
41ranked-venue papers
10as first author
29since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 21 · 6 first-author · 14 since 2021Artificial intelligence and machine learning · 20 · 5 first-author · 15 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Dropout Prompt Learning: Towards Robust and Adaptive Vision-Language Models
abstract
Dropout is a widely used regularization technique which improves the generalization ability of a model by randomly dropping neurons. In light of this, we propose Dropout Prompt Learning, which aims for applying dropout to improve the robustness of the vision-language models. Different from the vanilla dropout, we apply dropout on the tokens of the textual and visual branches, where we evaluate the token significance considering both intra-modal context and inter-modal alignment, enabling flexible dropout probabilities for each token. Moreover, to maintain semantic alignment for general knowledge transfer while encouraging the diverse representations that dropout introduces, we further propose residual entropy regularization. Experiments on 11 benchmarks show our method's effectiveness in challenging scenarios like low-shot learning, long-tail classification, and out-of-distribution generalization. Notably, our method surpasses regularization-based methods including KgCoOp by 5.10% and PromptSRC by 2.13% in performance on base-to-novel generalization.
Lin Zuo, Mengmeng Jing, Kunbin He
AAAI3
2026 Instruction-Guided Cross-Modal Clustering for Training-Free Visual Token Pruning in Vision-Language Models
abstract
Large vision-language models (LVLMs) have demonstrated remarkable capabilities in understanding multimodal data such as images and text. However, the number of visual tokens in these models often far exceeds that of textual tokens, resulting in substantial redundancy and high inference costs. Existing pruning methods primarily rely on either unimodal information or cross-modal attention mechanisms. The former often overlooks the semantic alignment between instructions and visual representations in the multimodal space, while the latter is prone to attention drift and dispersion, leading to significant performance degradation under high pruning ratios. All the above issues stem from the lack of effective textual guidance during the pruning process. To identify effective informational cues for guiding pruning, we conduct an in-depth analysis of the interaction between language instructions and visual features based on the cross-modal information bottleneck attribution (CIBA) method, revealing the presence of noun anchors. Based on this analysis, we propose the Instruction-Guided Cross-Modal Clustering Token Pruning (ICCTP) method, a plug-and-play, training-free pruning paradigm. Specifically, ICCTP first leverages global attention to retain a small set of visual tokens that preserve global context. It then extracts nouns from the instruction as clustering centers to perform cross-modal clustering over the remaining visual tokens. To balance semantic diversity and global relevance while reducing intra-cluster redundancy, we design an importance scoring mechanism. Finally, visual tokens within each cluster are pruned according to a specified pruning ratio. We evaluate ICCTP on multiple VLM architectures, including LLaVA-1.5-7B, LLaVA-1.5-13B, and LLaVA-NeXT-7B. Experimental results show that ICCTP maintains strong performance across various pruning rates without requiring retraining. Notably, even under an extreme setting where 94.4% of visual tokens are removed, ICCTP retains 90.02% of the original accuracy while reducing TFLOPs by 82.36%.
Yunqian Yu, Yunya Zhang, Tonglan Xie, Mengmeng Jing, Lin Zuo
AAAI5
2026 Learning spatio-temporal consistency in spiking neural networks by self-distillation
Lin Zuo, Yongqi Ding, Mengmeng Jing, Kunshan Yang, Hanpu Deng
Pattern Recognit.3
2025 Prelude echoes Finale: Video Domain Adaptation with Fine-grained Temporal Consistency
abstract
The prelude and finale of one video usually share similar themes and content, with the finale often echoing or deepening the concepts introduced in the prelude. This temporal correlation enhances the understanding of video content. Existing video domain adaptation methods, however, ignore this fine-grained temporal structure in the clip level, focusing instead on coarser temporal consistency at the video level. To make the prelude echo the finale, we propose to optimize fine-grained temporal consistency for video domain adaptation. Specifically, we first extract the prelude (the first n clips) and the finale (the last n clips) from the same video. Then, we maximize the feature similarity of the prelude and the finale, which enhances the connections between clips and achieves the temporal consistency of each video across source and target domains. On the other hand, we generate the cross-domain fusion features and optimize their discriminability to make the source domain features echo the target. As a result, the source and target domains are aligned. Extensive experiments on video domain adaptation benchmarks demonstrate the effectiveness of our method.
Mengmeng Jing, Xianlong Tian, Yuguo Hu, Lin Zuo
ICASSP2
2025 Learning Latent Representations with Codebook Priors for Single-Image Flare Removal
Shimin Luo, Yunya Zhang, Hanpu Deng, Mengmeng Jing, Lin Zuo
ICIC (3)5
2025 Multi-bit Mechanism: Towards Ultra-Low Time Steps for Spiking Neural Networks
Yongjun Xiao, Pei He, Hanpu Deng, Tonglan Xie, Mengmeng Jing, Lin Zuo
ICIC (21)5
2025 Rethinking Spiking Neural Networks from an Ensemble Learning Perspective
abstract
Spiking neural networks (SNNs) exhibit superior energy efficiency but suffer from limited performance. In this paper, we consider SNNs as ensembles of temporal subnetworks that share architectures and weights, and highlight a crucial issue that affects their performance: excessive differences in initial states (neuronal membrane potentials) across timesteps lead to unstable subnetwork outputs, resulting in degraded performance. To mitigate this, we promote the consistency of the initial membrane potential distribution and output through membrane potential smoothing and temporally adjacent subnetwork guidance, respectively, to improve overall stability and performance. Moreover, membrane potential smoothing facilitates forward propagation of information and backward propagation of gradients, mitigating the notorious temporal gradient vanishing problem. Our method requires only minimal modification of the spiking neurons without adapting the network structure, making our method generalizable and showing consistent performance gains in 1D speech, 2D object, and 3D point cloud recognition tasks. In particular, on the challenging CIFAR10-DVS dataset, we achieved 83.20\% accuracy with only four timesteps. This provides valuable insights into unleashing the potential of SNNs.
Yongqi Ding, Lin Zuo, Mengmeng Jing, Pei He, Hanpu Deng
ICLR3
2025 Diffusion-Driven Source Consistency for Gradual Domain Adaptation
abstract
Unsupervised Domain Adaptation (UDA) aims to directly transfer models from source domain to target domain. When the domain shifts are large, UDA performs poorly. Gradual Domain Adaptation (GDA) alleviates this problem in a mild way by gradually adapting from source to target domain using multiple intermediate domains. In this paper, we propose Source Adaptive Diffusion Model (SADiff) to leverage intermediate domains to get more linear and effective performance across samples. SADiff employs diffusion model to simulate the feature distribution shift between domains, ensuring smoothness and continuity in latent space through regularization constraints on cross-domain mixup features. Additionally, we introduce an energy-based labeling strategy to improve label consistency. Finally, we gradually improved the adaptive ability of the model, achieving smooth and effective domain adaptation. Extensive experiments on four GDA benchmarks verified the effectiveness of the proposed method.
Wenwei Luo, Yuguo Hu, Jiafu Yan, Mengmeng Jing, Lin Zuo
ICME4
2025 Bayesian-Inspired Cross-Spectral Fusion Network for Robust Depth Estimation
abstract
Depth estimation is a critical task in multimedia, with widespread applications in autonomous driving, robotics, and 3D reconstruction. Traditional methods relying on single-spectral data suffer from low-light and nighttime conditions. In view of this, we propose BCSFDepth, a bayesian-inspired cross-spectral fusion network fuses infrared and visible images to enhance depth estimation. Our method incorporates a frequency-domain interaction enhancement module and a Bayesian uncertainty-weighted fusion mechanism to effectively integrate complementary information across modalities. The proposed framework dynamically adjusts modality weights, mitigating the impact of low-quality features. We collected and calibrated a real-world cross-spectral dataset, IVLD, to evaluate our method under diverse weather conditions. Experimental results demonstrate substantial improvements in depth estimation accuracy, particularly in noisy and complex environments.
Jiafu Yan, Wenwei Luo, Yuguo Hu, Changhua Zhang, Mengmeng Jing, Lin Zuo
ICME5
2025 A Semantic-Enhanced Heterogeneous Graph Learning Method for Flexible Objects Recognition
abstract
Flexible objects recognition remains a significant challenge due to its inherently diverse shapes and sizes, translucent attributes, and subtle inter-class differences. Graph-based models, such as graph convolution networks and graph vision models, are promising in flexible objects recognition due to their ability of capturing variable relations within the flexible objects. These methods, however, often focus on global visual relationships or fail to align semantic and visual information. To alleviate these limitations, we propose a semantic-enhanced heterogeneous graph learning method. First, an adaptive scanning module is employed to extract discriminative semantic context, facilitating the matching of flexible objects with varying shapes and sizes while aligning semantic and visual nodes to enhance cross-modal feature correlation. Second, a heterogeneous graph generation module aggregates global visual and local semantic node features, improving the recognition of flexible objects. Additionally, We introduce the FSCW, a large-scale flexible dataset curated from existing sources. We validate our method through extensive experiments on flexible datasets (FDA and FSCW), and challenge benchmarks (CIFAR-100 and ImageNet-Hard), demonstrating competitive performance.
Kunshan Yang, Wenwei Luo, Yuguo Hu, Jiafu Yan, Mengmeng Jing, Lin Zuo
ICME5
2025 Toward End-to-End Bearing Fault Diagnosis for Industrial Scenarios with Spiking Neural Networks
abstract
This paper explores the application of spiking neural networks (SNNs), known for their low-power binary spikes, to bearing fault diagnosis, bridging the gap between high-performance AI algorithms and real-world industrial scenarios.In particular, we identify two key limitations of existing SNN fault diagnosis methods: inadequate encoding capacity that necessitates cumbersome data preprocessing, and non-spike-oriented architectures that constrain the performance of SNNs.To alleviate these problems, we propose a Multi-scale Residual Attention SNN (MRA-SNN) to simultaneously improve the efficiency, performance, and robustness of SNN methods.By incorporating a lightweight attention mechanism, we have designed a multi-scale attention encoding module to extract multiscale fault features from vibration signals and encode them as spatio-temporal spikes, eliminating the need for complicated preprocessing.Then, the spike residual attention block extracts high-dimensional fault features and enhances the expressiveness of sparse spikes with the attention mechanism for end-to-end diagnosis.In addition, the performance and robustness of MRA-SNN is further enhanced by introducing the lightweight attention mechanism within the spiking neurons to simulate the biological dendritic filtering effect.Extensive experiments on MFPT, JNU, Bearing, and Gearbox benchmark datasets demonstrate that MRA-SNN significantly outperforms existing methods in terms of accuracy, energy
Lin Zuo, Yongqi Ding, Mengmeng Jing, Kunshan Yang, Yunqian Yu
KDD (2)3
2025 Chain-of-Thought Guided Semantic Debiasing for Low-Shot Vision-Language Tasks
Kunbin He, Zhikun Zheng, Mengmeng Jing, Lin Zuo
ACM Multimedia4
2025 Bridging Inter-Class Ambiguity and Spatial Variability in Flexible Object Recognition via Graph Distillation
abstract
Flexible object recognition remains challenging in multimedia scenarios due to inherently diverse shapes and sizes, and subtle inter-class differences. Graph-based vision models show promise in flexible objects recognition by capturing variable relationships. However, they suffer from two problems: (1) inter-class ambiguity hinders model discrimination and (2) frequent scale changes degrade model generalization. To address these limitations, we propose a unified graph distillation framework that enhances inter-class discrimination and spatial generalization while maintaining computational efficiency. For inter-class ambiguity problem, we introduce a virtual prototype module that dynamically generates learnable class prototypes via clustering intermediate features. These prototypes are incorporated into the distillation loss to sharpen decision boundaries. A global-local distillation mechanism further capture both image-level global semantics and patch-level local details, enhancing inter-class discrimination. For frequent scale changes problem, we design a patch-aware distillation strategy that transfers knowledge across multiple patch scales, strengthening the student model's spatial generalization to match various shapes and sizes of flexible objects, thus alleviate generalization degradation. Extensive experiments on flexible-object datasets (FDA, FSCW, CCSN) and challenging benchmarks (CIFAR-100, Mini-ImageNet) confirm effectiveness and efficiency of our method.
Lin Zuo, Kunshan Yang, Mengmeng Jing, Xiangxu Zhao, Jiaqiao Chen
ACM Multimedia3
2025 Synergy Between the Strong and the Weak: Spiking Neural Networks are Inherently Self-Distillers
abstract
Brain-inspired spiking neural networks (SNNs) promise to be a low-power alternative to computationally intensive artificial neural networks (ANNs), although performance gaps persist. Recent studies have improved the performance of SNNs through knowledge distillation, but rely on large teacher models or introduce additional training overhead. In this paper, we show that SNNs can be naturally deconstructed into multiple submodels for efficient self-distillation. We treat each timestep instance of the SNN as a submodel and evaluate its output confidence, thus efficiently identifying the strong and the weak. Based on this strong and weak relationship, we propose two efficient self-distillation schemes: (1) Strong2Weak: During training, the stronger "teacher" guides the weaker "student", effectively improving overall performance. (2) Weak2Strong: The weak serve as the "teacher", distilling the strong in reverse with underlying dark knowledge, again yielding significant performance gains. For both distillation schemes, we offer flexible implementations such as ensemble, simultaneous, and cascade distillation. Experiments show that our method effectively improves the discriminability and overall performance of the SNN, while its adversarial robustness is also enhanced, benefiting from the stability brought by self-distillation. This ingeniously exploits the temporal properties of SNNs and provides insight into how to efficiently train high-performance SNNs.
Yongqi Ding, Lin Zuo, Mengmeng Jing, Kunshan Yang, Pei He, Tonglan Xie
NeurIPS3
2025 Visual residual aggregation network for visual-language prompt tuning
Yunqian Yu, Xianlong Tian, Mengmeng Jing, Lin Zuo
Appl. Intell.5
2025 Adaptive recognition of flexible objects via hierarchical deformable graph network
Kunshan Yang, Lin Zuo, Mengmeng Jing, Yongqi Ding, Xianlong Tian
Inf. Sci.3
2025 Flexible ViG: Learning the Self-Saliency for Flexible Object Recognition
abstract
Existing computer vision methods mainly focus on the recognition of rigid objects, whereas the recognition of flexible objects remains unexplored. Recognizing flexible objects poses significant challenges due to their inherently diverse shapes and sizes, translucent attributes, ambiguous boundaries, and subtle inter-class differences. In this paper, we claim that these problems primarily arise from the lack of object saliency. To this end, we propose the Flexible Vision Graph Neural Network (FViG) to optimize the self-saliency and thereby improve the discrimination of the representations for flexible objects. Specifically, on one hand, we propose to maximize the channel-aware saliency by extracting the weight of neighboring graph nodes, which is employed to identify flexible objects with minimal inter-class differences. On the other hand, we maximize the spatial-aware saliency based on clustering to aggregate neighborhood information for the centroid graph nodes. This introduces local context information and enables extracting of consistent representation, effectively adapting to the shape and size variations in flexible objects. To verify the performance of flexible objects recognition thoroughly, for the first time we propose the Flexible Dataset (FDA), which consists of various images of flexible objects collected from real-world scenarios or online. Extensive experiments evaluated on our FDA, FireNet, CIFAR-100 and ImageNet-Hard datasets demonstrate the effectiveness of our method on enhancing the discrimination of flexible objects.
Kunshan Yang, Lin Zuo, Mengmeng Jing, Xianlong Tian, Kunbin He, Yongqi Ding
IEEE Trans. Circuits Syst. Video Technol.3
2024 Shrinking Your TimeStep: Towards Low-Latency Neuromorphic Object Recognition with Spiking Neural Networks
abstract
Neuromorphic object recognition with spiking neural networks (SNNs) is the cornerstone of low-power neuromorphic computing. However, existing SNNs suffer from significant latency, utilizing 10 to 40 timesteps or more, to recognize neuromorphic objects. At low latencies, the performance of existing SNNs is drastically degraded. In this work, we propose the Shrinking SNN (SSNN) to achieve low-latency neuromorphic object recognition without reducing performance. Concretely, we alleviate the temporal redundancy in SNNs by dividing SNNs into multiple stages with progressively shrinking timesteps, which significantly reduces the inference latency. During timestep shrinkage, the temporal transformer smoothly transforms the temporal scale and preserves the information maximally. Moreover, we add multiple early classifiers to the SNN during training to mitigate the mismatch between the surrogate gradient and the true gradient, as well as the gradient vanishing/exploding, thus eliminating the performance degradation at low latency. Extensive experiments on neuromorphic datasets, CIFAR10-DVS, N-Caltech101, and DVS-Gesture have revealed that SSNN is able to improve the baseline accuracy by 6.55% ~ 21.41%. With only 5 average timesteps and without any data augmentation, SSNN is able to achieve an accuracy of 73.63% on CIFAR10-DVS. This work presents a heterogeneous temporal scale SNN and provides valuable insights into the development of high-performance, low-latency SNNs.
Yongqi Ding, Lin Zuo, Mengmeng Jing, Pei He, Yongjun Xiao
AAAI3
2024 Graph Neural Network-Based Structured Scene Graph Generation for Efficient Wildfire Detection
Yanning Ye, Shimin Luo, Mengmeng Jing, Yongqi Ding, Kunbin He, Lin Zuo
ICIC (3)3
2024 Visually Source-Free Domain Adaptation via Adversarial Style Matching
abstract
The majority of existing works explore Unsupervised Domain Adaptation (UDA) with an ideal assumption that samples in both domains are available and complete. In real-world applications, however, this assumption does not always hold. For instance, data-privacy is becoming a growing concern, the source domain samples may be not publicly available for training, leading to a typical Source-Free Domain Adaptation (SFDA) problem. Traditional UDA methods would fail to handle SFDA since there are two challenges in the way: the data incompleteness issue and the domain gaps issue. In this paper, we propose a visually SFDA method named Adversarial Style Matching (ASM) to address both issues. Specifically, we first train a style generator to generate source-style samples given the target images to solve the data incompleteness issue. We use the auxiliary information stored in the pre-trained source model to ensure that the generated samples are statistically aligned with the source samples, and use the pseudo labels to keep semantic consistency. Then, we feed the target domain samples and the corresponding source-style samples into a feature generator network to reduce the domain gaps with a self-supervised loss. An adversarial scheme is employed to further expand the distributional coverage of the generated source-style samples. The experimental results verify that our method can achieve comparative performance even compared with the traditional UDA methods with source samples for training.
Mengmeng Jing, Jingjing Li 0001, Ke Lu 0001, Lei Zhu 0002, Heng Tao Shen
IEEE Trans. Image Process.1
2023 Order-preserving Consistency Regularization for Domain Adaptation and Generalization
abstract
Deep learning models fail on cross-domain challenges if the model is oversensitive to domain-specific attributes, e.g., lightning, background, camera angle, etc. To alleviate this problem, data augmentation coupled with consistency regularization are commonly adopted to make the model less sensitive to domain-specific attributes. Consistency regularization enforces the model to output the same representation or prediction for two views of one image. These constraints, however, are either too strict or not order-preserving for the classification probabilities. In this work, we propose the Order-preserving Consistency Regularization (OCR) for cross-domain tasks. The order-preserving property for the prediction makes the model robust to task-irrelevant transformations. As a result, the model becomes less sensitive to the domain-specific attributes. The comprehensive experiments show that our method achieves clear advantages on five different cross-domain tasks.
Mengmeng Jing, Xiantong Zhen, Jingjing Li 0001, Cees Snoek
ICCV1
2023 Adversarial Mixup Ratio Confusion for Unsupervised Domain Adaptation
abstract
Multimedia applications often involve knowledge transfer across domains, e.g., from images to texts, where Unsupervised Domain Adaptation (UDA) can be used to reduce the domain shifts. Most of the UDA methods are based on adversarial learning. However, previous adversarial domain adaptation methods may suffer from three issues. First, although the features learned by previous methods could fool the domain classifier to make false classification predictions, they may not be domain-invariant. Second, the limited number of training samples make the latent space of features not smooth and continuous enough. Third, the target domain features may lack discriminability. In this paper, we propose a novel adversarial domain adaptation method named Adversarial Mixup Ratio Confusion (AMRC) to alleviate all the above issues. Specifically, we propose a new adversarial training pattern that uses mixup to generate multiple features with different mixup ratios, which represent different intermediate states between the source and target domain. Then, on one hand, we train an estimator to estimate the mixup ratio as accurately as possible. On the other hand, we train a generator to make the estimator be uncertain about the mixup ratio. In this way, our method could learn a continuous and domain-invariant latent space. Furthermore, we apply the intra-domain and cross-domain mixup regularizations to ensure the smoothness and continuity of the latent space, while making the classifier behave more linearly on in-between samples. At last, we exploit the sharpened pseudo-labels of the target samples for self-supervised learning to enhance the discriminability of the target features.The experimental results on 3 benchmarks verify the effectiveness of our method.
Mengmeng Jing, Lichao Meng, Jingjing Li 0001, Lei Zhu 0002, Heng Tao Shen
IEEE Trans. Multim.1
2023 Open Set Domain Adaptation via Joint Alignment and Category Separation
abstract
Prevalent domain adaptation approaches are suitable for a close-set scenario where the source domain and the target domain are assumed to share the same data categories. However, this assumption is often violated in real-world conditions where the target domain usually contains samples of categories that are not presented in the source domain. This setting is termed as open set domain adaptation (OSDA). Most existing domain adaptation approaches do not work well in this situation. In this article, we propose an effective method, named joint alignment and category separation (JACS), for OSDA. Specifically, JACS learns a latent shared space, where the marginal and conditional divergence of feature distributions for commonly known classes across domains is alleviated (Joint Alignment), the distribution discrepancy between the known classes and the unknown class is enlarged, and the distance between different known classes is also maximized (Category Separation). These two aspects are unified into an objective to reinforce the optimization of each part simultaneously. The classifier is achieved based on the learned new feature representations by minimizing the structural risk in the reproducing kernel Hilbert space. Extensive experiment results verify that our method outperforms other state-of-the-art approaches on several benchmark datasets.
Jieyan Liu, Mengmeng Jing, Jingjing Li 0001, Ke Lu 0001, Heng Tao Shen
IEEE Trans. Neural Networks Learn. Syst.2
2022 Variational Model Perturbation for Source-Free Domain Adaptation
abstract
We aim for source-free domain adaptation, where the task is to deploy a model pre-trained on source domains to target domains. The challenges stem from the distribution shift from the source to the target domain, coupled with the unavailability of any source data and labeled target data for optimization. Rather than fine-tuning the model by updating the parameters, we propose to perturb the source model to achieve adaptation to target domains. We introduce perturbations into the model parameters by variational Bayesian inference in a probabilistic framework. By doing so, we can effectively adapt the model to the target domain while largely preserving the discriminative ability. Importantly, we demonstrate the theoretical connection to learning Bayesian neural networks, which proves the generalizability of the perturbed model to target domains. To enable more efficient optimization, we further employ a parameter sharing strategy, which substantially reduces the learnable parameters compared to a fully Bayesian neural network. Our model perturbation provides a new probabilistic way for domain adaptation which enables efficient adaptation to target domains while maximally preserving knowledge in source models. Experiments on several source-free benchmarks under three different evaluation settings verify the effectiveness of the proposed variational model perturbation for source-free domain adaptation.
Mengmeng Jing, Xiantong Zhen, Jingjing Li 0001, Cees Snoek
NeurIPS1
2022 Investigating the Bilateral Connections in Generative Zero-Shot Learning
abstract
Zero-shot learning (ZSL) is a pretty intriguing topic in the computer vision community since it handles novel instances and unseen categories. In a typical ZSL setting, there is a main visual space and an auxiliary semantic space. Most existing ZSL methods handle the problem by learning either a visual-to-semantic mapping or a semantic-to-visual mapping. In other words, they investigate a unilateral connection from one end to the other. However, the connection between the visual space and the semantic space are bilateral in reality, that is, the visual space depicts the semantic space; the semantic space, on the other hand, describes the visual space. In this article, therefore, we investigate the bilateral connections in ZSL and present a novel model, called Boomerang-GAN, by taking advantage of conditional generative adversarial networks (GANs). Specifically, we generate unseen visual samples from their category semantic embeddings by a conditional GAN. Different from the existing generative ZSL methods that only consider generating visual features from class descriptions, our method also considers that the generated visual features can be translated back to their corresponding semantic embeddings by introducing a multimodal cycle-consistent loss. Extensive experiments of both ZSL and generalized ZSL on five widely used datasets verify that our method is able to outperform previous state-of-the-art approaches in both recognition and segmentation tasks.
Jingjing Li 0001, Mengmeng Jing, Ke Lu 0001, Lei Zhu 0002, Heng Tao Shen
IEEE Trans. Cybern.2
2022 Faster Domain Adaptation Networks
abstract
It is widely acknowledged that the success of deep learning is built upon large-scale training data and tremendous computing power. However, the data and computing power are not always available for many real-world applications. In this paper, we address the machine learning problem where it lacks training data and limits computing power. Specifically, we investigate domain adaptation which is able to transfer knowledge from one labeled source domain to an unlabeled target domain, so that we do not need much training data from the target domain. At the same time, we consider the situation that the running environment is confined, e.g., in edge computing the end device has very limited running resources. Technically, we present the Faster Domain Adaptation (FDA) protocol and further report two paradigms of FDA: early stopping and amid skipping. The former accelerates domain adaptation by multiple early exit points. The latter speeds up the adaptation by wisely skip several amid neural network blocks. Extensive experiments on standard benchmarks verify that our method is able to achieve the comparable and even better accuracy but employ much less computing resources. To the best of our knowledge, there are very few works which investigated accelerating knowledge adaptation in the community. This work is expected to inspire the topic for more discussion.
Jingjing Li 0001, Mengmeng Jing, Hongzu Su, Ke Lu 0001, Lei Zhu 0002, Heng Tao Shen
IEEE Trans. Knowl. Data Eng.2
2021 Balanced Open Set Domain Adaptation via Centroid Alignment
abstract
Open Set Domain Adaptation (OSDA) is a challenging domain adaptation setting which allows the existence of unknown classes on the target domain. Although existing OSDA methods are good at classifying samples of known classes, they ignore the classification ability for the unknown samples, making them unbalanced OSDA methods. To alleviate this problem, we propose a balanced OSDA methods which could recognize the unknown samples while maintain high classification performance for the known samples. Specifically, to reduce the domain gaps, we first project the features to a hyperspherical latent space. In this space, we propose to bound the centroid deviation angles to not only increase the intra-class compactness but also enlarge the inter-class margins. With the bounded centroid deviation angles, we employ the statistical Extreme Value Theory to recognize the unknown samples that are misclassified into known classes. In addition, to learn better centroids, we propose an improved centroid update strategy based on sample reweighting and adaptive update rate to cooperate with centroid alignment. Experimental results on three OSDA benchmarks verify that our method can significantly outperform the compared methods and reduce the proportion of the unknown samples being misclassified into known classes.
Mengmeng Jing, Jingjing Li 0001, Lei Zhu 0002, Zhengming Ding, Ke Lu 0001, Yang Yang 0002
AAAI1
2021 Challenging tough samples in unsupervised domain adaptation
Lin Zuo, Mengmeng Jing, Jingjing Li 0001, Lei Zhu 0002, Ke Lu 0001, Yang Yang 0002
Pattern Recognit.2
2021 Adaptive Component Embedding for Domain Adaptation
abstract
Domain adaptation is suitable for transferring knowledge learned from one domain to a different but related domain. Considering the substantially large domain discrepancies, learning a more generalized feature representation is crucial for domain adaptation. On account of this, we propose an adaptive component embedding (ACE) method, for domain adaptation. Specifically, ACE learns adaptive components across domains to embed data into a shared domain-invariant subspace, in which the first-order statistics is aligned and the geometric properties are preserved simultaneously. Furthermore, the second-order statistics of domain distributions is also aligned to further mitigate domain shifts. Then, the aligned feature representation is classified by optimizing the structural risk functional in the reproducing kernel Hilbert space (RKHS). Extensive experiments show that our method can work well on six domain adaptation benchmarks, which verifies the effectiveness of ACE.
Mengmeng Jing, Jidong Zhao, Jingjing Li 0001, Lei Zhu 0002, Yang Yang 0002, Heng Tao Shen
IEEE Trans. Cybern.1
2020 Incomplete Cross-modal Retrieval with Dual-Aligned Variational Autoencoders
abstract
Learning the relationship between the multi-modal data, e.g., texts, images and videos, is a classic task in the multimedia community. Cross-modal retrieval (CMR) is a typical example where the query and the corresponding results are in different modalities. Yet, a majority of existing works investigate CMR with an ideal assumption that the training samples in every modality are sufficient and complete. In real-world applications, however, this assumption does not always hold. Mismatch is common in multi-modal datasets. There is a high chance that samples in some modalities are either missing or corrupted. As a result, incomplete CMR has become a challenging issue. In this paper, we propose a Dual-Aligned Variational Autoencoders (DAVAE) to address the incomplete CMR problem. Specifically, we propose to learn modality-invariant representations for different modalities and use the learned representations for retrieval. We train multiple autoencoders, one for each modality, to learn the latent factors among different modalities. These latent representations are further dual-aligned at the distribution level and the semantic level to alleviate the modality gaps and enhance the discriminability of representations. For missing instances, we leverage generative models to synthesize latent representations for them. Notably, we test our method with different ratios of random incompleteness.Extensive experiments on three datasets verify that our method can consistently outperform the state-of-the-arts.
Mengmeng Jing, Jingjing Li 0001, Lei Zhu 0002, Ke Lu 0001, Yang Yang 0002, Zi Huang
ACM Multimedia1
2020 Learning Modality-Invariant Latent Representations for Generalized Zero-shot Learning
abstract
Recently, feature generating methods have been successfully applied to zero-shot learning (ZSL). However, most previous approaches only generate visual representations for zero-shot recognition. In fact, typical ZSL is a classic multi-modal learning protocol which consists of a visual space and a semantic space. In this paper, therefore, we present a new method which can simultaneously generate both visual representations and semantic representations so that the essential multi-modal information associated with unseen classes can be captured. Specifically, we address the most challenging issue in such a paradigm, i.e., how to handle the domain shift and thus guarantee that the learned representations are modality-invariant. To this end, we propose two strategies: 1) leveraging the mutual information between the latent visual representations and the semantic representations; 2) maximizing the entropy of the joint distribution of the two latent representations. By leveraging the two strategies, we argue that the two modalities can be well aligned. At last, extensive experiments on five widely used datasets verify that the proposed method is able to significantly outperform previous the state-of-the-arts.
Jingjing Li 0001, Mengmeng Jing, Lei Zhu 0002, Zhengming Ding, Ke Lu 0001, Yang Yang 0002
ACM Multimedia2
2020 Multi-source domain adaptation with graph embedding and adaptive label prediction
Ao Ma 0001, Fuming You, Mengmeng Jing, Jingjing Li 0001, Ke Lu 0001
Inf. Process. Manag.3
2020 Joint metric and feature representation learning for unsupervised domain adaptation
Zhekai Du, Jingjing Li 0001, Mengmeng Jing, Erpeng Chen, Ke Lu 0001
Knowl. Based Syst.4
2020 Learning explicitly transferable representations for domain adaptation
Mengmeng Jing, Jingjing Li 0001, Ke Lu 0001, Lei Zhu 0002, Yang Yang 0002
Neural Networks1
2019 From Zero-Shot Learning to Cold-Start Recommendation
abstract
Zero-shot learning (ZSL) and cold-start recommendation (CSR) are two challenging problems in computer vision and recommender system, respectively. In general, they are independently investigated in different communities. This paper, however, reveals that ZSL and CSR are two extensions of the same intension. Both of them, for instance, attempt to predict unseen classes and involve two spaces, one for direct feature representation and the other for supplementary description. Yet there is no existing approach which addresses CSR from the ZSL perspective. This work, for the first time, formulates CSR as a ZSL problem, and a tailor-made ZSL method is proposed to handle CSR. Specifically, we propose a Lowrank Linear Auto-Encoder (LLAE), which challenges three cruxes, i.e., domain shift, spurious correlations and computing efficiency, in this paper. LLAE consists of two parts, a low-rank encoder maps user behavior into user attributes and a symmetric decoder reconstructs user behavior from user attributes. Extensive experiments on both ZSL and CSR tasks verify that the proposed method is a win-win formulation, i.e., not only can CSR be handled by ZSL models with a significant performance improvement compared with several conventional state-of-the-art methods, but the consideration of CSR can benefit ZSL as well.
Jingjing Li 0001, Mengmeng Jing, Ke Lu 0001, Lei Zhu 0002, Yang Yang 0002, Zi Huang
AAAI2
2019 Leveraging the Invariant Side of Generative Zero-Shot Learning
abstract
Conventional zero-shot learning (ZSL) methods generally learn an embedding, e.g., visual-semantic mapping, to handle the unseen visual samples via an indirect manner. In this paper, we take the advantage of generative adversarial networks (GANs) and propose a novel method, named leveraging invariant side GAN (LisGAN), which can directly generate the unseen features from random noises which are conditioned by the semantic descriptions. Specifically, we train a conditional Wasserstein GANs in which the generator synthesizes fake unseen features from noises and the discriminator distinguishes the fake from real via a minimax game. Considering that one semantic description can correspond to various synthesized visual samples, and the semantic description, figuratively, is the soul of the generated features, we introduce soul samples as the invariant side of generative zero-shot learning in this paper. A soul sample is the meta-representation of one class. It visualizes the most semantically-meaningful aspects of each sample in the same category. We regularize that each generated sample (the varying side of generative ZSL) should be close to at least one soul sample (the invariant side) which has the same class label with it. At the zero-shot recognition stage, we propose to use two classifiers, which are deployed in a cascade way, to achieve a coarse-to-fine result. Experiments on five popular benchmarks verify that our proposed approach can outperform state-of-the-art methods with significant improvements.
Jingjing Li 0001, Mengmeng Jing, Ke Lu 0001, Zhengming Ding, Lei Zhu 0002, Zi Huang
CVPR2
2019 Adaptive Component Embedding for Unsupervised Domain Adaptation
abstract
Domain adaptation has obtained considerable interest from the literatures of multimedia, especially in cross-domain knowledge transfer problems. In this paper, we propose an effective yet time-saving approach, named Adaptive Component Embedding (ACE), for unsupervised domain adaptation. Specifically, ACE learns adaptive components across domains to embed all data in a shared subspace where the distribution divergence is mitigated and the underlying geometric structures in the local manifold are preserved. Then, an adaptive classifier is learned by using Representer Theorem in the Reproducing Kernel Hilbert Space (RKHS). The objective of our method can be efficiently solved in a closed form. Comprehensive experiments on both standard and large-scale datasets verify that ACE significantly outperforms previous state-of-the-art methods in terms of the classification accuracy and training time.
Mengmeng Jing, Jingjing Li 0001, Ke Lu 0001, Jieyan Liu, Zi Huang
ICME1
2019 Agile Domain Adaptation
abstract
Domain adaptation investigates the problem of leveraging knowledge from a well-labeled source domain to an unlabeled target domain, where the two domains are drawn from different data distributions. Because of the distribution shifts, different target samples have distinct degrees of difficulty in adaptation. However, existing domain adaptation approaches overwhelmingly neglect the degrees of difficulty and deploy exactly the same framework for all of the target samples. Generally, a simple or shadow framework is fast but rough. A sophisticated or deep framework, on the contrary, is accurate but slow. In this paper, we aim to challenge the fundamental contradiction between the accuracy and speed in domain adaptation tasks. We propose a novel approach, named agile domain adaptation, which agilely applies optimal frameworks to different target samples and classifies the target samples according to their adaptation difficulties. Specifically, we propose a paradigm which performs several early detections before the final classification. If a sample can be classified at one of the early stage with enough confidence, the sample would exit without the subsequent processes. Notably, the proposed method can significantly reduce the running cost of domain adaptation approaches, which can extend the application scenarios of domain adaptation to even mobile devices and real-time systems. Extensive experiments on two open benchmarks verify the effectiveness and efficiency of the proposed method.
Jingjing Li 0001, Mengmeng Jing, Ke Lu 0001, Zi Huang
IJCNN2
2019 Alleviating Feature Confusion for Generative Zero-shot Learning
abstract
Lately, generative adversarial networks (GANs) have been successfully applied to zero-shot learning (ZSL) and achieved state-of-the-art performance. By synthesizing virtual unseen visual features, GAN-based methods convert the challenging ZSL task into a supervised learning problem. However, since real unseen visual features are not available at the training stage, GAN-based ZSL methods have to train the GAN generator on the seen categories and further apply it to unseen instances. An inevitable issue of such a paradigm is that the synthesized unseen features are prone to seen references and incapable to reflect the novelty and diversity of real unseen instances. In a nutshell, the synthesized features are confusing. One cannot tell unseen categories from seen ones using the synthesized features. As a result, the synthesized features are too subtle to be classified in generalized zero-shot learning (GZSL) which involves both seen and unseen categories at the test stage. In this paper, we first introduce the feature confusion issue. Then, we propose a new feature generating network, named alleviating feature confusion GAN (AFC-GAN), to challenge the issue. Specifically, we present a boundary loss which maximizes the decision boundary of seen categories and unseen ones. Furthermore, a novel metric named feature confusion score (FCS) is proposed to quantify the feature confusion. Extensive experiments on five widely used datasets verify that our method is able to outperform previous state-of-the-arts under both ZSL and GZSL protocols.
Jingjing Li 0001, Mengmeng Jing, Ke Lu 0001, Lei Zhu 0002, Yang Yang 0002, Zi Huang
ACM Multimedia2
2019 Locality Preserving Joint Transfer for Domain Adaptation
abstract
Domain adaptation aims to leverage knowledge from a well-labeled source domain to a poorly labeled target domain. A majority of existing works transfer the knowledge at either feature level or sample level. Recent studies reveal that both of the paradigms are essentially important, and optimizing one of them can reinforce the other. Inspired by this, we propose a novel approach to jointly exploit feature adaptation with distribution matching and sample adaptation with landmark selection. During the knowledge transfer, we also take the local consistency between the samples into consideration so that the manifold structures of samples can be preserved. At last, we deploy label propagation to predict the categories of new instances. Notably, our approach is suitable for both homogeneous- and heterogeneous-domain adaptations by learning domain-specific projections. Extensive experiments on five open benchmarks, which consist of both standard and large-scale datasets, verify that our approach can significantly outperform not only conventional approaches but also end-to-end deep models. The experiments also demonstrate that we can leverage handcrafted features to promote the accuracy on deep features by heterogeneous adaptation.
Jingjing Li 0001, Mengmeng Jing, Ke Lu 0001, Lei Zhu 0002, Heng Tao Shen
IEEE Trans. Image Process.2
2018 Learning Distribution-Matched Landmarks for Unsupervised Domain Adaptation
Mengmeng Jing, Jingjing Li 0001, Jidong Zhao, Ke Lu 0001
DASFAA (2)1