VLDB 2026 Research / reviewers in the wild / expert
Zhiheng Ma
dblp:173/9652
· DBLP profile ↗
34ranked-venue papers
6as first author
29since 2021 · last 2026
0000-0002-0034-2065ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 25 · 3 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 6 first-author · 20 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Sample-Aware Knowledge Association and Enhancement for Open-Vocabulary Continual Learning
Zhilin Zhu 0001, Zhiheng Ma, Yabin Wang 0001, Yaguang Song, Yaowei Wang 0001, Xiaopeng Hong |
Int. J. Comput. Vis. | 2 |
| 2026 | Dual-Attention based prompt generation and catalyzing for instance-wise continual learning
Xiaopeng Hong, Yabin Wang 0001, Zhiheng Ma, Jinfeng Yang, Dongmei Jiang, Yaowei Wang 0001 |
Pattern Recognit. | 4 |
| 2026 | Linguistic profiling of deepfakes: An open database for next-Generation deepfake detection
Yabin Wang 0001, Xiaopeng Hong, Zhiheng Ma, Zhiwu Huang |
Pattern Recognit. | 4 |
| 2026 | Asymmetric modal fusion for multi-modal crowd counting
Xiaopeng Hong, Zhiheng Ma, Yabin Wang 0001 |
Pattern Recognit. | 3 |
| 2026 | Continual Conceptual Entity Learning for Text-to-Image Generative ModelsabstractCurrent Text-to-Image generative models struggle to continuously learn multiple distinct entities or concepts, limiting their scalability and hindering practical deployment in dynamic environments. We formulate this task as Continual Conceptual Entity Learning (CEL) and propose a novel framework called Continual Entity Adapter Learning (CEAL). CEAL leverages a compact set of tunable parameters, termed SuperLoRA, to efficient and scalable learning of new entities. We propose a dynamic rank-increasing strategy to train the SuperLoRA, balancing computational efficiency with performance. To evaluate our method, we create three benchmarks encompassing generic objects, human faces, and artistic styles. Experimental results demonstrate that CEAL effectively learns new entities while preserving prior knowledge, outperforming existing methods in both entity fidelity and parameter efficiency. Yabin Wang 0001, Xiaopeng Hong, Zhiheng Ma, Zhou Su 0001, Zhiwu Huang |
IEEE Trans. Multim. | 3 |
| 2025 | ComprehendEdit: A Comprehensive Dataset and Evaluation Framework for Multimodal Knowledge EditingabstractLarge multimodal language models (MLLMs) have revolutionized natural language processing and visual understanding, but often contain outdated or inaccurate information. Current multimodal knowledge editing evaluations are limited in scope and potentially biased, focusing on narrow tasks and failing to assess the impact on in-domain samples. To address these issues, we introduce ComprehendEdit, a comprehensive benchmark comprising eight diverse tasks from multiple datasets. We propose two novel metrics: Knowledge Generalization Index (KGI) and Knowledge Preservation Index (KPI), which evaluate editing effects on in-domain samples without relying on AI-synthetic samples. Based on insights from our framework, we establish Hierarchical In-Context Editing (HICE), a baseline method employing a two-stage approach that balances performance across all metrics. This study provides a more comprehensive evaluation framework for multimodal knowledge editing, reveals unique challenges in this field, and offers a baseline method demonstrating improved performance. Our work opens new perspectives for future research and provides a foundation for developing more robust and effective editing techniques for MLLMs. Yaohui Ma, Xiaopeng Hong, Shizhou Zhang, Huiyun Li, Zhilin Zhu 0001, Wei Luo 0014, Zhiheng Ma |
AAAI | 7 |
| 2025 | FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-trainingabstractLanguage-image pre-training faces significant challenges due to limited data in specific formats and the constrained capacities of text encoders. While prevailing methods attempt to address these issues through data augmentation and architecture modifications, they continue to struggle with processing long-form text inputs, and the inherent limitations of traditional CLIP text encoders lead to suboptimal downstream generalization. In this paper, we propose FLAME (Frozen Large lAnguage Models Enable data-efficient language-image pre-training) that leverages frozen large language models as text encoders, naturally processing long text inputs and demonstrating impressive multilingual generalization. FLAME comprises two key components: 1) a multifaceted prompt distillation technique for extracting diverse semantic representations from long captions, which better aligns with the multifaceted nature of images, and 2) a facet-decoupled attention mechanism, complemented by an offline embedding strategy, to ensure efficient computation. Extensive empirical evaluations demonstrate FLAME’s superior performance. When trained on CC3M, FLAME surpasses the previous state-of-the-art by 4.9% in ImageNet top-1 accuracy. On YFCC15M, FLAME surpasses the WIT-400M-trained CLIP by 44.4% in average image-to-text recall@1 across 36 languages, and by 34.6% in text-to-image recall@1 for long-context retrieval on Urban-1k. Code is available at https://github.com/MIV-XJTU/FLAME. Anjia Cao, Xing Wei 0001, Zhiheng Ma |
CVPR | 3 |
| 2025 | Semi-Supervised Counting via Pixel-by-Pixel Density Distribution ModelingabstractThis paper focuses on semi-supervised crowd counting, where only a small portion of the training data are labeled. We formulate the pixel-wise density value to regress as a probability distribution, instead of a single deterministic value. On this basis, we propose a semi-supervised crowd counting model. First, we design a pixel-wise distribution matching loss to measure the differences in the pixel-wise density distributions between the prediction and the ground-truth; Second, we enhance the transformer decoder by using density tokens to specialize the forwards of decoders w.r.t. different density intervals; Third, we design the interleaving consistency self-supervised learning mechanism to learn from unlabeled data efficiently. Extensive experiments on four datasets are performed to show that our method clearly outperforms the competitors by a large margin under various labeled ratio settings. Zhiheng Ma, Rongrong Ji, Yaowei Wang 0001, Zhou Su 0001, Xiaopeng Hong, Deyu Meng |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Joint Memory Optimization for Continual LearningabstractContinual learning, focusing on sequential knowledge acquisition and retention, necessitates efficient memory management. This paper introduces a holistic approach, diverging from traditional methods that separately optimize neural network and replay buffer memory. We aim to enhance overall memory efficiency, addressing neural network parameters and replay buffer concurrently within strict memory constraints. This is achieved by harnessing neural network parameter redundancies and employing compression techniques like pruning and quantization, allowing data replay storage without extra memory overhead. Balancing memory use across components is challenging due to the complex search space of combined tasks. We tackle this by conceptualizing it as a bi-level optimization problem, integrating all tasks under a single objective, thus optimizing memory use and managing the interplay between different components. We employ a synergy of optimization techniques to solve this challenging bi-level optimization problem. Our experimental findings affirm the superior performance of our proposed method, outperforming existing techniques such as prompt-based, feature-replay, exemplar-replay, and regularization-based methods under stringent memory constraints, consistently across various datasets and neural network architectures. Zhiheng Ma, Yaohui Ma, Xiaopeng Hong, Huiyun Li, Shizhou Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Curriculum Dataset DistillationabstractMost dataset distillation methods struggle to accommodate large-scale datasets due to their substantial computational and memory requirements. Recent research has begun to explore scalable disentanglement methods. However, there are still performance bottlenecks and room for optimization in this direction. In this paper, we present a curriculum-based dataset distillation framework aiming to harmonize performance and scalability. This framework strategically distills synthetic images, adhering to a curriculum that transitions from simple to complex. By incorporating curriculum evaluation, we address the issue of previous methods generating images that tend to be homogeneous and simplistic, doing so at a manageable computational cost. Furthermore, we introduce adversarial optimization towards synthetic images to further improve their representativeness and safeguard against their overfitting to the neural network involved in distilling. This enhances the generalization capability of the distilled images across various neural network architectures and also increases their robustness to noise. Extensive experiments demonstrate that our framework sets new benchmarks in large-scale dataset distillation, achieving substantial improvements of 11.1% on Tiny-ImageNet, 9.0% on ImageNet-1K, and 7.3% on ImageNet-21K. Our distilled datasets and code are available at https://github.com/MIV-XJTU/CUDD. Zhiheng Ma, Anjia Cao, Funing Yang, Yihong Gong, Xing Wei 0001 |
IEEE Trans. Image Process. | 1 |
| 2025 | Multidimensional Measure Matching for Crowd CountingabstractThis article addresses the challenge of scale variations in crowd-counting problems from a multidimensional measure-theoretic perspective. We start by formulating crowd counting as a measure-matching problem, based on the assumption that discrete measures can express the scattered ground truth and the predicted density map. In this context, we introduce the Sinkhorn counting loss and extend it to the semi-balanced form, which alleviates the problems including entropic bias, distance destruction, and amount constraints. We then model the measure matching under the multidimensional space, in order to learn the counting from both location and scale. To achieve this, we extend the traditional 2-D coordinate support to 3-D, incorporating an additional axis to represent scale information, where a pyramid-based structure will be leveraged to learn the scale value for the predicted density. Extensive experiments on four challenging crowd-counting datasets, namely, ShanghaiTech A, UCF-QNRF, JHU++, and NWPU have validated the proposed method. Code is released at https://github.com/LoraLinH/Multidimensional-Measure-Matching-for-Crowd-Counting. Xiaopeng Hong, Zhiheng Ma, Yaowei Wang 0001, Deyu Meng |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Evolving Parameterized Prompt Memory for Continual LearningabstractRecent studies have demonstrated the potency of leveraging prompts in Transformers for continual learning (CL). Nevertheless, employing a discrete key-prompt bottleneck can lead to selection mismatches and inappropriate prompt associations during testing. Furthermore, this approach hinders adaptive prompting due to the lack of shareability among nearly identical instances at more granular level. To address these challenges, we introduce the Evolving Parameterized Prompt Memory (EvoPrompt), a novel method involving adaptive and continuous prompting attached to pre-trained Vision Transformer (ViT), conditioned on specific instance. We formulate a continuous prompt function as a neural bottleneck and encode the collection of prompts on network weights. We establish a paired prompt memory system consisting of a stable reference and a flexible working prompt memory. Inspired by linear mode connectivity, we progressively fuse the working prompt memory and reference prompt memory during inter-task periods, resulting in continually evolved prompt memory. This fusion involves aligning functionally equivalent prompts using optimal transport and aggregating them in parameter space with an adjustable bias based on prompt node attribution. Additionally, to enhance backward compatibility, we propose compositional classifier initialization, which leverages prior prototypes from pre-trained models to guide the initialization of new classifiers in a subspace-aware manner. Comprehensive experiments validate that our approach achieves state-of-the-art performance in both class and domain incremental learning scenarios. Muhammad Rifki Kurniawan, Xiang Song 0005, Zhiheng Ma, Yuhang He 0001, Yihong Gong, Xing Wei 0001 |
AAAI | 3 |
| 2024 | Gramformer: Learning Crowd Counting via Graph-Modulated TransformerabstractTransformer has been popular in recent crowd counting work since it breaks the limited receptive field of traditional CNNs. However, since crowd images always contain a large number of similar patches, the self-attention mechanism in Transformer tends to find a homogenized solution where the attention maps of almost all patches are identical. In this paper, we address this problem by proposing Gramformer: a graph-modulated transformer to enhance the network by adjusting the attention and input node features respectively on the basis of two different types of graphs. Firstly, an attention graph is proposed to diverse attention maps to attend to complementary information. The graph is building upon the dissimilarities between patches, modulating the attention in an anti-similarity fashion. Secondly, a feature-based centrality encoding is proposed to discover the centrality positions or importance of nodes. We encode them with a proposed centrality indices scheme to modulate the node features and similarity relationships. Extensive experiments on four challenging crowd counting datasets have validated the competitiveness of the proposed method. Code is available at https://github.com/LoraLinH/Gramformer. Zhiheng Ma, Xiaopeng Hong, Qinnan Shangguan, Deyu Meng |
AAAI | 2 |
| 2024 | Multi-modal Crowd Counting via Modal Emulation
Xiaopeng Hong, Zhiheng Ma, Yabin Wang 0001, Xiaopeng Fan 0001 |
BMVC | 3 |
| 2024 | Reshaping the Online Data Buffering and Organizing Mechanism for Continual Test-Time Adaptation
Zhilin Zhu 0001, Xiaopeng Hong, Zhiheng Ma, Weijun Zhuang, Yaohui Ma, Yaowei Wang 0001 |
ECCV (82) | 3 |
| 2024 | Few-shot online anomaly detection and segmentation
Shenxing Wei, Xing Wei 0001, Zhiheng Ma, Songlin Dong, Shaochen Zhang, Yihong Gong |
Knowl. Based Syst. | 3 |
| 2024 | Knowledge Synergy Learning for Multi-Modal TrackingabstractBenefiting from the rich information provided by different modalities, multi-modal tracking has shown significant improvements compared to single-modal tracking. However, in practical applications, multi-modal tracking still faces two major challenges. Firstly, it is crucial to effectively integrate the complementary information from different modalities in order to improve tracking performance. Secondly, as trackers are often deployed in dynamic environments, it is difficult to ensure complete multi-modal data. Thus, handling modal-missing issues is essential to achieve robust and reliable tracking. To address these challenges, this paper proposes a Knowledge Synergy Network (KSNet) that integrates multi-modal features into a comprehensive representation and incorporates a modal compensation mechanism to handle modal-missing issues. With this framework, a multi-modal tracker (KSTrack) is built and trained using multi-modal data. KSTrack is capable of handling both complete and incomplete multi-modal data during inference. Comprehensive experiments on four large-scale RGB-Thermal (RGB-T) and RGB-Depth (RGB-D) benchmarks show that KSTrack surpasses state-of-the-art multi-modal trackers when using multi-modal data and outperforms single-modal trackers by a large margin when using single-modal data. Yuhang He 0001, Zhiheng Ma, Xing Wei 0001, Yihong Gong |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Isolation and Impartial Aggregation: A Paradigm of Incremental Learning without InterferenceabstractThis paper focuses on the prevalent stage interference and stage performance imbalance of incremental learning. To avoid obvious stage learning bottlenecks, we propose a new incremental learning framework, which leverages a series of stage-isolated classifiers to perform the learning task at each stage, without interference from others. To be concrete, to aggregate multiple stage classifiers as a uniform one impartially, we first introduce a temperature-controlled energy metric for indicating the confidence score levels of the stage classifiers. We then propose an anchor-based energy self-normalization strategy to ensure the stage classifiers work at the same energy level. Finally, we design a voting-based inference augmentation strategy for robust inference. The proposed method is rehearsal-free and can work for almost all incremental learning scenarios. We evaluate the proposed method on four large datasets. Extensive results demonstrate the superiority of the proposed method in setting up new state-of-the-art overall performance. Code is available at https://github.com/iamwangyabin/ESN. Yabin Wang 0001, Zhiheng Ma, Zhiwu Huang, Yaowei Wang 0001, Zhou Su 0001, Xiaopeng Hong |
AAAI | 2 |
| 2023 | Sparse Parameterization for Epitomic Dataset DistillationabstractThe success of deep learning relies heavily on large and diverse datasets, but the storage, preprocessing, and training of such data present significant challenges. To address these challenges, dataset distillation techniques have been proposed to obtain smaller synthetic datasets that capture the essential information of the originals. In this paper, we introduce a Sparse Parameterization for Epitomic datasEt Distillation (SPEED) framework, which leverages the concept of dictionary learning and sparse coding to distill epitomes that represent pivotal information of the dataset. SPEED prioritizes proper parameterization of the synthetic dataset and introduces techniques to capture spatial redundancy within and between synthetic images. We propose Spatial-Agnostic Epitomic Tokens (SAETs) and Sparse Coding Matrices (SCMs) to efficiently represent and select significant features. Additionally, we build a Feature-Recurrent Network (FReeNet) to generate hierarchical features with high compression and storage efficiency. Experimental results demonstrate the superiority of SPEED in handling high-resolution datasets, achieving state-of-the-art performance on multiple benchmarks and downstream applications. Our framework is compatible with a variety of dataset matching approaches, generally enhancing their performance. This work highlights the importance of proper parameterization in epitomic dataset distillation and opens avenues for efficient representation learning. Source code is available at https://github.com/MIV-XJTU/SPEED. Xing Wei 0001, Anjia Cao, Funing Yang, Zhiheng Ma |
NeurIPS | 4 |
| 2023 | Topology-preserving transfer learning for weakly-supervised anomaly detection and segmentation
Shenxing Wei, Xing Wei 0001, Muhammad Rifki Kurniawan, Zhiheng Ma, Yihong Gong |
Pattern Recognit. Lett. | 4 |
| 2023 | Semi-Supervised Crowd Counting via Multiple Representation LearningabstractThere has been a growing interest in counting crowds through computer vision and machine learning techniques in recent years. Despite that significant progress has been made, most existing methods heavily rely on fully-supervised learning and require a lot of labeled data. To alleviate the reliance, we focus on the semi-supervised learning paradigm. Usually, crowd counting is converted to a density estimation problem. The model is trained to predict a density map and obtains the total count by accumulating densities over all the locations. In particular, we find that there could be multiple density map representations for a given image in a way that they differ in probability distribution forms but reach a consensus on their total counts. Therefore, we propose multiple representation learning to train several models. Each model focuses on a specific density representation and utilizes the count consistency between models to supervise unlabeled data. To bypass the explicit density regression problem, which makes a strong parametric assumption on the underlying density distribution, we propose an implicit density representation method based on the kernel mean embedding. Extensive experiments demonstrate that our approach outperforms state-of-the-art semi-supervised methods significantly. Xing Wei 0001, Yunfeng Qiu, Zhiheng Ma, Xiaopeng Hong, Yihong Gong |
IEEE Trans. Image Process. | 3 |
| 2022 | Boosting Crowd Counting via Multifaceted AttentionabstractThis paper focuses on the challenging crowd counting task. As large-scale variations often exist within crowd images, neither fixed-size convolution kernel of CNN nor fixed-size attention of recent vision transformers can well handle this kind of variations. To address this problem, we propose a Multifaceted Attention Network (MAN) to improve transformer models in local spatial relation encoding. MAN incorporates global attention from vanilla transformer, learnable local attention, and instance attention into a counting model. Firstly, the local Learnable Region Attention (LRA) is proposed to assign attention exclusive for each feature location dynamically. Secondly, we design the Local Attention Regularization to supervise the training of LRA by minimizing the deviation among the attention for different feature locations. Finally, we provide an Instance Attention mechanism to focus on the most important instances dynamically during training. Extensive experiments on four challenging crowd counting datasets namely ShanghaiTech, UCF-QNRF, JHU++, and NWPU have validated the proposed method. Code: https://github.com/LoraLinH/Boosting-Crowd-Counting-via-Multifaceted-Attention. Zhiheng Ma, Rongrong Ji, Yaowei Wang 0001, Xiaopeng Hong |
CVPR | 2 |
| 2022 | Semi-supervised Crowd Counting via Density AgencyabstractIn this paper, we propose a new agency-guided semi-supervised counting approach. First, we build a learnable auxiliary structure, namely the density agency to bring the recognized foreground regional features close to corresponding density sub-classes (agents) and push away background ones. Second, we propose a density-guided contrastive learning loss to consolidate the backbone feature extractor. Third, we build a regression head by using a transformer structure to refine the foreground features further. Finally, an efficient noise depression loss is provided to minimize the negative influence of annotation noises. Extensive experiments on four challenging crowd counting datasets demonstrate that our method achieves superior performance to the state-of-the-art semi-supervised counting methods by a large margin. The code is available at https://github.com/LoraLinH/Semi-supervised-Crowd-Counting-via-Density-Agency. Zhiheng Ma, Xiaopeng Hong, Yaowei Wang 0001, Zhou Su 0001 |
ACM Multimedia | 2 |
| 2022 | ECCNAS: Efficient Crowd Counting Neural Architecture SearchabstractRecent solutions to crowd counting problems have already achieved promising performance across various benchmarks. However, applying these approaches to real-world applications is still challenging, because they are computation intensive and lack the flexibility to meet various resource budgets. In this article, we propose an efficient crowd counting neural architecture search (ECCNAS) framework to search efficient crowd counting network structures, which can fill this research gap. A novel search from pre-trained strategy enables our cross-task NAS to explore the significantly large and flexible search space with less search time and get more proper network structures. Moreover, our well-designed search space can intrinsically provide candidate neural network structures with high performance and efficiency. In order to search network structures according to hardwares with different computational performance, we develop a novel latency cost estimation algorithm in our ECCNAS. Experiments show our searched models get an excellent trade-off between computational complexity and accuracy and have the potential to deploy in practical scenarios with various resource budgets. We reduce the computational cost, in terms of multiply-and-accumulate (MACs), by up to 96% with comparable accuracy. And we further designed experiments to validate the efficiency and the stability improvement of our proposed search from pre-trained strategy. Yabin Wang 0001, Zhiheng Ma, Xing Wei 0001, Yaowei Wang 0001, Xiaopeng Hong |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2021 | Error-Aware Density Isomorphism Reconstruction for Unsupervised Cross-Domain Crowd CountingabstractThis paper focuses on the unsupervised domain adaptation problem for video-based crowd counting, in which we use labeled data as source domain and unlabelled video data as target domain. It is challenging as there is a huge gap between the source and the target domain and no annotations of samples are available in the target domain. The key issue is how to utilize unlabelled videos in the target domain for knowledge learning and transferring from the source domain. To tackle this problem, we propose a novel Error-aware Density Isomorphism REConstruction Network (EDIREC-Net) for cross-domain crowd counting. EDIREC-Net jointly transfers a pre-trained counting model to target domains using a density isomorphism reconstruction objective and models the reconstruction erroneousness by error reasoning. Specifically, as crowd flows in videos are consecutive, the density maps in adjacent frames turn out to be isomorphic. On this basis, we regard the density isomorphism reconstruction error as a self-supervised signal to transfer the pre-trained counting models to different target domains. Moreover, we leverage an estimation-reconstruction consistency to monitor the density reconstruction erroneousness and suppress unreliable density reconstructions during training. Experimental results on four benchmark datasets demonstrate the superiority of the proposed method and ablation studies investigate the efficiency and robustness. The source code is available at https://github.com/GehenHe/EDIREC-Net. Yuhang He 0001, Zhiheng Ma, Xing Wei 0001, Xiaopeng Hong, Wei Ke 0003, Yihong Gong |
AAAI | 2 |
| 2021 | Learning to Count via Unbalanced Optimal TransportabstractCounting dense crowds through computer vision technology has attracted widespread attention. Most crowd counting datasets use point annotations. In this paper, we formulate crowd counting as a measure regression problem to minimize the distance between two measures with different supports and unequal total mass. Specifically, we adopt the unbalanced optimal transport distance, which remains stable under spatial perturbations, to quantify the discrepancy between predicted density maps and point annotations. An efficient optimization algorithm based on the regularized semi-dual formulation of UOT is introduced, which alternatively learns the optimal transportation and optimizes the density regressor. The quantitative and qualitative results illustrate that our method achieves state-of-the-art counting and localization performance. Zhiheng Ma, Xing Wei 0001, Xiaopeng Hong, Yunfeng Qiu, Yihong Gong |
AAAI | 1 |
| 2021 | Towards A Universal Model for Cross-Dataset Crowd CountingabstractThis paper proposes to handle the practical problem of learning a universal model for crowd counting across scenes and datasets. We dissect that the crux of this problem is the catastrophic sensitivity of crowd counters to scale shift, which is very common in the real world and caused by factors such as different scene layouts and image resolutions. Therefore it is difficult to train a universal model that can be applied to various scenes. To address this problem, we propose scale alignment as a prime module for establishing a novel crowd counting framework. We derive a closed-form solution to get the optimal image rescaling factors for alignment by minimizing the distances between their scale distributions. A novel neural network together with a loss function based on an efficient sliced Wasserstein distance is also proposed for scale distribution estimation. Benefiting from the proposed method, we have learned a universal model that generally works well on several datasets where can even outperform state-of-the-art models that are particularly fine-tuned for each dataset significantly. Experiments also demonstrate the much better generalizability of our model to unseen scenes. Zhiheng Ma, Xiaopeng Hong, Xing Wei 0001, Yunfeng Qiu, Yihong Gong |
ICCV | 1 |
| 2021 | Anomaly Detection Via Self-Organizing MapabstractAnomaly detection plays a key role in industrial manufacturing for product quality control. Traditional methods for anomaly detection are rule-based with limited generalization ability. Recent methods based on supervised deep learning are more powerful but require large-scale annotated datasets for training. In practice, abnormal products are rare thus it is very difficult to train a deep model in a fully supervised way. In this paper, we propose a novel unsupervised anomaly detection approach based on Self-organizing Map (SOM). Our method, Self-organizing Map for Anomaly Detection (SOMAD) maintains normal characteristics by using topological memory based on multi-scale features. SOMAD achieves state-of-the-art performance on unsupervised anomaly detection and localization on the MVTec dataset. Kaitao Jiang, Zhiheng Ma, Xing Wei 0001, Xiaopeng Hong, Yihong Gong |
ICIP | 3 |
| 2021 | Direct Measure Matching for Crowd CountingabstractTraditional crowd counting approaches usually use Gaussian assumption to generate pseudo density ground truth, which suffers from problems like inaccurate estimation of the Gaussian kernel sizes. In this paper, we propose a new measure-based counting approach to regress the predicted density maps to the scattered point-annotated ground truth directly. First, crowd counting is formulated as a measure matching problem. Second, we derive a semi-balanced form of Sinkhorn divergence, based on which a Sinkhorn counting loss is designed for measure matching. Third, we propose a self-supervised mechanism by devising a Sinkhorn scale consistency loss to resist scale changes. Finally, an efficient optimization method is provided to minimize the overall loss function. Extensive experiments on four challenging crowd counting datasets namely ShanghaiTech, UCF-QNRF, JHU++ and NWPU have validated the proposed method. Xiaopeng Hong, Zhiheng Ma, Xing Wei 0001, Yunfeng Qiu, Yaowei Wang 0001, Yihong Gong |
IJCAI | 3 |
| 2020 | Superpixel Masking and Inpainting for Self-Supervised Anomaly Detection
Kaitao Jiang, Zhiheng Ma, Xing Wei 0001, Xiaopeng Hong, Yihong Gong |
BMVC | 4 |
| 2020 | Learning Scales from Points: A Scale-aware Probabilistic Model for Crowd CountingabstractCounting people automatically through computer vision technology is a challenging task. Recently, convolution neural network (CNN) based methods have made significant progress. Nonetheless, large scale variations of instances caused by, for example, perspective effects remain unsolved. Moreover, it is problematic to estimate scales with only point annotations. In this paper, we propose a scale-aware probabilistic model to handle this problem. Unlike previous methods that generate a single density map where instances of various scales are processed indiscriminately, we propose a density pyramid network (DPN), where each pyramid level handles instances within a particular scale range. Furthermore, we propose a scale distribution estimator (SDE) to learn scales of people from input data, under the weak supervision of point annotations. Finally, we adopt an instance-level probabilistic scale-aware model (IPSM) to guide the multi-scale training of DPN explicitly. Qualitative and quantitative experimental results demonstrate the effectiveness of the proposed method, which achieves competitive results on four widely used benchmarks. Zhiheng Ma, Xing Wei 0001, Xiaopeng Hong, Yihong Gong |
ACM Multimedia | 1 |
| 2020 | Transductive semi-supervised metric learning for person re-identification
Xinyuan Chang, Zhiheng Ma, Xing Wei 0001, Xiaopeng Hong, Yihong Gong |
Pattern Recognit. | 2 |
| 2019 | Bayesian Loss for Crowd Count Estimation With Point SupervisionabstractIn crowd counting datasets, each person is annotated by a point, which is usually the center of the head. And the task is to estimate the total count in a crowd scene. Most of the state-of-the-art methods are based on density map estimation, which convert the sparse point annotations into a “ground truth” density map through a Gaussian kernel, and then use it as the learning target to train a density map estimator. However, such a "ground-truth" density map is imperfect due to occlusions, perspective effects, variations in object shapes, etc. On the contrary, we propose Bayesian loss, a novel loss function which constructs a density contribution probability model from the point annotations. Instead of constraining the value at every pixel in the density map, the proposed training loss adopts a more reliable supervision on the count expectation at each annotated point. Without bells and whistles, the loss function makes substantial improvements over the baseline loss on all tested datasets. Moreover, our proposed loss function equipped with a standard backbone network, without using any external detectors or multi-scale architectures, plays favourably against the state of the arts. Our method outperforms previous best approaches by a large margin on the latest and largest UCF-QNRF dataset. Zhiheng Ma, Xing Wei 0001, Xiaopeng Hong, Yihong Gong |
ICCV | 1 |
| 2018 | Transductive Semi-Supervised Deep Learning Using Min-Max Features
Weiwei Shi 0003, Yihong Gong, Chris Ding, Zhiheng Ma, Nanning Zheng 0001 |
ECCV (5) | 4 |