VLDB 2026 Research / reviewers in the wild / expert
Jia Wan 0001
dblp:13/6504-1
· DBLP profile ↗
30ranked-venue papers
12as first author
22since 2021 · last 2026
0000-0001-8198-1629ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 8 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 7 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adaptive momentum weight averaging reduces initialization noise
Jia Wan 0001, Ziquan Liu, Junyu Gao 0001, Antoni B. Chan |
Pattern Recognit. | 1 |
| 2025 | Diffusion-based Data Augmentation for Object Counting ProblemsabstractCrowd counting, an important problem in computer vision, is commonly solved with deep learning approaches, such as convolutional networks and transformers. However, Deep networks often overfit when the available labeled crowd data is scarce. To overcome this, we have designed a pipeline that utilizes a diffusion model to generate extensive training data. We pioneer using diffusion models to generate images from high-density head location dot maps (a binary dot map that specifies the location of human heads) and are the first to use these diverse synthetic data to augment the crowd counting models. Our proposed smoothed density map input for ControlNet significantly improves ControlNet’s performance in generating crowds in the correct locations. Also, our proposed counting loss and guidance sampling for the diffusion model effectively minimize the discrepancies between the location dot map and the crowd images generated. Moreover, our versatile framework can be easily adapted to all kinds of counting problems. Extensive experiments demonstrate that our framework improves the counting performance on the ShanghaiTech, NWPU-Crowd, UCF-QNRF, and TRANCOS datasets. Yuelei Li, Jia Wan 0001, Nuno Vasconcelos |
ICASSP | 3 |
| 2025 | Video Individual Counting for Moving DronesabstractVideo Individual Counting (VIC) has received increasing attention for its importance in intelligent video surveillance. Existing works are limited in two aspects, i.e., dataset and method. Previous datasets are captured with fixed or rarely moving cameras with relatively sparse individuals, restricting evaluation for a highly varying view and time in crowded scenes. Existing methods rely on localization followed by association or classification, which struggle under dense and dynamic conditions due to inaccurate localization of small targets. To address these issues, we introduce the MovingDroneCrowd Dataset, featuring videos captured by fast-moving drones in crowded scenes under diverse illuminations, shooting heights and angles. We further propose a Shared Density map-guided Network (SDNet) using a Depth-wise Cross-Frame Attention (DCFA) module to directly estimate shared density maps between consecutive frames, from which the inflow and outflow density maps are derived by subtracting the shared density maps from the global density maps. The inflow density maps across frames are summed up to obtain the number of unique pedestrians in a video. Experiments on our datasets and publicly available ones show the superiority of our method over the state of the arts in highly dynamic and complex crowded scenes. Our dataset and codes have been released publicly. Yaowu Fan, Jia Wan 0001, Tao Han 0002, Antoni B. Chan, Andy Jinhua Ma |
ICCV | 2 |
| 2025 | Temporal Unlearnable Examples: Preventing Personal Video Data from Unauthorized Exploitation by Object TrackingabstractWith the rise of social media, vast amounts of user-uploaded videos (e.g., YouTube) are utilized as training data for Visual Object Tracking (VOT). However, the VOT community has largely overlooked video data-privacy issues, as many private videos have been collected and used for training commercial models without authorization. To alleviate these issues, this paper presents the first investigation on preventing personal video data from unauthorized exploitation by deep trackers. Existing methods for preventing unauthorized data use primarily focus on image-based tasks (e.g., image classification), directly applying them to videos reveals several limitations, including inefficiency, limited effectiveness, and poor generalizability. To address these issues, we propose a novel generative framework for generating Temporal Unlearnable Examples (TUEs), and whose efficient computation makes it scalable for usage on large-scale video datasets. The trackers trained w/ TUEs heavily rely on unlearnable noises for temporal matching, ignoring the original data structure and thus ensuring training video data-privacy. To enhance the effectiveness of TUEs, we introduce a temporal contrastive loss, which further corrupts the learning of existing trackers when using our TUEs for training. Extensive experiments demonstrate that our approach achieves state-of-the-art performance in video data-privacy protection, with strong transferability across VOT models, datasets, and temporal matching tasks. Qiangqiang Wu, Yi Yu 0011, Chenqi Kong, Ziquan Liu, Jia Wan 0001, Haoliang Li, Alex Chichung Kot, Antoni B. Chan |
ICCV | 5 |
| 2025 | Proximal Mapping Loss: Understanding Loss Functions in Crowd Counting & LocalizationabstractCrowd counting and localization involve extracting the number and distribution of crowds from images or videos using computer vision techniques. Most counting methods are based on density regression and are based on an ``intersection'' hypothesis, *i.e.*, one pixel is influenced by multiple points in the ground truth, which is inconsistent with reality since one pixel would not contain two objects. This paper proposes Proximal Mapping Loss (PML), a density regression method that eliminates this hypothesis. {PML} divides the predicted density map into multiple point-neighbor cases through the nearest neighbor, and then dynamically constructs a learning target for each sub-case via proximal mapping, leading to more robust and accurate training. {Furthermore}, PML is theoretically linked to various existing loss functions, such as Gaussian-blurred L2 loss, Bayesian loss, and the training schemes in P2PNet and DMC, demonstrating its versatility and adaptability. Experimentally, PML significantly improves the performance of crowd counting and localization, and illustrates the robustness against annotation noise. The code is available at [https://github.com/Elin24/pml](https://github.com/Elin24/pml). Wei Lin 0018, Jia Wan 0001, Antoni B. Chan |
ICLR | 2 |
| 2025 | Embodied Crowd CountingabstractOcclusion is one of the fundamental challenges in crowd counting. In the community, various data-driven approaches have been developed to address this issue, yet their effectiveness is limited. This is mainly because most existing crowd counting datasets on which the methods are trained are based on passive cameras, restricting their ability to fully sense the environment.
Recently, embodied navigation methods have shown significant potential in precise object detection in interactive scenes. These methods incorporate active camera settings, holding promise in addressing the fundamental issues in crowd counting. However, most existing methods are designed for indoor navigation, showing unknown performance in analyzing complex object distribution in large-scale scenes, such as crowds. Besides, most existing embodied navigation datasets are indoor scenes with limited scale and object quantity, preventing them from being introduced into dense crowd analysis. Based on this, a novel task, Embodied Crowd Counting (ECC), is proposed to count the number of persons in a large-scale scene actively. We then build up an interactive simulator, the Embodied Crowd Counting Dataset (ECCD), which enables large-scale scenes and large object quantities. A prior probability distribution approximating a realistic crowd distribution is introduced to generate crowds. Then, a zero-shot navigation method (ZECC) is proposed as a baseline. This method contains an MLLM-driven coarse-to-fine navigation mechanism, enabling active Z-axis exploration, and a normal-line-based crowd distribution analysis method for fine counting. Experimental results show that the proposed method achieves the best trade-off between counting accuracy and navigation cost. Code can be found at https://github.com/longrunling/ECC?. Runling Long, Jia Wan 0001, Xiang Deng 0002, Xinting Zhu, Weili Guan, Antoni B. Chan, Liqiang Nie |
NeurIPS | 3 |
| 2025 | Learning Crowd Scale and Distribution for Weakly Supervised Crowd Counting and LocalizationabstractThe count supervision used in weakly-supervised crowd counting is derived from the number of point annotations, which means that the labeling cost is not effectively reduced. Moreover, due to the lack of spatial information about the pedestrians during training, previous works struggle to accurately learn the positions of individuals. To address these challenges, we propose a crowd counting and localization method based on scene-specific synthetic data for surveillance scenarios, which can accurately predict the number and location of person without any manually labeled point-wise or count-wise annotations. Our method dynamically adjust scene-specific synthetic data to minimize domain differences from surveillance scenes by learning the crowd scale and distribution. Specifically, based on realistic synthetic data, the models learn precise location and scale information, which can then regenerate new synthetic data with a more reasonable pedestrian distribution and scale and generate high-quality pseudo point-wise annotations. Subsequently, the counter is trained using our proposed robust soft-weighted loss function, under the joint supervision of auto-generated point-wise annotations on synthetic data and pseudo point-wise annotations on real data in an end-to-end manner. Our proposed loss function, based on the designed weighted optimal transport, effectively mitigates noise in pseudo point-wise labels and is not only insensitive to hyperparemeters but also exhibits superior generalization ability on real data. We conduct comprehensive experiments across multiple scene-specific datasets, demonstrating our method’s superiority in counting and localization performance over count-supervised, fully-supervised, and state-of-the-art domain adaption algorithms. Code is available athttps://github.com/fyw1999/LCSD. Yaowu Fan, Jia Wan 0001, Andy Jinhua Ma |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | TriSAM: Tri-Plane SAM for Zero-Shot Cortical Blood Vessel Segmentation in VEM ImagesabstractWhile imaging techniques at macro and mesoscales have garnered substantial attention and resources, microscale Volume Electron Microscopy (vEM) imaging, capable of revealing intricate vascular details, has lacked the necessary benchmarking infrastructure. In this paper, we address a significant gap in this field of neuroimaging by introducing the first-in-class public benchmark, BvEM, designed specifically for cortical blood vessel segmentation in vEM images. Our BvEM benchmark is based on vEM image volumes from three mammals: adult mouse, macaque, and human. We standardized the resolution, addressed imaging variations, and meticulously annotated blood vessels through semi-automatic, manual, and quality control processes, ensuring high-quality 3D segmentation. Furthermore, we developed a zero-shot cortical blood vessel segmentation method named TriSAM, which leverages the powerful segmentation model SAM for 3D segmentation. To extend SAM from 2D to 3D volume segmentation, TriSAM employs a multi-seed tracking framework, leveraging the reliability of certain image planes for tracking while using others to identify potential turning points. This approach effectively achieves long-term 3D blood vessel segmentation without model training or fine-tuning. Experimental results show that TriSAM achieved superior performances on the BvEM benchmark across three species. Jia Wan 0001, Wanhua Li 0001, Jason Ken Adhinarta, Atmadeep Banerjee, Evelina Sjöstedt, Jingpeng Wu, Jeff Lichtman, Hanspeter Pfister, Donglai Wei 0001 |
IEEE J. Biomed. Health Informatics | 1 |
| 2024 | Learning a Dynamic Privacy-Preserving Camera Robust to Inversion Attacks
Jia Wan 0001, Nicholas Antipa, Nuno Vasconcelos |
ECCV (70) | 3 |
| 2024 | Robust Zero-Shot Crowd Counting and Localization With Adaptive Resolution SAM
Jia Wan 0001, Qiangqiang Wu, Wei Lin 0018, Antoni B. Chan |
ECCV (57) | 1 |
| 2024 | Boosting 3D Single Object Tracking with 2D Matching Distillation and 3D Pre-training
Qiangqiang Wu, Yan Xia 0003, Jia Wan 0001, Antoni B. Chan |
ECCV (12) | 3 |
| 2024 | Generalized Characteristic Function Loss for Crowd Analysis in the Frequency DomainabstractTypical approaches that learn crowd density maps are limited to extracting the supervisory information from the loosely organized spatial information in the crowd dot/density maps. This paper tackles this challenge by performing the supervision in the frequency domain. More specifically, we devise a new loss function for crowd analysis called generalized characteristic function loss (GCFL). This loss carries out two steps: 1) transforming the spatial information in density or dot maps to the frequency domain; 2) calculating a loss value between their frequency contents. For step 1, we establish a series of theoretical fundaments by extending the definition of the characteristic function for probability distributions to density maps, as well as proving some vital properties of the extended characteristic function. After taking the characteristic function of the density map, its information in the frequency domain is well-organized and hierarchically distributed, while in the spatial domain it is loose-organized and dispersed everywhere. In step 2, we design a loss function that can fit the information organization in the frequency domain, allowing the exploitation of the well-organized frequency information for the supervision of crowd analysis tasks. The loss function can be adapted to various crowd analysis tasks through the specification of its window functions. In this paper, we demonstrate its power in three tasks: Crowd Counting, Crowd Localization and Noisy Crowd Counting. We show the advantages of our GCFL compared to other SOTA losses and its competitiveness to other SOTA methods by theoretical analysis and empirical results on benchmark datasets. Our codes are available at https://github.com/wbshu/Crowd_Counting_in_the_Frequency_Domain. Weibo Shu, Jia Wan 0001, Antoni B. Chan |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Weakly-Supervised Scene-Specific Crowd Counting Using Real-Synthetic Hybrid DataabstractDue to the domain gap between the public large-scale datasets and actual scenes, the crowd counting models trained on the common datasets have a significant performance degradation when applying in practical applications. To address the above issue, one of the solution is to label additional data from the novel scenes, which is time-consuming and impractical for multiple scenes. Another solution is to utilize domain adaptation approaches to adapt a well-trained model to novel scenes. However, most of these approaches focus on appearance adaptation while the background and the crowd distribution is not adapted. In this paper, we propose a weakly-supervised method with real-synthetic hybrid data which only requires a small portion of unlabelled real images and auto-generated synthetic labelled images for training. First, the hybrid data is generated based on background from the real scene and random distributed synthetic persons. Second, an initialized counter is trained based on the hybrid data and the crowd distribution is predicted based on the predictions on real images. Then, a better crowd counter is trained based on new hybrid data generated from updated crowd distribution. The process is iterated until convergence. Extensive experiments demonstrate the effectiveness of the proposed method. Yaowu Fan, Jia Wan 0001, Yuan Yuan 0001, Qi Wang 0009 |
ICASSP | 2 |
| 2023 | Modeling Noisy Annotations for Point-Wise SupervisionabstractPoint-wise supervision is widely adopted in computer vision tasks such as crowd counting and human pose estimation. In practice, the noise in point annotations may affect the performance and robustness of algorithm significantly. In this paper, we investigate the effect of annotation noise in point-wise supervision and propose a series of robust loss functions for different tasks. In particular, the point annotation noise includes spatial-shift noise, missing-point noise, and duplicate-point noise. The spatial-shift noise is the most common one, and exists in crowd counting, pose estimation, visual tracking, etc, while the missing-point and duplicate-point noises usually appear in dense annotations, such as crowd counting. In this paper, we first consider the shift noise by modeling the real locations as random variables and the annotated points as noisy observations. The probability density function of the intermediate representation (a smooth heat map generated from dot annotations) is derived and the negative log likelihood is used as the loss function to naturally model the shift uncertainty in the intermediate representation. The missing and duplicate noise are further modeled by an empirical way with the assumption that the noise appears at high density region with a high probability. We apply the method to crowd counting, human pose estimation and visual tracking, propose robust loss functions for those tasks, and achieve superior performance and robustness on widely used datasets. Jia Wan 0001, Qiangqiang Wu, Antoni B. Chan |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Crowd Counting in the Frequency DomainabstractThis paper investigates crowd counting in the frequency domain, which is a novel direction compared to the traditional view in the spatial domain. By transforming the density map into the frequency domain and using the properties of the characteristic function, we propose a novel method that is simple, effective, and efficient. The solid theoretical analysis ends up as an implementation-friendly loss function, which requires only standard tensor operations in the training process. We prove that our loss function is an upper bound of the pseudo sup norm metric between the ground truth and the prediction density map (over all of their sub-regions), and demonstrate its efficacy and efficiency versus other loss functions. The experimental results also show its competitiveness to the state-of-the-art on five benchmark data sets: ShanghaiTech A & B, UCF-QNRF, JHU++, and NWPU. Our codes will be available at: wb-shu/Crowd_Couniing_in_the_Frequency_Domain Weibo Shu, Jia Wan 0001, Kay Chen Tan, Sam Kwong, Antoni B. Chan |
CVPR | 2 |
| 2022 | Kernel-Based Density Map Generation for Dense Object CountingabstractCrowd counting is an essential topic in computer vision due to its practical usage in surveillance systems. The typical design of crowd counting algorithms is divided into two steps. First, the ground-truth density maps of crowd images are generated from the ground-truth dot maps (density map generation), e.g., by convolving with a Gaussian kernel. Second, deep learning models are designed to predict a density map from an input image (density map estimation). The density map based counting methods that incorporate density map as the intermediate representation have improved counting performance dramatically. However, in the sense of end-to-end training, the hand-crafted methods used for generating the density maps may not be optimal for the particular network or dataset used. To address this issue, we propose an adaptive density map generator, which takes the annotation dot map as input, and learns a density map representation for a counter. The counter and generator are trained jointly within an end-to-end framework. We also show that the proposed framework can be applied to general dense object counting tasks. Extensive experiments are conducted on 10 datasets for 3 applications: crowd counting, vehicle counting, and general object counting. The experiment results on these datasets confirm the effectiveness of the proposed learnable density map representations. Jia Wan 0001, Qingzhong Wang, Antoni B. Chan |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | On Diversity in Image Captioning: Metrics and MethodsabstractDiversity is one of the most important properties in image captioning, as it reflects various expressions of important concepts presented in an image. However, the most popular metrics cannot well evaluate the diversity of multiple captions. In this paper, we first propose a metric to measure the diversity of a set of captions, which is derived from latent semantic analysis (LSA), and then kernelize LSA using CIDEr (R. Vedantam et al., 2015) similarity. Compared with mBLEU (R. Shetty et al., 2017), our proposed diversity metrics show a relatively strong correlation to human evaluation. We conduct extensive experiments, finding there is a large gap between the performance of the current state-of-the-art models and human annotations considering both diversity and accuracy; the models that aim to generate captions with higher CIDEr scores normally obtain lower diversity scores, which generally learn to describe images using common words. To bridge this "diversity" gap, we consider several methods for training caption models to generate diverse captions. First, we show that balancing the cross-entropy loss and CIDEr reward in reinforcement learning during training can effectively control the tradeoff between diversity and accuracy of the generated captions. Second, we develop approaches that directly optimize our diversity metric and CIDEr score using reinforcement learning. These proposed approaches using reinforcement learning (RL) can be unified into a self-critical (S. J. Rennie et al., 2017) framework with new RL baselines. Third, we combine accuracy and diversity into a single measure using an ensemble matrix, and then maximize the determinant of the ensemble matrix via reinforcement learning to boost diversity and accuracy, which outperforms its counterparts on the oracle test. Finally, inspired by determinantal point processes (DPP), we develop a DPP selection algorithm to select a subset of captions from a large number of candidate captions. The experimental results show that maximizing the determinant of the ensemble matrix outperforms other methods considerably improving diversity and accuracy. Qingzhong Wang, Jia Wan 0001, Antoni B. Chan |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | A Generalized Loss Function for Crowd Counting and LocalizationabstractPrevious work [40] shows that a better density map representation can improve the performance of crowd counting. In this paper, we investigate learning the density map representation through an unbalanced optimal transport problem, and propose a generalized loss function to learn density maps for crowd counting and localization. We prove that pixel-wise L2 loss and Bayesian loss [29] are special cases and suboptimal solutions to our proposed loss function. A perspective-guided transport cost function is further proposed to better handle the perspective transformation in crowd images. Since the predicted density will be pushed toward annotation positions, the density map prediction will be sparse and can naturally be used for localization. Finally, the proposed loss outperforms other losses on four large-scale datasets for counting, and achieves the best localization performance on NWPU-Crowd and UCF-QNRF. Jia Wan 0001, Ziquan Liu, Antoni B. Chan |
CVPR | 1 |
| 2021 | Progressive Unsupervised Learning for Visual Object TrackingabstractIn this paper, we propose a progressive unsupervised learning (PUL) framework, which entirely removes the need for annotated training videos in visual tracking. Specifically, we first learn a background discrimination (BD) model that effectively distinguishes an object from back-ground in a contrastive learning way. We then employ the BD model to progressively mine temporal corresponding patches (i.e., patches connected by a track) in sequential frames. As the BD model is imperfect and thus the mined patch pairs are noisy, we propose a noise-robust loss function to more effectively learn temporal correspondences from this noisy data. We use the proposed noise robust loss to train backbone networks of Siamese trackers. Without online fine-tuning or adaptation, our unsupervised real-time Siamese trackers can outperform state-of-the-art unsupervised deep trackers and achieve competitive results to the supervised baselines. Qiangqiang Wu, Jia Wan 0001, Antoni B. Chan |
CVPR | 2 |
| 2021 | Dynamic Momentum Adaptation for Zero-Shot Cross-Domain Crowd CountingabstractZero-shot cross-domain crowd counting is a challenging task where a crowd counting model is trained on a source domain (i.e., training dataset) and no additional labeled or unlabeled data is available for fine-tuning the model when testing on an unseen target domain (i.e., a different testing dataset). The generalisation performance of existing crowd counting methods is typically limited due to the large gap between source and target domains. Here, we propose a novel Crowd Counting framework built upon an external Momentum Template, termed C2MoT, which enables the encoding of domain specific information via an external template representation. Specifically, the Momentum Template (MoT) is learned in a momentum updating way during offline training, and then is dynamically updated for each test image in online cross-dataset evaluation. Thanks to the dynamically updated MoT, our C2MoT effectively generates dense target correspondences that explicitly accounts for head regions, and then effectively predicts the density map based on the normalized correspondence map. Experiments on large scale datasets show that our proposed C2MoT achieves leading zero-shot cross-domain crowd counting performance without model fine-tuning, while also outperforming domain adaptation methods that use fine-tuning on target domain data. Moreover, C2MoT also obtains state-of-the-art counting performance on the source domain. Qiangqiang Wu, Jia Wan 0001, Antoni B. Chan |
ACM Multimedia | 2 |
| 2021 | Fine-Grained Crowd CountingabstractCurrent crowd counting algorithms are only concerned about the number of people in an image, which lacks low-level fine-grained information of the crowd. For many practical applications, the total number of people in an image is not as useful as the number of people in each sub-category. For example, knowing the number of people waiting inline or browsing can help retail stores; knowing the number of people standing/sitting can help restaurants/cafeterias; knowing the number of violent/non-violent people can help police in crowd management. In this article, we propose fine-grained crowd counting, which differentiates a crowd into categories based on the low-level behavior attributes of the individuals (e.g. standing/sitting or violent behavior) and then counts the number of people in each category. To enable research in this area, we construct a new dataset of four real-world fine-grained counting tasks: traveling direction on a sidewalk, standing or sitting, waiting in line or not, and exhibiting violent behavior or not. Since the appearance features of different crowd categories are similar, the challenge of fine-grained crowd counting is to effectively utilize contextual information to distinguish between categories. We propose a two branch architecture, consisting of a density map estimation branch and a semantic segmentation branch. We propose two refinement strategies for improving the predictions of the two branches. First, to encode contextual information, we propose feature propagation guided by the density map prediction, which eliminates the effect of background features during propagation. Second, we propose a complementary attention model to share information between the two branches. Experiment results confirm the effectiveness of our method. Jia Wan 0001, Nikil Senthil Kumar, Antoni B. Chan |
IEEE Trans. Image Process. | 1 |
| 2021 | Angular-Driven Feedback Restoration Networks for Imperfect Sketch RecognitionabstractAutomatic hand-drawn sketch recognition is an important task in computer vision. However, the vast majority of prior works focus on exploring the power of deep learning to achieve better accuracy on complete and clean sketch images, and thus fail to achieve satisfactory performance when applied to incomplete or destroyed sketch images. To address this problem, we first develop two datasets that contain different levels of scrawl and incomplete sketches. Then, we propose an angular-driven feedback restoration network (ADFRNet), which first detects the imperfect parts of a sketch and then refines them into high quality images, to boost the performance of sketch recognition. By introducing a novel "feedback restoration loop" to deliver information between the middle stages, the proposed model can improve the quality of generated sketch images while avoiding the extra memory cost associated with popular cascading generation schemes. In addition, we also employ a novel angular-based loss function to guide the refinement of sketch images and learn a powerful discriminator in the angular space. Extensive experiments conducted on the proposed imperfect sketch datasets demonstrate that the proposed model is able to efficiently improve the quality of sketch images and achieve superior performance over the current state-of-the-art methods. Jia Wan 0001, Kaihao Zhang, Hongdong Li, Antoni B. Chan |
IEEE Trans. Image Process. | 1 |
| 2020 | Modeling Noisy Annotations for Crowd CountingabstractThe annotation noise in crowd counting is not modeled in traditional crowd counting algorithms based on crowd density maps. In this paper, we first model the annotation noise using a random variable with Gaussian distribution, and derive the pdf of the crowd density value for each spatial location in the image. We then approximate the joint distribution of the density values (i.e., the distribution of density maps) with a full covariance multivariate Gaussian density, and derive a low-rank approximate for tractable implementation. We use our loss function to train a crowd density map estimator and achieve state-of-the-art performance on three large-scale crowd counting datasets, which confirms its effectiveness. Examination of the predictions of the trained model shows that it can correctly predict the locations of people in spite of the noisy training data, which demonstrates the robustness of our loss function to annotation noise. Jia Wan 0001, Antoni B. Chan |
NeurIPS | 1 |
| 2019 | Residual Regression With Semantic Prior for Crowd CountingabstractCrowd counting is a challenging task due to factors such as large variations in crowdedness and severe occlusions. Although recent deep learning based counting algorithms have achieved a great progress, the correlation knowledge among samples and the semantic prior have not yet been fully exploited. In this paper, a residual regression framework is proposed for crowd counting utilizing the correlation information among samples. By incorporating such information into our network, we discover that more intrinsic characteristics can be learned by the network which thus generalizes better to unseen scenarios. Besides, we show how to effectively leverage the semantic prior to improve the performance of crowd counting. We also observe that the adversarial loss can be used to improve the quality of predicted density maps, thus leading to an improvement in crowd counting. Experiments on public datasets demonstrate the effectiveness and generalization ability of the proposed method. Jia Wan 0001, Wenhan Luo, Baoyuan Wu, Antoni B. Chan, Wei Liu 0005 |
CVPR | 1 |
| 2019 | Adaptive Density Map Generation for Crowd CountingabstractThe following topics are dealt with: learning (artificial intelligence); feature extraction; neural nets; image representation; object detection; convolutional neural nets; image segmentation; image classification; computer vision; video signal processing. Jia Wan 0001, Antoni B. Chan |
ICCV | 1 |
| 2019 | Hierarchical Feature Selection for Random ProjectionabstractRandom projection is a popular machine learning algorithm, which can be implemented by neural networks and trained in a very efficient manner. However, the number of features should be large enough when applied to a rather large-scale data set, which results in slow speed in testing procedure and more storage space under some circumstances. Furthermore, some of the features are redundant and even noisy since they are randomly generated, so the performance may be affected by these features. To remedy these problems, an effective feature selection method is introduced to select useful features hierarchically. Specifically, a novel criterion is proposed to select useful neurons for neural networks, which establishes a new way for network architecture design. The testing time and accuracy of the proposed method are improved compared with traditional methods and some variations on both classification and regression tasks. Extensive experiments confirm the effectiveness of the proposed method. Qi Wang 0009, Jia Wan 0001, Feiping Nie 0001, Bo Liu 0006, Chenggang Yan 0001, Xuelong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2018 | Locality constraint distance metric learning for traffic congestion detection
Qi Wang 0009, Jia Wan 0001, Yuan Yuan 0001 |
Pattern Recognit. | 2 |
| 2018 | Deep Metric Learning for Crowdedness RegressionabstractCross-scene regression tasks, such as congestion level detection and crowd counting, are useful but challenging. There are two main problems, which limit the performance of existing algorithms. The first one is that no appropriate congestion-related feature can reflect the real density in scenes. Though deep learning has been proved to be capable of extracting high level semantic representations, it is hard to converge on regression tasks, since the label is too weak to guide the learning of parameters in practice. Thus, many approaches utilize additional information, such as a density map, to guide the learning, which increases the effort of labeling. Another problem is that most existing methods are composed of several steps, for example, feature extraction and regression. Since the steps in the pipeline are separated, these methods face the problem of complex optimization. To remedy it, a deep metric learning-based regression method is proposed to extract density related features, and learn better distance measurement simultaneously. The proposed networks trained end-to-end for better optimization can be used for crowdedness regression tasks, including congestion level detection and crowd counting. Extensive experiments confirm the effectiveness of the proposed method. Qi Wang 0009, Jia Wan 0001, Yuan Yuan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | Traffic congestion analysis: A new PerspectiveabstractIn this paper, a new perspective of congestion is presented to promote the development of traffic video analysis. Our main contributions are threefold: a) An unified and quantifiable definition of congestion is proposed to describe the traffic state in video. b) Based on the definition, a congestion dataset which contains multiple traffic scenes is constructed as a platform for the research community. At the same time, a precise labeling method is introduced to get the ground truth of congestion level accurately. c) An algorithm based on Inverse Perspective Mapping (IPM) and pairwise regression is proposed to analyze traffic videos and serves as a baseline. We further compare the proposed method with two deep learning methods. Intensive experiments justify the effectiveness of the proposed method. Jia Wan 0001, Yuan Yuan 0001, Qi Wang 0009 |
ICASSP | 1 |
| 2016 | Congested scene classification via efficient unsupervised feature learning and density estimation
Yuan Yuan 0001, Jia Wan 0001, Qi Wang 0009 |
Pattern Recognit. | 2 |