EDBT 2026 Demo / reviewers in the wild / expert
Chao Zhang 0030
dblp:94/3019-30
· DBLP profile ↗
33ranked-venue papers
2as first author
27since 2021 · last 2026
0000-0002-0845-9217ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 1 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Community-level competitive influence payoff maximization
Jun Yu 0012, Chunzhi Gu, Takuya Akashi, Chao Zhang 0030 |
Knowl. Inf. Syst. | 5 |
| 2026 | Few-shot human action anomaly detection via a unified contrastive learning frameworkabstractHuman Action Anomaly Detection (HAAD) aims to identify anomalous actions using only normal-action data during training. Most prior work follows a one-model-per-category paradigm, requiring separate training for each action category and numerous normal samples. These requirements hinder scalability and limit applicability in real-world settings, where data are often scarce or novel categories frequently appear. To address these limitations, we propose a unified HAAD framework that supports few-shot settings. Our method learns a category-agnostic representation via contrastive learning and detects anomalies by comparing a test sample with a small support set of normal examples. To improve inter-category generalization and intra-category robustness, we introduce diffusion-based generative motion augmentation to synthesize diverse, realistic training samples. To the best of our knowledge, this is the first work to tailor diffusion-driven motion augmentation specifically to contrastive learning for action anomaly detection. Our few-shot design is particularly useful for monitoring applications, where normality often varies across individuals and contexts. Experiments on HumanAct12 demonstrate state-of-the-art performance on both seen and unseen categories, while improving training efficiency and scalability for few-shot HAAD. Koichiro Kamide, Shunsuke Sakai, Shun Maeda, Chunzhi Gu, Chao Zhang 0030 |
Knowl. Based Syst. | 5 |
| 2026 | Consensus-aware sparse subspace clustering via coreset selectionabstractSubspace clustering aims to recover low-dimensional subspaces from high-dimensional data and assign each data point to its corresponding subspace. State-of-the-art subspace clustering methods rely on self-expressive models to represent each data point as a linear combination of other data points. However, these methods typically suffer from scalability challenges when dealing with large-scale datasets. In particular, the limited capacity to capture similarities among data points within the same subspace leads to poor connectivity of the affinity matrix, often resulting in over-segmentation. In this paper, we propose a scalable self-expressive model based on consensus learning with selectively sampled subsets to address the connectivity issue. Our core insight is that an ideal alignment between global and local solutions across multiple small subsets plays a key role in promoting dense connections in the affinity matrix for large-scale datasets. To this end, our model is designed to flexibly fuse sparse local solutions obtained from small subsets in a consensus-aware manner to derive a global solution that captures richer pairwise relationships within each subspace. Moreover, we introduce a selective subsampling strategy to generate subsets that effectively approximate the spatial support of the original dataset. This contributes to the reliability of solved local solutions, which further enables robust clustering even in imbalanced datasets. Extensive experiments on synthetic and five real-world datasets show that our method achieves state-of-the-art performance, in terms of clustering accuracy and connectivity. Katsuya Hotta, Chunzhi Gu, Chao Zhang 0030 |
Pattern Recognit. | 3 |
| 2026 | Benchmarking real-world medical image classification with noisy labels: Challenges, practice, and outlookabstractLearning from noisy labels remains a major challenge in medical image analysis, where annotation demands expert knowledge and substantial inter-observer variability often leads to inconsistent or erroneous labels. Despite extensive research on learning with noisy labels (LNL), the robustness of existing methods in medical imaging has not been systematically assessed. To address this gap, we introduce LNMBench, a comprehensive benchmark for Label Noise in Medical imaging. LNMBench encompasses \textbf{10} representative methods evaluated across 7 datasets, 6 imaging modalities, and 3 noise patterns, establishing a unified and reproducible framework for robustness evaluation under realistic conditions. Comprehensive experiments reveal that the performance of existing LNL methods degrades substantially under high and real-world noise, highlighting the persistent challenges of class imbalance and domain variability in medical data. Motivated by these findings, we further propose a simple yet effective improvement to enhance model robustness under such conditions. The LNMBench codebase is publicly released to facilitate standardized evaluation, promote reproducible research, and provide practical insights for developing noise-resilient algorithms in both research and real-world medical applications.The codebase is publicly available on https://github.com/myyy777/LNMBench. Junlin Hou, Chao Zhang 0030, ZongYuan Ge, Haoran Xie 0002, Lie Ju |
Pattern Recognit. | 3 |
| 2026 | Frequency-guided multi-level human action anomaly detection with normalizing flowsabstractWe introduce the task of human action anomaly detection (HAAD), which aims to identify anomalous motions in an unsupervised manner given only the pre-determined normal category of training action samples. Compared to prior human-related anomaly detection tasks which primarily focus on unusual events from videos, HAAD involves the learning of specific action labels to recognize semantically anomalous human behaviors. To address this task, we propose a normalizing flow (NF)-based detection framework where the sample likelihood is effectively leveraged to indicate anomalies. As action anomalies often occur in some specific body parts, in addition to the full-body action feature learning, we incorporate extra encoding streams into our framework for finer modeling of body subsets. Our framework is thus multi-level to jointly discover global and local motion anomalies. Furthermore, to show awareness of the potentially jittery data during recording, we resort to discrete cosine transformation by converting the action samples from the temporal to the frequency domain to mitigate the issue of data instability. Extensive experimental results on two human action datasets demonstrate that our method outperforms the baselines formed by adapting state-of-the-art human activity AD approaches to our task of HAAD. Shun Maeda, Chunzhi Gu, Jun Yu 0012, Shogo Tokai, Shangce Gao, Chao Zhang 0030 |
Pattern Recognit. | 6 |
| 2025 | Leveraging Inter-Generational Knowledge Transfer in Large-Scale Global OptimizationabstractLarge-scale global optimization (LSGO) presents significant challenges due to the high dimensionality and complexity of the search space. We propose an IGKT (Inter-Generational Knowledge Transfer) optimization method incorporating a novel knowledge transfer mechanism to address these challenges. The proposed mechanism enables the algorithm to reduce reliance on stochastic exploration and enhance convergence efficiency by transferring information from the best-performing individuals across generations, guiding the population toward promising regions in the search space. Experimental results on the CEC2013 LSGO benchmark suite demonstrate that IGKT outperforms several state-of-the-art algorithms across various tested functions, achieving superior convergence speed and solution quality. Additionally, the IGKT framework handles both separable and non-separable functions and tasks, including those with overlapping and highly coupled variables. In summary, IGKT represents a powerful tool for addressing complex, high-dimensional optimization problems, providing a robust and adaptable solution for LSGO. Yuefeng Xu, Rui Zhong 0004, Chong Zhou, Chao Zhang 0030, Jun Yu 0012 |
CEC | 4 |
| 2025 | Dataset Distillation Via Vision-Language Category PrototypeabstractDataset distillation (DD) condenses large datasets into compact yet informative substitutes, preserving performance comparable to the original dataset while reducing storage, transmission costs, and computational consumption. However, previous DD methods mainly focus on distilling information from images, often overlooking the semantic information inherent in the data. The disregard for context hinders the model's generalization ability, particularly in tasks involving complex datasets, which may result in illogical outputs or the omission of critical objects. In this study, we integrate vision-language methods into DD by introducing text prototypes to distill language information and collaboratively synthesize data with image prototypes, thereby enhancing dataset distillation performance. Notably, the text prototypes utilized in this study are derived from descriptive text information generated by an open-source large language model. This framework demonstrates broad applicability across datasets without pre-existing text descriptions, expanding the potential of dataset distillation beyond traditional image-based approaches. Compared to other methods, the proposed approach generates logically coherent images containing target objects, achieving state-of-the-art validation performance and demonstrating robust generalization. Source code and generated data are available in https://github.com/zou-yawen/Dataset-Distillation-via-Vision-Language-Category-Prototype/ Yawen Zou, Guang Li 0008, Duo Su, Jun Yu 0012, Chao Zhang 0030 |
ICCV | 6 |
| 2025 | Subspace-Guided Feature Reconstruction for Unsupervised Anomaly LocalizationabstractABSTRACT Unsupervised anomaly localization aims to identify anomalous regions that deviate from normal sample patterns. Most recent methods perform feature matching or reconstruction for the target sample with pre‐trained deep neural networks. However, they still struggle to address challenging anomalies because the deep embeddings stored in the memory bank can be less powerful and informative. Specifically, prior methods often overly rely on the finite resources stored in the memory bank, which leads to low robustness to unseen targets. In this paper, we propose a novel subspace‐guided feature reconstruction framework to pursue adaptive feature approximation for anomaly localization. It first learns to construct low‐dimensional subspaces from the given nominal samples, and then learns to reconstruct the given deep target embedding by linearly combining the subspace basis vectors using the self‐expressive model. Our core is that, despite the limited resources in the memory bank, the out‐of‐bank features can be alternatively “mimicked” to adaptively model the target. Moreover, we propose a sampling method that leverages the sparsity of subspaces and allows the feature reconstruction to depend only on a small resource subset, contributing to less memory overhead. Extensive experiments on three benchmark datasets demonstrate that our approach generally achieves state‐of‐the‐art anomaly localization performance. Katsuya Hotta, Chao Zhang 0030, Yoshihiro Hagihara, Takuya Akashi |
IET Image Process. | 2 |
| 2025 | Incremental pseudo-labeling for black-box unsupervised domain adaptationabstractBlack-Box unsupervised domain adaptation (BBUDA) learns knowledge only with the prediction of target data from the source model without access to the source data and source model, which attempts to alleviate concerns about the privacy and security of data. However, incorrect pseudo-labels are prevalent in the prediction generated by the source model due to the cross-domain discrepancy, which may substantially degrade the performance of the target model. To address this problem, we propose a novel approach that incrementally selects high-confidence pseudo-labels to improve the generalization ability of the target model. Specifically, we first generate pseudo-labels using a source model and train a crude target model by a vanilla BBUDA method. Second, we iteratively select high-confidence data from the low-confidence data pool by thresholding the softmax probabilities, prototype labels, and intra-class similarity. Then, we iteratively train a stronger target network based on the crude target model to correct the wrongly labeled samples to improve the accuracy of the pseudo-label. Experimental results demonstrate that the proposed method achieves state-of-the-art black-box unsupervised domain adaptation performance on three benchmark datasets. Yawen Zou, Chunzhi Gu, Jun Yu 0012, Shangce Gao, Chao Zhang 0030 |
J. Vis. Commun. Image Represent. | 5 |
| 2025 | Crested ibis algorithm and its application in human-powered aircraft design
Yuefeng Xu, Rui Zhong 0004, Chao Zhang 0030, Jun Yu 0012 |
Knowl. Based Syst. | 3 |
| 2025 | Learning to Discriminate While Contrasting: Combating False Negative Pairs With Coupled Contrastive Learning for Incomplete Multi-View ClusteringabstractThe task of incomplete multi-view clustering (IMvC) aims to partition multi-view data with a lack of completeness into different clusters. The incompleteness can be typically categorized into the case of instance-missing and view-unaligned MvC. However, prior methods either consider each of them or struggle to pursue consistent latent representations among views. In this paper, we propose two forms of contrastive learning paradigms to jointly handle both cases for IMvC. Specifically, we design an instance-oriented contrastive (IOC) learning strategy to achieve intra-class consistency. As negative samples within different datasets can exhibit diverse distributions, we formulate a parameterized boundary for IOC learning to flexibly deal with such differing data modes. To preserve inter-view consistency, we further devise category-oriented contrastive (COC) learning such that data from different views can be seamlessly integrated into a combined semantic space. We also recover the missing instances with the learned latent representations in a reconstructing manner for realigning the incomplete multi-view data to facilitate clustering. Our approach unifies the solution to both incomplete cases into one formulation. To demonstrate the effectiveness of our model, we conduct four types of MvC tasks on six benchmark multi-view datasets and compare our method against state-of the-art IMvC methods. Extensive experiments show that our method achieves state-of-the-art performance, quantitatively and qualitatively. Katsuya Hotta, Chunzhi Gu, Ao Li 0002, Jun Yu 0012, Chao Zhang 0030 |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2024 | Handling Class Imbalance in Black-Box Unsupervised Domain Adaptation with Synthetic Minority Over-SamplingabstractBlack-box unsupervised domain adaptation (BBUDA) is a challenging task that transfers knowledge from the source domain to the target domain without access to the source data and source model, thus alleviating public concerns about data security. However, BBUDA requires the source model to function as a black-box predictor for the target data, and the pseudo-labels often exhibit class imbalance, which degrades the performance. To tackle this problem, we propose employing the synthetic minority oversampling technique (SMOTE) and adaptive sampling to rebalance data. Given that predictions often contain errors, we first select reliable high-confidence data before using SMOTE to generate synthetic samples for the minority class. Second, we incrementally select high-confidence data from the remaining low-confidence data with an adaptive sampling rate for each class, in which the minority class (with the fewest samples) is assigned a higher sampling rate and the majority class (with the most samples) is assigned a lower sampling rate. The experimental results demonstrate that our method can mitigate the class imbalance and further improve the performance of the target model. Yawen Zou, Chunzhi Gu, Guang Li 0008, Jun Yu 0012, Chao Zhang 0030 |
VCIP | 6 |
| 2024 | Cooperative coati optimization algorithm with transfer functions for feature selection and knapsack problems
Rui Zhong 0004, Chao Zhang 0030, Jun Yu 0012 |
Knowl. Inf. Syst. | 2 |
| 2024 | A multi-in and multi-out dendritic neuron model and its optimization
Jun Yu 0012, Chunzhi Gu, Shangce Gao, Chao Zhang 0030 |
Knowl. Based Syst. | 5 |
| 2024 | SRIME: a strengthened RIME with Latin hypercube sampling and embedded distance-based selection for engineering optimization problems
Rui Zhong 0004, Jun Yu 0012, Chao Zhang 0030, Masaharu Munetomo |
Neural Comput. Appl. | 3 |
| 2024 | Learning disentangled representations for controllable human motion prediction
Chunzhi Gu, Jun Yu 0012, Chao Zhang 0030 |
Pattern Recognit. | 3 |
| 2024 | Orientation-aware leg movement learning for action-driven human motion prediction
Chunzhi Gu, Chao Zhang 0030, Shigeru Kuriyama |
Pattern Recognit. | 2 |
| 2023 | Black-Box Targeted Adversarial Attack Based on Multi-Population Genetic AlgorithmabstractThe fast gradient signed method (FGSM) is an efficient white-box attack method that uses the gradient information to generate adversarial examples. However, applying the classic FGSM to real-world applications is often difficult due to the challenge of obtaining the internal structure of the models. Therefore, we have made slight modifications to the conventional genetic algorithm (GA) to effectively optimize the gradient signed function of the classic FGSM and generate adversarial examples from the perspective of the black-box attack. To attack multiple given target classes simultaneously, we initialize multiple different subpopulations and ensure that each subpopulation attacks a specified target class. Additionally, we propose two different strategies to migrate successfully attacked subpopulations into unsuccessful ones to ramp up attacks on unsuccessful classes. To evaluate the performance of the proposed algorithm, we compare it with the conventional GA when attacking the well-trained VGG19_BN model on the CIFAR-10 database. Furthermore, we investigate the impact of the proposed strategies on performance and analyze their respective contributions. The experimental results confirm that the proposed algorithm can successfully attack a greater variety of classes at a faster rate. Yuuto Aiza, Chao Zhang 0030, Jun Yu 0012 |
SMC | 2 |
| 2023 | PMSSC: Parallelizable multi-subset based self-expressive model for subspace clusteringabstractSubspace clustering methods which embrace a self-expressive model that represents each data point as a linear combination of other data points in the dataset provide powerful unsupervised learning techniques. However, when dealing with large datasets, representation of each data point by referring to all data points via a dictionary suffers from high computational complexity. To alleviate this issue, we introduce a parallelizable multi-subset based self-expressive model (PMS) which represents each data point by combining multiple subsets, with each consisting of only a small proportion of the samples. The adoption of PMS in subspace clustering (PMSSC) leads to computational advantages because the optimization problems decomposed over each subset are small, and can be solved efficiently in parallel. Furthermore, PMSSC is able to combine multiple self-expressive coefficient vectors obtained from subsets, which contributes to an improvement in self-expressiveness. Extensive experiments on synthetic and real-world datasets show the efficiency and effectiveness of our approach in comparison to other methods. Katsuya Hotta, Takuya Akashi, Shogo Tokai, Chao Zhang 0030 |
Comput. Vis. Media | 4 |
| 2023 | Teacher-student network for 3D point cloud anomaly detection with few normal samples
Jianjian Qin, Chunzhi Gu, Jun Yu 0012, Chao Zhang 0030 |
Expert Syst. Appl. | 4 |
| 2022 | Accelerating Fireworks Algorithm with Adaptive Scouting StrategyabstractWe propose an adaptive scouting strategy which is a refinement of our previous work to further improve the performance of the fireworks algorithm (FWA). The proposed strategy makes full use of the currently obtained fitness landscape information to avoid inefficient searches. Specifically, we introduce two new modifications to our previously proposed scouting strategy to more quickly adjust the balance between exploration and exploitation in the face of various optimization scenarios. The first modification is that the initial explosion center migrates with better generated spark individual instead of being fixed on initial firework individual, i.e., the next round of the initial explosion center will move to the recently generated spark individual when the current tracing direction has no potential. The other is to actively reduce the explosion amplitude of subsequent explosion operation when a potential spark individual is generated. Otherwise, increase the explosion amplitude to escape from the trapped local area quickly. To evaluate the performance of the new proposed strategy, we designed a series of comparative experiments and used 28 functions from the CEC 2013 test suite as the benchmark. The experimental results confirmed that the adaptive scouting strategy shows better performance and faster convergence speed especially for complex multimodal optimization problems. Jun Yu 0012, Chao Zhang 0030 |
SMC | 2 |
| 2022 | Component-based nearest neighbour subspace clusteringabstractAbstract In this paper, the problem of clustering data points that lie near or on a union of independent low‐dimensional subspaces is addressed. To this end, the popular spectral clustering‐based algorithms usually follow a two‐stage strategy that initially builds an affinity matrix and then applies spectral clustering. However, an inappropriate affinity matrix that does not sufficiently connect data points lying on the same subspace will easily lead to the issue of over‐segmentation. To alleviate this issue, building the affinity matrix based on subspace hypotheses generated by an iterative sampling operation according to the Random Cluster Model under the framework of energy minimisation is proposed. Specifically, each hypothesis is generated from a large number of data points by sampling a component in a K ‐nearest neighbour graph. Extensive experiments on synthetic data and real‐world datasets show that the proposed method can improve the connectivity of the affinity matrix and provide competitive results against state‐of‐the‐art methods. Katsuya Hotta, Haoran Xie 0002, Chao Zhang 0030 |
IET Image Process. | 3 |
| 2022 | DS-SRI: Diversity similarity measure against scaling, rotation, and illumination change for robust template matchingabstractAbstract This paper presents a novel multi‐scale template matching method that can be applied in unconstrained environments. The key component behind this is a general similarity measure is referred to as the diversity similarity measure against scaling, rotation, and illumination (DS‐SRI). Specifically, DS‐SRI exploits bidirectional diversity calculated from the nearest neighbour matches between two sets of points. Scaling and rotation changes are taken into consideration by introducing normalisation term on the scale change, and geometric consistency term with respect to the polar coordinate system. Moreover, in order to deal with the illumination change and further deformation, illumination‐corrected local appearance and rank information are jointly exploited during the nearest neighbour search. All the features of DS‐SRI are statistically assessed, and the extensive visual and quantitative results on both synthetic and real‐world data show that DS‐SRI can significantly outperform state‐of‐the‐art methods. Yi Zhang 0083, Chao Zhang 0030, Takuya Akashi |
IET Image Process. | 2 |
| 2022 | Learning to predict diverse human motions from a single image via mixture density networks
Chunzhi Gu, Yan Zhao 0038, Chao Zhang 0030 |
Knowl. Based Syst. | 3 |
| 2022 | Example-based color transfer with Gaussian mixture modeling
Chunzhi Gu, Xuequan Lu, Chao Zhang 0030 |
Pattern Recognit. | 3 |
| 2021 | DeepfakeUCL: Deepfake Detection via Unsupervised Contrastive LearningabstractFace deepfake detection has seen impressive results recently. Nearly all existing deep learning techniques for face deepfake detection are fully supervised and require labels during training. In this paper, we design a novel deepfake detection method via unsupervised contrastive learning. We first generate two different transformed versions of an image and feed them into two sequential sub-networks, i.e., an encoder and a projection head. The unsupervised training is achieved by maximizing the correspondence degree of the outputs of the projection head. To evaluate the detection performance of our unsupervised method, we further use the unsupervised features to train an efficient linear classification network. Extensive experiments show that our unsupervised learning method enables comparable detection performance to state-of-the-art supervised techniques, in both the intra- and inter-dataset settings. We also conduct ablation studies for our method. Sheldon Fung, Xuequan Lu, Chao Zhang 0030, Chang-Tsun Li |
IJCNN | 3 |
| 2021 | Blur Removal Via Blurred-Noisy Image PairabstractComplex blur such as the mixup of space-variant and space-invariant blur, which is hard to model mathematically, widely exists in real images. In this article, we propose a novel image deblurring method that does not need to estimate blur kernels. We utilize a pair of images that can be easily acquired in low-light situations: (1) a blurred image taken with low shutter speed and low ISO noise; and (2) a noisy image captured with high shutter speed and high ISO noise. Slicing the blurred image into patches, we extend the Gaussian mixture model (GMM) to model the underlying intensity distribution of each patch using the corresponding patches in the noisy image. We compute patch correspondences by analyzing the optical flow between the two images. The Expectation Maximization (EM) algorithm is utilized to estimate the parameters of GMM. To preserve sharp features, we add an additional bilateral term to the objective function in the M-step. We eventually add a detail layer to the deblurred image for refinement. Extensive experiments on both synthetic and real-world data demonstrate that our method outperforms state-of-the-art techniques, in terms of robustness, visual quality, and quantitative metrics. Chunzhi Gu, Xuequan Lu, Ying He 0001, Chao Zhang 0030 |
IEEE Trans. Image Process. | 4 |
| 2020 | Deep Patch-Based Human Segmentation
Dongbo Zhang 0004, Zheng Fang 0008, Xuequan Lu, Hong Qin 0001, Antonio Robles-Kelly, Chao Zhang 0030, Ying He 0001 |
ICONIP (1) | 6 |
| 2020 | G2MF-WA: Geometric multi-model fitting with weakly annotated dataabstractIn this paper we address the problem of geometric multi-model fitting using a few weakly annotated data points, which has been little studied so far. In weak annotating (WA), most manual annotations are supposed to be correct yet inevitably mixed with incorrect ones. Such WA data can naturally arise through interaction in various tasks. For example, in the case of homography estimation, one can easily annotate points on the same plane or object with a single label by observing the image. Motivated by this, we propose a novel method to make full use of WA data to boost multi-model fitting performance. Specifically, a graph for model proposal sampling is first constructed using the WA data, given the prior that WA data annotated with the same weak label has a high probability of belonging to the same model. By incorporating this prior knowledge into the calculation of edge probabilities, vertices (i.e., data points) lying on or near the latent model are likely to be associated and further form a subset or cluster for effective proposal generation. Having generated proposals, o-expansion is used for labeling, and our method in return updates the proposals. This procedure works in an iterative way. Extensive experiments validate our method and show that it produces noticeably better results than state-of-the-art techniques in most cases. Chao Zhang 0030, Xuequan Lu, Katsuya Hotta, Xi Yang 0017 |
Comput. Vis. Media | 1 |
| 2019 | Multi-scale Template Matching with Scalable Diversity Similarity in an Unconstrained Environment
Yi Zhang 0083, Chao Zhang 0030, Takuya Akashi |
BMVC | 2 |
| 2019 | Vehicle Rear-Lamp Detection at Nighttime via Probabilistic Bitwise Genetic AlgorithmabstractRear-lamp detection of a vehicle at nighttime is an important technique for advanced driver-assistance systems. We present a detection method by employing a variant of genetic algorithm, which utilizes bitwise genetic operation instead of classic crossover and mutation. That is, the detection task is cast to a localization problem under an evolutionary optimization framework. Specifically, geometric parameters of a rectangle pair form a model to represent the detected rear-lamp pair. The fitness function for evaluating each candidate solution is combinatorial, which consists of multiple fitness functions designed under handcrafted rules from the observation. In addition, the solution space is narrowed down by extracting the red-light sources, which yields in more efficient solution exploration. Experiment with a publicly available dataset which involves images captured in various traffic situations shows the effectiveness of our method qualitatively and quantitatively. Takumi Nakane, Tatsuya Takeshita, Shogo Tokai, Chao Zhang 0030 |
CW | 4 |
| 2019 | Bird Species Classification with Audio-Visual Data using CNN and Multiple Kernel LearningabstractRecently, deep convolutional neural networks (CNN) have become a new standard in many machine learning applications not only in image but also in audio processing. However, most of the studies only explore a single type of training data. In this paper, we present a study on classifying bird species by combining deep neural features of both visual and audio data using kernel-based fusion method. Specifically, we extract deep neural features based on the activation values of an inner layer of CNN. We combine these features by multiple kernel learning (MKL) to perform the final classification. In the experiment, we train and evaluate our method on a CUB-200-2011 standard data set combined with our originally collected audio data set with respect to 200 bird species (classes). The experimental results indicate that our CNN+MKL method which utilizes the combination of both categories of data outperforms single-modality methods, some simple kernel combination methods, and the conventional early fusion method. Bold Naranchimeg, Chao Zhang 0030, Takuya Akashi |
CW | 2 |
| 2015 | Fast Affine Template Matching over Galois Field
Chao Zhang 0030, Takuya Akashi |
BMVC | 1 |