VLDB 2026 Research / reviewers in the wild / expert
Yu Wang 0018
dblp:02/5889-18
· DBLP profile ↗
16ranked-venue papers
3as first author
9since 2021 · last 2024
0000-0003-3008-8712ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | MT-ASM: a multi-task attention strengthening model for fine-grained object recognition
Dichao Liu, Yu Wang 0018, Kenji Mase, Jien Kato |
Multim. Syst. | 2 |
| 2023 | Spatial-temporal Concept based Explanation of 3D ConvNetsabstractConvolutional neural networks (CNNs) have shown remarkable performance on various tasks. Despite its widespread adoption, the decision procedure of the network still lacks transparency and interpretability, making it difficult to enhance the performance further. Hence, there has been considerable interest in providing explanation and interpretability for CNNs over the last few years. Explainable artificial intelligence (XAI) investigates the relationship between input images or videos and output predictions. Recent studies have achieved outstanding success in explaining 2D image classification ConvNets. On the other hand, due to the high computation cost and complexity of video data, the explanation of 3D video recognition ConvNets is relatively less studied. And none of them are able to produce a high-level explanation. In this paper, we propose a STCE (Spatial-temporal Concept-based Explanation) framework for interpreting 3D ConvNets. In our approach: (1) videos are represented with high-level supervoxels, similar supervoxels are clustered as a concept, which is straightforward for human to understand; and (2) the interpreting framework calculates a score for each concept, which reflects its significance in the ConvNet decision procedure. Experiments on diverse 3D ConvNets demonstrate that our method can identify global concepts with different importance levels, allowing us to investigate the impact of the concepts on a target task, such as action recognition, in-depth. The source codes are publicly available at https://github.com/yingji425/STCE. Ying Ji 0003, Yu Wang 0018, Jien Kato |
CVPR | 2 |
| 2023 | Long-Tailed Image Recognition with Dynamic Re-WeightingabstractFor long-tailed image recognition tasks, re-weighting is effective to alleviate data imbalance by assigning higher weights to tail categories. However, existing re-weighting methods typically adopt a static weighting scheme, which usually hurts the accuracy of head categories. To deal with this issue, this paper proposes a progress-relevant weighting scheme called dynamic re-weighting, in which the weight assigned to a particular category first increases and then decreases, proportional to the number of samples that have been used in that category. In addition, we introduce a head-to-tail loss to control the evolving of weights, which makes the model gradually transfer its attention from head categories to tail categories. We conduct experiments on long-tailed CIFAR/ImageNet datasets, and confirm that our method not only outperforms static re-weighting methods, but also improves the accuracy on tail categories without sacrificing the accuracy of head categories. Yu Wang 0018, Jien Kato |
ICASSP | 2 |
| 2023 | Using Classifier Discrepancy for Cross-Domain Image RetrievalabstractIn recent years, cross-domain image retrieval (CDIM) has garnered considerable interest. The primary difficulty of CDIM is the domain gap, which makes it hard for the system to retrieve two photos that belong to the same category but have distinct domains. In this paper, we provide a novel multi-branch network employing the quintuplet structure to minimize retrieval loss and classifier discrepancy to minimize domain loss. Using three public datasets, we test the proposed method for zero-shot sketch-based image retrieval, which is one of CDIM's application tasks. Experiments validated the proposed method's state-of-the-art performance on the majority of datasets. Longjiao Zhao, Yu Wang 0018, Jien Kato |
ICIP | 2 |
| 2023 | Learn from each other to Classify better: Cross-layer mutual attention learning for fine-grained visual classificationabstractFine-grained visual classification (FGVC) is valuable yet challenging. The difficulty of FGVC mainly lies in its intrinsic inter-class similarity, intra-class variation, and limited training data. Moreover, with the popularity of deep convolutional neural networks, researchers have mainly used deep, abstract, semantic information for FGVC, while shallow, detailed information has been neglected. This work proposes a cross-layer mutual attention learning network (CMAL-Net) to solve the above problems. Specifically, this work views the shallow to deep layers of CNNs as “experts” knowledgeable about different perspectives. We let each expert give a category prediction and an attention region indicating the found clues. Attention regions are treated as information carriers among experts, bringing three benefits: (i) helping the model focus on discriminative regions; (ii) providing more training data; (iii) allowing experts to learn from each other to improve the overall performance. CMAL-Net achieves state-of-the-art performance on three competitive datasets: FGVC-Aircraft, Stanford Cars, and Food-11. The source code is available at https://github.com/Dichao-Liu/CMAL Dichao Liu, Longjiao Zhao, Yu Wang 0018, Jien Kato |
Pattern Recognit. | 3 |
| 2023 | Toward Extremely Lightweight Distracted Driver Recognition With Distillation-Based Neural Architecture Search and Knowledge TransferabstractThe number of traffic accidents has been continuously increasing in recent years worldwide. Many accidents are caused by distracted drivers, who take their attention away from driving. Motivated by the success of Convolutional Neural Networks (CNNs) in computer vision, many researchers developed CNN-based algorithms to recognize distracted driving from a dashcam and warn the driver against unsafe behaviors. However, current models have too many parameters, which is unfeasible for vehicle-mounted computing. This work proposes a novel knowledge-distillation-based framework to solve this problem. The proposed framework first constructs a high-performance teacher network by progressively strengthening the robustness to illumination changes from shallow to deep layers of a CNN. Then, the teacher network is used to guide the architecture searching process of a student network through knowledge distillation. After that, we use the teacher network again to transfer knowledge to the student network by knowledge distillation. Experimental results on the Statefarm Distracted Driver Detection Dataset and AUC Distracted Driver Dataset show that the proposed approach is highly effective for recognizing distracted driving behaviors from photos: (i) the teacher network’s accuracy surpasses the previous best accuracy; (ii) the student network achieves very high accuracy with only 0.42M parameters (around 55% of the previous most lightweight model). Furthermore, the student network architecture can be extended to a spatial-temporal 3D CNN for recognizing distracted driving from video clips. The 3D student network largely surpasses the previous best accuracy with only 2.03M parameters on the Drive&Act Dataset. The source code is available athttps://github.com/Dichao-Liu/Lightweight_Distracted_Driver_Recognition_with_Distillation-Based_NAS_and_Knowledge_Transfer Dichao Liu, Toshihiko Yamasaki, Yu Wang 0018, Kenji Mase, Jien Kato |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | One-shot Network Pruning at Initialization with Discriminative Image Patches
Yinan Yang 0001, Yu Wang 0018, Ying Ji 0003, Heng Qi, Jien Kato |
BMVC | 2 |
| 2021 | Rotation Invariance Analysis of Local Convolutional Features in Image RetrievalabstractRecently, local features computed using convolutional neural networks (CNNs) show good performance to image retrieval. However, the local convolutional features obtained by the CNNs (LC features) are inherently sensitive to rotation perturbations. This leads to miss-judgements in retrieval tasks. In this work, our objective is to enhance the robustness of LC features against image rotation. To do this, we conduct a thorough experimental evaluation of two candidate anti-rotation strategies (in-model data augmentation, and post-model feature augmentation), over two kinds of rotation attacks (dataset attack and query attack). We end up a series of good practices with steady quantitative supports, which lead to the best strategy for computing LC features with high rotation invariance in image retrieval. Longjiao Zhao, Yu Wang 0018, Jien Kato |
ICASSP | 2 |
| 2021 | Attention-Based Multi-Task Learning For Fine-Grained Image ClassificationabstractFine-Grained Image Classification is an inherently challenging task because of its inter-class similarity and intra-class variance. Most existing studies solve this problem by localization-and-classification strategies, which, however, always causes the problem of information loss or heavy computational expenses. Instead of localization-and-classification strategy, we propose a novel end-to-end optimization procedure named Multi-Task Attention Learning (MTAL), which reinforces the neural network’ correspondence to attention regions. Experimental results on CUB-Birds and Stanford Cars show that our procedure distinctly outperforms the baselines and is comparable with state-of-the-art studies despite its simplicity*. Dichao Liu, Yu Wang 0018, Kenji Mase, Jien Kato |
ICIP | 2 |
| 2020 | Contrastively-reinforced Attention Convolutional Neural Network for Fine-grained Image Recognition
Dichao Liu, Yu Wang 0018, Jien Kato, Kenji Mase |
BMVC | 2 |
| 2019 | Visual Violence Rating with Pairwise ComparisonabstractChildren's exposure to violence has become a severe problem with the rapid development of Internet. Recognizing violent video and estimating violence extent become crucial. Most researches focus on violent scene or violent action detection, lacking overall violence extent information. In this paper, we propose a violence rating prediction approach and build a novel violent video dataset. Our proposed method has two advantages: (1) videos are represented by features extracted from a learned two-stream network; (2) relationship between different violence extent can be learned and utilized to predict violence rating. To demonstrate the effectiveness of our method, we created a well-labelled dataset which contains 1, 930 violent videos. Each video is labelled with 6 objective violent attributes. Furthermore, we employ pairwise comparison method to obtain ground-truth violence rating for each video. Our proposed approach was evaluated on our dataset. Its results showed that our proposed method outperforms the state-of-art video classification methods. Ying Ji 0003, Yu Wang 0018, Jien Kato |
ICIP | 2 |
| 2019 | Good Choices for Deep Convolutional Feature EncodingabstractDeep convolutional neural networks can be used to produce discriminative image level features. However, when they are used as the feature extractor in a feature encoding pipeline, there are many design choices that are need to be made. In this work, we conduct a comprehensive study on deep convolutional feature encoding, by paying a special attention on its feature extraction aspect. We mainly evaluated the choices of the encoding methods; the choices of the base DCNN models; and the choices of the data augmentation methods. We not only quantitatively confirmed some known and previously unknown good choices for deep convolutional feature encoding, but also found out that some known good choices tune out to be bad. Base on the observations in the experiments, we present a very simple deep feature encoding pipeline, and confirmed its state-of-the-art performances on multiple image recognition datasets. Yu Wang 0018, Jien Kato |
WACV | 1 |
| 2017 | Solving Occlusion Problem in Pedestrian Detection by Constructing Discriminative Part LayersabstractOcclusion handling is one of the most challenging issues for pedestrian detection, and no satisfactory achievement has been found in this issue yet. Using human body parts has been considered as a reasonable way to overcome such an issue. In this paper, we propose a brand new approach based on the fusion of Mid-level body part mining and Convolutional Neural Network (CNN) to solve this problem, named DP-CNN(Discriminative Parts CNN). Two main discussions are included in this paper. First, we take an exhaustive analysis on how to mine useful body parts that contribute to pedestrian detection. Multiple ingredients (e.g. feature representation, pedestrian attributes) are analyzed through a wide range of experiments. Second, we convert the part detectors to the middle layer of CNN and re-train the model to get a better adaption of the dataset. Compare to existing approaches based on fine-tuning CNN models, our method is not only robust to occlusion handling but also has a smaller computational cost. Yu Wang 0018, Jien Kato, Guanwen Zhang, Kenji Mase |
WACV | 2 |
| 2015 | Action recognition with approximate sparse codingabstractIn this paper, we present a novel feature encoding approach called Approximate Sparse Coding (ASC). ASC computes the sparse codes for a large collection of prototype descriptors in the off-line learning phase with Sparse Coding (SC); and look up the nearest prototype's sparse code for each to-be-encoded descriptor in the encoding phase with Approximate Nearest Neighbour (ANN) search. It shares the low dimensionality of SC and the fast speed of ANN, which are both desired properties for the human action recognition task. We excessively evaluated ASC on the popular HMDB51 dataset, and confirme it is able to encode large number of video features into discriminative low dimensional representations efficiently. Yu Wang 0018, Jien Kato |
ICIP | 1 |
| 2012 | Local Distance Comparison for Multiple-shot People Re-identification
Guanwen Zhang, Yu Wang 0018, Jien Kato, Takafumi Marutani, Kenji Mase |
ACCV (3) | 2 |
| 2012 | A distance metric learning based summarization system for nursery school surveillance videoabstractIn this paper, we present a system for summarizing nursery school surveillance video. The system takes full use of a learned distance metric, which can properly measure the similarity between videos. The metric is combined with supervised classification and unsupervised clustering, to categorize raw video materials into individual events. By selecting representative videos for each event, the system produces short video digests as the summarization output. The digests cover and reflect the children's activities on a daily basis. They are not only of interest to the parents, but also provide easy access to the mass quantity of daily surveillance video data. We implemented the proposed system in a real nursery school environment and confirmed its performance through both quantitative experiment and questionnaire survey. Yu Wang 0018, Jien Kato |
ICIP | 1 |