Yongqing Sun

dblp:55/5195 · DBLP profile ↗
← Back
25ranked-venue papers
9as first author
15since 2021 · last 2025
0000-0003-3116-2371ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 20 · 9 first-author · 10 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 DLP-YOLOv9: Model with Fewer Parameters and Higher Precision Based on Improved YOLOv9 in Drone-Captured-Scenarios
abstract
With the rapid development of unmanned aerial vehicles (UAVs) technology and the increasingly extensive application fields, the accurate detection function of the objects in UAV shooting has gradually become the core issue of the industry’s frontier exploration and public attention. This paper introduces a novel object detection algorithm, designated as DLP-YOLOv9, which not only significantly improves the detection accuracy in Drone-Captured-Scenarios, but also greatly reduces the number of parameters, and realizes the dual optimization of performance and efficiency. Based on YOLOv9, we flexibly adjust the sampling position of the irregular convolution kernel to improve the accuracy of feature extraction. Meanwhile, we cleverly apply overlapping dilated convolution to propose a detection head structure DRRN, which effectively reduces parameters in YOLOv9. Finally, considering the ability to focus and process regression samples of different difficulty levels, we introduce the Focaler-CIoU method, which further improves the accuracy and efficiency of object detection. Through extensive experiments, our proposed DLP-YOLOv9 model shows excellent performance on the VisDrone2019 dataset.
Yongqing Sun
ICIP3
2025 Deep Dual Internal Learning for Hyperspectral Image Super-Resolution
Yongqing Sun, Hong Liu 0009, Qiong Chang, Xianhua Han
MMM (1)1
2024 Deep Counterfactual Representation Learning for Visual Recognition Against Weather Corruptions
abstract
Deep learning has been widely studied for processing and understanding multimedia data, and it does help improve performance. Recent research has shown that deep models are vulnerable to images containing adverse weather corruptions, leading to a safety risk for numerous safety-critical systems (e.g., autonomous driving systems). There are two problems with the current situation. First, collecting data under different weather scenarios is highly difficult in practice. Second, the performance degrades significantly when the training and test data are from different distributions, as exemplified by the weather corrupted test data. As a result, it is challenging to train a model without access to the images containing variations of various weather conditions, and it is difficult to make trained model generalized to unknown data under different weather conditions. In this paper, we introduce aCounterfactual Representation Learning(CRL) method to address these problems. Without access to training data including weather condition variations, our CRL makes the model resistant to unseen test data that has been corrupted by weather condition variations. Our basic idea is inspired by the perspective of counterfactual regularization. We build a causal model that introduces a counterfactual variable to eliminate the unobserved characteristics brought about by weather conditions. In particular, such a counterfactual variable is approximated by randomly shuffled features, echoing the previous empirical observation that the shuffling technique can perturb the shape details while preserving the local textures. We use information theoretic representation learning to encourage the neural networks to learn more powerful and robust features, which consist of two components. We conduct experiments on five benchmark datasets, namely, CIFAR-100-C, ImageNet-C, KITTI-C, BDD100 k, and CityScapes-C, all of which contain weather corruption. The results of our experiments show that our proposed method can not only be a plug-and-play technique but also work nicely for both object recognition and detection.
Hong Liu 0009, Yongqing Sun, Yukihiro Bandoh, Masaki Kitahara, Shin'ichi Satoh 0001
IEEE Trans. Multim.2
2024 TinyStereo: A Tiny Coarse-to-Fine Framework for Vision-Based Depth Estimation on Embedded GPUs
abstract
Stereo vision, a popular depth estimation technology in computing vision, finds wide-ranging applications in embedded systems, including robotics vision and autonomous driving. These applications demand both high accuracy and fast processing speeds. To address hardware limitations, most current embedded systems rely on nonlearning algorithms for fast matching, sacrificing accuracy. Some recent studies have explored using convolutional neural networks (CNNs) to improve matching accuracy, but the computational load of existing learning-based systems hampers real-world applicability. This article presents significant contributions: 1) a novel stereo matching framework that greatly enhances accuracy on real-time embedded platforms and 2) a two-pronged approach combining a nonlearning-based algorithm and a lightweight super-resolution residual neural network (sRRNet). The nonlearning-based algorithm yields a low-resolution disparity map, while the lightweight sRRNet generates a high-resolution disparity map. Experimental results on benchmark data demonstrate that the proposed method achieves a low matching error rate of 5.17% and a real-time processing speed of 51 fps using the embedded Jetson AGX GPU. The proposed method outperforms all existing real-time embedded systems.
Qiong Chang, Aolong Zha, Meng Joo Er, Yongqing Sun, Yun Li 0015
IEEE Trans. Syst. Man Cybern. Syst.5
2023 Adaptive Training Strategies for Small Object Detection Using Anchor-Based Detectors
Shenmeng Zhang, Yongqing Sun, Guoxi Gan, Zonghui Wen
ICANN (7)2
2023 Deep Quantigraphic Image Enhancement via Comparametric Equations
abstract
Most recent methods of deep image enhancement can be generally classified into two types: decompose-and-enhance and illumination estimation-centric. The former is usually less efficient, and the latter is constrained by a strong assumption regarding image reflectance as the desired enhancement result. To alleviate this constraint while retaining high efficiency, we propose a novel trainable module that diversifies the conversion from the low-light image and illumination map to the enhanced image. It formulates image enhancement as a comparametric equation parameterized by a camera response function and an exposure compensation ratio. By incorporating this module in an illumination estimation-centric DNN, our method improves the flexibility of deep image enhancement, limits the computational burden to illumination estimation, and allows for fully unsupervised learning adaptable to the diverse demands of different tasks.
Xiaomeng Wu, Yongqing Sun, Akisato Kimura
ICASSP2
2023 Distance Constraint-Based Generative Adversarial Networks for Hyperspectral Image Classification
abstract
Hyperspectral image (HSI) classification suffers from two serious problems, one is the limited labeled pixels, and the other is the class imbalance problem. As a result, the number of labeled pixels in many categories is not sufficient to characterize the spectral-spatial information, and train a satisfying deep model. By making full use of the information of unlabeled pixels, semi-supervised methods can provide better classification performance in the case of limited labeled pixels. However, they do not take into account the imbalance in the HSI data. As a method of data enhancement, generative adversarial networks focus on the above two problems and have also been widely used for the task of the HSI classification. In this work, we propose a distance constraints-based generative adversarial networks (DGAN) method for HSI classification to address these two problems. The DGAN employs the convolution autoencoder (AE) to extract the latent features of the HSI samples, and considers the reconstructed samples from the AE as the real samples for the later classifier and discriminator. In addition, the DGAN uses two distance constraints to solve the problems of the few labeled samples and class imbalance, the one latent-data distance constraint enforcing the generator to generate HSI samples for each class (especially the minority class), another discriminator-score distance constraint guiding the generator to synthesize samples that resemble the real HSI samples. Finally, the generated samples are combined classwise with the reconstructed samples and the real HSI samples to learn the parameters of the classifier and discriminator. Experimental results show that our method achieves state-of-the-art performance in terms of overall accuracy (OA) when trained with only 0.5%-4% of data sets from Indian Pines, Pavia University, and Botswana. Specifically, our method demonstrates improvements of 5.48%, 8.79%, and 0.91% on these three datasets, respectively. It reveals the great potential of the DGAN model in generating the HSI samples for each class, which contributes to improving the classification performance of the HSI data.
Anyong Qin, Zhuolin Tan, Yongqing Sun, Feng Yang 0015, Yue Zhao 0012, Chenqiang Gao
IEEE Trans. Geosci. Remote. Sens.4
2022 Active Learning for Hyperspectral Image Classification via Hypergraph Neural Network
abstract
Graph convolution network (GCN) has been extensively applied to the area of hyperspectral image (HSI) classification. However, the graph can not effectively describe the complex relationships between HSI pixels and the GCN still faces the challenge of insufficient labeled pixels. In order to alleviate the above two issues faced by the GCN in HSI classification, we propose a novel framework that integrates the active learning and the hypergraph neural network. First, we construct a hypergraph that can reveal the complex non-pairwise relationships embedded in the hyperspectral images. Next, we train a semi-supervised hypergraph neural network (GNN) with the fewer labeled training set. Then, exploiting the local structural properties of the hypergraph, the most useful HSI pixels are actively selected for labeling. Finally, we fine-tune the GNN with original training set along with the newly labeled pixels. And the last three steps are iteratively carried on. Compared with the other traditional and active learning approaches of HSI classification, the proposed active hypergraph neural network (ACGNN) can achieve better performance on the three HSI datasets.
Yongqing Sun, Anyong Qin, Yukihiro Bandoh, Chenqiang Gao, Yusuke Hiwasaki
ICIP1
2022 Motor Learning based on Presentation of a Tentative Goal
abstract
This paper presents a motor learning method based on the presenting of a personalized target motion, which we call a tentative goal. While many prior studies have focused on helping users correct their motor skill motions, most of them present the reference motion to users regardless of whether the motion is attainable or not. This makes it difficult for users to appropriately modify their motion to the reference motion when the difference between their motion and the reference motion is too significant. This study aims to provide a tentative goal that maximizes performance within a certain amount of motion change. To achieve this, predicting the performance of any motion is necessary. However, it is challenging to estimate the performance of a tentative goal by building a general model because of the large variety of human motion. Therefore, we built an individual model that predicts performance from a small training dataset and implemented it using our proposed data augmentation method. Experiments with basketball free-throw data demonstrate the effectiveness of the proposed method.
Yongqing Sun, Mitsuhiro Goto, Shigekuni Kondo, Dan Mikami, Susumu Yamamoto
ICMR2
2022 Contrast enhancement based on reflectance-oriented probabilistic equalization
Xiaomeng Wu, Yongqing Sun, Akisato Kimura, Kunio Kashino
Signal Process.2
2021 Reflectance-Oriented Probabilistic Equalization for Image Enhancement
abstract
Despite recent advances in image enhancement, it remains difficult for existing approaches to adaptively improve the brightness and contrast for both low-light and normal-light images. To solve this problem, we propose a novel 2D histogram equalization approach. It assumes intensity occurrence and co-occurrence to be dependent on each other and derives the distribution of intensity occurrence (1D histogram) by marginalizing over the distribution of intensity co-occurrence (2D histogram). This scheme improves global contrast more effectively and reduces noise amplification. The 2D histogram is defined by incorporating the local pixel value differences in image reflectance into the density estimation to alleviate the adverse effects of dark lighting conditions. Over 500 images were used for evaluation, demonstrating the superiority of our approach over existing studies. It can sufficiently improve the brightness of low-light images while avoiding over-enhancement in normal-light images.
Xiaomeng Wu, Yongqing Sun, Akisato Kimura, Kunio Kashino
ICASSP2
2021 Semantic Nighttime Image Segmentation Via Illumination and Position Aware Domain Adaptation
abstract
Due to the lack of the annotated nighttime images, general image segmentation models trained on the daytime image dataset do not perform well in nighttime scenes. The difference of the illumination condition and the difficulty to obtain the position information between daytime and nighttime makes the nighttime image segmentation tough. As a consequence, this paper proposes an end-to-end nighttime segmentation network based on the following two points: 1) Utilizing illumination adaptation with the different illumination condition on the daytime or nighttime to close the distribution gap at the feature map level; 2) With the prior information about the position of each object in the outdoor scene, some classification errors could be corrected by incorporating the self-attention mechanism. The scheme is tested on the open-source nighttime dataset Dark Zurich and night driving, with a 2.5% improvement compared to the base segmentation network.
Junhan Peng, Yongqing Sun, Zheng Wang 0007, Chia-Wen Lin
ICIP3
2021 Distribution Preserving Deep Semi-Nonnegative Matrix Factorization
abstract
Deep semi-nonnegative matrix factorization can obtain the hidden hierarchical representations according to the unknown attributes of the given data. On the other hand, the inherent structure of the each data cluster can be described by the distribution of the intra-class data. Then one hopes to learn a new low dimensional representation which can preserve the intrinsic structure embedded in the original high dimensional data space perfectly. Here we propose a novel distribution preserving deep semi-nonnegative matrix factorization method (DPNMF) to achieve this goal. As a result, the manifold structures in the raw data are well preserved in the feature space being from the top layer. The experimental results on the real-world datasets show that the proposed algorithm has good performance in terms of cluster accuracy and normalized mutual information (NMI).
Zhuolin Tan, Anyong Qin, Yongqing Sun, Yuan Yan Tang
SMC3
2021 MSLPNet: multi-scale location perception network for dental panoramic X-ray image segmentation
Qiaoyi Chen, Yue Zhao 0012, Yang Liu 0157, Yongqing Sun, Chongshi Yang, Pengcheng Li 0017, Chenqiang Gao
Neural Comput. Appl.4
2021 Infrared and Visible Cross-Modal Image Retrieval Through Shared Features
abstract
Image retrieval is one of the key techniques of computer vision, and has been studied for a long time. Nevertheless, little attention is paid to infrared and visible cross-modal retrieval which can be widely used in various applications, e.g., infrared and visible surveillance systems. In this paper, we propose a shared features based infrared-visible cross-modal image retrieval method. The similar visual features are extracted from infrared and visible images as the shared features, and the Euclidean distance is used to measure the similarity between these features. The core of the proposed method comes from three aspects: 1) Feature separation network can separate image features into shared features and exclusive features; 2) Maximum Mean Discrepancy (MMD) loss is employed to constrain the distribution of shared features, which can reduce the retrieval error caused by different imaging angles and similarity of infrared images. 3) The cross-layer fusion encoder compensates for the context loss in the convolution of infrared images. Experimental results on the Infrared-Visible dataset demonstrate the proposed method is effective and outperforms the state-of-the-art approaches.
Fangcen Liu, Chenqiang Gao, Yongqing Sun, Yue Zhao 0012, Feng Yang 0015, Anyong Qin, Deyu Meng
IEEE Trans. Circuits Syst. Video Technol.3
2020 Activity Normalization for Activity Detection in Surveillance Videos
abstract
A framework for activity detection in surveillance videos generally involves activity proposal generation and activity classification. An activity proposal is a spatial and temporal candidate region for an arbitrary activity, and an activity classifier identifies the activity class for activity proposals. One of the difficulties in activity classification is the variation in the number of activity appearances due to the diversity of the object moving directions and inter-object-positional relationships. To solve this problem, we propose an activity normalization method for rotating activity proposals so that the object-movement direction and inter-object-positional relationship are constant among all activity proposals before classification. The experimental results indicate that activity classification accuracy improves by adding our method to a general activity detection framework using the ActEV/VIRAT dataset.
Takashi Hosono, Kiyohito Sawada, Yongqing Sun, Kazuya Hayase, Jun Shimamura
ICIP3
2019 Weakly Supervised Instance Segmentation Using Hybrid Networks
abstract
Weakly-supervised instance segmentation, which could greatly save labor and time cost of pixel mask annotation, has attracted increasing attention in recent years. The commonly used pipeline firstly utilizes conventional image segmentation methods to automatically generate initial masks and then use them to train an off-the-shelf segmentation network in an iterative way. However, the initial generated masks usually contains a notable proportion of invalid masks which are mainly caused by small object instances. Directly using these initial masks to train segmentation models is harmful for the performance. To address this problem, we propose a kind of hybrid networks in this paper. In our architecture, there is a principle segmentation network which is used to handle the normal samples with valid generated masks. In addition, a complementary branch is added to handle the small and dim objects without valid masks. Experimental results indicate that our method can achieve significantly performance improvement both on the small object instances and large ones, and outperforms all state-of-the-art methods.
Shisha Liao, Yongqing Sun, Chenqiang Gao, Pranav Shenoy K. P, Song Mu, Jun Shimamura, Atsushi Sagata
ICASSP2
2016 Exploiting Objects with LSTMs for Video Categorization
abstract
Temporal dynamics play an important role for video classification. In this paper, we propose to leverage high-level semantic features to open the "black box" of the state-of-the-art temporal model, Long Short Term Memory (LSTM), with an aim to understand what is learned. More specifically, we first extract object features from a state-of-the-art CNN model that is trained to recognize 20K objects. Then we leverage LSTM with the extracted features as inputs to capture the temporal dynamics in videos. In combination with spatial and motion information, we achieve improvements for supervised video categorization. Furthermore, by masking the inputs, we demonstrate what is learned by LSTM, namely (i) which objects are crucial for recognizing a class-of-interest; (ii) how the LSTM model could assist the temporal localization of these detected objects.
Yongqing Sun, Zuxuan Wu, Xi Wang 0008, Hiroyuki Arai, Tetsuya Kinebuchi, Yu-Gang Jiang 0001
ACM Multimedia1
2016 Attribute Discovery for Person Re-Identification
Takayuki Umeda, Yongqing Sun, Go Irie, Kyoko Sudo, Tetsuya Kinebuchi
MMM (2)2
2016 Visual concept detection of web images based on group sparse ensemble learning
Yongqing Sun, Kyoko Sudo, Yukinobu Taniguchi
Multim. Tools Appl.1
2015 Cross-Domain Concept Detection with Dictionary Coherence by Leveraging Web Images
Yongqing Sun, Kyoko Sudo, Yukinobu Taniguchi
MMM (2)1
2013 Sampling of Web Images with Dictionary Coherence for Cross-Domain Concept Detection
Yongqing Sun, Kyoko Sudo, Yukinobu Taniguchi, Masashi Morimoto
MMM (2)1
2011 A novel method for semantic video concept learning using web images
abstract
In recent years, exploring the rich web image resources has been offering promising solutions to the problem of how to perform low-manual-cost concept learning. However, concept classifiers trained using web images perform poorly when they are directly applied to video concept detection. We propose a novel scheme to address video concept learning using web images, one that includes the selection of web training data and the transfer of subspace learning within a unified framework. Starting with a small set of video keyframes related to a video concept, we select web training data of good quality from the web by referring to the content of video keyframes. Then, by exploiting both the selected dataset and video keyframes, we train a robust concept classifier by means of a transfer subspace learning method. Experiment results demonstrate the robustness and effectiveness of our method.
Yongqing Sun, Akira Kojima
ACM Multimedia1
2008 A novel region-based approach to visual concept modeling using web images
abstract
A novel region-based approach is proposed to model semantic concepts using web images. Web images are mined to obtain multiple visual patterns automatically that then are used to model a semantic concept. First, the salient region groups corresponding to the representative visual patterns of a concept are mined and selected as positive samples. Next, a representative visual pattern is built in each salient region group by using a BDA classifier. Finally all the visual patterns are aggregated to describe the concept by using a BDA ensemble approach. Because the proposed method models a semantic concept utilizing multiple visual patterns, it enhances the visual variability of a visual model when learning from diverse web images and improves the robustness of the visual model in handling segmentation-related uncertainties. Experiment results demonstrate our method performs well on generic images including not only "object" concepts, but also complex "scene" concepts.
Yongqing Sun, Satoshi Shimada, Yukinobu Taniguchi, Akira Kojima
ACM Multimedia1
2005 HIRBIR: A hierarchical approach to region-based image retrieval
Yongqing Sun, Shinji Ozawa
Multim. Syst.1