Runqi Wang

dblp:266/9915 · DBLP profile ↗
← Back
26ranked-venue papers
9as first author
25since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 8 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 7 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Embedded intelligence in agriculture: A lightweight deep learning network for low-cost tomato leaf disease detection
Zhaolei Yang, Runqi Wang, Xueke An, Dexin Ma, Yuliang Yun
Eng. Appl. Artif. Intell.3
2025 VLMInferSlow: Evaluating the Efficiency Robustness of Large Vision-Language Models as a Service
abstract
Vision-Language Models (VLMs) have demonstrated great potential in real-world applications. While existing research primarily focuses on improving their accuracy, the efficiency remains underexplored. Given the real-time demands of many applications and the high inference overhead of VLMs, efficiency robustness is a critical issue. However, previous studies evaluate efficiency robustness under unrealistic assumptions, requiring access to the model architecture and parameters-an impractical scenario in ML-as-a-service settings, where VLMs are deployed via inference APIs. To address this gap, we propose VLMInferSlow, a novel approach for evaluating VLM efficiency robustness in a realistic black-box setting. VLMInferSlow incorporates fine-grained efficiency modeling tailored to VLM inference and leverages zero-order optimization to search for adversarial examples. Experimental results show that VLMInferSlow generates adversarial images with imperceptible perturbations, increasing the computational cost by up to 128.47%. We hope this research raises the community's awareness about the efficiency robustness of VLMs.
Xiasi Wang, Tianliang Yao, Runqi Wang, Kuofeng Gao, Yi Huang 0035
ACL (1)4
2025 SET: Spectral Enhancement for Tiny Object Detection
abstract
Deep learning has significantly advanced the object detection field. However, tiny object detection (TOD) remains a challenging problem. We provide a new analysis method to examine the TOD challenge through occlusion-based attribution analysis in the frequency domain. We observe that tiny objects become less distinct after feature encoding and can benefit from the removal of high-frequency information. In this paper, we propose a novel approach named Spectral Enhancement for Tiny object detection (SET), which amplifies the frequency signatures of tiny objects in a heterogeneous architecture. SET includes two modules. The Hierarchical Background Smoothing (HBS) module suppresses high-frequency noise in the background through adaptive smoothing operations. The Adversarial Perturbation Injection (API) module leverages adversarial perturbations to increase feature saliency in critical regions and prompt the refinement of object features during training. Extensive experiments on four datasets demonstrate the effectiveness of our method. Especially, SET boosts the prior art RFLA by 3.2% AP on the AI-TOD dataset.
Huixin Sun, Runqi Wang, Yanjing Li, Linlin Yang 0001, Shaohui Lin, Xianbin Cao 0001, Baochang Zhang 0001
CVPR2
2025 DFM: Differentiable Feature Matching for Anomaly Detection
abstract
Feature matching methods for unsupervised anomaly detection have demonstrated impressive performance. Existing methods primarily rely on self-supervised training and handcrafted matching schemes for task adaptation. However, they can only achieve an inferior feature representation for anomaly detection because the feature extraction and matching modules are separately trained. To address these issues, we propose a Differentiable Feature Matching (DFM) framework for joint optimization of the feature extractor and the matching head. DFM transforms nearest-neighbor matching into a pooling-based module and embeds it within a Feature Matching Network (FMN). This design enables end-to-end feature extraction and feature matching module training, thus providing better feature representation for anomaly detection tasks. DFM is generic and can be incorporated into existing feature-matching methods. We implement DFM with various backbones and conduct extensive experiments across various tasks and datasets, demonstrating its effectiveness. Notably, we achieve state-of-the-art results in the continual anomaly detection task with instance-AUROC improvement of up to 3.9% and pixel-AP improvement of up to 5.5%.
Yimi Wang, Yuguang Yang 0007, Runqi Wang, Guodong Guo, David S. Doermann, Baochang Zhang 0001
CVPR5
2025 DynamicFace: High-Quality and Consistent Face Swapping for Image and Video Using Composable 3D Facial Priors
Runqi Wang, Sijie Xu, Tianyao He, Dejia Song, Nemo Chen, Xu Tang 0007, Yao Hu 0002
ICCV1
2025 You Think, You ACT: the New Task of Arbitrary Text to Motion Generation
Runqi Wang, Caoyuan Ma, Hanrui Xu, Zheng Wang 0007
ICCV1
2025 WaveMamba: Wavelet-Driven Mamba Fusion for RGB-Infrared Object Detection
abstract
Leveraging the complementary characteristics of visible (RGB) and infrared (IR) imagery offers significant potential for improving object detection. In this paper, we propose WaveMamba, a cross-modality fusion method that efficiently integrates the unique and complementary frequency features of RGB and IR decomposed by Discrete Wavelet Transform (DWT). An improved detection head incorporating the Inverse Discrete Wavelet Transform (IDWT) is also proposed to reduce information loss and produce the final detection results. The core of our approach is the introduction of WaveMamba Fusion Block (WMFB), which facilitates comprehensive fusion across low-/high-frequency sub-bands. Within WMFB, the Low-frequency Mamba Fusion Block (LMFB), built upon the Mamba framework, first performs initial low-frequency feature fusion with channel swapping, followed by deep fusion with an advanced gated attention mechanism for enhanced integration. High-frequency features are enhanced using a strategy that applies an ``absolute maximum" fusion approach. These advancements lead to significant performance gains, with our method surpassing state-of-the-art approaches and achieving average mAP improvements of 4.5% on four benchmarks.
Haodong Zhu, Linlin Yang 0001, Hong Li 0016, Yuguang Yang 0007, Yangyang Ren, Qingcheng Zhu, Zichao Feng, Changbai Li, Shaohui Lin, Runqi Wang, Xiaoyan Luo, Baochang Zhang 0001
ICCV11
2025 Maximum Redundancy Pruning: A Principle-Driven Layerwise Sparsity Allocation for LLMs
abstract
Large language models (LLMs) have demonstrated impressive capabilities, but their enormous size poses significant challenges for deployment in real-world applications. To address this issue, researchers have sought to apply network pruning techniques to LLMs. A critical challenge in pruning is the allocation of sparsity for each layer. Recent sparsity allocation methods are often based on heuristics or search that can easily lead to suboptimal performance. In this paper, we conducted an extensive investigation into various LLMs and revealed three significant discoveries: (1) the Layerwise Pruning Sensitivity (LPS) of LLMs is highly non-uniform, (2) the choice of pruning metric affects LPS, and (3) the performance of a sparse model is related to the uniformity of its layerwise redundancy level. Based on these discoveries, we propose that the layerwise sparsity of LLMs should adhere to three principles: non-uniformity, pruning metric dependency, and uniform layerwise redundancy level in the pruned model. To this end, we proposed Maximum Redundancy Pruning (MRP), an iterative pruning algorithm that prunes in the most redundant layers (i.e., those with the highest non-outlier ratio) at each iteration. The achieved layerwise sparsity aligns with the outlined principles. We conducted extensive experiments on publicly available LLMs, including LLaMA2 and OPT, on various benchmarks. The experimental results validate the effectiveness of MRP, demonstrating its superiority over previous methods.
Chang Gao 0007, Runqi Wang, Jianfei Chen 0001, Liping Jing
ACM Multimedia3
2025 Leader is Guided: Interactive Motion Generation via Lead-Follow Paradigm and Trajectory Guidance
abstract
Generating interactive motion from texts has garnered significant attention in recent years. While text inputs offer greater flexibility, in many practical applications, there is a need to controllably impose strict constraints on the motion range or trajectory of virtual characters. However, existing trajectory-based methods are designed for single-actor scenarios and lack support for interactivity in interactive motions. Moreover, text-only methods struggle to accurately convey user-intended trajectories. The distribution shift between training and inference often leads to trajectory deviation and physical interpenetration. To address the questions mentioned, we introduce two key concepts: (1) Lead-Follow Paradigm: Inspired by role allocation in partner dancing, we decompose complex interactive motion tasks into a Lead-Follow paradigm. The leader's path is optimized first, and the follower's motion is subsequently adjusted for coherence and alignment. (2) Trajectory Guidance: We highlight the pivotal role of 3D trajectory guidance in interactive motion generation and accurately reflect user intentions. Through 3D trajectory control, we can more controllably generate the desired motion while avoiding physical interpenetration. In addition, we further investigate the refinement of motion scopes for interactive agents and propose an effective optimization strategy to enhance motion coherence and controllability. Experimental results show that the proposed approach, by more effectively using trajectory, outperforms existing methods in both realism and accuracy.
Runqi Wang, Caoyuan Ma, Jian Zhao 0013, Hanrui Xu, Dongfang Sun, Zheng Wang 0007, Xuelong Li 0001
ACM Multimedia1
2025 Single Trajectory Distillation for Accelerating Image and Video Style Transfer
abstract
Trajectory distillation based on consistency models (CMs) provides an effective framework for accelerating diffusion models by reducing inference steps. However, we find that existing CMs degrade style similarity and compromise aesthetic quality in stylization tasks-especially when handling image-to-image or video-to-video transformations that start denoising from partially noised inputs. The core limitation stems from existing methods enforcing initial-step alignment between the probability flow ODE (PF-ODE) trajectories of student models and their imperfect teacher models. This partial alignment strategy inevitably fails to guarantee full trajectory consistency, thereby compromising the overall generation quality. To address this issue, we propose Single Trajectory Distillation (STD), a training framework initiated from partial noise states. To counteract the additional time overhead introduced by STD, we design a trajectory bank that pre-stores intermediate states of the teacher model's PF-ODE trajectories, effectively offsetting the computational cost during student model training. This mechanism ensures STD maintains equivalent training efficiency compared to conventional consistency models. Furthermore, we incorporate an asymmetric adversarial loss to explicitly enhance style consistency and perceptual quality in generated outputs. Extensive experiments on image and video stylization demonstrate that our method surpasses existing acceleration models in terms of style similarity and aesthetic evaluations. Our code and results are available on the project page: https://single-trajectory-distillation.github.io/.
Sijie Xu, Runqi Wang, Dejia Song, Nemo Chen, Xu Tang 0007, Yao Hu 0002
ACM Multimedia2
2025 Enhancing multi-task performance through associative adversarial learning based on selective attacks
Yuanglong Yang, Bilang Zhang, Runqi Wang, Liping Jing, Baochang Zhang 0001
Neurocomputing4
2025 Modulated Convolutional Networks
abstract
While the deep convolutional neural network (DCNN) has achieved overwhelming success in various vision tasks, its heavy computational and storage overhead hinders the practical use of resource-constrained devices. Recently, compressing DCNN models has attracted increasing attention, where binarization-based schemes have generated great research popularity due to their high compression rate. In this article, we propose modulated convolutional networks (MCNs) to obtain binarized DCNNs with high performance. We lead a new architecture in MCNs to efficiently fuse the multiple features and achieve a similar performance as the full-precision model. The calculation of MCNs is theoretically reformulated as a discrete optimization problem to build binarized DCNNs, for the first time, which jointly consider the filter loss, center loss, and softmax loss in a unified framework. Our MCNs are generic and can decompose full-precision filters in DCNNs, e.g., conventional DCNNs, VGG, AlexNet, ResNets, or Wide-ResNets, into a compact set of binarized filters which are optimized based on a projection function and a new updated rule during the backpropagation. Moreover, we propose modulation filters (M-Filters) to recover filters from binarized ones, which lead to a specific architecture to calculate the network model. Our proposed MCNs substantially reduce the storage cost of convolutional filters by a factor of 32 with a comparable performance to the full-precision counterparts, achieving much better performance than other state-of-the-art binarized models.
Baochang Zhang 0001, Runqi Wang, Jungong Han, Rongrong Ji
IEEE Trans. Neural Networks Learn. Syst.2
2024 AQ-DETR: Low-Bit Quantized Detection Transformer with Auxiliary Queries
abstract
DEtection TRansformer (DETR)-based models have achieved remarkable performance. However, they are accompanied by a large computation overhead cost, which significantly prevents their applications on resource-limited devices. Prior arts attempt to reduce the computational burden of DETR using low-bit quantization, while these methods sacrifice a severe significant performance on weight-activation-attention low-bit quantization. We observe that the number of matching queries and positive samples affect much on the representation capacity of queries in DETR, while quantifying queries of DETR further reduces its representational capacity, thus leading to a severe performance drop. We introduce a new quantization strategy based on Auxiliary Queries for DETR (AQ-DETR), aiming to enhance the capacity of quantized queries. In addition, a layer-by-layer distillation is proposed to reduce the quantization error between quantized attention and full-precision counterpart. Through our extensive experiments on large-scale open datasets, the performance of the 4-bit quantization of DETR and Deformable DETR models is comparable to full-precision counterparts.
Runqi Wang, Huixin Sun, Linlin Yang 0001, Shaohui Lin, Chuanjian Liu, Yan Gao 0017, Yao Hu 0002, Baochang Zhang 0001
AAAI1
2024 Learning a Dynamic Neural Human via Poses Guided Dual Spaces Feature
abstract
Learning human representations from video is becoming increasingly important in various applications. However, due to the limited information in videos and the complexity of human deformation, existing methods cannot faithfully reconstruct the image representation of humans, including clothing folds and light and shadow. Our method is built upon a deformation-based approach, which uses pose-guided joint learning to derive human representations in both canonical space and observation space, thereby enhancing the model’s performance in human details. We conducted several experiments on publicly available datasets using our approach, achieving highly realistic reconstruction results that are difficult to distinguish from real frames. Our approach also showed improved overall evaluation metrics for video frames that were not visible in the original view angle.
Caoyuan Ma, Runqi Wang, Wu Liu 0005, Ziqiao Zhou, Zheng Wang 0007
AVSS2
2024 Generating synthetic computed tomography for radiotherapy: SynthRAD2023 challenge report
abstract
Radiation therapy plays a crucial role in cancer treatment, necessitating precise delivery of radiation to tumors while sparing healthy tissues over multiple days. Computed tomography (CT) is integral for treatment planning, offering electron density data crucial for accurate dose calculations. However, accurately representing patient anatomy is challenging, especially in adaptive radiotherapy, where CT is not acquired daily. Magnetic resonance imaging (MRI) provides superior soft-tissue contrast. Still, it lacks electron density information, while cone beam CT (CBCT) lacks direct electron density calibration and is mainly used for patient positioning. Adopting MRI-only or CBCT-based adaptive radiotherapy eliminates the need for CT planning but presents challenges. Synthetic CT (sCT) generation techniques aim to address these challenges by using image synthesis to bridge the gap between MRI, CBCT, and CT. The SynthRAD2023 challenge was organized to compare synthetic CT generation methods using multi-center ground truth data from 1080 patients, divided into two tasks: (1) MRI-to-CT and (2) CBCT-to-CT. The evaluation included image similarity and dose-based metrics from proton and photon plans. The challenge attracted significant participation, with 617 registrations and 22/17 valid submissions for tasks 1/2. Top-performing teams achieved high structural similarity indices (≥0.87/0.90) and gamma pass rates for photon (≥98.1%/99.0%) and proton (≥97.3%/97.0%) plans. However, no significant correlation was found between image similarity metrics and dose accuracy, emphasizing the need for dose evaluation when assessing the clinical applicability of sCT. SynthRAD2023 facilitated the investigation and benchmarking of sCT generation techniques, providing insights for developing MRI-only and CBCT-based adaptive radiotherapy. It showcased the growing capacity of deep learning to produce high-quality sCT, reducing reliance on conventional CT for treatment planning.
Evi M. C. Huijben, Maarten L. Terpstra, Arthur Jr Galapon, Suraj Pai, Adrian Thummerer, Peter J. Koopmans, Manya Afonso, Maureen van Eijnatten, Oliver J. Gurney-Champion, Zeli Chen, Kaiyi Zheng, Chuanpu Li, Haowen Pang, Chuyang Ye, Runqi Wang, Fuxin Fan, Jingna Qiu, Yixing Huang, Juhyung Ha, Jong Sung Park, Alexandra Alain-Beaudoin, Silvain Bériault, Pengxin Yu, Zhanyao Huang, Gengwan Li, Xueru Zhang, Yubo Fan, Bowen Xin, Aaron Nicolson, Lujia Zhong, Zhiwei Deng, Gustav Mueller-Franzes, Firas Khader, Xia Li 0005, Ye Zhang 0039, Cédric Hémon, Valentin Boussot, Shaobin Wang, Derk Mus, Bram Kooiman, Chelsea A. H. Sargeant, Edward G. A. Henderson, Satoshi Kondo, Satoshi Kasai, Reza Karimzadeh, Bulat Ibragimov, Thomas Helfer, Jessica Dafflon, Enpei Wang, Zoltán Perkó, Matteo Maspero
Medical Image Anal.16
2023 AttriCLIP: A Non-Incremental Learner for Incremental Knowledge Learning
abstract
Continual learning aims to enable a model to incrementally learn knowledge from sequentially arrived data. Previous works adopt the conventional classification architecture, which consists of a feature extractor and a classifier. The feature extractor is shared across sequentially arrived tasks or classes, but one specific group of weights of the classifier corresponding to one new class should be incrementally expanded. Consequently, the parameters of a continual learner gradually increase. Moreover, as the classifier contains all historical arrived classes, a certain size of the memory is usually required to store rehearsal data to mitigate classifier bias and catastrophic forgetting. In this paper, we propose a non-incremental learner, named AttriCLIP, to incrementally extract knowledge of new classes or tasks. Specifically, AttriCLIP is built upon the pre-trained visual-language model CLIP. Its image encoder and text encoder are fixed to extract features from both images and text. Text consists of a category name and a fixed number of learnable parameters which are selected from our designed attribute word bank and serve as attributes. As we compute the visual and textual similarity for classification, AttriCLIP is a non-incremental learner. The attribute prompts, which encode the common knowledge useful for classification, can effectively mitigate the catastrophic forgetting and avoid constructing a replay memory. We evaluate our AttriCLIP and compare it with CLIP-based and previous state-of-the-art continual learning methods in realistic settings with domain-shift and long-sequence learning. The results show that our method performs favorably against previous state-of-the-arts. The implementation code will be available at https://gitee.com/mindspore/models/tree/master/research/cv/AttriCLIP.
Runqi Wang, Xiaoyue Duan, Guoliang Kang, Jianzhuang Liu, Shaohui Lin, Songcen Xu, Jinhu Lü 0001, Baochang Zhang 0001
CVPR1
2023 Few-Shot Learning with Visual Distribution Calibration and Cross-Modal Distribution Alignment
abstract
Pre-trained vision-language models have inspired much research on few-shot learning. However, with only a few training images, there exist two crucial problems: (1) the visual feature distributions are easily distracted by class-irrelevant information in images, and (2) the alignment between the visual and language feature distributions is difficult. To deal with the distraction problem, we propose a Selective Attack module, which consists of trainable adapters that generate spatial attention maps of images to guide the attacks on class-irrelevant image areas. By messing up these areas, the critical features are captured and the visual distributions of image features are calibrated. To better align the visual and language feature distributions that describe the same object class, we propose a cross-modal distribution alignment module, in which we introduce a vision-language prototype for each class to align the distributions, and adopt the Earth Mover's Distance (EMD) to optimize the prototypes. For efficient computation, the upper bound of EMD is derived. In addition, we propose an augmentation strategy to increase the diversity of the images and the text prompts, which can reduce overfitting to the few-shot training images. Extensive experiments on 11 datasets demonstrate that our method consistently outperforms prior arts in few-shot learning. The implementation code will be available at https://gitee.com/mindspore/models/tree/master/research/cv/SADA.
Runqi Wang, Xiaoyue Duan, Jianzhuang Liu, Yuning Lu, Tian Wang 0002, Songcen Xu, Baochang Zhang 0001
CVPR1
2023 Cross-Level Distillation and Feature Denoising for Cross-Domain Few-Shot Classification
Runqi Wang, Jianzhuang Liu, Asako Kanezaki
ICLR2
2023 Few-Shot Learning with Complex-Valued Neural Networks and Dependable Learning
Runqi Wang, Baochang Zhang 0001, Guodong Guo, David S. Doermann
Int. J. Comput. Vis.1
2023 Anti-Bandit for Neural Architecture Search
Runqi Wang, Linlin Yang 0001, Wei Wang 0016, David S. Doermann, Baochang Zhang 0001
Int. J. Comput. Vis.1
2022 Anti-retroactive Interference for Lifelong Learning
Runqi Wang, Yuxiang Bao, Baochang Zhang 0001, Jianzhuang Liu, Wentao Zhu 0001, Guodong Guo
ECCV (24)1
2022 MSPCD: predicting circRNA-disease associations via integrating multi-source data and hierarchical neural network
abstract
BACKGROUND: Increasing evidence shows that circRNA plays an essential regulatory role in diseases through interactions with disease-related miRNAs. Identifying circRNA-disease associations is of great significance to precise diagnosis and treatment of diseases. However, the traditional biological experiment is usually time-consuming and expensive. Hence, it is necessary to develop a computational framework to infer unknown associations between circRNA and disease. RESULTS: In this work, we propose an efficient framework called MSPCD to infer unknown circRNA-disease associations. To obtain circRNA similarity and disease similarity accurately, MSPCD first integrates more biological information such as circRNA-miRNA associations, circRNA-gene ontology associations, then extracts circRNA and disease high-order features by the neural network. Finally, MSPCD employs DNN to predict unknown circRNA-disease associations. CONCLUSIONS: Experiment results show that MSPCD achieves a significantly more accurate performance compared with previous state-of-the-art methods on the circFunBase dataset. The case study also demonstrates that MSPCD is a promising tool that can effectively infer unknown circRNA-disease associations.
Lei Deng 0002, Dayun Liu, Yizhan Li, Runqi Wang, Hui Liu 0026
BMC Bioinform.4
2022 Citizen Participation in the Co-Production of Urban Natural Resource Assets: Analysis Based on Social Media Big Data
abstract
Abundant natural resources are the basis of urbanisation and industrialisation. Citizens are the key factor in promoting a sustainable supply of natural resources and the high-quality development of urban areas. This study focuses on the co-production behaviours of citizens regarding urban natural resource assets in the age of big data, and uses the latent Dirichlet allocation algorithm and the stepwise regression analysis method to evaluate citizens’ experiences and feelings related to the urban capitalisation of natural resources. Results show that, firstly, the machine learning algorithm based on natural language processing can effectively identify and deal with the demands of urban natural resource assets. Secondly, in the experience of urban natural resources, citizens pay more attention to the combination of history, culture, infrastructure and natural landscape. Unique natural resource can enhance citizens’ sense of participation. Finally, the scenery, entertainment and quality and value of urban natural resources are the influencing factors of citizens’ satisfaction.
Shaojun Ma, Runqi Wang, Yilin Zheng
J. Glob. Inf. Manag.3
2022 Data-adaptive binary neural networks for efficient object detection and recognition
Junhe Zhao, Sheng Xu 0007, Runqi Wang, Baochang Zhang 0001, Guodong Guo, David S. Doermann, Dianmin Sun
Pattern Recognit. Lett.3
2021 IDARTS: Interactive Differentiable Architecture Search
abstract
Differentiable Architecture Search (DARTS) improves the efficiency of architecture search by learning the architecture and network parameters end-to-end. However, the intrinsic relationship between the architecture’s parameters is neglected, leading to a sub-optimal optimization process. The reason lies in the fact that the gradient descent method used in DARTS ignores the coupling relationship of the parameters and therefore degrades the optimization. In this paper, we address this issue by formulating DARTS as a bi-linear optimization problem and introducing an Interactive Differentiable Architecture Search (IDARTS). We first develop a backtracking backpropagation process, which can decouple the relationships of different kinds of parameters and train them in the same framework. The backtracking method coordinates the training of different parameters that fully explore their interaction and optimize training. We present experiments on the CIFAR10 and ImageNet datasets that demonstrate the efficacy of the IDARTS approach by achieving a top-1 accuracy of 76.52% on ImageNet without additional search cost vs. 75.8% with the state-of-the-art PC-DARTS.
Runqi Wang, Baochang Zhang 0001, Tian Wang 0002, Guodong Guo, David S. Doermann
ICCV2
2020 Visual exploration of latent space for traditional Chinese music
abstract
Generating compact and effective numerical representations of data is a fundamental step for many machine learning tasks. Traditionally, handcrafted features are used but as deep learning starts to show its potential, using deep learning models to extract compact representations becomes a new trend. Among them, adopting vectors from the model’s latent space is the most popular. There are several studies focused on visual analysis of latent space in NLP and computer vision . However, relatively little work has been done for music information retrieval (MIR) especially incorporating visualization. To bridge this gap, we propose a visual analysis system utilizing Autoencoders to facilitate analysis and exploration of traditional Chinese music. Due to the lack of proper traditional Chinese music data, we construct a labeled dataset from a collection of pre-recorded audios and then convert them into spectrograms . Our system takes music features learned from two deep learning models (a fully-connected Autoencoder and a Long Short-Term Memory (LSTM) Autoencoder) as input. Through interactive selection, similarity calculation, clustering and listening, we show that the latent representations of the encoded data allow our system to identify essential music elements, which lay the foundation for further analysis and retrieval of Chinese music in the future.
Jingyi Shen, Runqi Wang, Han-Wei Shen
Vis. Informatics2