Danyang Cao

dblp:201/5398 · DBLP profile ↗
← Back
13ranked-venue papers
13as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 4 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Improving local association in end-to-end tracking via Hard Negative Confusion-aware Denoising
Danyang Cao
Eng. Appl. Artif. Intell.1
2026 LG-Fast-ACDTW: a fast adaptive dynamic time warping algorithm
Danyang Cao, Hanyu Cui, Di Liu 0014, Zifeng Lin
Knowl. Inf. Syst.1
2026 Chinese poetry generation using a lightweight model with attention-assisted loss function
Danyang Cao, Jingke Tan
Multim. Syst.1
2025 G-WVDTW: A Generalised Weighted Variance Dynamic Time Warping Algorithm for Subsequence Matching in Multivariate Time Series
abstract
ABSTRACT Dynamic time warping (DTW) is an algorithm used to measure the similarity between sequences, with widespread applications in domains such as speech recognition, image processing and video synchronisation. However, when matching a shorter multivariate time subsequence to a longer time series containing a similar subsequence, existing DTW variants struggle to accurately determine the matching path. To address this issue, we propose an improved algorithm, generalised weighted variance DTW (G‐WVDTW). We extend the DTW algorithm to multivariate time series and introduce a weighted variance‐based approach to calculate local distances. This allows the algorithm to better assess the distance between different time points in multivariate time series. Additionally, we modify the algorithm's boundary conditions, enabling it to handle subsequence matching tasks in multivariate time series. We conducted similarity retrieval experiments using public datasets and evaluated the algorithm's performance with the AUC metric, achieving up to a 19% improvement on certain datasets. Furthermore, we performed alignment experiments on industrial data, where we artificially generated aligned sequences and quantitatively assessed the alignment errors, which were lower than those produced by other DTW variants. Finally, we validated the algorithm's superior performance in multivariate time series subsequence matching tasks using a synthetic dataset and showcased its use in motif detection using a wind power generation dataset. The algorithm can be applied in fields such as industrial, meteorological and electrocardiogram (ECG) signal analysis for tasks like time series retrieval, matching and data labelling.
Danyang Cao, ZiFeng Lin, Xiaoyuan Chai
Expert Syst. J. Knowl. Eng.1
2025 YOLOv7-DBV: A Lightweight Object Detection Algorithm Based on Cross-Stage Local Network
abstract
Object detection algorithms that achieve high accuracy in complex application scenarios often result in complex network structures and exhibit poor detection performance for deformed objects and densely distributed small objects. These complex networks have high hardware requirements for mobile terminals or devices that are resource-constrained, significantly limiting their practical applications. To balance detection accuracy and network complexity, we propose a lightweight object detection algorithm based on a cross-stage local network, termed YOLOv7-DBV. First, a lightweight VoV-GSDSBCSP structure is designed to optimize the Head network structure of YOLOv7. Next, DCNv3 is introduced to enhance the YOLOv7 backbone network, improving the extraction of feature information for deformed objects. Finally, a Bi-level Routing Attention (BRA) mechanism is incorporated to further optimize the backbone network of YOLOv7, emphasizing the feature information of small objects. Experiments on a public underwater object detection dataset demonstrate that the YOLOv7-DBV model increases mAP0.5 by 1.05% and reduces the number of parameters by 15.31% compared to the original YOLOv7 model. The proposed algorithm achieves a better balance between detection accuracy and network complexity, outperforming YOLOv7, YOLOv6l, YOLOv5l, and other algorithms.
Danyang Cao, Yongfu Wong
Int. J. Pattern Recognit. Artif. Intell.1
2025 Compact and efficient language modeling for classical poetry: generation and interpretation
Danyang Cao, Guanmin Zhang
Multim. Syst.1
2025 Improve the image caption generation on out-domain dataset by external knowledge augmented
Danyang Cao
Multim. Syst.1
2024 Multidimensional time series motif group discovery based on matrix profile
Danyang Cao, Zifeng Lin
Knowl. Based Syst.1
2024 Emotional voice conversion using DBiLSTM-NN with MFCC and LogF0 features
Danyang Cao, Chengzhi Miao
Multim. Tools Appl.1
2024 NeuralVC: Any-to-Any Voice Conversion Using Neural Networks Decoder for Real-Time Voice Conversion
abstract
With the advancement of Automatic Speech Recognition (ASR) and Text-to-Speech (TTS) technologies, high-quality speech conversion can now be achieved by extracting source speech content and reconstructing waveforms based on target speaker information. However, current methods still require improvement in terms of inference speed, especially when running on CPUs with slow speeds, greatly limiting the real-time application of speech conversion. To address this issue, we propose a real-time speech conversion model called NeuralVC. Our model is based on the VITS architecture, where the decoder for synthesizing speech plays a crucial role in influencing speech synthesis speed. To obtain speaker-independent content information, we introduce pre-trained HuBERT for extracting speech content features. To improve synthesis speed, we integrate a lightweight neural decoder based on SEANet for synthesizing speech, and modify it to receive and understand speaker information, significantly enhancing the speaker similarity of the converted speech. Additionally, we introduce a pre-trained speaker encoder and combine it with speaker consistency loss to improve the model's conversion ability in unseen scenarios, achieving any-to-any speech conversion. Experimental results demonstrate that our proposed model can achieve high-quality real-time speech conversion and maintain good performance in unseen scenarios.
Danyang Cao
IEEE Signal Process. Lett.1
2024 AD-autoformer: decomposition transformers with attention distilling for long sequence time-series forecasting
Danyang Cao
J. Supercomput.1
2021 Video object detection algorithm based on dynamic combination of sparse feature propagation and dense feature aggregation
Danyang Cao, Jinfeng Ma, Zhixin Chen
Multim. Tools Appl.1
2019 An image caption method based on object detection
Danyang Cao, Menggui Zhu
Multim. Tools Appl.1