Xinyi Tong 0001

dblp:171/0531-1 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0002-3269-3628ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Video Echoed in Music: Semantic, Temporal, and Rhythmic Alignment for Video-to-Music Generation
abstract
Video-to-Music generation seeks to generate musically appropriate background music that enhances audiovisual immersion for videos. However, current approaches suffer from two critical limitations: 1) incomplete representation of video details, leading to weak alignment, and 2) inadequate temporal and rhythmic correspondence, particularly in achieving precise beat synchronization. To address the challenges, we propose Video Echoed in Music (VeM), a latent music diffusion that generates high-quality soundtracks with semantic, temporal, and rhythmic alignment for input videos. To capture video details comprehensively, VeM employs a hierarchical video parsing that acts as a music conductor, orchestrating multi-level information across modalities. Modality-specific encoders, coupled with a storyboard-guided cross-attention mechanism (SG-CAtt), integrate semantic cues while maintaining temporal coherence through position and duration encoding. For rhythmic precision, the frame-level transition-beat aligner and adapter (TB-As) dynamically synchronize visual scene transitions with music beats. We further contribute a novel video-music paired dataset sourced from e-commerce advertisements and video-sharing platforms, which imposes stricter transition-beat synchronization requirements. Meanwhile, we introduce novel metrics tailored to the task. Experimental results demonstrate superiority, particularly in semantic relevance and rhythmic precision.
Xinyi Tong 0001, Yiran Zhu, Jishang Chen, Chunru Zhan, Tianle Wang 0007, Sirui Zhang, Nian Liu 0003, Tiezheng Ge, Duo Xu 0004, Xin Jin 0015, Feng Yu 0032, Song-Chun Zhu
AAAI1
2025 Learning Uniformly Distributed Embedding Clusters of Stylistic Skills for Physically Simulated Characters
Nian Liu 0003, Zi Wang 0014, Tengyu Liu, Hongzhao Xie, Xinyi Tong 0001, Libin Liu 0002, Yaodong Yang 0001, Zhaofeng He 0001
ACM Multimedia6
2025 Video Echoed in Harmony: Learning and Sampling Video-Integrated Chord Progression Sequences for Controllable Video Background Music Generation
abstract
Automatically generating video background music mitigates the inefficiency and time-consuming drawbacks of current manual video editing. Two key challenges hinder the expansion of the inception of video-to-music tasks. 1) Limited availability of high-quality video–music datasets and annotations. 2) Absence of music generation methods that consider actual musicality, which are controlled by interpretable factors based on music theory. In the article, we propose video echoed in harmony (VEH), a method for learning and sampling video-integrated chord progression sequences. Our approach adopts harmony, represented by chord progressions that are aligned with various music formats [musical instrument digital interface (MIDI), audio, and score], imitating chord precedence in human music composition. Visual-language models link visual features to chord progressions through genre labels and descriptive words in generated textualized videos. The two aforementioned features collectively obviate the necessity of extensive video–music paired data. Besides, an energy-based chord progression learning and sampling algorithm quantifies abstract harmony impressions to statistical features, serving as interpretable factors for the controllable music generation based on music theory. Experimental results demonstrate that the proposed method outperforms the state-of-the-art, producing a superior music alignment for the given video.
Xinyi Tong 0001, Peiyang Yu, Nian Liu 0003, Hui Qv, Tao Ma 0008, Bo Zheng 0007, Feng Yu 0032, Song-Chun Zhu
IEEE Trans. Comput. Soc. Syst.1
2022 RecDis-SNN: Rectifying Membrane Potential Distribution for Directly Training Spiking Neural Networks
abstract
The brain-inspired and event-driven Spiking Neural Network (SNN) aiming at mimicking the synaptic activity of biological neurons has received increasing attention. It transmits binary spike signals between network units when the membrane potential exceeds the firing threshold. This biomimetic mechanism of SNN appears energy-efficiency with its power sparsity and asynchronous operations on spike events. Unfortunately, with the propagation of binary spikes, the distribution of membrane potential will shift, leading to degeneration, saturation, and gradient mismatch problems, which would be disadvantageous to the network optimization and convergence. Such undesired shifts would prevent the SNN from performing well and going deep. To tackle these problems, we attempt to rectify the membrane potential distribution (MPD) by designing a novel distribution loss, MPD-Loss, which can explicitly penalize the un-desired shifts without introducing any additional operations in the inference phase. Moreover, the proposed method can also mitigate the quantization error in SNNs, which is usually ignored in other works. Experimental results demonstrate that the proposed method can directly train a deeper, larger, and better-performing SNN within fewer timesteps.
Yufei Guo 0001, Xinyi Tong 0001, Yuanpei Chen, Liwen Zhang 0001, Xiaode Liu, Zhe Ma 0001, Xuhui Huang
CVPR2
2022 Reducing Information Loss for Spiking Neural Networks
Yufei Guo 0001, Yuanpei Chen, Liwen Zhang 0001, YingLei Wang, Xiaode Liu, Xinyi Tong 0001, Yuanyuan Ou, Xuhui Huang, Zhe Ma 0001
ECCV (11)6
2022 Real Spike: Learning Real-Valued Spikes for Spiking Neural Networks
Yufei Guo 0001, Liwen Zhang 0001, Yuanpei Chen, Xinyi Tong 0001, Xiaode Liu, YingLei Wang, Xuhui Huang, Zhe Ma 0001
ECCV (12)4
2022 Clustering by centroid drift and boundary shrinkage
Hui Qv, Tao Ma 0008, Xinyi Tong 0001, Xuhui Huang, Zhe Ma 0001, Jiehong Feng
Pattern Recognit.3
2020 Few-Shot Learning With Attention-Weighted Graph Convolutional Networks For Hyperspectral Image Classification
abstract
In this paper, to alleviate the demand for enormous labeled data in the classification task, an Attention-weighted Graph Convolutional Networks (AwGCN) model for hyperspectral image (HSI) few-shot classification is proposed, which aims to explore the internal relationships of data for semi-supervised label propagation. To be specific, the attention-weighted graph is exploited to fully quantify the relationships of all samples, which is potential to solve the HSI few-shot learning problems. Subsequently, Graph Convolutional Networks (GCN) are applied to spread the labels, which ascertain the categories of samples based on the trained attention-weighted graph. The robust prediction of our proposed approach is validated on the real HSI and the experimental results show a competitive good performance, which demonstrates the superior ability of AwGCN in HSI few-shot classification.
Xinyi Tong 0001, Jihao Yin, Bingnan Han, Hui Qv
ICIP1
2019 Global Self-Labeled Distribution Analysis for Hyperspectral Band Selection
abstract
A global self-labeled distribution analysis (GSLDA) for hyperspectral image (HSI) band selection is proposed in this paper, which focuses on an unsupervised method to ascertain the band discrimination. In order to generate the band labels for further analysis, the concept of the local minimum spanning forest (LMSF) is introduced into the construction of the global self-labeled band partitions based on graph theory. Meanwhile, the novel scoring strategy of triple-density indexes is applied to analyze the labeled-band distribution for determining the selected band subset with clear discrimination. The feasibility of the proposed method is evaluated on real hyperspectral data and the experiment results show a competitive good performance, which demonstrates that the selected bands hold apparent global discrimination and robust noise immunity.
Xinyi Tong 0001, Jihao Yin, Limin Wu, Hui Qv
IGARSS1