EDBT 2026 Demo / reviewers in the wild / expert
Qin Zhang 0009
dblp:45/47-9
· DBLP profile ↗
27ranked-venue papers
0as first author
14since 2021 · last 2025
0009-0001-0205-6986ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 7 since 2021Artificial intelligence and machine learning · 8 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Pop-Diffuseq: Controllable Symbolic Music Multi-Instrument Infilling and Accompaniment Generation with Long-Axis AttentionabstractControllability is a major challenge in music infilling and accompaniment tasks. Solutions based on transformer decoders have been widely adopted, while data-driven approaches with full self-attention result in high costs and unsatisfied outcomes for fine-grained control. Existing diffusion methods rely on trained classifiers, unconditional frameworks, or solo track, etc. To address these issues, we explore novel methods to enhance the controllability and quality of music model while reducing computational complexity. Firstly, we improve the classifier-free diffusion for multi-instrumental pop music. Secondly, we design a long-axis attention algorithm that combines long with axial attention to acquire the feature correlations of multi-dimensional attributes. Additionally, we contribute a pop band dataset with melody, style and mood labels handcrafted by musicians. After experiments on the benchmark dataset, our method demonstrates high-quality controllable results and outperforms existing state-of-the-art models. The GPU memory of our model is 26.9% lower than Diffuseq under the same hyperparameters. Haonan Cheng, Long Ye, Qin Zhang 0009 |
ICME | 4 |
| 2024 | Coarse-to-Fine Domain Adaptation for Cross-Subject EEG Emotion Recognition with Contrastive Learning
Shuang Ran, Wei Zhong 0001, Long Ye, Qin Zhang 0009 |
PRCV (15) | 5 |
| 2024 | Mind to Music: An EEG Signal-Driven Real-Time Emotional Music Generation SystemabstractMusic is an important way for emotion expression, and traditional manual composition requires a solid knowledge of music theory. It is needed to find a simple but accurate method to express personal emotions in music creation. In this paper, we propose and implement an EEG signal‐driven real‐time emotional music generation system for generating exclusive emotional music. To achieve real‐time emotion recognition, the proposed system can obtain the model suitable for a newcomer quickly through short‐time calibration. And then, both the recognized emotion state and music structure features are fed into the network as the conditional inputs to generate exclusive music which is consistent with the user’s real emotional expression. In the real‐time emotion recognition module, we propose an optimized style transfer mapping algorithm based on simplified parameter optimization and introduce the strategy of instance selection into the proposed method. The module can obtain and calibrate a suitable model for a new user in short‐time, which achieves the purpose of real‐time emotion recognition. The accuracies have been improved to 86.78% and 77.68%, and the computing time is just to 7 s and 10 s on the public SEED and self‐collected datasets, respectively. In the music generation module, we propose an emotional music generation network based on structure features and embed it into our system, which breaks the limitation of the existing systems by calling third‐party software and realizes the controllability of the consistency of generated music with the actual one in emotional expression. The experimental results show that the proposed system can generate fluent, complete, and exclusive music consistent with the user’s real‐time emotion recognition results. Shuang Ran, Wei Zhong 0001, Danting Duan, Long Ye, Qin Zhang 0009 |
Int. J. Intell. Syst. | 6 |
| 2024 | Interaction Between Dynamic Affection and Arithmetic Cognitive Ability: A Practical Investigation With EEG MeasurementabstractEmotions play an essential role in affecting the performance of cognitive abilities in continuous cognitive tasks. Most previous studies share a common issue in that the evoked emotions are simply presumed to be real emotions, without taking into account the observation that emotions may be changed when carrying out cognitive activities. This may lead to the inaccurate detection of true emotions, which further adversely affects the investigation of interactions between emotion and cognition. To address this challenging problem, the present work develops an innovative study using EEG measurement to investigate the interaction between dynamic affection and cognitive ability. In particular, a real-time emotion detection model by the use of physiological signals (i.e., EEG) is constructed, to dynamically monitor the current emotional state. Given the observed emotion, the analysis of the interaction between cognitive abilities and dynamic emotions is undertaken from the perspectives of both behavioral performance and brain mechanisms. Research outcomes indicate that emotions are not stable, and are indeed dynamically changed by cognitive performance. Meanwhile, cognitive activities also influence the brain activation pattern revealed under different emotions, which validates the necessity of introducing the dynamic emotion monitoring model. In addition, the best performance has been found when the emotional state is neutral in terms of accuracy and response time. The results of this study provide a potential basis for assessing the cognitive abilities of individuals with different emotions in a variety of applications of cognitive scenarios. Yilu Peng, Qin Zhang 0009, Xia Wu 0001 |
IEEE Trans. Affect. Comput. | 5 |
| 2024 | MusicECAN: An Automatic Denoising Network for Music Recordings With Efficient Channel AttentionabstractIn this work, we address the long-standing problem of automatic recorded music denoising. In previous audio denoising research, the primary focus has been on speech, and music denoising works only considered noise types in indoor conversation scenarios or old gramophone recordings, neglecting the amateur music recording scenario. To this end, we first propose MusicECAN, an automatic music denoising method designed to filter out additional noise components in recorded music. The novel architecture comprises two key components, namely, a feature learning module and a noise filtering module, which can efficiently but effectively model, refine and denoise the noisy input. Specifically, in order to capture sufficient noisy music information, an ECA-U-SAM based feature learning module is designed by incorporating an efficient channel attention (ECA) mechanism into the traditional U-Net model with a supervised attention module (SAM). To train our MusicECAN, we collect M&N, a dataset containing various clean music and noise recordings. Through the combination of different clean-noise recording pairs, we can effectively simulate possible music performance environments with various background noise. Extensive quantitative and qualitative comparisons demonstrate that our MusicECAN outperforms the state-of-the-art audio denoising methods. Haonan Cheng, Zhicheng Lian, Long Ye, Qin Zhang 0009 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2024 | A Dataset and Benchmark for 3D Scene Plausibility AssessmentabstractThe surge in popularity of 3D scene synthesis has driven the development of diverse methods for assessing the quality of synthesized scenes. While subjective assessment methods are widespread, their time-consuming and labor-intensive nature prompts exploration into more efficient objective alternatives. This paper introduces an objective approach to evaluating scene plausibility, aiming to overcome the limitations associated with subjective methods. To underpin our objective evaluation, we present the 3D-SPAD dataset, comprising plausibility scores for 3000 scenes across 46 object categories. Leveraging this dataset, we propose a graph attention-based network designed to accurately estimate scene plausibility. A comprehensive evaluation of our network is conducted through a series of experiments, showcasing its feasibility and reliability. For additional details and access to the code, please refer to our GitHub repository athttps://github.com/Mayibo-cuc/3D-SPAN. Yibo Ma, Wei Zhong 0001, Long Ye, Xinyan Yang, Qin Zhang 0009 |
IEEE Trans. Multim. | 7 |
| 2024 | Interpretability Diversity for Decision-Tree-Initialized Dendritic Neuron Model EnsembleabstractTo construct a strong classifier ensemble, base classifiers should be accurate and diverse. However, there is no uniform standard for the definition and measurement of diversity. This work proposes a learners' interpretability diversity (LID) to measure the diversity of interpretable machine learners. It then proposes a LID-based classifier ensemble. Such an ensemble concept is novel because: 1) interpretability is used as an important basis for diversity measurement and 2) before its training, the difference between two interpretable base learners can be measured. To verify the proposed method's effectiveness, we choose a decision-tree-initialized dendritic neuron model (DDNM) as a base learner for ensemble design. We apply it to seven benchmark datasets. The results show that the DDNM ensemble combined with LID obtains superior performance in terms of accuracy and computational efficiency compared to some popular classifier ensembles. A random-forest-initialized dendritic neuron model (RDNM) combined with LID is an outstanding representative of the DDNM ensemble. Xudong Luo 0003, Long Ye, Xiaohao Wen, MengChu Zhou, Qin Zhang 0009 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2023 | Pruning of Dendritic Neuron Model with Significance Constraints for ClassificationabstractThe dendritic neural model (DNM) simulates the information-processing mechanisms and procedures of neurons by mimicking the nonlinearity of synapses in the human brain. This allows for a better understanding of biological nervous systems and has been used to a wonderful effect in various fields. However, there are some problems with the existing DNM, such as the high complexity and the limited generalization capability. Pruning is one of the most common and crucial approaches to model compression. It is a model optimization technique that involves removing redundant values from a weight tensor to develop smaller and more efficient neural networks. A compressed neural network can enable faster running and reduced computational costs in network training. To improve the model performance, this work proposes a DNM pruning method with dendrite layer significance constraints. Our proposed method not only calculates the significance of dendrite layers but also makes the significance of a few dendrite layers in the trained model concentrated in a few dendrite layers so that the dendrite layers with low significance can be pruned away. The results of the simulation experiment on classification problems show that our proposed method outperforms the existing pruning methods in terms of network size and generalization performance. Xudong Luo 0003, Long Ye, Xiaohao Wen, Qin Zhang 0009 |
IJCNN | 5 |
| 2023 | Gender-Sensitive EEG Channel Selection for Emotion Recognition Using Enhanced Genetic AlgorithmabstractEEG channel selection aims to choose informative and representative channels to reduce data redundancy. It is very beneficial for improving the utility and efficiency of emotion recognition. Previous studies on EEG channel selection have not considered the influence of genders despite long-standing belief in gender differences with respect to emotion analysis. In this paper, we collected EEG signals from 20 subjects containing 10 males and 10 females by letting them watch short emotional videos. Then, to reduce data redundancy, we propose an enhanced genetic algorithm to select the optimal channel subsets separately for male and female subjects by incorporating a novel evolution operation. Experimental results show that the proposed algorithm achieves higher accuracy in terms of emotion recognition than several compared methods with a smaller channel subset. Besides, experimental results also indicate that the gender differences in neural patterns indeed exist. Through this study, the gender-sensitive channel selection offers a new avenue for further development of EEG based emotion recognition. Danting Duan, Qiang Yang 0008, Wei Zhong 0001, Long Ye, Qin Zhang 0009, Jun Zhang 0003 |
SMC | 6 |
| 2022 | Unsupervised Quantized Prosody Representation for Controllable Speech SynthesisabstractIn this paper, we propose a novel prosody disentangle method for prosodic Text-to-Speech (TTS) model, which introduces the vector quantization (VQ) method to the auxiliary prosody encoder to obtain the decomposed prosody representations in an unsupervised manner. Rely on its advantages, the speaking styles, such as pitch, speaking velocity, local pitch variance, etc., are decomposed automatically into the latent quantize vectors. We also investigate the internal mechanism of VQ disentangle process by means of a latent variables counter and find that higher value dimensions usually represent prosody information. Experiments show that our model can control the speaking styles of synthesis results by directly manipulating the latent variables. The objective and subjective evaluations illustrated that our model outperforms the popular models. Yuankun Xie, Hui Wang 0070, Qin Zhang 0009 |
ICME | 5 |
| 2022 | Multi-source Information-Shared Domain Adaptation for EEG Emotion Recognition
Wei Zhong 0001, Long Ye, Qin Zhang 0009 |
PRCV (2) | 5 |
| 2022 | Visually aligned sound generation via sound-producing motion parsing
Wei Zhong 0001, Long Ye, Qin Zhang 0009 |
Neurocomputing | 4 |
| 2021 | MovieREP: A New Movie Reproduction Framework for Film SoundtrackabstractFilm sound reproduction is the process of converting the image-form film soundtrack to wave-form movie sound. In this paper, a novel optical imaging based reproduction framework is proposed with the basic idea that restoring film audio damage in the image domain. In traditional reproduction method, the scanning light emitted by film projector causes inversible physical damage to the flammable film soundtrack (made of Nitrate compounds). By using optical imaging method in film soundtrack capturing, our framework can avoid the damage and the self-ignition problem. Experiment results show that our framework can improve the reproduction speed to 2 times while maintaining equal sound quality. Also, the sound sampling rate can be enhanced to 162.08%. Long Ye, Qin Zhang 0009 |
ACM Multimedia | 3 |
| 2021 | A two-stage complex network using cycle-consistent generative adversarial networks for speech enhancement
Guochen Yu, Hui Wang 0070, Qin Zhang 0009, Chengshi Zheng |
Speech Commun. | 4 |
| 2020 | Global Affective Video Content Regression Based on Complementary Audio-Visual Features
Xiaona Guo, Wei Zhong 0001, Long Ye, Yan Heng, Qin Zhang 0009 |
MMM (2) | 6 |
| 2019 | Semantic based autoencoder-attention 3D reconstruction network
Long Ye, Wei Zhong 0001, Tie Yun, Qin Zhang 0009 |
Graph. Model. | 6 |
| 2019 | Target Disassembly Sequencing and Scheme Evaluation for CNC Machine Tools Using Improved Multiobjective Ant Colony Algorithm and Fuzzy IntegralabstractDisassembly planning aims to perform the optimal disassembly sequence given a used or obsolete product in terms of cost and environmental impact. This paper presents a new multiobjective programming model for the target disassembly sequencing. It proposes an improved multiobjective ant colony algorithm to derive optimal target disassembly sequences. This work also establishes some indices on disassembly scheme evaluation and a fuzzy integral method to evaluate the obtained disassembly scheme. A CNC machine tool example is given to illustrate the proposed models and the effectiveness of the proposed algorithm. Both theoretical and simulation results demonstrate that the proposed approach can perform the quantitative analysis of a disassembly process effectively. Such results can help decision makers select the best plans and sequences when executing a disassembly process of a product. Yixiong Feng, MengChu Zhou, Guangdong Tian, Zhiwu Li 0001, Qin Zhang 0009, Jianrong Tan |
IEEE Trans. Syst. Man Cybern. Syst. | 6 |
| 2018 | Design of linear-phase nonsubsampled nonuniform directional filter bank with arbitrary directional partitioning
Long Ye, Tie Yun, Wei Zhong 0001, Qin Zhang 0009 |
J. Vis. Commun. Image Represent. | 5 |
| 2017 | Online Multi-threshold Learning with Imbalanced Data Stream
Xufen Cai, Long Ye, Qin Zhang 0009 |
ISNN (1) | 6 |
| 2017 | A novel image compression framework at edgesabstractFor the new developed area of edge computing, the traditional image coding schemes always have poor adaptability on the requirements of high-compression and low-delay coding, due to the fact that they do not consider the correlations between encoding image and external images. Aiming at this problem, we propose a novel image coding framework based on the content similarity analysis. Unlike the current state-of-the-art systems, our method compresses all the images stored in the edge computing terminal as a whole. The images are represented by their pHash values and classified into many groups, then each group is compressed in a quasi inter-prediction coding way. With our coding scheme, the redundancies among the similar images could be removed to achieve more coding gain. Experimental results demonstrate the superiorities of our method in the edge computing environment compared with current static image coding solutions. Long Ye, Qianhan Liu, Wei Zhong 0001, Qin Zhang 0009 |
VCIP | 4 |
| 2016 | Speech enhancement using magnitude and phase spectrum compensationabstractBackground noise is a severe problem in speech related systems. In order to solve this problem, it is important to eliminate the noise from the noisy speech, which is called speech enhancement. Typical speech enhancement algorithms only operate on the short-time magnitude spectrum, while keeping the short-time phase spectrum unchanged for synthesis. Or only compensate the phase spectrum while keeping the magnitude spectrum unchanged. In this paper, we present a novel method by changing both magnitude and phase spectra to produce a modified complex spectrum. The test of an objective speech quality measure PESQ, and spectrogram analysis had showed that the proposed method can obtain better enhancement performance. Wenjin Wu, Qin Zhang 0009, Shilei Bai |
ICIS | 3 |
| 2016 | Design of M-channel linear-phase non-uniform filter banks with arbitrary rational sampling factorsabstractThe majority of the existing work on designing non‐uniform filter banks (NUFBs) cannot realise arbitrary rational frequency partitioning, due to the ineliminable large aliasing caused by the non‐feasible permutation of rational sampling factors. In this study, the authors extend the efficient modulation technique to non‐feasible partitioned NUFBs and then generalise the phase modification structure to the case of rational decimators to achieve the linear‐phase (LP) property. Through a step‐by‐step analysis on the aliasing elimination of both feasible and non‐feasible subbands, it is proved that when the symmetry of filters and phase modification factors are chosen to meet the derived matching conditions, the non‐LP characteristic of shifted filters can be transferred to LP one and further the non‐feasible partitioned NUFB is obtained with highly desired LP property. Compared with the existing typical NUFB designs, the proposed approach can achieve comparable performance with much lower system delay and implementation complexity. Wei Zhong 0001, Qin Zhang 0009 |
IET Signal Process. | 3 |
| 2013 | Stereo matching using belief propagation with spatiotemporal consistencyabstractIn this paper, we propose a stereo matching approach using belief propagation for video disparity estimation by establishing a novel spatiotemporal belief propagation model. The proposed model extends 2D belief propagation algorithm to 3D mode by propagating the belief of preceding frame to the following frame. Additionally, the propagating messages of the preceding frame are translated through referring to motion vector and then used as the initial values of message for the current frame. Meanwhile, the consistency of the motion vector is incorporated to the smoothness constraint for the current frame. The proposed spatiotemporal model of belief propagation has more systematic and comprehensive combination of temporal correlation compared to previous works. The experimental results show that it outperforms the algorithms based on 2D belief propagation especially for non-deformation motion in middle-low speed. Yingyun Yang, Xie Song, Qin Zhang 0009 |
ICMV | 3 |
| 2012 | Distributed Markov Chain Monte Carlo kernel based particle filtering for object tracking
Danling Wang, Qin Zhang 0009, John Morris |
Multim. Tools Appl. | 2 |
| 2010 | Combined just noticeable difference model guided image watermarkingabstractPerceptual Watermarking should take full advantage of the results from human visual system (HVS) studies. Just noticeable difference (JND), which refers to the maximum difference that the HVS does not perceive, gives us a way to model the HVS accurately. In this paper, we exploit a combined JND model which represents additional accurate perceptual visibility threshold profile to guide watermarking for digital images. The proposed combined JND model guided watermarking scheme, where visual models are fully used to determine image dependent upper bounds on watermark insertion, allows us to provide the maximum strength transparent watermark. Experimental results confirm the improved performance of our combined JND model. Our combined JND model is capable of yielding higher injected-watermark energy without introducing noticeable difference to the original image and outperforms the relevant existing visual models. Robustness results show the proposed JND model guided watermarking scheme performs much better than other algorithms based on Watson's perceptual model. Yaqing Niu, Sridhar Krishnan 0001, Qin Zhang 0009 |
ICME | 4 |
| 2008 | Image Restoration Using Piecewise Iterative Curve Fitting and Texture Synthesis
Yingyun Yang, Long Ye, Qin Zhang 0009 |
ICIC (1) | 4 |
| 2008 | Use hierarchical genetic particle filter to figure articulated human trackingabstractUsing particle filter to track human movement, a key problem is how to draw samples in high-dimensional state space. In this paper, we present a novel framework of particle filtering, namely Hierarchical Genetic Particle Filter (HGPF), to improve the efficiency of samples by a hierarchical evolutionary detection. As a result, we can obtain reasonably distributed samples thus translating into reliable tracking performance. Finally, we apply the technique to 2D articulated human movement tracking. Result demonstrates the effectiveness of HGPF in solving the tracking problem like self-occlusion and cluttered background. Long Ye, Qin Zhang 0009, Ling Guan |
ICME | 2 |