EDBT 2026 Demo / reviewers in the wild / expert
Long Ye
dblp:86/2624
· DBLP profile ↗
65ranked-venue papers
3as first author
55since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 41 · 3 first-author · 35 since 2021Artificial intelligence and machine learning · 24 · 22 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Detect All-Type Deepfake Audio: Wavelet Prompt Tuning for Enhanced Auditory PerceptionabstractThe rapid advancement of audio generation technologies has escalated the risks of malicious deepfake audio across speech, sound, singing voice, and music, threatening multimedia security and trust. While existing countermeasures (CMs) perform well in single-type audio deepfake detection (ADD), their performance declines in cross-type scenarios. This paper is dedicated to studying the all-type ADD task. We are the first to comprehensively establish an all-type ADD benchmark to evaluate current CMs, incorporating cross-type deepfake detection across speech, sound, singing voice, and music. Then, we introduce the prompt tuning self-supervised learning (PT-SSL) training paradigm, which optimizes SSL front-end by learning specialized prompt tokens for ADD, requiring 458× fewer trainable parameters than fine-tuning (FT). Considering the auditory perception of different audio types, we propose the wavelet prompt tuning (WPT)-SSL method to capture type-invariant auditory deepfake information from the frequency domain without requiring additional training parameters, thereby enhancing performance over FT in the all-type ADD task. To achieve an universally CM, we utilize all types of deepfake audio for co-training. Experimental results demonstrate that WPT-XLSR-AASIST achieved the best performance, with an average EER of 3.58% across all evaluation sets. Yuankun Xie, Ruibo Fu, Songjun Cao, Haonan Cheng, Long Ye |
AAAI | 8 |
| 2026 | PSRNet: Progressive Semantic Refinement for Human Parsing via Text Conditioning and Embedding-Based CalibrationabstractHuman parsing requires fine-grained, pixel-level delineation of human parts and accessories, yet visually correlated categories and long-tailed parts often cause semantic confusion and boundary ambiguity. We propose PSRNet, a Progressive Semantic Refinement Network that exploits fixed class semantic embeddings derived from category text to refine parsing in a coarse-to-fine manner. First, a Text-Conditioned Feature Modulation (TCFM) module injects class semantics to modulate encoder features, enhancing low-level discriminability for confusing parts. Second, a Semantic-Embedding Calibration and Fusion (SECF) module combines a conventional linear classifier with an embedding-similarity head to calibrate category logits, effectively reducing misclassification among semantically close classes. Third, we introduce a Morphology-Guided Boundary Refinement loss (MGBR) that constructs boundary supervision via dilation–erosion operations on predictions and ground truth, encouraging sharper and more consistent part boundaries. Extensive experiments on two widely used benchmarks, LIP and CIHP, demonstrate that PSRNet consistently improves both parsing accuracy and boundary quality over strong baselines, with particularly notable gains on confusing part pairs. Xingxing Xiang, Long Ye, Lei Zhang 0218, Zhaoxin Fan |
ICMR | 4 |
| 2026 | Differential impact of evaluative vs. non-evaluative AI feedback on employee safety performance and affective wellbeing
Chaorui Shen, Long Ye |
Decis. Support Syst. | 2 |
| 2026 | A condition-driven hybrid GRU-GMCVAE framework for dynamic anomaly detection: Application to industrial petrochemical processes
Qiulei Xue, Long Ye, Yanxia Xu |
Expert Syst. Appl. | 3 |
| 2026 | Interpretable semi-supervised 3D deep anomaly detection for surface defect localization in wire-laser directed energy deposition using point cloudsabstractWire-laser directed energy deposition (WL-DED) enables the fabrication of large-scale metallic components but frequently suffers from process-induced surface defects that hinder part quality and increase post-processing costs. Automated inspection is challenging because defects are diverse and rare, making large labelled datasets impractical. This paper proposes an interpretable semi-supervised framework for surface detection and localization on WL-DED components using high-density 3D point clouds acquired by laser scanning. The workflow includes point-cloud preprocessing, patch-based segmentation, voxelization, and semi-supervised representation learning of defect-free surface morphology. Two 3D deep autoencoder models, i.e., a convolutional autoencoder (CAE) and a variational autoencoder (VAE), are trained exclusively on normal patches and detect anomalies through voxel-wise reconstruction errors. Defects are localized by mapping reconstruction-error heatmaps back onto the original surface, enabling quantitative visualization of defect severity. Experimental results on WL-DED thin-wall samples show that the optimized CAE achieves 86.09% precision, while the VAE reaches 86.43% precision with improved defect localization (mIoU up to 0.7234). Activation-map analysis provides interpretability by highlighting geometric regions that drive anomaly responses. A hyperparameter study demonstrates that lower voxel resolutions and smaller patch sizes improve robustness and reduce false positives. The proposed framework generalizes to more complex multi-bead, multi-layer structures with minimal retraining, supporting practical deployment for intelligent inspection and decision-making in additive manufacturing quality assurance. Long Ye |
Expert Syst. Appl. | 3 |
| 2026 | OpenST: Toward open-set source tracing for neural codec deepfake audio
Yuankun Xie, Ruibo Fu, Zhengqi Wen, Songjun Cao, Chenxing Li, Haonan Cheng, Long Ye |
Neurocomputing | 10 |
| 2026 | Audio-visual perceptual quality measurement via multi-perspective spatio-temporal EEG analysis
Shuzhan Hu, Weiwei Jiang 0003, Bingrui Geng, Wei Zhong 0001, Long Ye |
Pattern Recognit. | 7 |
| 2026 | Visual Label Augmentation-Driven Multimodal Emotion Recognition
Qinglan Wei, Yaqi Zhou, Junzhe Zhou, Long Ye |
Pattern Recognit. | 4 |
| 2026 | UNQA: Unified No-Reference Quality Assessment for Audio, Image, Video, and Audio-Visual ContentabstractAs multimedia data flourishes on the Internet, quality assessment (QA) of multimedia data becomes paramount for digital media applications. Since multimedia data includes multiple modalities including audio, image, video, and audio-visual (A/V) content, researchers have developed a range of QA methods to evaluate the quality of different modality data. While they exclusively focus on addressing the single modality QA issues, a unified QA model that can handle diverse media across multiple modalities is still missing, whereas the latter can better resemble human perception behaviour and also have a wider range of applications. In this paper, we propose the Unified No-reference Quality Assessment model (UNQA) for audio, image, video, and A/V content, which tries to train a single QA model across different media modalities. To tackle the issue of inconsistent quality scales among different QA databases, we develop a multi-modality strategy to jointly train UNQA on multiple QA databases. Based on the input modality, UNQA selectively extracts the spatial features, motion features, and audio features, and calculates a final quality score via the four corresponding modality regression modules. Compared with existing QA methods, UNQA has two advantages: 1) the multi-modality training strategy makes the QA model learn more general and robust quality-aware feature representation as evidenced by the superior performance of UNQA compared to state-of-the-art QA methods. 2) UNQA reduces the number of models required to assess multimedia data across different modalities. and is friendly to deploy to practical applications. Code are available at https://github.com/charlotte9524/UNQA. Yuqin Cao, Xiongkuo Min, Wei Sun 0029, Long Ye, Weisi Lin, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | RD-VTA: Rule-Data Guided Video-to-Audio Generation for Fine-Grained Footstep SoundabstractIt is challenging to implement visually guided fine-grained footstep sounds based on a limited number of samples in complex scenes. This is due to the interference of redundant information in the complex background of the visual scene for audiovisual mapping. As well as the complex coupling of sound features makes audiovisual fine-grained linear mapping difficult. To address the mentioned problems, we propose an automated video-to-audio generation method (RD-VTA) for footstep sound that incorporates data-driven and rule-based modelling approaches. First, we design a data-driven masked footstep sound generation network (DM-AFSG) to acquire audiovisual temporal concordance. The network is capable of separating visual sound objects, reducing background redundant interference, and generating initial target sounds that capture temporal cues. Secondly, a rule-based fine-grained footstep sound adjustment method (RT-AFSG) is designed based on visual guides such as material, motion type and displacement distance. The proposed RT-AFSG effectively achieves diverse sounds with a limited number of sound samples through sound texture analysis and modification. Moreover, it constructs the mapping relationship between different visual cues and footstep sounds, and realizes the fine variation of footstep sounds. To adequately validate the effectiveness of the method in terms of audiovisual temporal consistency and content granularity, we perform objective synchronization metrics and subjective human evaluation on the footsteps audiovisual dataset VAFoot. The experimental results show that the method obtains an average of 5% improvement in sound synchronization performance and significantly outperforms several existing methods in terms of sound content granularity. We encourage readers to watch and listen to the footstep sound results on our demo website:https://quinntt.github.io/RD-VTA/. Qiutang Qi, Haonan Cheng, Hengyan Huang, Long Ye, Shaobin Li |
IEEE Trans. Multim. | 4 |
| 2026 | VGL-DPO: Vision-Guided Lexical Direct Preference Optimization for Mitigating Hallucination in Multimodal Large Language ModelsabstractMultimodal large language models (MLLMs) have achieved significant advancements in multimodal understanding, reasoning, and interaction. However, they still suffer from hallucination, where the generated text often deviates from the factual content of the input image. To mitigate this issue, prior studies have primarily employed direct preference optimization (DPO) for human preference alignment. However, these approaches treat all textual words equally, neglecting the varying significance of individual words in grounding text generation to image content. This limitation hinders fine-grained semantic alignment and consequently constrains their effectiveness in hallucination suppression. To address this limitation, we propose a vision-guided lexical DPO method, called VGL-DPO. Specifically, we quantify the significance of words in positive preference data based on their relevance to the visual input and dynamically assign different weights to different words during training. This facilitates more precise optimization by emphasizing critical words that contribute to factual grounding. Additionally, we leverage the importance differences between high-significance words in positive and negative preference data to adaptively adjust the weight of the negative preference loss. This dynamic reweighting mechanism further refines the model’s ability to suppress hallucinated content while reinforcing factual accuracy. Extensive experiments across various models demonstrate that our method outperforms existing state-of-the-art methods in reducing hallucination and enhancing factual accuracy. Siyuan Li 0001, Feng Wang 0063, Simeng Qin, Ranjie Duan, Haonan Cheng, Long Ye |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2025 | Depth-Guided Bundle Sampling for Efficient Generalizable Neural Radiance Field ReconstructionabstractRecent advancements in generalizable novel view synthesis have achieved impressive quality through interpolation between nearby views. However, rendering highresolution images remains computationally intensive due to the need for dense sampling of all rays. Recognizing that natural scenes are typically piecewise smooth and sampling all rays is often redundant, we propose a novel depth- guided bundle sampling strategy to accelerate rendering. By grouping adjacent rays into a bundle and sampling them collectively, a shared representation is generated for decoding all rays within the bundle. To further optimize efficiency, our adaptive sampling strategy dynamically allocates samples based on depth confidence, concentrating more samples in complex regions while reducing them in smoother areas. When applied to ENeRF, our method achieves up to a 1.27 dB PSNR improvement and a 47% increase in FPS on the DTU dataset. Extensive experiments on synthetic and real-world datasets demonstrate state-of-the-art rendering quality and up to 2× faster rendering compared to existing generalizable methods. Code is available at https://github.com/KLMAV-CUC/GDB-NeRF. Longlong Chen, Long Ye, Zhan Ma 0001 |
CVPR | 5 |
| 2025 | GoLF-NRT: Integrating Global Context and Local Geometry for Few-Shot View SynthesisabstractNeural Radiance Fields (NeRF) have transformed novel view synthesis by modeling scene-specific volumetric representations directly from images. While generalizable NeRF models can generate novel views across unknown scenes by learning latent ray representations, their performance heavily depends on a large number of multi-view observations. However, with limited input views, these methods experience significant degradation in rendering quality. To address this limitation, we propose GoLF-NRT: a Global and Local feature Fusion-based Neural Rendering Transformer. GoLF-NRT enhances generalizable neural rendering from few input views by leveraging a 3D transformer with efficient sparse attention to capture global scene context. In parallel, it integrates local geometric features extracted along the epipolar line, enabling high-quality scene reconstruction from as few as 1 to 3 input views. Furthermore, we introduce an adaptive sampling strategy based on attention weights and kernel regression, improving the accuracy of transformer-based neural rendering. Extensive experiments on public datasets show that GoLF-NRT achieves state-of-the-art performance across varying numbers of input views, highlighting the effectiveness and superiority of our approach. Code is available at https://github.com/KLMAV-CUC/GoLF-NRT. Long Ye, Zhan Ma 0001 |
CVPR | 5 |
| 2025 | Deciphering the Visual Style of China's Hit Short Videos Through Computer Vision
Qinglan Wei, Shenlian Xiang, Chen Zhang 0049, Long Ye |
ICIG (2) | 5 |
| 2025 | MixLGN: Mixed Local-Global Network for 3D Human Pose GenerationabstractAccurate 3D human pose generation plays a vital role in human-centric AI tasks. Although condition-guided (e.g. 2D pose) solutions have obtained suitable 3D human poses, the performance is insufficient while facing complex scenarios, e.g. occlusion or complicated actions. In this paper, we propose a novel Mixed Local-Global Network (MixLGN) that takes both intactness and plausibility into account for 3D human pose generation. Considering the local connectivity of adjacent keypoints, a novel dual convolution module is introduced to enhance the local ability for better spatial feature extraction. To guarantee the intactness of the 3D pose, we propose a part-to-body integrating mechanism that fuses local and global features together by taking guidance from human body hierarchy properties. Specifically, a part-aware pose feature encoding module is employed to learn structure-aware representation for each body part, and then put them together for global localization. We have verified the proposed MixLGN on two popular benchmark datasets, i.e., Human3.6M and MPI-INF-3DHP. The experimental results show that our model achieves the best performance both on single-hypothesis and multi-hypothesis. Sanyi Zhang, Chixuan Wei, Yinghao Yang 0002, Long Ye |
ICME | 5 |
| 2025 | GE-Talker: Generalizable and Efficient Neural Rendering for Talking Head GenerationabstractTalking head generation aims to create videos that preserve a source character’s identity while replicating synchronized lip movements, facial expressions, and head gestures from audio inputs. While personalized methods have achieved impressive results by training neural radiance fields (NeRFs) for specific identities, they often face limitations in generalization and incur high computational costs due to identity-specific retraining. We propose GE-Talker, a Generalizable and Efficient NeRF for talking head generation, which addresses these challenges through two key innovations. First, we leverage FLAME-based full-head modeling as intermediate representations, conditioned on audio features, to achieve precise lip synchronization and natural facial movements. Second, we introduce a semantic-aware depth-guided sampling strategy that uses FLAME-generated depth maps to restrict sampling ranges and semantic segmentation to focus on human regions, improving rendering quality and efficiency. GE-Talker achieves high-quality outputs for unseen speaker identities with a 109% speed-up over baseline methods and enables rapid fine-tuning on new speakers within 14 minutes, establishing it as a powerful and adaptable solution for talking head generation. Long Ye |
ICME | 4 |
| 2025 | Pop-Diffuseq: Controllable Symbolic Music Multi-Instrument Infilling and Accompaniment Generation with Long-Axis AttentionabstractControllability is a major challenge in music infilling and accompaniment tasks. Solutions based on transformer decoders have been widely adopted, while data-driven approaches with full self-attention result in high costs and unsatisfied outcomes for fine-grained control. Existing diffusion methods rely on trained classifiers, unconditional frameworks, or solo track, etc. To address these issues, we explore novel methods to enhance the controllability and quality of music model while reducing computational complexity. Firstly, we improve the classifier-free diffusion for multi-instrumental pop music. Secondly, we design a long-axis attention algorithm that combines long with axial attention to acquire the feature correlations of multi-dimensional attributes. Additionally, we contribute a pop band dataset with melody, style and mood labels handcrafted by musicians. After experiments on the benchmark dataset, our method demonstrates high-quality controllable results and outperforms existing state-of-the-art models. The GPU memory of our model is 26.9% lower than Diffuseq under the same hyperparameters. Haonan Cheng, Long Ye, Qin Zhang 0009 |
ICME | 3 |
| 2025 | FG-Midiformer: A Symbolic Music Understanding Model towards Fine-Grained Learning of Multi-Attributes
Haonan Cheng, Hengyan Huang, Long Ye |
ACM Multimedia | 4 |
| 2025 | Event Chain-Driven Communication Strategy Generation for News Videos
Qinglan Wei, Ruiqi Xue, Mingyue Liao, Long Ye |
ACM Multimedia | 4 |
| 2025 | Domain knowledge integrated CAM system based on multi-objective path optimal planning and deep convolutional neural network
Kangsen Li, Long Ye, Feng Gong |
Expert Syst. Appl. | 4 |
| 2025 | Full-Body Pose Motion Tracking From Sparse Data via Morphology-Aware Constraints
Yinghao Yang 0002, Sanyi Zhang, Chixuan Wei, Chenxi Feng, Long Ye |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2025 | LCIQA: A Lightweight Contrastive-Learning Framework for Image Quality Assessment via Cross-Scale Consistency MinimizationabstractBlind image quality assessment (BIQA), which functions without the need for a reference image, is a challenging yet essential task in various image processing systems and downstream vision applications, ranging from semantic recognition to image enhancement. Traditionally, numerous BIQA models have been developed using supervised learning methodologies, which rely heavily on the availability and quality of ground truth data. To improve the generalization capability and robustness of these models, recent studies have explored the application of contrastive learning, aiming to enhance the quality representation capacity of model backbones through a self-supervised approach. However, the training process for contrastive learning is computationally intensive, posing significant challenges in resource-constrained environments. To mitigate this issue, we propose a Lightweight Contrastive-learning-based IQA (LCIQA) framework, designed to be efficiently trained on a single GPU without relying on ground truth data. This framework maintains a fixed vision backbone and focuses on optimizing the parameters of subsequent IQA heads through contrastive learning. To accommodate a lightweight framework, we incorporate a quality task adapter to eliminate semantic biases introduced by the features extracted from the fixed-parameter backbone. A coarse-to-fine contrastive learning strategy is then employed to train the quality regression module. Extensive experiments demonstrate the superior performance of our model in terms of both accuracy and complexity. In addition, ablation studies validate the effectiveness of each component within the proposed framework. Chenxi Feng, Xiongkuo Min, Long Ye, Yinghao Yang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Point Evolution Hierarchy Network for Weak Single-Point Human ParsingabstractSparsely single-point human parsing aims at segmenting the human body into fine-grained categories via weak point-level labels (e.g., point-level, scribble-level, or image-level, etc). The point-level label, especially single-point supervision, can simultaneously preserve spatial positions as well as take light annotation time, which is particularly advantageous in alleviating the human labeling burden. However, how to obtain satisfactory parsing performance under limited sparse point annotations is challenging, which requires further investigation. In this paper, we propose a novel end-to-end Point Evolution Hierarchy human parsing Network (PEHNet) for fine-grained human parsing task that just leverages single-point supervision. Motivated by the concept of a divide-and-conquer strategy, we partition all pixels into three distinct groups, i.e., single-point labels, pseudo-region labels, and unlabeled pixels, then optimize each group with suitable mechanisms. To expand the coverage of single-point labels, we introduce a point dissemination module that generates high-quality pseudo-region labels. Furthermore, the point-level spatial position information inherently preserves the structural characteristics of the human body. Inspired by this hierarchical property, we devise a point-level human hierarchy-wise constraint that guides the prediction probabilities to align with the inherent hierarchy of the human body. Experimental results demonstrate that the proposed PEHNet outperforms state-of-the-art parsing methods on two popular human parsing benchmark datasets (LIP and ATR) and one semantic segmentation dataset (Pascal VOC 2012). Sanyi Zhang, Xiaochun Cao, Long Ye, Zhanjie Song, Guo-Jun Qi, Jie Zhou 0001 |
IEEE Trans. Image Process. | 3 |
| 2024 | DNIT: Enhancing Day-Night Image-to-Image Translation through Fine-Grained Feature Handling (Student Abstract)abstractExisting image-to-image translation methods perform less satisfactorily in the "day-night" domain due to insufficient scene feature study. To address this problem, we propose DNIT, which performs fine-grained handling of features by a nighttime image preprocessing (NIP) module and an edge fusion detection (EFD) module. The NIP module enhances brightness while minimizing noise, facilitating extracting content and style features. Meanwhile, the EFD module utilizes two types of edge images as additional constraints to optimize the generator. Experimental results show that we can generate more realistic and higher-quality images compared to other methods, proving the effectiveness of our DNIT. Haonan Cheng, Long Ye |
AAAI | 3 |
| 2024 | Multivariate Time-Series Representation Learning for Continuous Medical DiagnosisabstractMultivariate time series (MTS) data in electronic health records (EHR) pose unique challenges due to their sparsity and irregular time intervals. Existing methods often tend towards imputation or isolated encoding, resulting in suboptimal representation learning. Moreover, most existing research solely focuses on single-shot diagnosis, neglecting the importance of continuous diagnosis, particularly for critically ill patients. Continuous diagnosis provides significant opportunities for timely intervention and rational resource allocation. To address these challenges, we propose an innovative multivariate time-series representation learning for continuous medical diagnosis. Specifically, we first address sparsity issues by combining feature names and record values encoded by time. Then, we utilize a transformer variant with gated units to extract contextual features. Additionally, we introduce Time Update Block, a component that combines the strengths of long short-term memory and attention mechanisms, aimed at improving the model's ability for continuous diagnosis. Based on extensive experimental evaluations on real-world medical datasets, we demonstrate the superior performance of the proposed method. Xiongjun Zhao, Linzhuang Zou, Long Ye, Shaoliang Peng |
BIBM | 3 |
| 2024 | Binauralmusic: A Diverse Dataset for Improving Cross-Modal Binaural Audio GenerationabstractCross-modal binaural audio generation is an important task and has broad applications such as game sound development and auditory assistance for the visually impaired. However, existing datasets lack binaural samples with abundant visual venues. As a consequence, state-of-the-art cross-modal binaural audio generation methods have weak generalization. To support research on building robust binaural audio generation, we construct BinauralMusic dataset consisting of 5,462 performance video clips with binaural audio from 9 musical instrument categories. The performance venues involve indoor closed places such as shopping mall, hotel, bedroom, as well as outdoor open areas such as field, garden and seashore. Experiments show that the performance of the cross-modal binaural audio generation model can be significantly improved by 10.62% by using the BinauralMusic dataset as training material. Moreover, different from previous datasets, the BinauralMusic dataset can also support other audio-visual cross-modal learning tasks, including visually guided sound source localization and separation. Haonan Cheng, Long Ye |
ICASSP | 4 |
| 2024 | An Efficient Temporary Deepfake Location Approach Based Embeddings for Partially Spoofed Audio DetectionabstractPartially spoofed audio detection is a challenging task, lying in the need to accurately locate the authenticity of audio at the frame level. To address this issue, we propose a fine-grained partially spoofed audio detection method, namely Temporal Deepfake Location (TDL), which can effectively capture information of both features and locations. Specifically, our approach involves two novel parts: embedding similarity module and temporal convolution operation. To enhance the identification between the real and fake features, the embedding similarity module is designed to generate an embedding space that can separate the real frames from fake frames. To effectively concentrate on the position information, temporal convolution operation is proposed to calculate the frame-specific similarities among neighboring frames, and dynamically select informative neighbors to convolution. Extensive experiments show that our method outperform baseline models in ASVspoof2019 Partial Spoof dataset and demonstrate superior performance even in the cross-dataset scenario. Yuankun Xie, Haonan Cheng, Long Ye |
ICASSP | 4 |
| 2024 | FSD: An Initial Chinese Dataset for Fake Song DetectionabstractSinging voice synthesis and singing voice conversion have significantly advanced, revolutionizing musical experiences. However, the rise of "Deepfake Songs" generated by these technologies raises concerns about authenticity. Unlike Audio DeepFake Detection (ADD), the field of song deepfake detection lacks specialized datasets or methods for song authenticity verification. In this paper, we initially construct a Chinese Fake Song Detection (FSD) dataset to investigate the field of song deepfake detection. The fake songs in the FSD dataset are generated by five state-of-the-art singing voice synthesis and singing voice conversion methods. Our initial experiments on FSD revealed the ineffectiveness of existing speech-trained ADD models for the task of song deepfake detection. Thus, we employ the FSD dataset for the training of ADD models. We subsequently evaluate these models under two scenarios: one with the original songs and another with separated vocal tracks. Experiment results show that song-trained ADD models exhibit a 38.58% reduction in average equal error rate compared to speech-trained ADD models on the FSD test set. Yuankun Xie, Xiaolin Lu, Zhenghao Jiang, Haonan Cheng, Long Ye |
ICASSP | 7 |
| 2024 | HMDST: A Hybrid Model-Data Driven Approach for Spatio-Temporally Consistent Video InpaintingabstractVideo inpainting fills in the missing regions in videos, which should be coherent and natural. Conventional model-driven methods with manual priors may produce blurry, distorted or inconsistent results due to the absence of high-level semantics. While data-driven approaches like deep learning directly learn mappings from observations to target videos, they may face challenges in quality, generalization, and robustness. Existing methods use optical flow to model the motion and context between frames, but the flow estimation may be inaccurate or unstable in the missing regions. This paper introduces a hybrid model-data driven approach for spatio-temporally consistent video inpainting. It combines prior knowledge of imaging and deep prior learned through training, and employs elaborately designed modules to accurately model the motion and contextual information. Our network can be trained end-to-end, leading to a more efficient and effective inpainting process. Extensive experiments demonstrate the superiority of our method qualitatively and quantitatively. Kaijun Zou, Zhiye Chen, Long Ye |
ICME | 4 |
| 2024 | Generalized Source Tracing: Detecting Novel Audio Deepfake Algorithm with Real Emphasis and Fake Dispersion Strategy
Yuankun Xie, Ruibo Fu, Zhengqi Wen, Haonan Cheng, Long Ye, Jianhua Tao 0001 |
INTERSPEECH | 7 |
| 2024 | Coarse-to-Fine Domain Adaptation for Cross-Subject EEG Emotion Recognition with Contrastive Learning
Shuang Ran, Wei Zhong 0001, Long Ye, Qin Zhang 0009 |
PRCV (15) | 4 |
| 2024 | PoseVR: Structure-Aware Hybrid Full-Body Pose Estimation in Virtual Reality
Yinghao Yang 0002, Sanyi Zhang, Long Ye, Neng Rao |
PRCV (11) | 3 |
| 2024 | Some notes on the pan-integrals of set-valued functions
Tong Kang, Leifan Yan, Long Ye, Jun Li 0014 |
Fuzzy Sets Syst. | 3 |
| 2024 | Mind to Music: An EEG Signal-Driven Real-Time Emotional Music Generation SystemabstractMusic is an important way for emotion expression, and traditional manual composition requires a solid knowledge of music theory. It is needed to find a simple but accurate method to express personal emotions in music creation. In this paper, we propose and implement an EEG signal‐driven real‐time emotional music generation system for generating exclusive emotional music. To achieve real‐time emotion recognition, the proposed system can obtain the model suitable for a newcomer quickly through short‐time calibration. And then, both the recognized emotion state and music structure features are fed into the network as the conditional inputs to generate exclusive music which is consistent with the user’s real emotional expression. In the real‐time emotion recognition module, we propose an optimized style transfer mapping algorithm based on simplified parameter optimization and introduce the strategy of instance selection into the proposed method. The module can obtain and calibrate a suitable model for a new user in short‐time, which achieves the purpose of real‐time emotion recognition. The accuracies have been improved to 86.78% and 77.68%, and the computing time is just to 7 s and 10 s on the public SEED and self‐collected datasets, respectively. In the music generation module, we propose an emotional music generation network based on structure features and embed it into our system, which breaks the limitation of the existing systems by calling third‐party software and realizes the controllability of the consistency of generated music with the actual one in emotional expression. The experimental results show that the proposed system can generate fluent, complete, and exclusive music consistent with the user’s real‐time emotion recognition results. Shuang Ran, Wei Zhong 0001, Danting Duan, Long Ye, Qin Zhang 0009 |
Int. J. Intell. Syst. | 5 |
| 2024 | Deep generative network for image inpainting with gradient semantics and spatial-smooth attention
Ziqi Sheng, Cong Lin 0003, Wei Lu 0001, Long Ye |
J. Vis. Commun. Image Represent. | 5 |
| 2024 | DiffuseRoll: multi-track multi-attribute music generation based on diffusion model
Haonan Cheng, Long Ye |
Multim. Syst. | 4 |
| 2024 | MusicECAN: An Automatic Denoising Network for Music Recordings With Efficient Channel AttentionabstractIn this work, we address the long-standing problem of automatic recorded music denoising. In previous audio denoising research, the primary focus has been on speech, and music denoising works only considered noise types in indoor conversation scenarios or old gramophone recordings, neglecting the amateur music recording scenario. To this end, we first propose MusicECAN, an automatic music denoising method designed to filter out additional noise components in recorded music. The novel architecture comprises two key components, namely, a feature learning module and a noise filtering module, which can efficiently but effectively model, refine and denoise the noisy input. Specifically, in order to capture sufficient noisy music information, an ECA-U-SAM based feature learning module is designed by incorporating an efficient channel attention (ECA) mechanism into the traditional U-Net model with a supervised attention module (SAM). To train our MusicECAN, we collect M&N, a dataset containing various clean music and noise recordings. Through the combination of different clean-noise recording pairs, we can effectively simulate possible music performance environments with various background noise. Extensive quantitative and qualitative comparisons demonstrate that our MusicECAN outperforms the state-of-the-art audio denoising methods. Haonan Cheng, Zhicheng Lian, Long Ye, Qin Zhang 0009 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2024 | Domain Generalization via Aggregation and Separation for Audio Deepfake DetectionabstractIn this paper, we propose an Aggregation and Separation Domain Generalization (ASDG) method for Audio DeepFake Detection (ADD). Fake speech generated from different methods exhibits varied amplitude and frequency distributions rather than genuine speech. In addition, the spoofing attacks in training sets may not keep pace with the evolving diversity of real-world deepfake distributions. In light of this, we attempt to learn an ideal feature space that can aggregate real speech and separate fake speech to achieve better generalizability in the detection of unseen target domains. Specifically, we first propose a feature generator based on Lightweight Convolutional Neural Networks (LCNN), which is employed for generating a feature space and categorizing the feature into real and fake. Meanwhile, single-side domain adversarial learning is leveraged to make only the real speech from different domains indistinguishable, which enables the distribution of real speech to be aggregated in the feature space. Furthermore, a triplet loss is adopted to separate the distribution of fake speech while aggregating the distribution of real speech. Finally, in order to test the generalizability of the model, we train it with three different English datasets and evaluate in harsh conditions: cross-language and noisy datasets. The extensive experiments show that ASDG outperforms the baseline models in cross-domain tasks and decreases Equal Error Rate (EER) by up to 39.24% when compared to that of RawNet2. It is proved that the proposed Aggregation and Separation Domain Generalization method can be an effective strategy to improve the model generalizability. Yuankun Xie, Haonan Cheng, Long Ye |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | A Dataset and Benchmark for 3D Scene Plausibility AssessmentabstractThe surge in popularity of 3D scene synthesis has driven the development of diverse methods for assessing the quality of synthesized scenes. While subjective assessment methods are widespread, their time-consuming and labor-intensive nature prompts exploration into more efficient objective alternatives. This paper introduces an objective approach to evaluating scene plausibility, aiming to overcome the limitations associated with subjective methods. To underpin our objective evaluation, we present the 3D-SPAD dataset, comprising plausibility scores for 3000 scenes across 46 object categories. Leveraging this dataset, we propose a graph attention-based network designed to accurately estimate scene plausibility. A comprehensive evaluation of our network is conducted through a series of experiments, showcasing its feasibility and reliability. For additional details and access to the code, please refer to our GitHub repository athttps://github.com/Mayibo-cuc/3D-SPAN. Yibo Ma, Wei Zhong 0001, Long Ye, Xinyan Yang, Qin Zhang 0009 |
IEEE Trans. Multim. | 4 |
| 2024 | Disjoint Masking With Joint Distillation for Efficient Masked Image ModelingabstractMasked image modeling (MIM) has shown great promise for self-supervised learning (SSL) yet been criticized for learning inefficiency. We believe the insufficient utilization of training signals should be responsible. To alleviate this issue, we introduce a conceptually simple yet learning-efficient MIM training scheme, termedDisjointMasking withJointDistillation (DMJD). For disjoint masking (DM), we sequentially sample multiple masked views per image in a mini-batch with the disjoint regulation to raise the usage of tokens for reconstruction in each image while keeping the masking rate of each view. For joint distillation (JD), we adopt a dual branch architecture to respectively predict invisible (masked) and visible (unmasked) tokens with superior learning targets. Rooting in orthogonal perspectives for training efficiency improvement, DM and JD cooperatively accelerate the training convergence yet not sacrificing the model generalization ability. Concretely, DM can train ViT with less effective training epochs (at most$3.7\times$less time-consuming) to report competitive performance. With JD, our DMJD clearly improves the linear probing classification accuracy, up to 3.4$\%$. On fine-grained downstream tasks like semantic segmentation, object detection,etc., our DMJD also presents superior generalization compared with state-of-the-art SSL methods. Xin Ma 0019, Chang Liu 0047, Chunyu Xie, Long Ye, Yafeng Deng, Xiangyang Ji |
IEEE Trans. Multim. | 4 |
| 2024 | Cross-Modal Quantization for Co-Speech Gesture GenerationabstractLearning proper representations for speech and gesture is essential for co-speech gesture generation. Existing approaches either utilize direct representations or independently encode the speech and gesture, which neglect the joint representation to highlight the interplay between these two modalities. In this work, we propose a novel Cross-modal Quantization (CMQ) to jointly learn the quantized codes for speech and gesture together. Such representation highlights the speech-gesture interaction before actually learning the complex mapping, and thus better suits the intricate mapping between speech and gesture. Specifically, the Cross-modal Quantizer jointly encodes speech and gesture as discrete codebooks, enabling better cross-modal interaction. Cross-modal Predictor subsequently utilizes the learned codebooks to autoregressively predict the next-step gesture. With cross-modal quantization, our approach yields much higher codebook usage and generates more realistic and diverse gestures in practice. Extensive experiments are conducted on both 3D and 2D datasets as well as the subjective user study, demonstrating a clear performance gain compared to several baseline models in terms of audio-visual alignment and gesture diversity. In particular, our method demonstrates a three-fold improvement in diversity compared to baseline models, while simultaneously maintaining high motion fidelity. Zheng Wang 0059, Wei Zhang 0031, Long Ye, Dan Zeng 0001, Tao Mei 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | Interpretability Diversity for Decision-Tree-Initialized Dendritic Neuron Model EnsembleabstractTo construct a strong classifier ensemble, base classifiers should be accurate and diverse. However, there is no uniform standard for the definition and measurement of diversity. This work proposes a learners' interpretability diversity (LID) to measure the diversity of interpretable machine learners. It then proposes a LID-based classifier ensemble. Such an ensemble concept is novel because: 1) interpretability is used as an important basis for diversity measurement and 2) before its training, the difference between two interpretable base learners can be measured. To verify the proposed method's effectiveness, we choose a decision-tree-initialized dendritic neuron model (DDNM) as a base learner for ensemble design. We apply it to seven benchmark datasets. The results show that the DDNM ensemble combined with LID obtains superior performance in terms of accuracy and computational efficiency compared to some popular classifier ensembles. A random-forest-initialized dendritic neuron model (RDNM) combined with LID is an outstanding representative of the DDNM ensemble. Xudong Luo 0003, Long Ye, Xiaohao Wen, MengChu Zhou, Qin Zhang 0009 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Pruning of Dendritic Neuron Model with Significance Constraints for ClassificationabstractThe dendritic neural model (DNM) simulates the information-processing mechanisms and procedures of neurons by mimicking the nonlinearity of synapses in the human brain. This allows for a better understanding of biological nervous systems and has been used to a wonderful effect in various fields. However, there are some problems with the existing DNM, such as the high complexity and the limited generalization capability. Pruning is one of the most common and crucial approaches to model compression. It is a model optimization technique that involves removing redundant values from a weight tensor to develop smaller and more efficient neural networks. A compressed neural network can enable faster running and reduced computational costs in network training. To improve the model performance, this work proposes a DNM pruning method with dendrite layer significance constraints. Our proposed method not only calculates the significance of dendrite layers but also makes the significance of a few dendrite layers in the trained model concentrated in a few dendrite layers so that the dendrite layers with low significance can be pruned away. The results of the simulation experiment on classification problems show that our proposed method outperforms the existing pruning methods in terms of network size and generalization performance. Xudong Luo 0003, Long Ye, Xiaohao Wen, Qin Zhang 0009 |
IJCNN | 2 |
| 2023 | Learning A Self-Supervised Domain-Invariant Feature Representation for Generalized Audio Deepfake Detection
Yuankun Xie, Haonan Cheng, Long Ye |
INTERSPEECH | 4 |
| 2023 | RD-FGFS: A Rule-Data Hybrid Framework for Fine-Grained Footstep Sound Synthesis from Visual GuidanceabstractExisting methods are difficult to synthesize fine-grained footsteps based on video frames only. This is due to the complicated nonlinear mapping relationships between motion states, spatial locations and different footstep sounds. Aiming to address this issue, we propose a Rule-Data guided Fine-Grained Footstep Sound (RD-FGFS) synthesis method. To the best of our knowledge, our work takes the first step in integrating data-driven and rule modeling approaches for visually aligned footstep sound synthesis. Firstly, we design a learning-based footstep sound generation network (FSGN) architecture driven by pose and flow features. The FSGN is proposed for generating an initial target sound which captures timing cues. Secondly, a rule-based fine-grained footstep sound adjustment (FGFSA) method is designed based on the visual guidance, namely ground material, movement type, and displacement distance. The proposed FGFSA effectively constructs a mapping relationship between different visual cues and footstep sounds, enabling fine-grained variations of footstep sounds. Experimental results show that our method improves the visual and sound synchronization results of footsteps and achieves impressive performance in footstep sound fine-grained control. Qiutang Qi, Haonan Cheng, Yang Wang 0191, Long Ye, Shaobin Li |
ACM Multimedia | 4 |
| 2023 | Gender-Sensitive EEG Channel Selection for Emotion Recognition Using Enhanced Genetic AlgorithmabstractEEG channel selection aims to choose informative and representative channels to reduce data redundancy. It is very beneficial for improving the utility and efficiency of emotion recognition. Previous studies on EEG channel selection have not considered the influence of genders despite long-standing belief in gender differences with respect to emotion analysis. In this paper, we collected EEG signals from 20 subjects containing 10 males and 10 females by letting them watch short emotional videos. Then, to reduce data redundancy, we propose an enhanced genetic algorithm to select the optimal channel subsets separately for male and female subjects by incorporating a novel evolution operation. Experimental results show that the proposed algorithm achieves higher accuracy in terms of emotion recognition than several compared methods with a smaller channel subset. Besides, experimental results also indicate that the gender differences in neural patterns indeed exist. Through this study, the gender-sensitive channel selection offers a new avenue for further development of EEG based emotion recognition. Danting Duan, Qiang Yang 0008, Wei Zhong 0001, Long Ye, Qin Zhang 0009, Jun Zhang 0003 |
SMC | 5 |
| 2023 | EEG Feature Selection via Global Redundancy Minimization for Emotion RecognitionabstractA common drawback of EEG-based emotion recognition is that volume conduction effects of the human head introduce interchannel dependence and result in highly correlated information among most EEG features. These highly correlated EEG features cannot provide extra useful information, and they actually reduce the performance of emotion recognition. However, the existing feature selection methods, commonly used to remove redundant EEG features for emotion recognition, ignore the correlation between the EEG features or utilize a greedy strategy to evaluate the interdependence, which leads to the algorithms retaining the correlated and redundant features with similar feature scores in the EEG feature subset. To solve this problem, we propose a novel EEG feature selection method for emotion recognition, termed global redundancy minimization in orthogonal regression (GRMOR). GRMOR can effectively evaluate the dependence among all EEG features from a global view and then select a discriminative and nonredundant EEG feature subset for emotion recognition. To verify the performance of GRMOR, we utilized three EEG emotional data sets (DEAP, SEED, and HDED) with different numbers of channels (32, 62, and 128). The experimental results demonstrate that GRMOR is a promising tool for redundant feature removal and informative feature selection from highly correlated EEG features. Xueyuan Xu, Tianyuan Jia, Qing Li 0027, Fulin Wei, Long Ye, Xia Wu 0001 |
IEEE Trans. Affect. Comput. | 5 |
| 2022 | Global-Local Similarity Function for Automatic Playlist GenerationabstractThis paper proposes the Global-Local Similarity Function (GLSF) to exploit the multi-scale cues in track sequences for automatic playlist generation (APG). Unlike previous neighborhood-based methods only looking on local similarities for a given playlist, GLTS is constructed by first modeling the fine-grained audio features of each track, then capturing the long-term relations among consecutive tracks. Specifically, the fine-grained audio features are captured beat-by-beat to represent the rhythmic variation of music. The long- term relations are modeled by a designed track distance constraint (TD-constraint) to alleviate the incoherences and un- smooth transition in track sequences. The fine-grained audio features and TD-constraint are aggregated as the final GLSF by a simple distance function. Objective and subjective evaluations show that GLSF-based APG achieves better smooth transition and ensure the long-term content consistency among the tracks. Furthermore, GLSF yields a better understanding of the sequential relationship between tracks and propose a promising way to improve APG algorithms.1 Haonan Cheng, Ruyu Zhang, Long Ye |
ICME | 4 |
| 2022 | W-Hash: A Novel Word Hash Clustering Algorithm for Large-Scale Chinese Short Text Analysis
Yaofeng Chen, Long Ye, Xiaogang Peng, Meikang Qiu, Weipeng Cao |
KSEM (3) | 3 |
| 2022 | Multi-source Information-Shared Domain Adaptation for EEG Emotion Recognition
Wei Zhong 0001, Long Ye, Qin Zhang 0009 |
PRCV (2) | 4 |
| 2022 | Visually aligned sound generation via sound-producing motion parsing
Wei Zhong 0001, Long Ye, Qin Zhang 0009 |
Neurocomputing | 3 |
| 2021 | Few-shot object detection with anti-confusion groupingabstractRecent approaches have achieved excellent results on few-shot object detection. However, most detectors are easily confused by visually similar classes, leading to misclassification of interesting objects. In this work, we introduce an anti-confusion grouping mechanism for this problem. Our model can refine the results of the major multi-class classifier of the few-shot object detector with an anti-confusion module. Instead of maximizing the feature distribution distance of similar classes in the feature space, our approach uses additional auxiliary grouping module to distinguish similar classes on the same feature space as in base training phase. Concretely, the class groups are obtained according to the class visual similarity, and then they are utilized to train the auxiliary module. The main classifier, regressor and auxiliary anti-confusion module are end-to-end trained based on a multi-task loss. In the test phase, the auxiliary module is combined with the main classifier to provide the final classification result. Through extensive experiments, we demonstrate that our model outperforms well-established baselines for few-shot object detection. We also present analysis on various aspects of our model, aiming to provide some inspiration for future few-shot detection works. Long Ye |
ICMV | 1 |
| 2021 | MovieREP: A New Movie Reproduction Framework for Film SoundtrackabstractFilm sound reproduction is the process of converting the image-form film soundtrack to wave-form movie sound. In this paper, a novel optical imaging based reproduction framework is proposed with the basic idea that restoring film audio damage in the image domain. In traditional reproduction method, the scanning light emitted by film projector causes inversible physical damage to the flammable film soundtrack (made of Nitrate compounds). By using optical imaging method in film soundtrack capturing, our framework can avoid the damage and the self-ignition problem. Experiment results show that our framework can improve the reproduction speed to 2 times while maintaining equal sound quality. Also, the sound sampling rate can be enhanced to 162.08%. Long Ye, Qin Zhang 0009 |
ACM Multimedia | 2 |
| 2021 | Text to Scene: A System of Configurable 3D Indoor Scene SynthesisabstractIn this work, we show the Text to Scene system, which can configure 3D indoor scene from natural language. Given a text, the system will organize inclusive semantic message to a graph template, complete the graph with a novel graph-based contextual completion method Contextual ConvE(CConvE) and visulize the graph by arranging 3D models under an object location protocol. In the experiments, qualitative results obtained by the Text to Scene(T2S) system and quantitative evaluation of CConvE compared with other state-of-the-art approaches are reported. Xinyan Yang, Long Ye |
ACM Multimedia | 3 |
| 2021 | A hybrid deep-learning approach for complex biochemical named entity recognition
Lei Gao 0002, Sujie Guo, Long Ye, Qinghua Meng, Asef Nazari, Dhananjay R. Thiruvady |
Knowl. Based Syst. | 6 |
| 2020 | Global Affective Video Content Regression Based on Complementary Audio-Visual Features
Xiaona Guo, Wei Zhong 0001, Long Ye, Yan Heng, Qin Zhang 0009 |
MMM (2) | 3 |
| 2020 | Concurrent optimization of multiple base learners in neural network ensembles: An adaptive niching differential evolution approach
Ting Huang 0001, Danting Duan, Yue-Jiao Gong, Long Ye, Wing W. Y. Ng, Jun Zhang 0003 |
Neurocomputing | 4 |
| 2019 | M2-VISD: A Visual Intelligence Evaluation Dataset Based on Virtual SceneabstractMost of the existing visual intelligence evaluation datasets like ImageNet or ActivityNet share a common property, namely consisting of many images or videos captured from the real world. This property makes this evaluation datasets suitable for actual applications, but has limitations in achieving the diversity and interactivity of evaluation environments. The currently available datasets do not systematically provide data of different scales and different angles. In order to solve the above problems, this paper constructs a multi-angle and multi-scale data set based on UE4 platform, which can evaluate the performance of the algorithm from both scale and angle in an all-round way. Our experiments show that the algorithms detection performance varies greatly under different scale and different angle. Particularly, in scale-data, when the distance between the camera and the object is less than 50cm or greater than 3200cm, algorithm performance is poor, and in angle-data, the camera’s pitch angle is also poor when it is 18 degree and 0 degree. This result has a guiding role for the correct evaluation of the algorithm performance and provides help for the innovation and optimization of the algorithm. Yaning Tan, Xinyan Yang, Tianyi Feng, Jianbiao Wang, Long Ye |
ICIS | 6 |
| 2019 | Semantic based autoencoder-attention 3D reconstruction network
Long Ye, Wei Zhong 0001, Tie Yun, Qin Zhang 0009 |
Graph. Model. | 2 |
| 2018 | Design of linear-phase nonsubsampled nonuniform directional filter bank with arbitrary directional partitioning
Long Ye, Tie Yun, Wei Zhong 0001, Qin Zhang 0009 |
J. Vis. Commun. Image Represent. | 2 |
| 2018 | Identifying facial expression using adaptive sub-layer compensation based feature extraction
Xin Guo 0005, Tie Yun, Long Ye, Jinyao Yan |
J. Vis. Commun. Image Represent. | 3 |
| 2017 | Online Multi-threshold Learning with Imbalanced Data Stream
Xufen Cai, Long Ye, Qin Zhang 0009 |
ISNN (1) | 5 |
| 2017 | A novel image compression framework at edgesabstractFor the new developed area of edge computing, the traditional image coding schemes always have poor adaptability on the requirements of high-compression and low-delay coding, due to the fact that they do not consider the correlations between encoding image and external images. Aiming at this problem, we propose a novel image coding framework based on the content similarity analysis. Unlike the current state-of-the-art systems, our method compresses all the images stored in the edge computing terminal as a whole. The images are represented by their pHash values and classified into many groups, then each group is compressed in a quasi inter-prediction coding way. With our coding scheme, the redundancies among the similar images could be removed to achieve more coding gain. Experimental results demonstrate the superiorities of our method in the edge computing environment compared with current static image coding solutions. Long Ye, Qianhan Liu, Wei Zhong 0001, Qin Zhang 0009 |
VCIP | 1 |
| 2008 | Image Restoration Using Piecewise Iterative Curve Fitting and Texture Synthesis
Yingyun Yang, Long Ye, Qin Zhang 0009 |
ICIC (1) | 3 |
| 2008 | Use hierarchical genetic particle filter to figure articulated human trackingabstractUsing particle filter to track human movement, a key problem is how to draw samples in high-dimensional state space. In this paper, we present a novel framework of particle filtering, namely Hierarchical Genetic Particle Filter (HGPF), to improve the efficiency of samples by a hierarchical evolutionary detection. As a result, we can obtain reasonably distributed samples thus translating into reliable tracking performance. Finally, we apply the technique to 2D articulated human movement tracking. Result demonstrates the effectiveness of HGPF in solving the tracking problem like self-occlusion and cluttered background. Long Ye, Qin Zhang 0009, Ling Guan |
ICME | 1 |