Yuanzhe Chen

dblp:59/11438 · DBLP profile ↗
← Back
25ranked-venue papers
5as first author
10since 2021 · last 2025
0009-0000-1129-214XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 19 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Sound-VECaps: Improving Audio Generation with Visually Enhanced Captions
abstract
Generative models have shown significant achievements in audio generation tasks. However, existing models struggle with complex and detailed prompts, leading to potential performance degradation. We hypothesize that this problem stems from the simplicity and scarcity of the training data. This work aims to create a large-scale audio dataset with rich captions for improving audio generation models. We first develop an automated pipeline to generate detailed captions by transforming predicted visual captions, audio captions, and tagging labels into comprehensive descriptions using a Large Language Model (LLM). The resulting dataset, Sound-VECaps, comprises 1.66M high-quality audio-caption pairs with enriched details including audio event orders, occurred places and environment information. We then demonstrate that training the text-to-audio generation models with Sound-VECaps significantly improves the performance on complex prompts. Furthermore, we conduct ablation studies of the models on several downstream audio-language tasks, showing the potential of Sound-VECaps in advancing audio-text representation learning.Dataset and demos are available at https://yyua8222.github.io/Sound-VECaps-demo/.
Dongya Jia, Xiaobin Zhuang, Yuanzhe Chen, Zhuo Chen 0006, Yuping Wang 0005, Yuxuan Wang 0002, Xubo Liu 0001, Xiyuan Kang, Mark D. Plumbley, Wenwu Wang 0001
ICASSP4
2025 MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
abstract
We introduce MMAR, a new benchmark designed to evaluate the deep reasoning capabilities of Audio-Language Models (ALMs) across massive multi-disciplinary tasks. MMAR comprises 1,000 meticulously curated audio-question-answer triplets, collected from real-world internet videos and refined through iterative error corrections and quality checks to ensure high quality. Unlike existing benchmarks that are limited to specific domains of sound, music, or speech, MMAR extends them to a broad spectrum of real-world audio scenarios, including mixed-modality combinations of sound, music, and speech. Each question in MMAR is hierarchically categorized across four reasoning layers: Signal, Perception, Semantic, and Cultural, with additional sub-categories within each layer to reflect task diversity and complexity. To further foster research in this area, we annotate every question with a Chain-of-Thought (CoT) rationale to promote future advancements in audio reasoning. Each item in the benchmark demands multi-step deep reasoning beyond surface-level understanding. Moreover, a part of the questions requires graduate-level perceptual and domain-specific knowledge, elevating the benchmark's difficulty and depth. We evaluate MMAR using a broad set of models, including Large Audio-Language Models (LALMs), Large Audio Reasoning Models (LARMs), Omni Language Models (OLMs), Large Language Models (LLMs), and Large Reasoning Models (LRMs), with audio caption inputs. The performance of these models on MMAR highlights the benchmark's challenging nature, and our analysis further reveals critical limitations of understanding and reasoning capabilities among current models. These findings underscore the urgent need for greater research attention in audio-language reasoning, including both data and algorithm innovation. We hope MMAR will serve as a catalyst for future advances in this important but little-explored area.
Ziyang Ma 0001, Yinghao Ma, Yanqiao Zhu 0003, Yi-Wen Chao, Yuanzhe Chen, Zhuo Chen 0006, Jian Cong, Keliang Li, Siyou Li, Xinfeng Li, Xiquan Li, Zheng Lian 0004, Yuzhe Liang, Minghao Liu 0003, Zhikang Niu, Tianrui Wang, Yuping Wang 0005, Yuxuan Wang 0002, Guanrou Yang, Jianwei Yu 0001, Ruibin Yuan, Zhisheng Zheng, Ziya Zhou, Haina Zhu, Wei Xue 0002, Emmanouil Benetos, Kai Yu 0004, Chng Eng Siong, Xie Chen 0001
NeurIPS8
2024 StreamVoice: Streamable Context-Aware Language Modeling for Real-time Zero-Shot Voice Conversion
abstract
Recent language model (LM) advancements have showcased impressive zero-shot voice conversion (VC) performance.However, existing LM-based VC models usually apply offline conversion from source semantics to acoustic features, demanding the complete source speech and limiting their deployment to realtime applications.In this paper, we introduce StreamVoice, a novel streaming LM-based model for zero-shot VC, facilitating real-time conversion given arbitrary speaker prompts and source speech.Specifically, to enable streaming capability, StreamVoice employs a fully causal context-aware LM with a temporalindependent acoustic predictor, while alternately processing semantic and acoustic features at each time step of autoregression which eliminates the dependence on complete source speech.To address the potential performance degradation from the incomplete context in streaming processing, we enhance the contextawareness of the LM through two strategies: 1) teacher-guided context foresight, using a teacher model to summarize the present and future semantic context during training to guide the model's forecasting for missing context; 2) semantic masking strategy, promoting acoustic prediction from preceding corrupted semantic and acoustic input, enhancing context-learning ability.Notably, StreamVoice is the first LMbased streaming zero-shot VC model without any future look-ahead.Experiments demonstrate StreamVoice's streaming conversion capability while achieving zero-shot performance comparable to non-streaming VC systems.
Zhichao Wang 0002, Yuanzhe Chen, Lei Xie 0001, Yuping Wang 0005
ACL (1)2
2024 StreamVoice+: Evolving Into End-to-End Streaming Zero-Shot Voice Conversion
abstract
StreamVoice has recently pushed the boundaries of zero-shot voice conversion (VC) in the streaming domain. It uses a streamable language model (LM) with a context-aware approach to convert semantic features from automatic speech recognition (ASR) into acoustic features with the desired speaker timbre. Despite its innovations, StreamVoice faces challenges due to its dependency on a streaming ASR within a cascaded framework, which complicates system deployment and optimization, affects VC system's design and performance based on the choice of ASR, and struggles with conversion stability when faced with low-quality semantic inputs. To overcome these limitations, we introduce StreamVoice+, an enhanced LM-based end-to-end streaming framework that operates independently of streaming ASR. StreamVoice+ integrates a semantic encoder and a connector with the original StreamVoice framework, now trained using a non-streaming ASR. This model undergoes a two-stage training process: initially, the StreamVoice backbone is pre-trained for voice conversion and the semantic encoder for robust semantic extraction. Subsequently, the system is fine-tuned end-to-end, incorporating a LoRA matrix to activate comprehensive streaming functionality. Furthermore, StreamVoice+ mainly introduces two strategic enhancements to boost conversion quality: a residual compensation mechanism in the connector to ensure effective semantic transmission and a self-refinement strategy that leverages pseudo-parallel speech pairs generated by the conversion backbone to improve speech decoupling. Experiments demonstrate that StreamVoice+ not only achieves higher naturalness and speaker similarity in voice conversion than its predecessor but also provides versatile support for both streaming and non-streaming conversion scenarios.
Zhichao Wang 0002, Yuanzhe Chen, Lei Xie 0001, Yuping Wang 0005
IEEE Signal Process. Lett.2
2024 Multi-Level Temporal-Channel Speaker Retrieval for Zero-Shot Voice Conversion
abstract
Zero-shot voice conversion (VC) converts source speech into the voice of any desired speaker using only one utterance of the speaker without requiring additional model updates. Typical methods use a speaker representation from a pre-trained speaker verification (SV) model or learn speaker representation during VC training to achieve zero-shot VC. However, existing speaker modeling methods overlook the variation of speaker information richness in temporal and frequency channel dimensions of speech. This insufficient speaker modeling hampers the ability of the VC model to accurately represent unseen speakers who are not in the training dataset. In this study, we present a robust zero-shot VC model withmulti-leveltemporal-channelretrieval, referred to as MTCR-VC. Specifically, to flexibly adapt to the dynamic-variant speaker characteristic in the temporal and channel axis of the speech, we propose a novel fine-grained speaker modeling method, calledtemporal-channelretrieval (TCR), to find outwhenandwherespeaker information appears in speech. It retrieves variable-length speaker representation from both temporal and channel dimensions under the guidance of a pre-trained SV model. Besides, inspired by the hierarchical process of human speech production, the MTCR speaker module stacks several TCR blocks to extract speaker representations from multi-granularity levels. Furthermore, we introduce a cycle-based training strategy to simulate zero-shot inference recurrently to achieve better speech disentanglement and reconstruction. To drive this process, we adopt perceptual constraints on three aspects: content, style, and speaker. Experiments demonstrate that MTCR-VC is superior to the previous zero-shot VC methods in modeling speaker timbre while maintaining good speech naturalness.
Zhichao Wang 0002, Liumeng Xue, Qiuqiang Kong, Lei Xie 0001, Yuanzhe Chen, Qiao Tian 0001, Yuping Wang 0005
IEEE ACM Trans. Audio Speech Lang. Process.5
2023 Streaming Voice Conversion via Intermediate Bottleneck Features and Non-Streaming Teacher Guidance
abstract
Streaming voice conversion (VC) is the task of converting the voice of one person to another in real-time. Previous streaming VC methods use phonetic posteriorgrams (PPGs) extracted from automatic speech recognition (ASR) systems to represent speaker-independent information. However, PPGs lack the prosody and phonation information of the source speaker, and streaming PPGs contain undesired leaked timbre of the source speaker. In this paper, we propose to use intermediate bottleneck features (IBFs) to replace PPGs. VC systems trained with IBFs retain more prosody and phonation information of the source speaker. Furthermore, we propose a non-streaming teacher guidance (TG) framework that addresses the timbre leakage problem. Experiments show that our proposed IBFs and the TG framework achieve a state-of-the-art streaming VC naturalness of 3.85, a content consistency of 3.77, and a timbre similarity of 3.77 under a future receptive field of 160 ms which significantly outperform previous streaming VC systems.
Yuanzhe Chen, Tang Li 0001, Qiuqiang Kong, Zhichao Wang 0002, Qiao Tian 0001, Yuping Wang 0005, Yuxuan Wang 0002
ICASSP1
2023 Delivering Speaking Style in Low-Resource Voice Conversion with Multi-Factor Constraints
abstract
Conveying the linguistic content and maintaining the source speech’s speaking style, such as intonation and emotion, is essential in voice conversion (VC). However, in a low-resource situation, where only limited utterances from the target speaker are accessible, existing VC methods are hard to meet this requirement and capture the target speaker’s timber. In this work, a novel VC model, referred to as MFC-StyleVC, is proposed for the low-resource VC task. Specifically, speaker timbre constraint generated by clustering method is newly proposed to guide target speaker timbre learning in different stages. Meanwhile, to prevent over-fitting to the target speaker’s limited data, perceptual regularization constraints explicitly maintain model performance on specific aspects, including speaking style, linguistic content, and speech quality. Besides, a simulation mode is introduced to simulate the inference process to alleviate the mis-match between training and inference. Extensive experiments performed on highly expressive speech demonstrate the superiority of the proposed method in low-resource VC.
Zhichao Wang 0002, Lei Xie 0001, Yuanzhe Chen, Qiao Tian 0001, Yuping Wang 0005
ICASSP4
2023 Zero-Shot Accent Conversion using Pseudo Siamese Disentanglement Network
Dongya Jia, Qiao Tian 0001, Kainan Peng, Yuanzhe Chen, Mingbo Ma, Yuping Wang 0005, Yuxuan Wang 0002
INTERSPEECH5
2023 LM-VC: Zero-Shot Voice Conversion via Speech Generation Based on Language Models
abstract
Language model (LM) based audio generation frameworks, e.g., AudioLM, have recently achieved new state-of-the-art performance in zero-shot audio generation. In this paper, we explore the feasibility of LMs forzero-shot voice conversion. An intuitive approach is to follow AudioLM – Tokenizing speech into semantic and acoustic tokens respectively by HuBERT and SoundStream, and converting source semantic tokens to target acoustic tokens conditioned on acoustic tokens of the target speaker. However, such an approach encounters several issues: 1) the linguistic content contained in semantic tokens may get dispersed during multi-layer modeling while the lengthy speech input in the voice conversion task makes contextual learning even harder; 2) the semantic tokens still contain speaker-related information, which may be leaked to the target speech, lowering the target speaker similarity; 3) the generation diversity in the sampling of the LM can lead to unexpected outcomes during inference, leading to unnatural pronunciation and speech quality degradation. To mitigate these problems, we proposeLM-VC, a two-stage language modeling approach that generates coarse acoustic tokens for recovering the source linguistic content and target speaker's timbre, and then reconstructs the fine for acoustic details as converted speech. Specifically, to enhance content preservation and facilitates better disentanglement, a masked prefix LM with a mask prediction strategy is used for coarse acoustic modeling. This model is encouraged to recover the masked content from the surrounding context and generate target speech based on the target speaker's utterance and corrupted semantic tokens. Besides, to further alleviate the sampling error in the generation, an external LM, which employs window attention to capture the local acoustic relations, is introduced to participate in the coarse acoustic modeling through shallow fusion. Finally, a prefix LM reconstructs fine acoustic tokens from the coarse and results in the converted speech. Experiments demonstrate that LM-VC outperforms competitive systems in speech naturalness and speaker similarity.
Zhichao Wang 0002, Yuanzhe Chen, Lei Xie 0001, Qiao Tian 0001, Yuping Wang 0005
IEEE Signal Process. Lett.2
2022 Cloning One's Voice Using Very Limited Data in the Wild
abstract
With the increasing popularity of speech synthesis products, the industry has put forward more requirements for personalized speech synthesis: (1) How to use low-resource, easily accessible data to clone a person’s voice. (2) How to clone a person’s voice while controlling the style and prosody. To solve the above two problems, we proposed the Hieratron model framework in which the prosody and timbre are modeled separately using two modules, therefore, the independent control of timbre and the other characteristics of audio can be achieved while generating speech. The practice shows that, for very limited target speaker data in the wild, Hieratron has obvious advantages over the traditional method, in addition to controlling the style and language of the generated speech, the mean opinion score on speech quality of the generated speech has also been improved by more than 0.2 points.
Dongyang Dai, Yuanzhe Chen, Qiao Tian 0001, Yuping Wang 0005, Yuxuan Wang 0002
ICASSP2
2020 DFSeer: A Visual Analytics Approach to Facilitate Model Selection for Demand Forecasting
abstract
Selecting an appropriate model to forecast product demand is critical to the manufacturing industry. However, due to the data complexity, market uncertainty and users' demanding requirements for the model, it is challenging for demand analysts to select a proper model. Although existing model selection methods can reduce the manual burden to some extent, they often fail to present model performance details on individual products and reveal the potential risk of the selected model. This paper presents DFSeer, an interactive visualization system to conduct reliable model selection for demand forecasting based on the products with similar historical demand. It supports model comparison and selection with different levels of details. Besides, it shows the difference in model performance on similar products to reveal the risk of model selection and increase users' confidence in choosing a forecasting model. Two case studies and interviews with domain experts demonstrate the effectiveness and usability of DFSeer.
Dong Sun 0001, Zezheng Feng, Yuanzhe Chen, Yong Wang 0021, Mingxuan Yuan, Ting-Chuen Pong, Huamin Qu
CHI3
2020 ViSeq: Visual Analytics of Learning Sequence in Massive Open Online Courses
abstract
The research on massive open online courses (MOOCs) data analytics has mushroomed recently because of the rapid development of MOOCs. The MOOC data not only contains learner profiles and learning outcomes, but also sequential information about when and which type of learning activities each learner performs, such as reviewing a lecture video before undertaking an assignment. Learning sequence analytics could help understand the correlations between learning sequences and performances, which further characterize different learner groups. However, few works have explored the sequence of learning activities, which have mostly been considered aggregated events. A visual analytics system called ViSeq is introduced to resolve the loss of sequential information, to visualize the learning sequence of different learner groups, and to help better understand the reasons behind the learning behaviors. The system facilitates users in exploring learning sequences from multiple levels of granularity. ViSeq incorporates four linked views: the projection view to identify learner groups, the pattern view to exhibit overall sequential patterns within a selected group, the sequence view to illustrate the transitions between consecutive events, and the individual view with an augmented sequence chain to compare selected personal learning sequences. Case studies and expert interviews were conducted to evaluate the system.
Qing Chen 0001, Xuanwu Yue, Xavier Plantaz, Yuanzhe Chen, Conglei Shi, Ting-Chuen Pong, Huamin Qu
IEEE Trans. Vis. Comput. Graph.4
2020 PlanningVis: A Visual Analytics Approach to Production Planning in Smart Factories
abstract
Production planning in the manufacturing industry is crucial for fully utilizing factory resources (e.g., machines, raw materials and workers) and reducing costs. With the advent of industry 4.0, plenty of data recording the status of factory resources have been collected and further involved in production planning, which brings an unprecedented opportunity to understand, evaluate and adjust complex production plans through a data-driven approach. However, developing a systematic analytics approach for production planning is challenging due to the large volume of production data, the complex dependency between products, and unexpected changes in the market and the plant. Previous studies only provide summarized results and fail to show details for comparative analysis of production plans. Besides, the rapid adjustment to the plan in the case of an unanticipated incident is also not supported. In this paper, we propose PlanningVis, a visual analytics system to support the exploration and comparison of production plans with three levels of details: a plan overview presenting the overall difference between plans, a product view visualizing various properties of individual products, and a production detail view displaying the product dependency and the daily production details in related factories. By integrating an automatic planning algorithm with interactive visual explorations, PlanningVis can facilitate the efficient optimization of daily production planning as well as support a quick response to unanticipated incidents in manufacturing. Two case studies with real-world data and carefully designed interviews with domain experts demonstrate the effectiveness and usability of PlanningVis.
Dong Sun 0001, Renfei Huang, Yuanzhe Chen, Yong Wang 0021, Mingxuan Yuan, Ting-Chuen Pong, Huamin Qu
IEEE Trans. Vis. Comput. Graph.3
2018 StageMap: Extracting and Summarizing Progression Stages in Event Sequences
abstract
Temporal event sequences are becoming increasingly important in many application domains such as website click streams, user interaction logs, electronic health records and car service records. A real-world dataset with a large number of event sequences of varying lengths is complex and difficult to analyze. To support visual exploration of the data, it is desirable yet challenging to provide a concise and meaningful overview of sequences. In this paper, we focus on the stage, that is, a frequently occurring subsequence in the dataset. We introduce StageMap, a novel visualization technique to summarize event sequence data into a set of stage progression patterns. The resulting overview is more concise compared with event-level summarization and supports level-of-detail exploration. We further present a visual analytics system with four linked views, which are overview, tree view, stage view and sequences view. We also present case studies and discuss advantages and limitations of applying StageMap to real-world scenarios.
Yuanzhe Chen, Abishek Puri, Linping Yuan, Huamin Qu
IEEE BigData1
2018 Sequence Synopsis: Optimize Visual Summary of Temporal Event Data
abstract
Event sequences analysis plays an important role in many application domains such as customer behavior analysis, electronic health record analysis and vehicle fault diagnosis. Real-world event sequence data is often noisy and complex with high event cardinality, making it a challenging task to construct concise yet comprehensive overviews for such data. In this paper, we propose a novel visualization technique based on the minimum description length (MDL) principle to construct a coarse-level overview of event sequence data while balancing the information loss in it. The method addresses a fundamental trade-off in visualization design: reducing visual clutter vs. increasing the information content in a visualization. The method enables simultaneous sequence clustering and pattern extraction and is highly tolerant to noises such as missing or additional events in the data. Based on this approach we propose a visual analytics framework with multiple levels-of-detail to facilitate interactive data exploration. We demonstrate the usability and effectiveness of our approach through case studies with two real-world datasets. One dataset showcases a new application domain for event sequence visualization, i.e., fault development path analysis in vehicles for predictive maintenance. We also discuss the strengths and limitations of the proposed method based on user feedback.
Yuanzhe Chen, Liu Ren 0001
IEEE Trans. Vis. Comput. Graph.1
2016 Visual Analytics in Urban Computing: An Overview
abstract
Nowadays, various data collected in urban context provide unprecedented opportunities for building a smarter city through urban computing. However, due to heterogeneity, high complexity and large volumes of these urban data, analyzing them is not an easy task, which often requires integrating human perception in analytical process, triggering a broad use of visualization. In this survey, we first summarize frequently used data types in urban visual analytics, and then elaborate on existing visualization techniques for time, locations and other properties of urban data. Furthermore, we discuss how visualization can be combined with automated analytical approaches. Existing work on urban visual analytics is categorized into two classes based on different outputs of such combinations: 1) For data exploration and pattern interpretation, we describe representative visual analytics tools designed for better insights of different types of urban data. 2) For visual learning, we discuss how visualization can help in three major steps of automated analytical approaches (i.e., cohort construction; feature selection & model construction; result evaluation & tuning) for a more effective machine learning or data mining process, leading to sort of artificial intelligence, such as a classifier, a predictor or a regression model. Finally, we outlook the future of urban visual analytics, and conclude the survey with potential research directions.
Yixian Zheng, Yuanzhe Chen, Huamin Qu, Lionel M. Ni
IEEE Trans. Big Data3
2016 PeakVizor: Visual Analytics of Peaks in Video Clickstreams from Massive Open Online Courses
abstract
Massive open online courses (MOOCs) aim to facilitate open-access and massive-participation education. These courses have attracted millions of learners recently. At present, most MOOC platforms record the web log data of learner interactions with course videos. Such large amounts of multivariate data pose a new challenge in terms of analyzing online learning behaviors. Previous studies have mainly focused on the aggregate behaviors of learners from a summative view; however, few attempts have been made to conduct a detailed analysis of such behaviors. To determine complex learning patterns in MOOC video interactions, this paper introduces a comprehensive visualization system called PeakVizor. This system enables course instructors and education experts to analyze the "peaks" or the video segments that generate numerous clickstreams. The system features three views at different levels: the overview with glyphs to display valuable statistics regarding the peaks detected; the flow view to present spatio-temporal information regarding the peaks; and the correlation view to show the correlation between different learner groups and the peaks. Case studies and interviews conducted with domain experts have demonstrated the usefulness and effectiveness of PeakVizor, and new findings about learning behaviors in MOOC platforms have been reported.
Qing Chen 0001, Yuanzhe Chen, Dongyu Liu, Conglei Shi, Yingcai Wu, Huamin Qu
IEEE Trans. Vis. Comput. Graph.2
2014 Finding Coherent Motions and Semantic Regions in Crowd Scenes: A Diffusion and Clustering Approach
Weiyue Wang 0002, Weiyao Lin, Yuanzhe Chen, Jianxin Wu 0001, Jingdong Wang 0001, Bin Sheng 0001
ECCV (1)3
2014 A New Network-Based Algorithm for Human Activity Recognition in Videos
abstract
In this paper, a new network-transmission-based (NTB) algorithm is proposed for human activity recognition in videos. The proposed NTB algorithm models the entire scene as an error-free network. In this network, each node corresponds to a patch of the scene and each edge represents the activity correlation between the corresponding patches. Based on this network, we further model people in the scene as packages, while human activities can be modeled as the process of package transmission in the network. By analyzing these specific package transmission processes, various activities can be effectively detected. The implementation of our NTB algorithm into abnormal activity detection and group activity recognition are described in detail in this paper. Experimental results demonstrate the effectiveness of our proposed algorithm.
Weiyao Lin, Yuanzhe Chen, Jianxin Wu 0001, Hanli Wang, Bin Sheng 0001, Hongxiang Li 0001
IEEE Trans. Circuits Syst. Video Technol.2
2013 A New Network-Based Algorithm for Human Group Activity Recognition in Videos
Gaojian Li, Weiyao Lin, Jianxin Wu 0001, Yuanzhe Chen, Hui Wei 0001
MMM (1)5
2013 Intra-and-Inter-Constraint-Based Video Enhancement Based on Piecewise Tone Mapping
abstract
Video enhancement plays an important role in various video applications. In this paper, we propose a new intra-and-inter-constraint-based video enhancement approach aiming to: 1) achieve high intraframe quality of the entire picture where multiple regions-of-interest (ROIs) can be adaptively and simultaneously enhanced, and 2) guarantee the interframe quality consistencies among video frames. We first analyze features from different ROIs and create a piecewise tone mapping curve for the entire frame such that the intraframe quality can be enhanced. We further introduce new interframe constraints to improve the temporal quality consistency. Experimental results show that the proposed algorithm obviously outperforms the state-of-the-art algorithms.
Yuanzhe Chen, Weiyao Lin, Zhenzhong Chen 0001, Ning Xu 0007
IEEE Trans. Circuits Syst. Video Technol.1
2012 A new heat-map-based algorithm for human group activity recognition
abstract
In this paper, a new heat-map-based (HMB) algorithm is proposed for human group activity recognition. The proposed algorithm first models people trajectories as series of "heat sources" and then applies a thermal diffusion process to create a heat map (HM) for representing the group activities. Based on this heat map, a new surface-fitting (SF) method is also proposed for recognizing human group activities. Our proposed HM feature can efficiently keep the temporal motion information of the group activities while the proposed SF method can effectively catch the characteristics of the heat map for activity recognition. Experimental results demonstrate the effectiveness of our proposed algorithm.
Hang Chu, Weiyao Lin, Jianxin Wu 0001, Xingtong Zhou, Yuanzhe Chen, Hongxiang Li 0001
ACM Multimedia5
2012 A patch-based framework for detecting abnormal activities with a PTZ camera
abstract
In this paper, a novel patch-based (PB) framework is proposed for detecting abnormal activities using a Pan-Tilt-Zoom (PTZ) camera. We first propose a new scene-patch-based (SSB) algorithm which can efficiently extract the target object's global trajectory from the PTZ camera. Furthermore, we propose an extended network-based (ENB) algorithm for detecting abnormal activities. The proposed ENB algorithm models the entire scene as a network where each node in the network corresponds to a patch of the scene and each edge between nodes corresponds to the activity correlation between the scene patchs. Based on this network, a recursive training strategy is proposed to train the edge weights in the network such that abnormal activities can be effectively detected through these trained edge weights. Experimental results demonstrate the effectiveness of our proposed framework.
Yisi Tao, Yuanzhe Chen, Weiyao Lin, Xintong Han, Hongxiang Li 0001, Zheng Lu 0003
VCIP2
2011 A new package-group-transmission-based algorithm for human activity recognition in videos
abstract
In this paper, a new package-group-transmission-based algorithm is proposed for human activity recognition in videos. The proposed algorithm first models the entire scene as a network where each node in the network corresponds to a segmentation of the scene. Based on this network, we further model people in the scene as groups of packages. Thus, various human activities can be modeled as the process of "package group transmission" in the network and these activities can be efficiently recognized by suitably analyzing the "package transmission" process. Our proposed algorithm can not only detect activities under the challenging multiple camera scenario, but also be able to recognize various complex group activities among people. Experimental results demonstrate the effectiveness of our proposed algorithm.
Yuanzhe Chen, Weiyao Lin, Hongxiang Li 0001, Hangzai Luo, Yisi Tao, Donghua Liu
VCIP1
2011 A new global-based video enhancement algorithm by fusing features of multiple region-of-interests
abstract
Video enhancement plays an important role in various video applications. It is desirable to achieve high visual quality of the entire picture where multiple region-of-interests (ROIs) within the frame can be adaptively and simultaneously enhanced. In this paper, a new global-based video enhancement algorithm is proposed. The proposed algorithm first analyzes features from different ROIs. Then, a 'global' tone mapping curve is created for the entire picture which can adaptively enhance different regions at the same time. According to the statistics of ROIs, two fusion strategies, i.e., piecewise-based and factor-based fusions, are proposed for creating the global tone mapping curve. Experimental results show that the proposed algorithm can obtain more appealing perceptual quality than the state-of-the-art algorithms.
Ning Xu 0007, Weiyao Lin, Yu Zhou 0015, Yuanzhe Chen, Zhenzhong Chen 0001, Hongxiang Li 0001
VCIP4