Wenmian Yang

dblp:205/4005 · DBLP profile ↗
← Back
22ranked-venue papers
6as first author
15since 2021 · last 2026
0000-0001-8493-4449ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 3 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ODUTQA-MDC: A Task for Open-Domain Underspecified Tabular QA with Multi-turn Dialogue-based Clarification
abstract
The advancement of large language models (LLMs) has enhanced tabular question answering (Tabular QA), yet they struggle with opendomain queries exhibiting underspecified or uncertain expressions.To address this, we introduce the ODUTQA-MDC task and the first comprehensive benchmark to tackle it.This benchmark includes: (1) a large-scale ODUTQA dataset with 209 tables and 25,105 QA pairs; (2) a fine-grained labeling scheme for detailed evaluation; and (3) a dynamic clarification interface that simulates user feedback for interactive assessment.We also propose MAIC-TQA, a multi-agent framework that excels at detecting ambiguities, clarifying them through dialogue, and refining answers.Experiments validate our benchmark and framework, establishing them as a key resource for advancing conversational, underspecification-aware Tabular QA research.The data and code are available at https://github.com/jensenw1/ ODUTQA-MDC.Q3: What is the greenery rate of the Vanke Light of the Future community?Q1: What is the greenery rate of the Vanke Light of the Future community in Bao'an District, Shenzhen?Q2: How is the environment of the Vanke Light of the Future community in Bao'an District, Shenzhen?Q4: What is the greenery rate of the Vanke community in Bao'an District, Shenzhen?
Zhensheng Wang 0001, ZhanTeng Lin, Wenmian Yang, Yiquan Zhang, Weijia Jia 0001
ACL (1)3
2026 ReCoQA: A Benchmark for Tool-Augmented and Multi-Step Reasoning in Real Estate Question and Answering
abstract
Developing agents capable of navigating fragmented, multi-source information remains challenging, primarily due to the scarcity of benchmarks reflecting hybrid workflows combining database querying with external APIs.To bridge this gap, we introduce ReCoQA, a largescale benchmark of 29,270 real-estate instances featuring machine-verifiable supervision for intermediate steps, including structured intent labels, SQL queries, and API calls.Complementarily, we propose HIRE-Agent, a hierarchical framework instantiating an understand-plan-execute architecture as a strong baseline.By orchestrating a Front-end parser, a planning Supervisor, and execution Specialists, HIRE-Agent effectively integrates heterogeneous evidence.Extensive experiments demonstrate that HIRE-Agent constitutes a strong baseline and substantiates the necessity of hierarchical collaboration for complex, real-world reasoning tasks.The benchmark and source code are available at: https://github.com/ Husky-989/ReCoQA
Yindong Zhang, Wenmian Yang, Yiquan Zhang, Weijia Jia 0001
ACL (1)2
2026 JUST: Cost-efficient joint cluster upgrade sequencing and task scheduling for containerized edge computing
Zhiqing Tang, Wenmian Yang, Jianxiong Guo, Tian Wang 0001, Weijia Jia 0001
Comput. Networks4
2026 GMDRA: Optimizing Edge-Cloud LLM Security and Efficiency via Game-Theoretic Detection Resource Allocation
abstract
Large language models (LLMs) have made notable advancements across various dimensions of human experience, while prompt engineering has further refined their capability and efficiency in generating relevant, high-quality outputs. Nonetheless, the rise of prompt engineering has also led to an increase in prompt attacks, resulting in critical issues such as privacy breaches, increased latency, and inefficient resource utilization. Existing security mechanisms based on Reinforcement Learning from Human Feedback often fall short in addressing the complexities posed by these diverse prompt attacks, emphasizing the pressing need for effective prompt security mechanisms. The prompt detection model serves as an effective supplementary measure to ensure the safety of large model outputs. However, the current detection mechanism checks all prompts, which increases system resource consumption and user latency. In this paper, we present a holistic analysis of prompt security, service latency, and resource optimization within Edge-Cloud LLM (EC-LLM) systems confronted with various prompt attack scenarios. To strengthen prompt security, we introduce a novel, lightweight attack detection mechanism that uses vector databases. We conceptualize the intricate interplay between prompt detection, latency, and resource optimization within a 2-stage dynamic Bayesian game framework. A reproducible equilibrium strategy is derived by estimating the potential number of malicious tasks and updating beliefs iteratively at each stage using Bayesian methods. Our proposed solution has been rigorously evaluated in a practical EC-LLM system, yielding results that highlight significant enhancements in security, reductions in service latency for legitimate users, and decreased overall resource consumption compared to current leading algorithms.
Jianxiong Guo, Tianhui Meng, Wenmian Yang, Tian Wang 0001, Weijia Jia 0001
IEEE Trans. Mob. Comput.4
2025 RETQA: A Large-Scale Open-Domain Tabular Question Answering Dataset for Real Estate Sector
abstract
The real estate market relies heavily on structured data, such as property details, market trends, and price fluctuations. However, the lack of specialized Tabular Question Answering datasets in this domain limits the development of automated question-answering systems. To fill this gap, we introduce RETQA, the first large-scale open-domain Chinese Tabular Question Answering dataset for Real Estate. RETQA comprises 4,932 tables and 20,762 question-answer pairs across 16 sub-fields within three major domains: property information, real estate company finance information and land auction information. Compared with existing tabular question answering datasets, RETQA poses greater challenges due to three key factors: long-table structures, open-domain retrieval, and multi-domain queries. To tackle these challenges, we propose the SLUTQA framework, which integrates large language models with spoken language understanding tasks to enhance retrieval and answering accuracy. Extensive experiments demonstrate that SLUTQA significantly improves the performance of large language models on RETQA by in-context learning. RETQA and SLUTQA provide essential resources for advancing tabular question answering research in the real estate domain, addressing critical challenges in open-domain and long-table question-answering.
Zhensheng Wang 0001, Wenmian Yang, Yiquan Zhang, Weijia Jia 0001
AAAI2
2025 TLCCSP: A Scalable Framework for Enhancing Time Series Forecasting with Time-Lagged Cross-Correlations
abstract
Time series forecasting is critical across various domains, such as weather, finance and real estate forecasting, as accurate forecasts support informed decision-making and risk mitigation. While recent deep learning models have improved predictive capabilities, they often overlook time-lagged cross-correlations between related sequences, which are crucial for capturing complex temporal relationships. To address this, we propose the Time-Lagged Cross-Correlations-based Sequence Prediction framework (TLCCSP), which enhances forecasting accuracy by effectively integrating time-lagged cross-correlated sequences. TLCCSP employs the Sequence Shifted Dynamic Time Warping (SSDTW) algorithm to capture lagged correlations and a contrastive learning-based encoder to efficiently approximate SSDTW distances. Experimental results on weather, finance and real estate time series datasets demonstrate the effectiveness of our framework. On the weather dataset, SSDTW reduces mean squared error (MSE) by 16.01% compared with single-sequence methods, while the contrastive learning encoder (CLE) further decreases MSE by 17.88%. On the stock dataset, SSDTW achieves a 9.95% MSE reduction, and CLE reduces it by 6.13%. For the real estate dataset, SSDTW and CLE reduce MSE by 21.29% and 8.62%, respectively. Additionally, the contrastive learning approach decreases SSDTW computational time by approximately 99%, ensuring scalability and real-time applicability across multiple time series forecasting tasks.
Jianfei Wu, Wenmian Yang, Bingning Liu, Weijia Jia 0001
CIKM2
2025 HiGarment: Cross-Modal Harmony Based Diffusion Model for Flat Sketch to Realistic Garment Image
abstract
Diffusion-based garment synthesis tasks primarily focus on the design phase in the fashion domain, while the garment production process remains largely underexplored. To bridge this gap, we introduce a new task: Flat Sketch to Realistic Garment Image (FS2RG), which generates realistic garment images by integrating flat sketches and textual guidance. FS2RG presents two key challenges: 1) fabric characteristics are solely guided by textual prompts, providing insufficient visual supervision for diffusion-based models, which limits their ability to capture fine-grained fabric details; 2) flat sketches and textual guidance may provide conflicting information, requiring the model to selectively preserve or modify garment attributes while maintaining structural coherence. To tackle this task, we propose HiGarment, a novel framework that comprises two core components: i) a multi-modal semantic enhancement mechanism that enhances fabric representation across textual and visual modalities, and ii) a harmonized cross-attention mechanism that dynamically balances information from flat sketches and text prompts, allowing controllable synthesis by generating either sketch-aligned (image-biased) or text-guided (text-biased) outputs. Furthermore, we collect Multi-modal Detailed Garment, the largest open-source dataset for garment generation. Experimental results and user studies demonstrate the effectiveness of HiGarment in garment synthesis. The code and dataset are available at https://github.com/Maple498/HiGarment.
Junyi Guo, Fangyu Wu 0001, Huanda Lu, Qiufeng Wang 0001, Wenmian Yang, Eng Gee Lim, Dongming Lu
ICCV6
2025 A Virtual-Label-Based Hierarchical Domain Adaptation Method for Time-Series Classification
abstract
Unsupervised domain adaptation (UDA) is becoming a prominent solution for the domain-shift problem in many time-series classification tasks. With sequence properties, time-series data contain both local and sequential features, and the domain shift exists in both features. However, conventional UDA methods usually cannot distinguish those two features but mix them into one variable for direct alignment, which harms the performance. To address this problem, we propose a novel virtual-label-based hierarchical domain adaptation (VLH-DA) approach for time-series classification. Specifically, we first slice the original time-series data and introduce virtual labels to represent the type of each slice (called local patterns). With the help of virtual labels, we decompose the end-to-end (i.e., signal to time-series label) time-series task into two parts, i.e., signal sequence to local pattern sequence and local pattern sequence to time-series label. By decomposing the complex time-series UDA task into two simpler subtasks, the local features and sequential features can be aligned separately, making it easier to mitigate distribution discrepancies. Experiments on four public time-series datasets demonstrate that our VLH-DA outperforms all state-of-the-art (SOTA) methods.
Wenmian Yang, Lizhi Cheng, Mohamed Ragab 0002, Min Wu 0008, Sinno Jialin Pan, Zhenghua Chen
IEEE Trans. Neural Networks Learn. Syst.1
2025 Enhancing LLM QoS Through Cloud-Edge Collaboration: A Diffusion-Based Multi-Agent Reinforcement Learning Approach
abstract
Large Language Models (LLMs) are widely used across various domains, but deploying them in cloud data centers often leads to significant response delays and high costs, undermining Quality of Service (QoS) at the network edge. Although caching LLM request results at the edge using vector databases can greatly reduce response times and costs for similar requests, this approach has been overlooked in prior research. To address this, we propose a novelVector database-assisted cloud-Edge collaborativeLLM QoSOptimization (VELO) framework that caches LLM request results at the edge using vector databases, thereby reducing response times for subsequent similar requests. Unlike methods that modify LLMs directly, VELO leaves the LLM's internal structure intact and is applicable to various LLMs. Building on VELO, we formulate the QoS optimization problem as a Markov Decision Process (MDP) and design an algorithm based on Multi-Agent Reinforcement Learning (MARL). Our algorithm employs a diffusion-based policy network to extract the LLM request features, determining whether to request the LLM in the cloud or retrieve results from the edge's vector database. Implemented in a real edge system, our experimental results demonstrate that VELO significantly enhances user satisfaction by simultaneously reducing delays and resource consumption for edge users of LLMs. Our DLRS algorithm improves performance by 15.0% on average for similar requests and by 14.6% for new requests compared to the baselines.
Zhi Yao, Zhiqing Tang, Wenmian Yang, Weijia Jia 0001
IEEE Trans. Serv. Comput.3
2023 A Scope Sensitive and Result Attentive Model for Multi-Intent Spoken Language Understanding
abstract
Multi-Intent Spoken Language Understanding (SLU), a novel and more complex scenario of SLU, is attracting increasing attention. Unlike traditional SLU, each intent in this scenario has its specific scope. Semantic information outside the scope even hinders the prediction, which tremendously increases the difficulty of intent detection. More seriously, guiding slot filling with these inaccurate intent labels suffers error propagation problems, resulting in unsatisfied overall performance. To solve these challenges, in this paper, we propose a novel Scope-Sensitive Result Attention Network (SSRAN) based on Transformer, which contains a Scope Recognizer (SR) and a Result Attention Network (RAN). SR assignments scope information to each token, reducing the distraction of out-of-scope tokens. RAN effectively utilizes the bidirectional interaction between SF and ID results, mitigating the error propagation problem. Experiments on two public datasets indicate that our model significantly improves SLU performance (5.4% and 2.1% on Overall accuracy) over the state-of-the-art baseline.
Lizhi Cheng, Wenmian Yang, Weijia Jia 0001
AAAI2
2023 Mitigating Exposure Bias in Grammatical Error Correction with Data Augmentation and Reweighting
abstract
The most popular approach in grammatical error correction (GEC) is based on sequence-tosequence (seq2seq) models.Similar to other autoregressive generation tasks, seq2seq GEC also faces the exposure bias problem, i.e., the context tokens are drawn from different distributions during training and testing, caused by the teacher forcing mechanism.In this paper, we propose a novel data manipulation approach to overcome this problem, which includes a data augmentation method during training to mimic the decoder input at inference time, and a data reweighting method to automatically balance the importance of each kind of augmented samples.Experimental results on benchmark GEC datasets show that our method achieves significant improvements compared to prior approaches. 1
Hannan Cao, Wenmian Yang, Hwee Tou Ng
EACL2
2023 Capture Salient Historical Information: A Fast and Accurate Non-autoregressive Model for Multi-turn Spoken Language Understanding
abstract
Spoken Language Understanding (SLU), a core component of the task-oriented dialogue system, expects a shorter inference facing the impatience of human users. Existing work increases inference speed by designing non-autoregressive models for single-turn SLU tasks but fails to apply to multi-turn SLU in confronting the dialogue history. The intuitive idea is to concatenate all historical utterances and utilize the non-autoregressive models directly. However, this approach seriously misses the salient historical information and suffers from the uncoordinated-slot problems. To overcome those shortcomings, we propose a novel model for multi-turn SLU named Salient History Attention with Layer-Refined Transformer (SHA-LRT), which comprises a SHA module, a Layer-Refined Mechanism (LRM), and a Slot Label Generation (SLG) task. SHA captures salient historical information for the current dialogue from both historical utterances and results via a well-designed history-attention mechanism. LRM predicts preliminary SLU results from Transformer’s middle states and utilizes them to guide the final prediction, and SLG obtains the sequential dependency information for the non-autoregressive encoder. Experiments on public datasets indicate that our model significantly improves multi-turn SLU performance (17.5% on Overall) with accelerating (nearly 15 times) the inference process over the state-of-the-art baseline as well as effective on the single-turn SLU tasks.
Lizhi Cheng, Weijia Jia 0001, Wenmian Yang
ACM Trans. Inf. Syst.3
2022 Representation and Reinforcement Learning for Task Scheduling in Edge Computing
abstract
Recently, many deep reinforcement learning (DRL)-based task scheduling algorithms have been widely used in edge computing (EC) to reduce energy consumption. Unlike the existing algorithms considering fixed and fewer edge nodes (servers) and tasks, in this article, a representation model with a DRL based algorithm is proposed to adapt the dynamic change of nodes and tasks and solve the dimensional disaster in DRL caused by a massive scale. Specifically, 1) we apply the representation learning models to describe the different nodes and tasks in EC, i.e., nodes and tasks are mapped to corresponding vector sub-spaces to reduce the dimensions and store the vector space efficiently. 2) With the space after dimensionality reduction, a DRL-based algorithm is employed to learn the vector representations of nodes and tasks and make scheduling decisions. 3) The experiments are conducted with the real-world data set, and the results show that the proposed representation model with DRL-based algorithm outperforms the baselines 18.04 and 9.94 percent on average regarding energy consumption and service level agreement violation (SLAV), respectively.
Zhiqing Tang, Wei Jia 0001, Wenmian Yang, Yongjian You
IEEE Trans. Big Data4
2021 An Effective Non-Autoregressive Model for Spoken Language Understanding
abstract
Spoken Language Understanding (SLU), a core component of the task-oriented dialogue system, expects a shorter inference latency due to the impatience of humans. Non-autoregressive SLU models clearly increase the inference speed but suffer uncoordinated-slot problems caused by the lack of sequential dependency information among each slot chunk. To gap this shortcoming, in this paper, we propose a novel non-autoregressive SLU model named Layered-Refine Transformer, which contains a Slot Label Generation (SLG) task and a Layered Refine Mechanism (LRM). SLG is defined as generating the next slot label with the token sequence and generated slot labels. With SLG, the non-autoregressive model can efficiently obtain dependency information during training and spend no extra time in inference. LRM predicts the preliminary SLU results from Transformer's middle states and utilizes them to guide the final prediction. Experiments on two public datasets indicate that our model significantly improves SLU performance (1.5% on Overall accuracy) while substantially speed up (more than 10 times) the inference process over the state-of-the-art baseline.
Lizhi Cheng, Weijia Jia 0001, Wenmian Yang
CIKM3
2021 A Result Based Portable Framework for Spoken Language Understanding
abstract
Spoken language understanding (SLU), which is a core component of the task-oriented dialogue system, has made substantial progress in the research of single-turn dialogue. However, the performance in multi-turn dialogue is still not satisfactory in the sense that the existing multi-turn SLU methods have low portability and compatibility for other single-turn SLU models. Further, existing multi-turn SLU methods do not exploit the historical predicted results when predicting the current utterance, which wastes helpful information. To gap those shortcomings, in this paper, we propose a novel Result-based Portable Framework for SLU (RPFSLU). RPFSLU allows most existing single-turn SLU models to obtain the contextual information from multi-turn dialogues and takes full advantage of predicted results in the dialogue history during the current prediction. Experimental results on the public dataset KVRET have shown that all SLU models in baselines acquire enhancement by RPFSLU on multi-turn SLU tasks.
Lizhi Cheng, Wenmian Yang, Weijia Jia 0001
ICME2
2020 MedDialog: Large-scale Medical Dialogue Datasets
abstract
Guangtao Zeng, Wenmian Yang, Zeqian Ju, Yue Yang, Sicheng Wang, Ruisi Zhang, Meng Zhou, Jiaqi Zeng, Xiangyu Dong, Ruoyu Zhang, Hongchao Fang, Penghui Zhu, Shu Chen, Pengtao Xie. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Guangtao Zeng, Wenmian Yang, Zeqian Ju, Sicheng Wang 0001, Ruisi Zhang, Jiaqi Zeng, Xiangyu Dong 0002, Hongchao Fang, Penghui Zhu, Pengtao Xie
EMNLP (1)2
2019 Improving Abstractive Document Summarization with Salient Information Modeling
abstract
Comprehensive document encoding and salient information selection are two major difficulties for generating summaries with adequate salient information.To tackle the above difficulties, we propose a Transformerbased encoder-decoder framework with two novel extensions for abstractive document summarization.Specifically, (1) to encode the documents comprehensively, we design a focus-attention mechanism and incorporate it into the encoder.This mechanism models a Gaussian focal bias on attention scores to enhance the perception of local context, which contributes to producing salient and informative summaries.(2) To distinguish salient information precisely, we design an independent saliency-selection network which manages the information flow from encoder to decoder.This network effectively reduces the influences of secondary information on the generated summaries.Experimental results on the popular CNN/Daily Mail benchmark demonstrate that our model outperforms other state-of-the-art baselines on the ROUGE metrics.
Yongjian You, Weijia Jia 0001, Wenmian Yang
ACL (1)4
2019 Interactive Variance Attention based Online Spoiler Detection for Time-Sync Comments
abstract
Nowadays, time-sync comment (TSC), a new form of interactive comments, has become increasingly popular on Chinese video websites. By posting TSCs, people can easily express their feelings and exchange their opinions with others when watching online videos. However, some spoilers appear among the TSCs. These spoilers reveal crucial plots in videos that ruin people's surprise when they first watch the video. In this paper, we proposed a novel Similarity-Based Network with Interactive Variance Attention (SBN-IVA) to classify comments as spoilers or not. In this framework, we firstly extract textual features of TSCs through the word-level attentive encoder. We design Similarity-Based Network (SBN) to acquire neighbor and keyframe similarity according to semantic similarity and timestamps of TSCs. Then, we implement Interactive Variance Attention (IVA) to eliminate the impact of noise comments. Finally, we obtain the likelihood of spoiler based on the difference between the neighbor and keyframe similarity. Experiments show SBN-IVA is on average 11.2% higher than the state-of-the-art method on F1-score in baselines.
Wenmian Yang, Weijia Jia 0001, Wenyuan Gao, Yutao Luo
CIKM1
2019 Herding Effect Based Attention for Personalized Time-Sync Video Recommendation
abstract
Time-sync comment (TSC) is a new form of user-interaction review associated with real-time video contents, which contains a user's preferences for videos and therefore well suited as the data source for video recommendations. However, existing review-based recommendation methods ignore the context-dependent (generated by user-interaction), real-time, and time-sensitive properties of TSC data. To bridge the above gaps, in this paper, we use video images and users' TSCs to design an Image-Text Fusion model with a novel Herding Effect Attention mechanism (called ITF-HEA), which can predict users' favorite videos with model-based collaborative filtering. Specifically, in the HEA mechanism, we weight the context information based on the semantic similarities and time intervals between each TSC and its context, thereby considering influences of the herding effect in the model. Experiments show that ITF-HEA is on average 3.78% higher than the state-of-the-art method upon F1-score in baselines.
Wenmian Yang, Wenyuan Gao, Weijia Jia 0001, Yutao Luo
ICME1
2019 Legal Judgment Prediction via Multi-Perspective Bi-Feedback Network
abstract
The Legal Judgment Prediction (LJP) is to determine judgment results based on the fact descriptions of the cases. LJP usually consists of multiple subtasks, such as applicable law articles prediction, charges prediction, and the term of the penalty prediction. These multiple subtasks have topological dependencies, the results of which affect and verify each other. However, existing methods use dependencies of results among multiple subtasks inefficiently. Moreover, for cases with similar descriptions but different penalties, current methods cannot predict accurately because the word collocation information is ignored. In this paper, we propose a Multi-Perspective Bi-Feedback Network with the Word Collocation Attention mechanism based on the topology structure among subtasks. Specifically, we design a multi-perspective forward prediction and backward verification framework to utilize result dependencies among multiple subtasks effectively. To distinguish cases with similar descriptions but different penalties, we integrate word collocations features of fact descriptions into the network via an attention mechanism. The experimental results show our model achieves significant improvements over baselines on all prediction tasks.
Wenmian Yang, Weijia Jia 0001, Yutao Luo
IJCAI1
2019 Time-Sync Video Tag Extraction Using Semantic Association Graph
abstract
Time-sync comments (TSCs) reveal a new way of extracting the online video tags. However, such TSCs have lots of noises due to users’ diverse comments, introducing great challenges for accurate and fast video tag extractions. In this article, we propose an unsupervised video tag extraction algorithm named Semantic Weight-Inverse Document Frequency (SW-IDF). Specifically, we first generate corresponding semantic association graph (SAG) using semantic similarities and timestamps of the TSCs. Second, we propose two graph cluster algorithms, i.e., dialogue-based algorithm and topic center-based algorithm, to deal with the videos with different density of comments. Third, we design a graph iteration algorithm to assign the weight to each comment based on the degrees of the clustered subgraphs, which can differentiate the meaningful comments from the noises. Finally, we gain the weight of each word by combining Semantic Weight (SW) and Inverse Document Frequency (IDF). In this way, the video tags are extracted automatically in an unsupervised way. Extensive experiments have shown that SW-IDF (dialogue-based algorithm) achieves 0.4210 F1-score and 0.4932 MAP (Mean Average Precision) in high-density comments, 0.4267 F1-score and 0.3623 MAP in low-density comments; while SW-IDF (topic center-based algorithm) achieves 0.4444 F1-score and 0.5122 MAP in high-density comments, 0.4207 F1-score and 0.3522 MAP in low-density comments. It has a better performance than the state-of-the-art unsupervised algorithms in both F1-score and MAP.
Wenmian Yang, Kun Wang 0005, Na Ruan, Wenyuan Gao, Weijia Jia 0001, Wei Zhao 0001, Yunyong Zhang
ACM Trans. Knowl. Discov. Data1
2017 Crowdsourced time-sync video tagging using semantic association graph
abstract
Time-sync comments reveal a new way of extracting the online video tags. However, such time-sync comments have lots of noises due to users' diverse comments, introducing great challenges for accurate and fast video tag extractions. In this paper, we propose an unsupervised video tag extraction algorithm named Semantic Weight-Inverse Document Frequency (SW-IDF). SW-IDF first generates corresponding semantic association graph (SAG) using semantic similarities and timestamps of the time-sync comments. Then it clusters the comments into sub-graphs of different topics and assigns weight to each comment based on SAG. This can clearly differentiate the meaningful comments with the noises. In this way, the noises can be identified, and effectively eliminated. Extensive experiments have shown that SW-IDF can achieve 0.3045 precision and 0.6530 recall in high-density comments; 0.3800 precision and 0.4460 recall in low-density comments. It is the best performance among the existing unsupervised algorithms.
Wenmian Yang, Na Ruan, Wenyuan Gao, Kun Wang 0005, Wensheng Ran, Weijia Jia 0001
ICME1