EDBT 2026 Demo / reviewers in the wild / expert
Xiaoyang Wang 0001
dblp:81/1832-1
· DBLP profile ↗
34ranked-venue papers
12as first author
16since 2021 · last 2025
0000-0002-0746-1059ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 31 · 12 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 10 first-author · 1 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Router-Tuning: A Simple and Effective Approach for Dynamic DepthabstractThe Mixture of Depths (MoD) was introduced to improve computational efficiency by dynamically skipping less important layers, reducing redundant computation while maintaining model capacity.Despite its promise, existing MoD approaches remain under-explored and face two main challenges: (1) high training costs due to the need to train the entire model along with the routers that determine which layers to skip, and (2) performance degradation when important layers are bypassed.In response to the first issue, we propose Router-Tuning, which fine-tunes only the routers on a small dataset, drastically reducing the computational overhead associated with full model training.For the second challenge, we investigate Router-Tuning across different architectures and granularities, demonstrating its effectiveness on Attention layers and MoE layers.This method preserves the model's performance while significantly enhancing computational and memory efficiency.Extensive experiments demonstrate that our approach delivers competitive results while dramatically improving the computation efficiency, e.g., 21% speedup and only a 0.2% performance drop. Shwai He, Tao Ge 0001, Guoheng Sun, Bowei Tian, Xiaoyang Wang 0001, Dong Yu 0001 |
EMNLP | 5 |
| 2025 | DSBench: How Far Are Data Science Agents from Becoming Data Science Experts?abstractLarge Language Models (LLMs) and Large Vision-Language Models (LVLMs) have demonstrated impressive language/vision reasoning abilities, igniting the recent trend of building agents for targeted applications such as shopping assistants or AI software engineers. Recently, many data science benchmarks have been proposed to investigate their performance in the data science domain. However, existing data science benchmarks still fall short when compared to real-world data science applications due to their simplified settings. To bridge this gap, we introduce DSBench, a comprehensive benchmark designed to evaluate data science agents with realistic tasks. This benchmark includes 466 data analysis tasks and 74 data modeling tasks, sourced from Eloquence and Kaggle competitions. DSBench offers a realistic setting by encompassing long contexts, multimodal task backgrounds, reasoning with large data files and multi-table structures, and performing end-to-end data modeling tasks. Our evaluation of state-of-the-art LLMs, LVLMs, and agents shows that they struggle with most tasks, with the best agent solving only 34.12% of data analysis tasks and achieving a 34.74% Relative Performance Gap (RPG). These findings underscore the need for further advancements in developing more practical, intelligent, and autonomous data science agents. Liqiang Jing, Zhehui Huang, Xiaoyang Wang 0001, Wenlin Yao, Wenhao Yu 0002, Kaixin Ma, Hongming Zhang 0009, Xinya Du, Dong Yu 0001 |
ICLR | 3 |
| 2024 | SportsMetrics: Blending Text and Numerical Data to Understand Information Fusion in LLMsabstractYebowen Hu, Kaiqiang Song, Sangwoo Cho, Xiaoyang Wang, Hassan Foroosh, Dong Yu, Fei Liu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Yebowen Hu, Kaiqiang Song, Sangwoo Cho, Xiaoyang Wang 0001, Hassan Foroosh, Dong Yu 0001, Fei Liu 0004 |
ACL (1) | 4 |
| 2024 | When Reasoning Meets Information Aggregation: A Case Study with Sports NarrativesabstractReasoning is most powerful when an LLM accurately aggregates relevant information.We examine the critical role of information aggregation in reasoning by requiring the LLM to analyze sports narratives.To succeed at this task, an LLM must infer points from actions, identify related entities, attribute points accurately to players and teams, and compile key statistics to draw conclusions.We conduct comprehensive experiments with real NBA basketball data and present SPORTSGEN, a new method to synthesize game narratives.By synthesizing data, we can rigorously evaluate LLMs' reasoning capabilities under complex scenarios with varying narrative lengths and density of information.Our findings show that most models, including GPT-4o, often fail to accurately aggregate basketball scores due to frequent scoring patterns.Open-source models like Llama-3 further suffer from significant score hallucinations.Finally, the effectiveness of reasoning is influenced by narrative complexity, information density, and domain-specific terms, highlighting the challenges in analytical reasoning tasks. 1 * Work done during Yebowen Hu's internship; Kaiqiang Song and Sangwoo Cho were full-time researchers at Tencent AI Lab, Seattle, USA at the time of this work. https://github.com/YebowenHu/SportsGenAnalyze the team-player affiliations and play-by-play descriptions below to determine the total points scored by each team (player).Please explain your reasoning step by step and provide the final results in JSON format.Start with: {New York Knicks: 0, Denver Nuggets: 0} ({Andrea Bargnani: 0, Timofey Mozgov: 0, ... Yebowen Hu, Kaiqiang Song, Sangwoo Cho, Xiaoyang Wang 0001, Wenlin Yao, Hassan Foroosh, Dong Yu 0001, Fei Liu 0004 |
EMNLP | 4 |
| 2024 | Polarity Calibration for Opinion SummarizationabstractYuanyuan Lei, Kaiqiang Song, Sangwoo Cho, Xiaoyang Wang, Ruihong Huang, Dong Yu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Yuanyuan Lei 0001, Kaiqiang Song, Sangwoo Cho, Xiaoyang Wang 0001, Ruihong Huang, Dong Yu 0001 |
NAACL-HLT | 4 |
| 2024 | MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction TuningabstractFuxiao Liu, Xiaoyang Wang, Wenlin Yao, Jianshu Chen, Kaiqiang Song, Sangwoo Cho, Yaser Yacoob, Dong Yu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Fuxiao Liu, Xiaoyang Wang 0001, Wenlin Yao, Jianshu Chen, Kaiqiang Song, Sangwoo Cho, Yaser Yacoob, Dong Yu 0001 |
NAACL-HLT | 2 |
| 2024 | From Language Modeling to Instruction Following: Understanding the Behavior Shift in LLMs after Instruction TuningabstractXuansheng Wu, Wenlin Yao, Jianshu Chen, Xiaoman Pan, Xiaoyang Wang, Ninghao Liu, Dong Yu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Xuansheng Wu, Wenlin Yao, Jianshu Chen, Xiaoman Pan, Xiaoyang Wang 0001, Ninghao Liu 0001, Dong Yu 0001 |
NAACL-HLT | 5 |
| 2023 | Generating User-Engaging News HeadlinesabstractPengshan Cai, Kaiqiang Song, Sangwoo Cho, Hongwei Wang, Xiaoyang Wang, Hong Yu, Fei Liu, Dong Yu. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Pengshan Cai, Kaiqiang Song, Sangwoo Cho, Hongwei Wang 0010, Xiaoyang Wang 0001, Hong Yu 0001, Fei Liu 0004, Dong Yu 0001 |
ACL (1) | 5 |
| 2023 | DecipherPref: Analyzing Influential Factors in Human Preference Judgments via GPT-4abstractHuman preference judgments are pivotal in guiding large language models (LLMs) to produce outputs that align with human values.Human evaluations are also used in summarization tasks to compare outputs from various systems, complementing existing automatic metrics.Despite their significance, however, there has been limited research probing these pairwise or kwise comparisons.The collective impact and relative importance of factors such as output length, informativeness, fluency, and factual consistency are still not well understood.It is also unclear if there are other hidden factors influencing human judgments.In this paper, we conduct an in-depth examination of a collection of pairwise human judgments released by Ope-nAI.Utilizing the Bradley-Terry-Luce (BTL) model, we reveal the inherent preferences embedded in these human judgments.We find that the most favored factors vary across tasks and genres, whereas the least favored factors tend to be consistent, e.g., outputs are too brief, contain excessive off-focus content or hallucinated facts.Our findings have implications on the construction of balanced datasets in human preference evaluations, which is a crucial step in shaping the behaviors of future LLMs. Yebowen Hu, Kaiqiang Song, Sangwoo Cho, Xiaoyang Wang 0001, Hassan Foroosh, Fei Liu 0004 |
EMNLP | 4 |
| 2023 | More Than Spoken Words: Nonverbal Message Extraction and GenerationabstractNonverbal messages (NM) such as speakers' facial expressions and speed of speech are essential for face-to-face communication, and they can be regarded as implicit knowledge as they are usually not included in existing dialogue understanding or generation tasks.This paper introduces the task of extracting NMs in written text and generating NMs for spoken text.Previous studies merely focus on extracting NMs from relatively small-scale well-structured corpora such as movie scripts wherein NMs are enclosed in parentheses by scriptwriters, which greatly decreases the difficulty of extraction.To enable extracting NMs from unstructured corpora, we annotate the first NM extraction dataset for Chinese based on novels and develop three baselines to extract single-span or multi-span NM of a target utterance from its surrounding context.Furthermore, we use the extractors to extract 749K (context, utterance, NM) triples from Chinese novels and investigate whether we can use them to improve NM generation via semi-supervised learning.Experimental results demonstrate that the automatically extracted triples can serve as high-quality augmentation data of clean triples extracted from scripts to generate more relevant, fluent, valid, and factually consistent 1 NMs than the purely supervised generator, and the resulting generator can in turn help Chinese dialogue understanding tasks such as dialogue machine reading comprehension and emotion classification by simply adding the predicted "unspoken" NM to each utterance or narrative in inputs. Dian Yu 0001, Xiaoyang Wang 0001, Wanshun Chen, Longyue Wang, Haitao Mi, Dong Yu 0001 |
EMNLP | 2 |
| 2022 | Towards Abstractive Grounded Summarization of Podcast TranscriptsabstractPodcasts have shown a recent rise in popularity.Summarization of podcasts is of practical benefit to both content providers and consumers.It helps people quickly decide whether they will listen to a podcast and/or reduces the cognitive load of content providers to write summaries.Nevertheless, podcast summarization faces significant challenges including factual inconsistencies of summaries with respect to the inputs.The problem is exacerbated by speech disfluencies and recognition errors in transcripts of spoken language.In this paper, we explore a novel abstractive summarization method to alleviate these issues.Our approach learns to produce an abstractive summary while grounding summary segments in specific regions of the transcript to allow for full inspection of summary details.We conduct a series of analyses of the proposed approach on a large podcast dataset and show that the approach can achieve promising results.Grounded summaries bring clear benefits in locating the summary and transcript segments that contain inconsistent information, and hence improve summarization quality in terms of automatic and human evaluation. Kaiqiang Song, Chen Li 0003, Xiaoyang Wang 0001, Dong Yu 0001, Fei Liu 0004 |
ACL (1) | 3 |
| 2022 | Toward Unifying Text Segmentation and Long Document SummarizationabstractText segmentation is important for signaling a document's structure.Without segmenting a long document into topically coherent sections, it is difficult for readers to comprehend the text, let alone find important information.The problem is only exacerbated by a lack of segmentation in transcripts of audio/video recordings.In this paper, we explore the role that section segmentation plays in extractive summarization of written and spoken documents.Our approach learns robust sentence representations by performing summarization and segmentation simultaneously, which is further enhanced by an optimization-based regularizer to promote selection of diverse summary sentences.We conduct experiments on multiple datasets ranging from scientific articles to spoken transcripts to evaluate the model's performance.Our findings suggest that the model can not only achieve state-of-the-art performance on publicly available benchmarks, but demonstrate better crossgenre transferability when equipped with text segmentation.We perform a series of analyses to quantify the impact of section segmentation on summarizing written and spoken documents of substantial length and complexity. Sangwoo Cho, Kaiqiang Song, Xiaoyang Wang 0001, Fei Liu 0004, Dong Yu 0001 |
EMNLP | 3 |
| 2022 | Salience Allocation as Guidance for Abstractive SummarizationabstractFei Wang, Kaiqiang Song, Hongming Zhang, Lifeng Jin, Sangwoo Cho, Wenlin Yao, Xiaoyang Wang, Muhao Chen, Dong Yu. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Fei Wang 0060, Kaiqiang Song, Hongming Zhang 0009, Lifeng Jin, Sangwoo Cho, Wenlin Yao, Xiaoyang Wang 0001, Muhao Chen 0001, Dong Yu 0001 |
EMNLP | 7 |
| 2022 | Z-LaVI: Zero-Shot Language Solver Fueled by Visual ImaginationabstractLarge-scale pretrained language models have made significant advances in solving downstream language understanding tasks.However, they generally suffer from reporting bias, the phenomenon describing the lack of explicit commonsense knowledge in written text, e.g., "an orange is orange".To overcome this limitation, we develop a novel approach, Z-LaVI, to endow language models with visual imagination capabilities.Specifically, we leverage two complementary types of "imaginations": (i) recalling existing images through retrieval and (ii) synthesizing nonexistent images via text-toimage generation.Jointly exploiting the language inputs and the imagination, a pretrained vision-language model (e.g., CLIP) eventually composes a zero-shot solution to the original language tasks.Notably, fueling language models with imagination can effectively leverage visual knowledge to solve plain language tasks.In consequence, Z-LaVI consistently improves the zero-shot performance of existing language models across a diverse set of language tasks. 1 * Work was done during the internship at Tencent AI Lab.(a) Word Sense Disambiguation (b) Science Question Answering (c) Topic Classification sense1: bank(institute) sense2: bank(geography) Input: The species prefers the {bank} of pond. Wenlin Yao, Hongming Zhang 0009, Xiaoyang Wang 0001, Dong Yu 0001, Jianshu Chen |
EMNLP | 4 |
| 2022 | Meta-learning without data via Wasserstein distributionally-robust model fusionabstractExisting meta-learning works assume that each task has available training and testing data. However, there are many available pre-trained models without accessing their training data in practice. We often need a single model to solve different tasks simultaneously as this is much more convenient to deploy the models. Our work aims to meta-learn a model initialization from these pre-trained models without using corresponding training data. We name this challenging problem setting as Data-Free Learning To Learn (DFL2L). We propose a distributionally robust optimization (DRO) framework to learn a black-box model to fuse and compress all the pre-trained models into a single network to address this problem. To encourage good generalization to the unseen new tasks, the proposed DRO framework diversifies the learned task embedding associated with each pre-trained model to cover the diversity in the underlying training task distributions. A model initialization is sampled from the black-box network during meta-testing as the meta learned initialization. Extensive experiments on offline and online DFL2L settings and several real image datasets demonstrate the effectiveness of the proposed methods. Zhenyi Wang 0001, Xiaoyang Wang 0001, Li Shen 0008, Qiuling Suo, Kaiqiang Song, Dong Yu 0001, Yan Shen 0002, Mingchen Gao |
UAI | 2 |
| 2021 | NaturalConv: A Chinese Dialogue Dataset Towards Multi-turn Topic-driven ConversationabstractIn this paper, we propose a Chinese multi-turn topic-driven conversation dataset, NaturalConv, which allows the participants to chat anything they want as long as any element from the topic is mentioned and the topic shift is smooth. Our corpus contains 19.9K conversations from six domains, and 400K utterances with an average turn number of 20.1. These conversations contain in-depth discussions on related topics or widely natural transition between multiple topics. We believe either way is normal for human conversation. To facilitate the research on this corpus, we provide results of several benchmark models. Comparative results show that for this dataset, our current models are not able to provide significant improvement by introducing background knowledge/topic. Therefore, the proposed dataset should be a good benchmark for further research to evaluate the validity and naturalness of multi-turn conversation systems. Our dataset is available at https://ai.tencent.com/ailab/nlp/dialogue/#datasets. Xiaoyang Wang 0001, Chen Li 0003, Jianqiao Zhao, Dong Yu 0001 |
AAAI | 1 |
| 2020 | Towards Faithful Neural Table-to-Text Generation with Content-Matching ConstraintsabstractText generation from a knowledge base aims to translate knowledge triples to naturallanguage descriptions.Most existing methods ignore the faithfulness between a generated text description and the original table, leading to generated information that goes beyond the content of the table.In this paper, for the first time, we propose a novel Transformerbased generation framework to achieve the goal.The core techniques in our method to enforce faithfulness include a new table-text optimal-transport matching loss and a tabletext embedding similarity loss based on the Transformer model.Furthermore, to evaluate faithfulness, we propose a new automatic metric specialized to the table-to-text generation problem.We also provide detailed analysis on each component of our model in our experiments.Automatic and human evaluations show that our framework can significantly outperform state-of-the-art by a large margin. Zhenyi Wang 0001, Xiaoyang Wang 0001, Bang An 0001, Dong Yu 0001, Changyou Chen |
ACL | 2 |
| 2020 | Requet: Real-Time QoE Metric Detection for Encrypted YouTube TrafficabstractAs video traffic dominates the Internet, it is important for operators to detect video quality of experience (QoE) to ensure adequate support for video traffic. With wide deployment of end-to-end encryption, traditional deep packet inspection--based traffic monitoring approaches are becoming ineffective. This poses a challenge for network operators to monitor user QoE and improve upon their experience. To resolve this issue, we develop and present a system for RE al-time QU ality of experience metric detection for E ncrypted T raffic— Requet —which is suitable for network middlebox deployment. Requet uses a detection algorithm that we develop to identify video and audio chunks from the IP headers of encrypted traffic. Features extracted from the chunk statistics are used as input to a machine learning algorithm to predict QoE metrics, specifically buffer warning (low buffer, high buffer), video state (buffer increase, buffer decay, steady, stall), and video resolution. We collect a large YouTube dataset consisting of diverse video assets delivered over various WiFi and LTE network conditions to evaluate the performance. We compare Requet with a baseline system based on previous work and show that Requet outperforms the baseline system in accuracy of predicting buffer low warning, video state, and video resolution by 1.12×, 1.53×, and 3.14×, respectively. Craig Gutterman, Katherine Guo, Sarthak Arora, Trey Gilliland, Xiaoyang Wang 0001, Les Wu, Ethan Katz-Bassett, Gil Zussman |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2019 | Requet: real-time QoE detection for encrypted YouTube trafficabstractAs video traffic dominates the Internet, it is important for operators to detect video Quality of Experience (QoE) in order to ensure adequate support for video traffic. With wide deployment of end-to-end encryption, traditional deep packet inspection based traffic monitoring approaches are becoming ineffective. This poses a challenge for network operators to monitor user QoE and improve upon their experience. To resolve this issue, we develop and present a system for REal-time QUality of experience metric detection for Encrypted Traffic, Requet. Requet uses a detection algorithm we develop to identify video and audio chunks from the IP headers of encrypted traffic. Features extracted from the chunk statistics are used as input to a Machine Learning (ML) algorithm to predict QoE metrics, specifically, buffer warning (low buffer, high buffer), video state (buffer increase, buffer decay, steady, stall), and video resolution. We collect a large YouTube dataset consisting of diverse video assets delivered over various WiFi network conditions to evaluate the performance. We compare Requet with a baseline system based on previous work and show that Requet outperforms the baseline system in accuracy of predicting buffer low warning, video state, and video resolution by 1.12X, 1.53X, and 3.14X, respectively. Craig Gutterman, Katherine Guo, Sarthak Arora, Xiaoyang Wang 0001, Les Wu, Ethan Katz-Bassett, Gil Zussman |
MMSys | 4 |
| 2017 | Hierarchical Context Modeling for Video Event RecognitionabstractCurrent video event recognition research remains largely target-centered. For real-world surveillance videos, targetcentered event recognition faces great challenges due to large intra-class target variation, limited image resolution, and poor detection and tracking results. To mitigate these challenges, we introduced a context-augmented video event recognition approach. Specifically, we explicitly capture different types of contexts from three levels including image level, semantic level, and prior level. At the image level, we introduce two types of contextual features including the appearance context features and interaction context features to capture the appearance of context objects and their interactions with the target objects. At the semantic level, we propose a deep model based on deep Boltzmann machine to learn event object representations and their interactions. At the prior level, we utilize two types of prior-level contexts including scene priming and dynamic cueing. Finally, we introduce a hierarchical context model that systematically integrates the contextual information at different levels. Through the hierarchical context model, contexts at different levels jointly contribute to the event recognition. We evaluate the hierarchical context model for event recognition on benchmark surveillance video datasets. Results show that incorporating contexts in each level can improve event recognition performance, and jointly integrating three levels of contexts through our hierarchical model achieves the best performance. Xiaoyang Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2016 | Taxonomy augmented object recognitionabstractRealistic scene object recognition in computer vision still faces great challenges due to the large intra-class variation of object images caused by factors like object appearance variation and viewpoint change. To address this challenge, we propose to exploit the semantic relationships embedded in object taxonomy for improved object recognition. Specifically, we exploit the relationships in the object taxonomy to augment the learning of object classifiers. We utilize two types of relationships in the taxonomy, including the overall relationship and the local relationship. Our proposed approach jointly incorporates both the overall relationship as the loss function for classifier learning, and the local relationship as classifier learning constraints. Experiments on benchmark datasets demonstrate the effectiveness of our method in incorporating taxonomy for object recognition compared to the state-art-the-art methods. Xiaoyang Wang 0001, Yue Zhao 0013 |
ICPR | 1 |
| 2016 | Multilingual articulatory features augmentation learningabstractArticulatory features are used as an universal set of speech attributes shared across many different languages. Some multilingual and cross-language speech recognition systems using articulatory features have been shown to improve the performance. The existing articulatory features are defined by phonetician as a set of articulatory descriptions of phones, which represent some semantic information explaining how humans produce speech sounds via the interaction of different physiological structures. But these manually specified attributes suffer from the incomplete capturing articulation information of all languages and are not distinctive enough for accurate monolingual and multilingual phoneme recognition. In this paper, we are solving the problem of a more complete set of articulatory features representation by sparse coding methods. We learned the latent attributes that sparsely represent more speech articulation information sharing between English and Tibetan languages. Models based on the concatenated semantic and latent speech attributes performed the better accuracy over the existing methods in our experiments for English-Tibetan bilingual phone recognition. Yue Zhao 0013, Rui Zhao 0015, Xiaoyang Wang 0001 |
ICPR | 3 |
| 2016 | Object Recognition with Hidden Attributes
Xiaoyang Wang 0001 |
IJCAI | 1 |
| 2015 | Video event recognition with deep hierarchical context modelabstractVideo event recognition still faces great challenges due to large intra-class variation and low image resolution, in particular for surveillance videos. To mitigate these challenges and to improve the event recognition performance, various context information from the feature level, the semantic level, as well as the prior level is utilized. Different from most existing context approaches that utilize context in one of the three levels through shallow models like support vector machines, or probabilistic models like BN and MRF, we propose a deep hierarchical context model that simultaneously learns and integrates context at all three levels, and holistically utilizes the integrated contexts for event recognition. We first introduce two types of context features describing the event neighborhood, and then utilize the proposed deep model to learn the middle level representations and combine the bottom feature level, middle semantic level and top prior level contexts together for event recognition. The experiments on state of art surveillance video event benchmarks including VIRAT 1.0 Ground Dataset, VIRAT 2.0 Ground Dataset, and the UT-Interaction Dataset demonstrate that the proposed model is quite effective in utilizing the context information for event recognition. It outperforms the existing context approaches that also utilize multiple level contexts on these event benchmarks. Xiaoyang Wang 0001 |
CVPR | 1 |
| 2014 | A Hierarchical Context Model for Event Recognition in Surveillance VideoabstractDue to great challenges such as tremendous intra-class variations and low image resolution, context information has been playing a more and more important role for accurate and robust event recognition in surveillance videos. The context information can generally be divided into the feature level context, the semantic level context, and the prior level context. These three levels of context provide crucial bottom-up, middle level, and top down information that can benefit the recognition task itself. Unlike existing researches that generally integrate the context information at one of the three levels, we propose a hierarchical context model that simultaneously exploits contexts at all three levels and systematically incorporate them into event recognition. To tackle the learning and inference challenges brought in by the model hierarchy, we develop complete learning and inference algorithms for the proposed hierarchical context model based on variational Bayes method. Experiments on VIRAT 1.0 and 2.0 Ground Datasets demonstrate the effectiveness of the proposed hierarchical context model for improving the event recognition performance even under great challenges like large intra-class variations and low image resolution. Xiaoyang Wang 0001 |
CVPR | 1 |
| 2014 | Attribute Augmentation with Sparse CodingabstractThis work proposes a novel sparse coding based approach for augmenting attributes in both object recognition and facial expression recognition applications. Attributes are a set of manually specified binary descriptions of visual objects. Though playing an important role in different applications like zero-shot learning, image description and recognition, the manually specified attributes suffer from the incomplete capturing of the original image data. In this work, we propose to augment the original manually specified semantic attributes with the augmented attributes which are also sparse, based on the minimization of the reconstruction error between the original image and the concatenated semantic and augmented attributes. We propose to iteratively learn the dictionaries as well as recover the augmented attributes in the optimization. For our applications of object recognition and facial expression recognition, the augmented attributes combined with the predicted semantic attributes can improve the overall recognition rate. Also, our learned dictionaries show certain meanings captured by the attributes. Xiaoyang Wang 0001 |
ICPR | 1 |
| 2014 | Learning with Hidden InformationabstractIn many classification problems, there exists additional information which is available during training but not available during testing. In this paper we denote such information as hidden information, and study how to incorporate it to improve the learning performance. Despite its importance, learning with hidden information has not attracted enough attention from the field and existing work in this area remains limited. In this paper we make improvements from two perspectives. First, unlike the related work, we propose a general framework to capture hidden information, which is not limited to a specific type of classifier but is widely applicable to different classifiers. Second, borrowing the tool of Bootstrap widely used in statistics, we are able to numerically quantify the benefits and identify the most useful hidden information. Experiments on both digit and object recognition demonstrate the effectiveness of the proposed approach. Ziheng Wang 0001, Xiaoyang Wang 0001 |
ICPR | 2 |
| 2014 | Context augmented Dynamic Bayesian Networks for event recognition
Xiaoyang Wang 0001 |
Pattern Recognit. Lett. | 1 |
| 2013 | A Unified Probabilistic Approach Modeling Relationships between Attributes and ObjectsabstractThis paper proposes a unified probabilistic model to model the relationships between attributes and objects for attribute prediction and object recognition. As a list of semantically meaningful properties of objects, attributes generally relate to each other statistically. In this paper, we propose a unified probabilistic model to automatically discover and capture both the object-dependent and object-independent attribute relationships. The model utilizes the captured relationships to benefit both attribute prediction and object recognition. Experiments on four benchmark attribute datasets demonstrate the effectiveness of the proposed unified model for improving attribute prediction as well as object recognition in both standard and zero-shot learning cases. Xiaoyang Wang 0001 |
ICCV | 1 |
| 2012 | Incorporating contextual knowledge to Dynamic Bayesian Networks for event recognition
Xiaoyang Wang 0001 |
ICPR | 1 |
| 2012 | A novel probabilistic approach utilizing clip attribute as hidden knowledge for event recognition
Xiaoyang Wang 0001 |
ICPR | 1 |
| 2012 | Learning dynamic Bayesian network discriminatively for human activity recognition
Xiaoyang Wang 0001 |
ICPR | 1 |
| 2011 | AVSS 2011 demo session: A large-scale benchmark dataset for event recognition in surveillance videoabstractSummary form only given. We present a concept for automatic construction site monitoring by taking into account 4D information (3D over time), that is acquired from highly-overlapping digital aerial images. On the one hand today's maturity of flying micro aerial vehicles (MAVs) enables a low-cost and an efficient image acquisition of high-quality data that maps construction sites entirely from many varying viewpoints. On the other hand, due to low-noise sensors and high redundancy in the image data, recent developments in 3D reconstruction workflows have benefited the automatic computation of accurate and dense 3D scene information. Having both an inexpensive high-quality image acquisition and an efficient 3D analysis workflow enables monitoring, documentation and visualization of observed sites over time with short intervals. Relating acquired 4D site observations, composed of color, texture, geometry over time, largely supports automated methods toward full scene understanding, the acquisition of both the change and the construction site's progress. Sangmin Oh, Anthony Hoogs, A. G. Amitha Perera, Naresh P. Cuntoor, Chia-Chih Chen, Jong Taek Lee, Saurajit Mukherjee, Jake K. Aggarwal, Hyungtae Lee, Larry Davis 0001, Eran Swears, Xiaoyang Wang 0001, Kishore K. Reddy, Mubarak Shah, Carl Vondrick, Hamed Pirsiavash, Deva Ramanan, Jenny Yuen, Antonio Torralba 0001, Bi Song, Anesco Fong, Amit K. Roy-Chowdhury, Mita Desai |
AVSS | 12 |
| 2011 | A large-scale benchmark dataset for event recognition in surveillance videoabstractWe introduce a new large-scale video dataset designed to assess the performance of diverse visual event recognition algorithms with a focus on continuous visual event recognition (CVER) in outdoor areas with wide coverage. Previous datasets for action recognition are unrealistic for real-world surveillance because they consist of short clips showing one action by one individual [15, 8]. Datasets have been developed for movies [11] and sports [12], but, these actions and scene conditions do not apply effectively to surveillance videos. Our dataset consists of many outdoor scenes with actions occurring naturally by non-actors in continuously captured videos of the real world. The dataset includes large numbers of instances for 23 event types distributed throughout 29 hours of video. This data is accompanied by detailed annotations which include both moving object tracks and event examples, which will provide solid basis for large-scale evaluation. Additionally, we propose different types of evaluation modes for visual recognition tasks and evaluation metrics along with our preliminary experimental results. We believe that this dataset will stimulate diverse aspects of computer vision research and help us to advance the CVER tasks in the years ahead. Sangmin Oh, Anthony Hoogs, A. G. Amitha Perera, Naresh P. Cuntoor, Chia-Chih Chen, Jong Taek Lee, Saurajit Mukherjee, Jake K. Aggarwal, Hyungtae Lee, Larry Davis 0001, Eran Swears, Xiaoyang Wang 0001, Kishore K. Reddy, Mubarak Shah, Carl Vondrick, Hamed Pirsiavash, Deva Ramanan, Jenny Yuen, Antonio Torralba 0001, Bi Song, Anesco Fong, Amit K. Roy-Chowdhury, Mita Desai |
CVPR | 12 |