EDBT 2026 Demo / reviewers in the wild / expert
Ming-Wei Chang
dblp:69/4618
· DBLP profile ↗
63ranked-venue papers
15as first author
17since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 55 · 14 first-author · 17 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Instruct-Imagen: Image Generation with Multi-modal InstructionabstractThis paper presents Instruct-Imagen, a model that tackles heterogeneous image generation tasks and generalizes across unseen tasks. We introduce multi-modal in-struction for image generation, a task representation artic-ulating a range of generation intents with precision. It uses natural language to amalgamate disparate modalities (e.g., text, edge, style, subject, etc.), such that abundant generation intents can be standardized in a uniform format. We then build Instruct - Imagen by fine-tuning a pre-trained text-to-image diffusion model with two stages. First, we adapt the model using the retrieval-augmented training, to enhance model's capabilities to ground its generation on external multi-modal context. Subsequently, we fine-tune the adapted model on diverse image generation tasks that requires vision-language understanding (e.g., subject-driven generation, etc.), each paired with a multi-modal instruction encapsulating the task's essence. Human evaluation on various image generation datasets re-veals that Instruct-Imagen matches or surpasses prior task-specific models in-domain and demonstrates promising generalization to unseen and more complex tasks. Our evaluation suite will be made publicly available. Hexiang Hu, Kelvin C. K. Chan, Yu-Chuan Su, Wenhu Chen, Yandong Li, Kihyuk Sohn, Xue Ben, Boqing Gong, William W. Cohen, Ming-Wei Chang, Xuhui Jia |
CVPR | 11 |
| 2024 | MagicLens: Self-Supervised Image Retrieval with Open-Ended InstructionsabstractImage retrieval, i.e., finding desired images given a reference image, inherently encompasses rich, multi-faceted search intents that are difficult to capture solely using image-based measures. Recent works leverage text instructions to allow users to more freely express their search intents. However, they primarily focus on image pairs that are visually similar and/or can be characterized by a small set of pre-defined relations. The core thesis of this paper is that text instructions can enable retrieving images with richer relations beyond visual similarity. To show this, we introduce MagicLens, a series of self-supervised image retrieval models that support open-ended instructions. MagicLens is built on a key novel insight: image pairs that naturally occur on the same web pages contain a wide range of implicit relations (e.g., inside view of), and we can bring those implicit relations explicit by synthesizing instructions via foundation models. Trained on 36.7M (query image, instruction, target image) triplets with rich semantic relations mined from the web, MagicLens achieves results comparable with or better than prior best on eight benchmarks of various image retrieval tasks, while maintaining high parameter efficiency with a significantly smaller model size. Additional human analyses on a 1.4M-image unseen corpus further demonstrate the diversity of search intents supported by MagicLens. Code and models are publicly available at the https://open-vision-language.github.io/MagicLens/. Kai Zhang 0033, Yi Luan, Hexiang Hu, Kenton Lee, Siyuan Qiao, Wenhu Chen, Yu Su 0001, Ming-Wei Chang |
ICML | 8 |
| 2023 | QUEST: A Retrieval Dataset of Entity-Seeking Queries with Implicit Set OperationsabstractFormulating selective information needs results in queries that implicitly specify set operations, such as intersection, union, and difference.For instance, one might search for "shorebirds that are not sandpipers" or "science-fiction films shot in England".To study the ability of retrieval systems to meet such information needs, we construct QUEST, a dataset of 3357 natural language queries with implicit set operations, that map to a set of entities corresponding to Wikipedia documents.The dataset challenges models to match multiple constraints mentioned in queries with corresponding evidence in documents and correctly perform various set operations.The dataset is constructed semi-automatically using Wikipedia category names.Queries are automatically composed from individual categories, then paraphrased and further validated for naturalness and fluency by crowdworkers.Crowdworkers also assess the relevance of entities based on their documents and highlight attribution of query constraints to spans of document text.We analyze several modern retrieval systems, finding that they often struggle on such queries.Queries involving negation and conjunction are particularly challenging and systems are further challenged with combinations of these operations. 1 * Work done during an internship at Google.retrieving an exhaustive document set, instead lim-65 iting annotation to the top few results of a baseline 66 information retrieval system.67 To analyze how well retrieval systems handle 68 such queries, we present QUEST, a dataset with 69 natural language queries from four domains, that 70 are mapped to relatively comprehensive sets of en-71 tities corresponding to Wikipedia pages.We use 72 Wikipedia categories and their mapping to entities 73 in Wikipedia as a building block for our dataset 74 construction approach, but do not allow access to 75 this semi-structured data source at inference time, 76 to simulate text-based retrieval.Wikipedia cate-77 gories represent a broad set of natural language 78 descriptions of entity properties and often corre-79 spond to selective information need queries that 80 could be plausibly issued by a search engine user 81 ([At least 90% of the time based on our filtering?]).82 The correspondence between property names and 83 document text is also often subtle and requires so-84 phisticated reasoning to determine relevance, rep-85 resenting the natural language inference challenge 86 inherent in the task, while the knowledge of cate-87 gory membership allows us to construct relatively 88 comprehensive sets of candidate entities for atomic 89 categories and their combinations.90 Our dataset construction process is outlined in 91 Figure 1.The base queries in our dataset are 92 semi-automatically generated using Wikipedia cat-93 egory names.To construct queries, we sample 94 category names and compose them into complex 95 queries by using pre-defined templates (for exam-96 ple, A \ B \ C).Next, we ask crowdworkers to 97 paraphrase these automatically generated queries, 98 while ensuring that the paraphrased queries are 99 fluent and clearly describe what a user could be 00 looking for.These are then validated for natural-01 ness and fluency by a different set of crowdworkers, 02 and filtered according to those criteria.Finally, for 03 a large subset of our dataset, we collect scalar rel-04 evance labels based on the entity documents, and 05 textual attributions mapping query constraints to 06 spans of document text, to aid the development of 07 systems that can make precise inferences based on 08 trusted sources.09 Performing well on this dataset requires sys-10 tems that can match query constraints with cor-11 Chaitanya Malaviya, Peter Shaw 0004, Ming-Wei Chang, Kenton Lee, Kristina Toutanova |
ACL (1) | 3 |
| 2023 | Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?abstractPre-trained vision and language models (Chen et al., 2023b,a;Dai et al., 2023; Li et al., 2023b) have demonstrated state-of-the-art capabilities over existing tasks involving images and texts, including visual question answering.However, it remains unclear whether these models possess the capability to answer questions that are not only querying visual content but knowledge-intensive and informationseeking.In this study, we introduce INFOS-EEK 1 , a visual question answering dataset tailored for information-seeking questions that cannot be answered with only common sense knowledge.Using INFOSEEK, we analyze various pre-trained visual question answering models and gain insights into their characteristics.Our findings reveal that state-of-the-art pre-trained multi-modal models (e.g., PaLI-X, BLIP2, etc.) face challenges in answering visual information-seeking questions, but finetuning on the INFOSEEK dataset elicits models to use fine-grained knowledge that was learned during their pre-training.Furthermore, we show that accurate visual entity recognition can be used to improve performance on INFOSEEK by retrieving relevant documents, showing a significant space for improvement.* Work done when interned at Google 1 Our dataset is available at https:// open-vision-language.github.io/infoseek/.Dataset OK-VQA ViQuAE INFOSEEK PaLM (Q-only) 23.8 31.5 5.6 Current SotA 66.1 22.1 18.2 Require Knowledge † 29.2% 95.2% 95.6% † :% of questions that require knowledge to answer.PaLM (Q-only): a question-only baseline using PaLM. Yang Chen 0065, Hexiang Hu, Yi Luan, Haitian Sun, Soravit Changpinyo, Alan Ritter, Ming-Wei Chang |
EMNLP | 7 |
| 2023 | Open-domain Visual Entity Recognition: Towards Recognizing Millions of Wikipedia EntitiesabstractLarge-scale multi-modal pre-training models such as CLIP [30] and PaLI [8] exhibit strong generalization on various visual domains and tasks. However, existing image classification benchmarks often evaluate recognition on a specific domain (e.g., outdoor images) or a specific task (e.g., classifying plant species), which falls short of evaluating whether pre-trained foundational models are universal visual recognizers. To address this, we formally present the task of Open-domain Visual Entity recognitioN (Oven), where a model need to link an image onto a Wikipedia entity with respect to a text query. We construct Oven-Wiki‡by repurposing 14 existing datasets with all labels grounded onto one single label space: Wikipedia entities. Oven-Wiki challenges models to select among six million possible Wikipedia entities, making it a general visual recognition benchmark with the largest number of labels. Our study on state-ofthe-art pre-trained models reveals large headroom in generalizing to the massive-scale label space. We show that a PaLI-based auto-regressive visual recognition model performs surprisingly well, even on Wikipedia entities that have never been seen during fine-tuning. We also find existing pretrained models yield different strengths: while PaLI-based models obtain higher overall performance, CLIP-based models are better at recognizing tail entities. Hexiang Hu, Yi Luan, Yang Chen 0065, Urvashi Khandelwal, Mandar Joshi, Kenton Lee, Kristina Toutanova, Ming-Wei Chang |
ICCV | 8 |
| 2023 | Promptagator: Few-shot Dense Retrieval From 8 Examples
Zhuyun Dai, Vincent Y. Zhao, Ji Ma 0004, Yi Luan, Jianmo Ni, Jing Lu 0014, Anton Bakalov, Kelvin Guu, Keith B. Hall, Ming-Wei Chang |
ICLR | 10 |
| 2023 | Pix2Struct: Screenshot Parsing as Pretraining for Visual Language UnderstandingabstractVisually-situated language is ubiquitous---sources range from textbooks with diagrams to web pages with images and tables, to mobile apps with buttons and forms. Perhaps due to this diversity, previous work has typically relied on domain-specific recipes with limited sharing of the underlying data, model architectures, and objectives. We present Pix2Struct, a pretrained image-to-text model for purely visual language understanding, which can be finetuned on tasks containing visually-situated language. Pix2Struct is pretrained by learning to parse masked screenshots of web pages into simplified HTML. The web, with its richness of visual elements cleanly reflected in the HTML structure, provides a large source of pretraining data well suited to the diversity of downstream tasks. Intuitively, this objective subsumes common pretraining signals such as OCR, language modeling, and image captioning. In addition to the novel pretraining strategy, we introduce a variable-resolution input representation and a more flexible integration of language and vision inputs, where language prompts such as questions are rendered directly on top of the input image. For the first time, we show that a single pretrained model can achieve state-of-the-art results in six out of nine tasks across four domains: documents, illustrations, user interfaces, and natural images. Kenton Lee, Mandar Joshi, Iulia Turc, Hexiang Hu, Fangyu Liu 0001, Julian Martin Eisenschlos, Urvashi Khandelwal, Peter Shaw 0004, Ming-Wei Chang, Kristina Toutanova |
ICML | 9 |
| 2023 | Subject-driven Text-to-Image Generation via Apprenticeship LearningabstractRecent text-to-image generation models like DreamBooth have made remarkable progress in generating highly customized images of a target subject, by fine-tuning an ``expert model'' for a given subject from a few examples.
However, this process is expensive, since a new expert model must be learned for each subject.
In this paper, we present SuTI, a Subject-driven Text-to-Image generator that replaces subject-specific fine tuning with {in-context} learning.
Given a few demonstrations of a new subject, SuTI can instantly generate novel renditions of the subject in different scenes, without any subject-specific optimization.
SuTI is powered by {apprenticeship learning}, where a single apprentice model is learned from data generated by a massive number of subject-specific expert models.
Specifically, we mine millions of image clusters from the Internet, each centered around a specific visual subject. We adopt these clusters to train a massive number of expert models, each specializing in a different subject. The apprentice model SuTI then learns to imitate the behavior of these fine-tuned experts.
SuTI can generate high-quality and customized subject-specific images 20x faster than optimization-based SoTA methods. On the challenging DreamBench and DreamBench-v2, our human evaluation shows that SuTI significantly outperforms existing models like InstructPix2Pix, Textual Inversion, Imagic, Prompt2Prompt, Re-Imagen and DreamBooth. Wenhu Chen, Hexiang Hu, Yandong Li, Nataniel Ruiz, Xuhui Jia, Ming-Wei Chang, William W. Cohen |
NeurIPS | 6 |
| 2023 | Rethinking the Role of Token Retrieval in Multi-Vector RetrievalabstractMulti-vector retrieval models such as ColBERT [Khattab et al., 2020] allow token-level interactions between queries and documents, and hence achieve state of the art on many information retrieval benchmarks. However, their non-linear scoring function cannot be scaled to millions of documents, necessitating a three-stage process for inference: retrieving initial candidates via token retrieval, accessing all token vectors, and scoring the initial candidate documents. The non-linear scoring function is applied over all token vectors of each candidate document, making the inference process complicated and slow. In this paper, we aim to simplify the multi-vector retrieval by rethinking the role of token retrieval. We present XTR, ConteXtualized Token Retriever, which introduces a simple, yet novel, objective function that encourages the model to retrieve the most important document tokens first. The improvement to token retrieval allows XTR to rank candidates only using the retrieved tokens rather than all tokens in the document, and enables a newly designed scoring stage that is two-to-three orders of magnitude cheaper than that of ColBERT. On the popular BEIR benchmark, XTR advances the state-of-the-art by 2.8 nDCG@10 without any distillation. Detailed analysis confirms our decision to revisit the token retrieval stage, as XTR demonstrates much better recall of the token retrieval stage compared to ColBERT. Jinhyuk Lee, Zhuyun Dai, Sai Meher Karthik Duddu, Tao Lei 0001, Iftekhar Naim, Ming-Wei Chang, Vincent Y. Zhao |
NeurIPS | 6 |
| 2023 | Conditional Adapters: Parameter-efficient Transfer Learning with Fast InferenceabstractWe propose Conditional Adapter (CoDA), a parameter-efficient transfer learning method that also improves inference efficiency. CoDA generalizes beyond standard adapter approaches to enable a new way of balancing speed and accuracy using conditional computation.
Starting with an existing dense pretrained model, CoDA adds sparse activation together with a small number of new parameters and a light-weight training phase.
Our experiments demonstrate that the CoDA approach provides an unexpectedly efficient way to transfer knowledge.
Across a variety of language, vision, and speech tasks, CoDA achieves a 2x to 8x inference speed-up compared to the state-of-the-art Adapter approaches with moderate to no accuracy loss and the same parameter efficiency. Tao Lei 0001, Junwen Bai, Siddhartha Brahma, Joshua Ainslie, Kenton Lee, Yanqi Zhou, Nan Du 0002, Vincent Y. Zhao, Yuexin Wu, Bo Li 0028, Yu Zhang 0033, Ming-Wei Chang |
NeurIPS | 12 |
| 2022 | Meta-Learning Fast Weight Language ModelsabstractDynamic evaluation of language models (LMs) adapts model parameters at test time using gradient information from previous tokens and substantially improves LM performance.However, it requires over 3x more compute than standard inference.We present Fast Weight Layers (FWLs), a neural component that provides the benefits of dynamic evaluation much more efficiently by expressing gradient updates as linear attention.A key improvement over dynamic evaluation is that FWLs can also be applied at training time so the model learns to make good use of gradient updates.FWLs can easily be added on top of existing transformer models, require relatively little extra compute or memory to run, and significantly improve language modeling perplexity. Kevin Clark, Kelvin Guu, Ming-Wei Chang, Panupong Pasupat, Geoffrey E. Hinton, Mohammad Norouzi 0002 |
EMNLP | 3 |
| 2022 | Large Dual Encoders Are Generalizable RetrieversabstractJianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai, Gustavo Hernandez Abrego, Ji Ma, Vincent Zhao, Yi Luan, Keith Hall, Ming-Wei Chang, Yinfei Yang. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Jianmo Ni, Chen Qu 0001, Jing Lu 0014, Zhuyun Dai, Gustavo Hernández Ábrego, Ji Ma 0004, Vincent Y. Zhao, Yi Luan, Keith B. Hall, Ming-Wei Chang, Yinfei Yang |
EMNLP | 10 |
| 2022 | ASQA: Factoid Questions Meet Long-Form AnswersabstractAn abundance of datasets and availability of reliable evaluation metrics have resulted in strong progress in factoid question answering (QA).This progress, however, does not easily transfer to the task of long-form QA, where the goal is to answer questions that require in-depth explanations.The hurdles include (i) a lack of high-quality data, and (ii) the absence of a well-defined notion of the answer's quality.In this work, we address these problems by (i) releasing a novel dataset and a task that we call ASQA (Answer Summaries for Questions which are Ambiguous); and (ii) proposing a reliable metric for measuring performance on ASQA.Our task focuses on factoid questions that are ambiguous, that is, have different correct answers depending on interpretation.Answers to ambiguous questions should synthesize factual information from multiple sources into a long-form summary that resolves the ambiguity.In contrast to existing long-form QA tasks (such as ELI5), ASQA admits a clear notion of correctness: a user faced with a good summary should be able to answer different interpretations of the original ambiguous question.We use this notion of correctness to define an automated metric of performance for ASQA.Our analysis demonstrates an agreement between this metric and human judgments, and reveals a considerable gap between human performance and strong baselines. Ivan Stelmakh, Yi Luan, Bhuwan Dhingra, Ming-Wei Chang |
EMNLP | 4 |
| 2022 | FRUIT: Faithfully Reflecting Updated Information in TextabstractRobert Iv, Alexandre Passos, Sameer Singh, Ming-Wei Chang. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Robert L. Logan IV, Alexandre Tachard Passos, Sameer Singh 0001, Ming-Wei Chang |
NAACL-HLT | 4 |
| 2021 | Compositional Generalization and Natural Language Variation: Can a Semantic Parsing Approach Handle Both?abstractPeter Shaw, Ming-Wei Chang, Panupong Pasupat, Kristina Toutanova. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Peter Shaw 0004, Ming-Wei Chang, Panupong Pasupat, Kristina Toutanova |
ACL/IJCNLP (1) | 2 |
| 2021 | Joint Passage Ranking for Diverse Multi-Answer RetrievalabstractWe study multi-answer retrieval, an underexplored problem that requires retrieving passages to cover multiple distinct answers for a given question.This task requires joint modeling of retrieved passages, as models should not repeatedly retrieve passages containing the same answer at the cost of missing a different valid answer.In this paper, we introduce JPR, the first joint passage retrieval model for multi-answer retrieval.JPR makes use of an autoregressive reranker that selects a sequence of passages, each conditioned on previously selected passages.JPR is trained to select passages that cover new answers at each timestep and uses a tree-decoding algorithm to enable flexibility in the degree of diversity.Compared to prior approaches, JPR achieves significantly better answer coverage on three multianswer datasets.When combined with downstream question answering, the improved retrieval enables larger answer generation models since they need to consider fewer passages, establishing a new state-of-the-art. Sewon Min, Kenton Lee, Ming-Wei Chang, Kristina Toutanova, Hannaneh Hajishirzi |
EMNLP (1) | 3 |
| 2021 | Open Question Answering over Tables and Text
Wenhu Chen, Ming-Wei Chang, Eva Schlinger, William Yang Wang, William W. Cohen |
ICLR | 2 |
| 2020 | Probabilistic Assumptions Matter: Improved Models for Distantly-Supervised Document-Level Question AnsweringabstractWe address the problem of extractive question answering using document-level distant supervision, pairing questions and relevant documents with answer strings.We compare previously used probability space and distant supervision assumptions (assumptions on the correspondence between the weak answer string labels and possible answer mention spans).We show that these assumptions interact, and that different configurations provide complementary benefits.We demonstrate that a multiobjective model can efficiently combine the advantages of multiple assumptions and outperform the best individual formulation.Our approach outperforms previous state-of-the-art models by 4.3 points in F1 on TriviaQA-Wiki and 1.7 points in Rouge-L on NarrativeQA summaries.1 Hao Cheng 0002, Ming-Wei Chang, Kenton Lee, Kristina Toutanova |
ACL | 2 |
| 2020 | Exploring Unexplored Generalization Challenges for Cross-Database Semantic ParsingabstractWe study the task of cross-database semantic parsing (XSP), where a system that maps natural language utterances to executable SQL queries is evaluated on databases unseen during training. Recently, several datasets, including Spider, were proposed to support development of XSP systems. We propose a challenging evaluation setup for cross-database semantic parsing, focusing on variation across database schemas and in-domain language use. We re-purpose eight semantic parsing datasets that have been well-studied in the setting where in-domain training data is available, and instead use them as additional evaluation data for XSP systems instead. We build a system that performs well on Spider, and find that it struggles to generalize to our re-purposed set. Our setup uncovers several generalization challenges for cross-database semantic parsing, demonstrating the need to use and develop diverse training and evaluation datasets. Alane Suhr, Ming-Wei Chang, Peter Shaw 0004, Kenton Lee |
ACL | 2 |
| 2020 | CapWAP: Image Captioning with a PurposeabstractThe traditional image captioning task uses generic reference captions to provide textual information about images.Different user populations, however, will care about different visual aspects of images.In this paper, we propose a new task, Captioning with A Purpose (CAPWAP).Our goal is to develop systems that can be tailored to be useful for the information needs of an intended population, rather than merely provide generic information about an image.In this task, we use questionanswer (QA) pairs-a natural expression of information need-from users, instead of reference captions, for both training and postinference evaluation.We show that it is possible to use reinforcement learning to directly optimize for the intended information need, by rewarding outputs that allow a question answering model to provide correct answers to sampled user questions.We convert several visual question answering datasets into CAP-WAP datasets, and demonstrate that under a variety of scenarios our purposeful captioning system learns to anticipate and fulfill specific information needs better than its generic counterparts, as measured by QA performance on user questions from unseen images, when using the caption alone as context. Adam Fisch, Kenton Lee, Ming-Wei Chang, Jonathan H. Clark, Regina Barzilay |
EMNLP (1) | 3 |
| 2020 | Retrieval Augmented Language Model Pre-TrainingabstractLanguage model pre-training has been shown to capture a surprising amount of world knowledge, crucial for NLP tasks such as question answering. However, this knowledge is stored implicitly in the parameters of a neural network, requiring ever-larger networks to cover more facts. To capture knowledge in a more modular and interpretable way, we augment language model pre-training with a latent knowledge retriever, which allows the model to retrieve and attend over documents from a large corpus such as Wikipedia, used during pre-training, fine-tuning and inference. For the first time, we show how to pre-train such a knowledge retriever in an unsupervised manner, using masked language modeling as the learning signal and backpropagating through a retrieval step that considers millions of documents. We demonstrate the effectiveness of Retrieval-Augmented Language Model pre-training (REALM) by fine-tuning on the challenging task of Open-domain Question Answering (Open-QA). We compare against state-of-the-art models for both explicit and implicit knowledge storage on three popular Open-QA benchmarks, and find that we outperform all previous methods by a significant margin (4-16% absolute accuracy), while also providing qualitative benefits such as interpretability and modularity. Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, Ming-Wei Chang |
ICML | 5 |
| 2019 | Handling Divergent Reference Texts when Evaluating Table-to-Text GenerationabstractAutomatically constructed datasets for generating text from semi-structured data (tables), such as WikiBio (Lebret et al., 2016), often contain reference texts that diverge from the information in the corresponding semistructured data.We show that metrics which rely solely on the reference texts, such as BLEU and ROUGE, show poor correlation with human judgments when those references diverge.We propose a new metric, PAR-ENT, which aligns n-grams from the reference and generated texts to the semi-structured data before computing their precision and recall.Through a large scale human evaluation study of table-to-text models for WikiBio, we show that PARENT correlates with human judgments better than existing text generation metrics.We also adapt and evaluate the information extraction based evaluation proposed in Wiseman et al. (2017), and show that PAR-ENT has comparable correlation to it, while being easier to use.We show that PARENT is also applicable when the reference texts are elicited from humans using the data from the WebNLG challenge.1 * Work done during an internship at Google. Bhuwan Dhingra, Manaal Faruqui, Ankur P. Parikh, Ming-Wei Chang, Dipanjan Das 0001, William W. Cohen |
ACL (1) | 4 |
| 2019 | Latent Retrieval for Weakly Supervised Open Domain Question AnsweringabstractRecent work on open domain question answering (QA) assumes strong supervision of the supporting evidence and/or assumes a blackbox information retrieval (IR) system to retrieve evidence candidates.We argue that both are suboptimal, since gold evidence is not always available, and QA is fundamentally different from IR.We show for the first time that it is possible to jointly learn the retriever and reader from question-answer string pairs and without any IR system.In this setting, evidence retrieval from all of Wikipedia is treated as a latent variable.Since this is impractical to learn from scratch, we pre-train the retriever with an Inverse Cloze Task.We evaluate on open versions of five QA datasets.On datasets where the questioner already knows the answer, a traditional IR system such as BM25 is sufficient.On datasets where a user is genuinely seeking an answer, we show that learned retrieval is crucial, outperforming BM25 by up to 19 points in exact match. Kenton Lee, Ming-Wei Chang, Kristina Toutanova |
ACL (1) | 2 |
| 2019 | Zero-Shot Entity Linking by Reading Entity DescriptionsabstractWe present the zero-shot entity linking task, where mentions must be linked to unseen entities without in-domain labeled data.The goal is to enable robust transfer to highly specialized domains, and so no metadata or alias tables are assumed.In this setting, entities are only identified by text descriptions, and models must rely strictly on language understanding to resolve the new entities.First, we show that strong reading comprehension models pre-trained on large unlabeled data can be used to generalize to unseen entities.Second, we propose a simple and effective adaptive pre-training strategy, which we term domainadaptive pre-training (DAP), to address the domain shift problem associated with linking unseen entities in a new domain.We present experiments on a new dataset that we construct for this task and show that DAP improves over strong pre-training baselines, including BERT. Lajanugen Logeswaran, Ming-Wei Chang, Kenton Lee, Kristina Toutanova, Jacob Devlin, Honglak Lee |
ACL (1) | 2 |
| 2019 | Natural Questions: a Benchmark for Question Answering ResearchabstractWe present the Natural Questions corpus, a question answering data set. Questions consist of real anonymized, aggregated queries issued to the Google search engine. An annotator is presented with a question along with a Wikipedia page from the top 5 search results, and annotates a long answer (typically a paragraph) and a short answer (one or more entities) if present on the page, or marks null if no long/short answer is present. The public release consists of 307,373 training examples with single annotations; 7,830 examples with 5-way annotations for development data; and a further 7,842 examples with 5-way annotated sequestered as test data. We present experiments validating quality of the data. We also describe analysis of 25-way annotations on 302 examples, giving insights into human variability on the annotation task. We introduce robust metrics for the purposes of evaluating question answering systems; demonstrate high human upper bounds on these metrics; and establish baseline results using competitive methods drawn from related literature. Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins 0001, Ankur P. Parikh, Christopher Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc V. Le, Slav Petrov |
Trans. Assoc. Comput. Linguistics | 14 |
| 2018 | A Knowledge-Grounded Neural Conversation ModelabstractNeural network models are capable of generating extremely natural sounding conversational interactions. However, these models have been mostly applied to casual scenarios (e.g., as “chatbots”) and have yet to demonstrate they can serve in more useful conversational applications. This paper presents a novel, fully data-driven, and knowledge-grounded neural conversation model aimed at producing more contentful responses. We generalize the widely-used Sequence-to-Sequence (Seq2Seq) approach by conditioning responses on both conversation history and external “facts”, allowing the model to be versatile and applicable in an open-domain setting. Our approach yields significant improvements over a competitive Seq2Seq baseline. Human judges found that our outputs are significantly more informative. Marjan Ghazvininejad, Chris Brockett, Ming-Wei Chang, William B. Dolan, Jianfeng Gao 0001, Scott Yih, Michel Galley |
AAAI | 3 |
| 2018 | Policy Shaping and Generalized Update Equations for Semantic Parsing from DenotationsabstractSemantic parsing from denotations faces two key challenges in model training: (1) given only the denotations (e.g., answers), search for good candidate semantic parses, and (2) choose the best model update algorithm.We propose effective and general solutions to each of them.Using policy shaping, we bias the search procedure towards semantic parses that are more compatible to the text, which provide better supervision signals for training.In addition, we propose an update equation that generalizes three different families of learning algorithms, which enables fast model exploration.When experimented on a recently proposed sequential question answering dataset, our framework leads to a new state-of-theart model that outperforms previous work by 5.0% absolute on exact match accuracy.Question: what nation scored the most points Dipendra Misra, Ming-Wei Chang, Xiaodong He 0001, Scott Yih |
EMNLP | 2 |
| 2017 | Search-based Neural Structured Learning for Sequential Question AnsweringabstractRecent work in semantic parsing for question answering has focused on long and complicated questions, many of which would seem unnatural if asked in a normal conversation between two humans.In an effort to explore a conversational QA setting, we present a more realistic task: answering sequences of simple but inter-related questions.We collect a dataset of 6,066 question sequences that inquire about semistructured tables from Wikipedia, with 17,553 question-answer pairs in total.To solve this sequential question answering task, we propose a novel dynamic neural semantic parsing framework trained using a weakly supervised reward-guided search.Our model effectively leverages the sequential context to outperform state-of-the-art QA systems that are designed to answer highly complex questions. Mohit Iyyer, Scott Yih, Ming-Wei Chang |
ACL (1) | 3 |
| 2017 | Annotating Derivations: A New Evaluation Strategy and Dataset for Algebra Word ProblemsabstractWe propose a new evaluation for automatic solvers for algebra word problems, which can identify mistakes that existing evaluations overlook.Our proposal is to evaluate such solvers using derivations, which reflect how an equation system was constructed from the word problem.To accomplish this, we develop an algorithm for checking the equivalence between two derivations, and show how derivation annotations can be semi-automatically added to existing datasets.To make our experiments more comprehensive, we include the derivation annotation for DRAW-1K, a new dataset containing 1000 general algebra word problems.In our experiments, we found that the annotated derivations enable a more accurate evaluation of automatic solvers than previously used metrics.We release derivation annotations for over 2300 algebra word problems for future evaluations. Shyam Upadhyay, Ming-Wei Chang |
EACL (1) | 2 |
| 2017 | Maximum Margin Reward Networks for Learning from Explicit and Implicit SupervisionabstractNeural networks have achieved state-ofthe-art performance on several structuredoutput prediction tasks, trained in a fully supervised fashion.However, annotated examples in structured domains are often costly to obtain, which thus limits the applications of neural networks.In this work, we propose Maximum Margin Reward Networks, a neural networkbased framework that aims to learn from both explicit (full structures) and implicit supervision signals (delayed feedback on the correctness of the predicted structure).On named entity recognition and semantic parsing, our model outperforms previous systems on the benchmark datasets, CoNLL-2003 and WebQuestionsSP. Haoruo Peng, Ming-Wei Chang, Scott Yih |
EMNLP | 2 |
| 2016 | Learning from Explicit and Implicit Supervision Jointly For Algebra Word ProblemsabstractAutomatically solving algebra word problems has raised considerable interest recently.Existing state-of-the-art approaches mainly rely on learning from human annotated equations.In this paper, we demonstrate that it is possible to efficiently mine algebra problems and their numerical solutions with little to no manual effort.To leverage the mined dataset, we propose a novel structured-output learning algorithm that aims to learn from both explicit (e.g., equations) and implicit (e.g., solutions) supervision signals jointly.Enabled by this new algorithm, our model gains 4.6% absolute improvement in accuracy on the ALG-514 benchmark compared to the one without using implicit supervision.The final model also outperforms the current state-of-the-art approach by 3%. Shyam Upadhyay, Ming-Wei Chang, Kai-Wei Chang 0001, Scott Yih |
EMNLP | 2 |
| 2016 | Toward Socially-Infused Information Extraction: Embedding Authors, Mentions, and EntitiesabstractEntity linking is the task of identifying mentions of entities in text, and linking them to entries in a knowledge base. This task is especially difficult in microblogs, as there is little additional text to provide disambiguating context; rather, authors rely on an implicit common ground of shared knowledge with their readers. In this paper, we attempt to capture some of this implicit context by exploiting the social network structure in microblogs. We build on the theory of homophily, which implies that socially linked individuals share interests, and are therefore likely to mention the same sorts of entities. We implement this idea by encoding authors, mentions, and entities in a continuous vector space, which is constructed so that socially-connected authors have similar vector representations. These vectors are incorporated into a neural structured prediction model, which captures structural constraints that are inherent in the entity linking task. Together, these design decisions yield F1 improvements of 1%-5% on benchmark datasets, as compared to the previous state-of-the-art. Yi Yang 0038, Ming-Wei Chang, Jacob Eisenstein |
EMNLP | 2 |
| 2015 | S-MART: Novel Tree-based Structured Learning Algorithms Applied to Tweet Entity LinkingabstractYi Yang, Ming-Wei Chang. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Yi Yang 0038, Ming-Wei Chang |
ACL (1) | 2 |
| 2015 | Semantic Parsing via Staged Query Graph Generation: Question Answering with Knowledge BaseabstractWen-tau Yih, Ming-Wei Chang, Xiaodong He, Jianfeng Gao. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Scott Yih, Ming-Wei Chang, Xiaodong He 0001, Jianfeng Gao 0001 |
ACL (1) | 2 |
| 2015 | Inferring Missing Entity Type Instances for Knowledge Base Completion: New Dataset and MethodsabstractMost of previous work in knowledge base (KB) completion has focused on the problem of relation extraction. In this work, we focus on the task of inferring missing entity type in-stances in a KB, a fundamental task for KB competition yet receives little attention. Due to the novelty of this task, we construct a large-scale dataset and design an automatic evaluation methodology. Our knowledge base completion method uses information within the existing KB and external information from Wikipedia. We show that individual methods trained with a global objective that consid-ers unobserved cells from both the entity and the type side gives consistently higher qual-ity predictions compared to baseline methods. We also perform manual evaluation on a small subset of the data to verify the effectiveness of our knowledge base completion methods and the correctness of our proposed automatic evaluation method. 1 Arvind Neelakantan, Ming-Wei Chang |
HLT-NAACL | 2 |
| 2015 | Placer++: Semantic place labels beyond the visitabstractPlace labeling is the process of giving semantic labels to locations, such as home, work, and school. For a particular person, these labels can be computed automatically based on features of that person's visits to these locations. A previous system called Placer used the person's demographic data and the timing of their visits to label places with a learned decision tree. We developed Placer++ as a more accurate labeler, augmenting Placer's features of individual visits with (1) labeled visits from other people and (2) features about the sequence of the individual's visits. In processing sequences, we adopt structural learning techniques to take into account the relationships between visits. Accuracy increased by 8.85 percentage points over the baseline of Placer. We describe and justify the features and present our experiments on government diary data. John Krumm, Dany Rouhana, Ming-Wei Chang |
PerCom | 3 |
| 2015 | Open Domain Question Answering via Semantic EnrichmentabstractMost recent question answering (QA) systems query large-scale knowledge bases (KBs) to answer a question, after parsing and transforming natural language questions to KBs-executable forms (e.g., logical forms). As a well-known fact, KBs are far from complete, so that information required to answer questions may not always exist in KBs. In this paper, we develop a new QA system that mines answers directly from the Web, and meanwhile employs KBs as a significant auxiliary to further boost the QA performance. Specifically, to the best of our knowledge, we make the first attempt to link answer candidates to entities in Freebase, during answer candidate generation. Several remarkable advantages follow: (1) Redundancy among answer candidates is automatically reduced. (2) The types of an answer candidate can be effortlessly determined by those of its corresponding entity in Freebase. (3) Capitalizing on the rich information about entities in Freebase, we can develop semantic features for each answer candidate after linking them to Freebase. Particularly, we construct answer-type related features with two novel probabilistic models, which directly evaluate the appropriateness of an answer candidate's types under a given question. Overall, such semantic features turn out to play significant roles in determining the true answers from the large answer candidate pool. The experimental results show that across two testing datasets, our QA system achieves an 18%~54% improvement under F_1 metric, compared with various existing QA systems. Huan Sun 0001, Hao Ma 0001, Scott Yih, Chen-Tse Tsai, Jingjing Liu 0001, Ming-Wei Chang |
WWW | 6 |
| 2014 | Virtual keyboard for head mounted display-based wearable devicesabstractWearable devices eliminate the need of physically taking out a mobile device before operating on it and are emerging as the next wave of mobile systems. Head-mounted display (HMD) is a key building block of wearable devices, and offers users immediate access to relevant information in a glance. However, most existing user input mechanisms accompanying HMDs are designed for interactive information exploration rather than for extended text entry. This paper describes the design, implementation and evaluation of a text input system for HMDs called Air Typing, which requires only a standard camera and is shown to be comparable in effectiveness to single-hand text input on tablet computers in a lab setting. Air Typing features a novel two-level virtual keyword layout, which substantially improves the typing speed by cutting down unnecessary hand movements during typing and greatly simplifies the associated image processing task by doing away with fine-grained matching between fingertips and keys. The current Air Typing prototype incorporates an OpenCV-based virtual key press detection algorithm that runs on the featured two-level virtual keyboard. In our tests, an experienced user's typing speeds of one-hand text input and of two-hand text input under Air Typing are 13 and 15 words per minute (WPM), respectively. Ming-Wei Chang, Tzi-cker Chiueh |
ICPADS | 1 |
| 2014 | ERD'14: entity recognition and disambiguation challengeabstractNo abstract available. David Carmel, Ming-Wei Chang, Evgeniy Gabrilovich, Bo-June Paul Hsu, Kuansan Wang |
SIGIR | 2 |
| 2014 | Modeling action-level satisfaction for search task satisfaction predictionabstractSearch satisfaction is a property of a user's search process. Understanding it is critical for search providers to evaluate the performance and improve the effectiveness of search engines. Existing methods model search satisfaction holistically at the search-task level, ignoring important dependencies between action-level satisfaction and overall task satisfaction. We hypothesize that searchers' latent action-level satisfaction (i.e., whether they believe they were satisfied with the results of a query or click) influences their observed search behaviors and contributes to overall search satisfaction. We conjecture that by modeling search satisfaction at the action level, we can build more complete and more accurate predictors of search-task satisfaction. To do this, we develop a latent structural learning method, whereby rich structured features and dependency relations unique to search satisfaction prediction are explored. Using in-situ search satisfaction judgments provided by searchers, we show that there is significant value in modeling action-level satisfaction in search-task satisfaction prediction. In addition, experimental results on large-scale logs from Bing.com demonstrate clear benefit from using inferred action satisfaction labels for other applications such as document relevance estimation and query suggestion. Hongning Wang, Yang Song 0008, Ming-Wei Chang, Xiaodong He 0001, Ahmed Awadallah 0001, Ryen W. White |
SIGIR | 3 |
| 2014 | Entity Linking on Microblogs with Spatial and Temporal SignalsabstractMicroblogs present an excellent opportunity for monitoring and analyzing world happenings. Given that words are often ambiguous, entity linking becomes a crucial step towards understanding microblogs. In this paper, we re-examine the problem of entity linking on microblogs. We first observe that spatiotemporal ( i.e., spatial and temporal) signals play a key role, but they are not utilized in existing approaches. Thus, we propose a novel entity linking framework that incorporates spatiotemporal signals through a weakly supervised process. Using entity annotations on real-world data, our experiments show that the spatiotemporal model improves F1 by more than 10 points over existing systems. Finally, we present a qualitative study to visualize the effectiveness of our approach. Ming-Wei Chang |
Trans. Assoc. Comput. Linguistics | 2 |
| 2013 | Question Answering Using Enhanced Lexical Semantic Models
Scott Yih, Ming-Wei Chang, Christopher Meek, Andrzej Pastusiak |
ACL (1) | 2 |
| 2013 | To Link or Not to Link? A Study on End-to-End Tweet Entity Linking
Stephen D. Guo, Ming-Wei Chang, Emre Kiciman |
HLT-NAACL | 2 |
| 2013 | Personalized ranking model adaptation for web searchabstractSearch engines train and apply a single ranking model across all users, but searchers' information needs are diverse and cover a broad range of topics. Hence, a single user-independent ranking model is insufficient to satisfy different users' result preferences. Conventional personalization methods learn separate models of user interests and use those to re-rank the results from the generic model. Those methods require significant user history information to learn user preferences, have low coverage in the case of memory-based methods that learn direct associations between query-URL pairs, and have limited opportunity to markedly affect the ranking given that they only re-order top-ranked items. Hongning Wang, Xiaodong He 0001, Ming-Wei Chang, Yang Song 0008, Ryen W. White |
SIGIR | 3 |
| 2013 | Learning to extract cross-session search tasksabstractSearch tasks, comprising a series of search queries serving the same information need, have recently been recognized as an accurate atomic unit for modeling user search intent. Most prior research in this area has focused on short-term search tasks within a single search session, and heavily depend on human annotations for supervised classification model learning. In this work, we target the identification of long-term, or cross-session, search tasks (transcending session boundaries) by investigating inter-query dependencies learned from users' searching behaviors. A semi-supervised clustering model is proposed based on the latent structural SVM framework, and a set of effective automatic annotation rules are proposed as weak supervision to release the burden of manual annotation. Experimental results based on a large-scale search log collected from Bing.com confirms the effectiveness of the proposed model in identifying cross-session search tasks and the utility of the introduced weak supervision signals. Our learned model enables a more comprehensive understanding of users' search behaviors via search logs and facilitates the development of dedicated search-engine support for long-term tasks. Hongning Wang, Yang Song 0008, Ming-Wei Chang, Xiaodong He 0001, Ryen W. White |
WWW | 3 |
| 2013 | Dual Coordinate Descent Algorithms for Efficient Large Margin Structured PredictionabstractDue to the nature of complex NLP problems, structured prediction algorithms have been important modeling tools for a wide range of tasks. While there exists evidence showing that linear Structural Support Vector Machine (SSVM) algorithm performs better than structured Perceptron, the SSVM algorithm is still less frequently chosen in the NLP community because of its relatively slow training speed. In this paper, we propose a fast and easy-to-implement dual coordinate descent algorithm for SSVMs. Unlike algorithms such as Perceptron and stochastic gradient descent, our method keeps track of dual variables and updates the weight vector more aggressively. As a result, this training process is as efficient as existing online learning methods, and yet derives consistently better models, as evaluated on four benchmark NLP datasets for part-of-speech tagging, named-entity recognition and dependency parsing. Ming-Wei Chang, Scott Yih |
Trans. Assoc. Comput. Linguistics | 1 |
| 2012 | Learning shared body plansabstractWe cast the problem of recognizing related categories as a unified learning and structured prediction problem with shared body plans. When provided with detailed annotations of objects and their parts, these body plans model objects in terms of shared parts and layouts, simultaneously capturing a variety of categories in varied poses. We can use these body plans to jointly train many detectors in a shared framework with structured learning, leading to significant gains for each supervised task. Using our model, we can provide detailed predictions of objects and their parts for both familiar and unfamiliar categories. Ian Endres, Vivek Srikumar, Ming-Wei Chang, Derek Hoiem |
CVPR | 3 |
| 2012 | Unified Expectation Maximization
Rajhans Samdani, Ming-Wei Chang, Dan Roth 0001 |
HLT-NAACL | 2 |
| 2012 | Structured learning with constrained conditional models
Ming-Wei Chang, Lev-Arie Ratinov, Dan Roth 0001 |
Mach. Learn. | 1 |
| 2010 | Driving Semantic Parsing from the World's Response
James Clarke, Dan Goldwasser, Ming-Wei Chang, Dan Roth 0001 |
CoNLL | 3 |
| 2010 | The Necessity of Combining Adaptation Methods
Ming-Wei Chang, Michael Connor, Dan Roth 0001 |
EMNLP | 1 |
| 2010 | Structured Output Learning with Indirect Supervision
Ming-Wei Chang, Vivek Srikumar, Dan Goldwasser, Dan Roth 0001 |
ICML | 1 |
| 2010 | Discriminative Learning over Constrained Latent Representations
Ming-Wei Chang, Dan Goldwasser, Dan Roth 0001, Vivek Srikumar |
HLT-NAACL | 1 |
| 2010 | Modified Frequency-Partitioned Spectrum Estimation for a Wireless Health Advanced Monitoring Bio-Diagnosis SystemabstractThis paper proposes a technique for frequency-partitioned spectrum estimation (FPSE), which is used in the National Taiwan University Wireless Health Advanced Monitoring Bio-Diagnosis System for electrocardiogram analysis. A process for analyzing the RR interval (which is a time series formed by the heat-beat duration that represents heart-rate variations) in conjunction with the fuzzy clustering technique is proposed for arrhythmia recognition. FPSE helps reduce data transmission errors and allows the computational load to be moved to a remote server; however, it suffers from waveform deterioration during reconstruction of the signal power spectrum. To compensate for this problem, this paper proposes a modified FPSE approach that imposes an additional boundary constraint to ensure that the estimated spectrum is smooth. The simulation results show that the proposed algorithm is more effective at recovering the original frequency information and achieves a globally asymptotic trend. The proposed arrhythmia recognition procedure was applied to the Massachusetts Institute of Technology-Boston's Beth Israel Hospital (MIT-BIH) database (developed by MIT and Boston's Beth Israel Deaconess Medical Center), which demonstrated that it is both very convenient and efficient. Ching-En Tseng, Jia-Yush Yen, Ming-Wei Chang, Wei-Chien Chang, Chih-Kung Lee |
IEEE Trans. Syst. Man Cybern. Part A | 3 |
| 2009 | Unsupervised Constraint Driven Learning For Transliteration Discovery
Ming-Wei Chang, Dan Goldwasser, Dan Roth 0001, Yuancheng Tu |
HLT-NAACL | 1 |
| 2008 | Learning and Inference with Constraints
Ming-Wei Chang, Lev-Arie Ratinov, Nick Rizzolo, Dan Roth 0001 |
AAAI | 1 |
| 2008 | Importance of Semantic Representation: Dataless Classification
Ming-Wei Chang, Lev-Arie Ratinov, Dan Roth 0001, Vivek Srikumar |
AAAI | 1 |
| 2008 | Partitioned logistic regression for spam filteringabstractNaive Bayes and logistic regression perform well in different regimes. While the former is a very simple generative model which is efficient to train and performs well empirically in many applications,the latter is a discriminative model which often achieves better accuracy and can be shown to outperform naive Bayes asymptotically. In this paper, we propose a novel hybrid model, partitioned logistic regression, which has several advantages over both naive Bayes and logistic regression. This model separates the original feature space into several disjoint feature groups. Individual models on these groups of features are learned using logistic regression and their predictions are combined using the naive Bayes principle to produce a robust final estimation. We show that our model is better both theoretically and empirically. In addition, when applying it in a practical application, email spam filtering, it improves the normalized AUC score at 10% false-positive rate by 28.8% and 23.6% compared to naive Bayes and logistic regression, when using the exact same training examples. Ming-Wei Chang, Scott Yih, Christopher Meek |
KDD | 1 |
| 2007 | Guiding Semi-Supervision with Constraint-Driven Learning
Ming-Wei Chang, Lev-Arie Ratinov, Dan Roth 0001 |
ACL | 1 |
| 2006 | A Pipeline Framework for Dependency Parsing
Ming-Wei Chang, Quang Do, Dan Roth 0001 |
ACL | 1 |
| 2006 | A Pipeline Model for Bottom-Up Dependency Parsing
Ming-Wei Chang, Quang Do, Dan Roth 0001 |
CoNLL | 1 |
| 2005 | Leave-One-Out Bounds for Support Vector Regression Model SelectionabstractMinimizing bounds of leave-one-out errors is an important and efficient approach for support vector machine (SVM) model selection. Past research focuses on their use for classification but not regression. In this letter, we derive various leave-one-out bounds for support vector regression (SVR) and discuss the difference from those for classification. Experiments demonstrate that the proposed bounds are competitive with Bayesian SVR for parameter selection. We also discuss the differentiability of leave-one-out bounds. Ming-Wei Chang, Chih-Jen Lin |
Neural Comput. | 1 |
| 2004 | Analysis of switching dynamics with competing support vector machinesabstractWe present a framework for the unsupervised segmentation of switching dynamics using support vector machines. Following the architecture by Pawelzik et al., where annealed competing neural networks were used to segment a nonstationary time series, in this paper, we exploit the use of support vector machines, a well-known learning technique. First, a new formulation of support vector regression is proposed. Second, an expectation-maximization step is suggested to adaptively adjust the annealing parameter. Results indicate that the proposed approach is promising. Ming-Wei Chang, Chih-Jen Lin, Ruby Chiu-Hsing Weng |
IEEE Trans. Neural Networks | 1 |