Yutaka Matsuo

dblp:m/YMatsuo · DBLP profile ↗
← Back
144ranked-venue papers
18as first author
56since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 110 · 12 first-author · 53 since 2021Databases, data management, data science and information retrieval · 39 · 7 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 4 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 5 first-authorHuman-computer interaction and ubiquitous computing · 12 · 2 first-authorSystems, architecture and hardware · 4 · 3 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Understanding Emergent Misalignment via Feature Superposition Geometry
abstract
Emergent misalignment, where fine-tuning on narrow, non-harmful tasks induces harmful behaviors, poses a key challenge for AI safety in LLMs.Despite growing empirical evidence, its underlying mechanism remains unclear.To uncover the reason behind this phenomenon, we propose a geometric account based on the geometry of feature superposition.Because features are encoded in overlapping representations, fine-tuning that amplifies a target feature also unintentionally strengthens nearby harmful features in accordance with their similarity.We give a simple gradient-level derivation of this effect and empirically test it in multiple LLMs (Gemma-2 2B/9B/27B, LLaMA-3.1 8B, gpt-oss 20B).Using sparse autoencoders (SAEs), we identify features tied to misalignment-inducing data and to harmful behaviors, and show that they are geometrically closer to each other than features derived from non-inducing data.This trend generalizes across domains (e.g., health, career, legal advice).Finally, we show that a geometry-aware approach-filtering training samples closest to toxic features-reduces misalignment by 34.5%, substantially outperforming random removal and achieving comparable or slightly lower misalignment than LLM-as-ajudge-based filtering.Our study links emergent misalignment to feature superposition, providing a basis for understanding and mitigating this phenomenon.This paper analyzes datasets and model outputs that include offensive language.
Gouki Minegishi, Hiroki Furuta, Takeshi Kojima, Yusuke Iwasawa, Yutaka Matsuo
ACL (1)5
2026 Automated Refinement of Essay Scoring Rubrics for Language Models via Reflect-and-Revise
abstract
Large Language Models (LLMs) are increasingly used for Automated Essay Scoring (AES), yet the scoring rubrics they rely on are typically designed for human raters and may not be optimal for LLMs.Inspired by the calibration process that human raters undergo before formal scoring, we propose Reflect-and-Revise, an iterative framework that refines scoring rubrics by prompting models to reflect on their own chain-of-thought rationales and score discrepancies with human labels.At each iteration, the model identifies scoring-error patterns from sampled mismatches and revises the rubric accordingly.Experiments on three essay scoring benchmarks (ASAP, ASAP 2.0, and TOEFL11) with three LLMs (GPT-5 mini, Gemini 3 Flash, and Qwen3-Next-80B-A3B-Instruct) demonstrate that our method yields improvements in Quadratic Weighted Kappa (QWK), achieving gains of up to +0.403 over human-authored rubrics.Starting from a minimal seed rubric that specifies only the score scale, our method matches or exceeds expert rubric performance in most dataset-model combinations, indicating that iterative refinement can reduce the manual effort of rubric authoring.Analysis of the refined rubrics reveals that the refinement process introduces explicit procedural structures, such as conditional gating rules and quantitative thresholds, that are absent from humanauthored rubrics, highlighting a gap between rubrics designed for human raters and those effective for LLMs. 1
Keno Harada, Lui Yoshida, Takeshi Kojima, Yusuke Iwasawa, Yutaka Matsuo
CoNLL5
2026 One-shot Portrait Stylization via Geometric Alignment
abstract
Portrait stylization casts vivid artistic style drawn from style examples to portrait photos. Although recently extensively studied with machine learning algorithms, existing methods still face challenges in stylizing portraits from a single style reference, severely limiting their potential for real-world applications. In this paper, we propose a portrait stylization method that learns style reference from a single artistic portrait image. Unlike previous StyleGAN based methods that heavily rely on the quality of GAN inversion or diffusion based methods that introduce computational expensive operations and fall short of precise control, our method achieves high-quality stylization with small computation and parameter budget. Specifically, we employ geometric alignment to build spatial correlation between content images and style reference. A geometry LoRA and a style LoRA are then jointly optimized based on a pre-trained diffusion backbone respectively, with orthogonal adaptation used to disentangle the geometry and style information. During inference, the style LoRA is integrated into the diffusion backbone and ControlNet is further combined to facilitate better spatial and identity control. We illustrate abundant stylized portraits with multiple styles. Qualitative comparison, quantitative validation and user study prove that our method outperforms existing methods, and ablation study demonstrates the effectiveness of each components.
Zilin Guo, Zhuoru Li, Yusuke Iwasawa, Yutaka Matsuo, Jiaxian Guo
WACV7
2026 A two-stage learning architecture for reliable financial time series forecasting
Albi Isufaj, Pablo Mollá, Jing Li 0111, Yutaka Matsuo, Helmut Prendinger
Neurocomputing4
2025 AGENTiGraph: A Multi-Agent Knowledge Graph Framework for Interactive, Domain-Specific LLM Chatbots
abstract
AGENTiGraph is a user-friendly, agent-driven system that enables intuitive interaction and management of domain-specific data through the manipulation of knowledge graphs in natural language. It gives non-technical users a complete, visual solution to incrementally build and refine their knowledge bases, allowing multi-round dialogues and dynamic updates without specialized query languages. The flexible design of AGENTiGraph, including intent classification, task planning, and automatic knowledge integration, ensures seamless reasoning between diverse tasks. Evaluated on a 3,500-query benchmark within an educational scenario, the system outperforms strong zero-shot baselines (achieving 95.12% classification accuracy, 90.45% execution success), indicating potential scalability to compliance-critical or multi-step queries in legal and medical domains, e.g., incorporating new statutes or research on the fly. Our open-source demo offers a powerful new paradigm for multi-turn enterprise knowledge management that bridges LLMs and structured graphs.
Xinjie Zhao 0004, Moritz Blum, Yingjian Chen, Boming Yang, Luis Marquez-Carpintero, Monica Pina-Navarro, Yanran Fu, So Morikawa, Yusuke Iwasawa, Yutaka Matsuo, Chanjun Park, Irene Li
CIKM11
2025 Slender-Mamba: Fully Quantized Mamba in 1.58 Bits From Head to Toe
abstract
Large language models (LLMs) have achieved significant performance improvements in natural language processing (NLP) domain. However, these models often require large computational resources for training and inference. Recently, Mamba, a language model architecture based on State-Space Models (SSMs), has achieved comparable performance to Transformer models while significantly reducing costs by compressing context windows during inference. We focused on the potential of the lightweight Mamba architecture by applying BitNet quantization method to the model architecture. In addition, while prior BitNet methods generally quantized only linear layers in the main body, we extensively quantized the embedding and projection layers considering their significant proportion of model parameters. In our experiments, we applied ternary quantization to the Mamba-2 (170M) architecture and pre-trained the model with 150 B tokens from scratch. Our method achieves approximately 90.0% reduction in the bits used by all parameters, achieving a significant improvement compared with a 48.4% reduction by the conventional BitNet quantization method. In addition, our method experienced minimal performance degradation in both the pre-training perplexity and downstream tasks. These findings demonstrate the potential of incorporating lightweight language models into edge devices, which will become more demanding in the future.
Zhenxuan Yu, Takeshi Kojima, Yutaka Matsuo, Yusuke Iwasawa
COLING3
2025 Image Referenced Sketch Colorization Based on Animation Creation Workflow
abstract
Sketch colorization plays an important role in animation and digital illustration production tasks. However, existing methods still meet problems in that text-guided methods fail to provide accurate color and style reference, hint-guided methods still involve manual operation, and image-referenced methods are prone to cause artifacts. To address these limitations, we propose a diffusion-based framework inspired by real-world animation production work-flows. Our approach leverages the sketch as the spatial guidance and an RGB image as the color reference, and separately extracts foreground and background from the reference image with spatial masks. Particularly, we introduce a split cross-attention mechanism with LoRA (Low-Rank Adaptation) modules. They are trained separately with foreground and background regions to control the corresponding embeddings for keys and values in cross-attention. This design allows the diffusion model to integrate information from foreground and background independently, preventing interference and eliminating the spatial artifacts. During inference, we design switchable inference modes for diverse use scenarios by changing modules activated in the framework. Extensive qualitative and quantitative experiments, along with user studies, demonstrate our advantages over existing methods in generating high-qualigy artifact-free results with geometric mismatched references. Ablation studies further confirm the effectiveness of each component. Codes are available at https://github.com/tellurion-kanata/colorizeDiffusion.
Dingkun Yan, Zhuoru Li, Suguru Saito, Yusuke Iwasawa, Yutaka Matsuo, Jiaxian Guo
CVPR6
2025 MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation
abstract
Weihao Xuan, Rui Yang, Heli Qi, Qingcheng Zeng, Yunze Xiao, Aosong Feng, Dairui Liu, Yun Xing, Junjue Wang, Fan Gao, Jinghui Lu, Yuang Jiang, Huitao Li, Xin Li, Kunyu Yu, Ruihai Dong, Shangding Gu, Yuekang Li, Xiaofei Xie, Felix Juefei-Xu, Foutse Khomh, Osamu Yoshie, Qingyu Chen, Douglas Teodoro, Nan Liu, Randy Goebel, Lei Ma, Edison Marrese-Taylor, Shijian Lu, Yusuke Iwasawa, Yutaka Matsuo, Irene Li. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Weihao Xuan, Rui Yang 0016, Heli Qi, Qingcheng Zeng, Yunze Xiao, Aosong Feng, Dairui Liu, Yun Xing 0001, Jinghui Lu, Yuang Jiang, Huitao Li, Xin Li 0079, Kunyu Yu, Ruihai Dong, Shangding Gu, Yuekang Li, Xiaofei Xie, Felix Juefei-Xu, Foutse Khomh, Osamu Yoshie, Qingyu Chen 0001, Douglas Teodoro, Nan Liu 0003, Randy Goebel, Lei Ma 0003, Edison Marrese-Taylor, Shijian Lu, Yusuke Iwasawa, Yutaka Matsuo, Irene Li
EMNLP31
2025 ReAgent: Reversible Multi-Agent Reasoning for Knowledge-Enhanced Multi-Hop QA
abstract
Zhao Xinjie, Fan Gao, Xingyu Song, Yingjian Chen, Rui Yang, Yanran Fu, Yuyang Wang, Yusuke Iwasawa, Yutaka Matsuo, Irene Li. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Xinjie Zhao 0004, Xingyu Song, Yingjian Chen, Rui Yang 0016, Yanran Fu, Yusuke Iwasawa, Yutaka Matsuo, Irene Li
EMNLP9
2025 CityNav: A Large-Scale Dataset for Real-World Aerial Navigation
Jungdae Lee, Taiki Miyanishi, Shuhei Kurita, Koya Sakamoto, Daichi Azuma, Yutaka Matsuo, Nakamasa Inoue
ICCV6
2025 GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields
abstract
The advancement of 3D language fields has enabled intuitive interactions with 3D scenes via natural language. However, existing approaches are typically limited to small-scale environments, lacking the scalability and compositional reasoning capabilities necessary for large, complex urban settings. To overcome these limitations, we propose GeoProg3D, a visual programming framework that enables natural language-driven interactions with city-scale high-fidelity 3D scenes. GeoProg3D consists of two key components: (i) a Geography-aware City-scale 3D Language Field (GCLF) that leverages a memory-efficient hierarchical 3D model to handle large-scale data, integrated with geographic information for efficiently filtering vast urban spaces using directional cues, distance measurements, elevation data, and landmark references; and (ii) Geographical Vision APIs (GV-APIs), specialized geographic vision tools such as area segmentation and object detection. Our framework employs large language models (LLMs) as reasoning engines to dynamically combine GV-APIs and operate GCLF, effectively supporting diverse geographic vision tasks. To assess performance in city-scale reasoning, we introduce GeoEval3D, a comprehensive benchmark dataset containing 952 query-answer pairs across five challenging tasks: grounding, spatial reasoning, comparison, counting, and measurement. Experiments demonstrate that GeoProg3D significantly outperforms existing 3D language fields and vision-language models across multiple tasks. To our knowledge, GeoProg3D is the first framework enabling compositional geographic reasoning in high-fidelity city-scale 3D environments via natural language. The code is available at https://snskysk.github.io/GeoProg3D/.
Shunsuke Yasuki, Taiki Miyanishi, Nakamasa Inoue, Shuhei Kurita, Koya Sakamoto, Daichi Azuma, Masato Taki, Yutaka Matsuo
ICCV8
2025 Near-Optimal Policy Identification in Robust Constrained Markov Decision Processes via Epigraph Form
abstract
Designing a safe policy for uncertain environments is crucial in real-world control systems. However, this challenge remains inadequately addressed within the Markov decision process (MDP) framework. This paper presents the first algorithm guaranteed to identify a near-optimal policy in a robust constrained MDP (RCMDP), where an optimal policy minimizes cumulative cost while satisfying constraints in the worst-case scenario across a set of environments. We first prove that the conventional policy gradient approach to the Lagrangian max-min formulation can become trapped in suboptimal solutions. This occurs when its inner minimization encounters a sum of conflicting gradients from the objective and constraint functions. To address this, we leverage the epigraph form of the RCMDP problem, which resolves the conflict by selecting a single gradient from either the objective or the constraints. Building on the epigraph form, we propose a bisection search algorithm with a policy gradient subroutine and prove that it identifies an $\varepsilon$-optimal policy in an RCMDP with $\widetilde{\mathcal{O}}(\varepsilon^{-4})$ robust policy evaluations.
Toshinori Kitamura, Tadashi Kozuno, Wataru Kumagai, Kenta Hoshino, Yohei Hosoe, Kazumi Kasaura, Masashi Hamaya, Paavo Parmas, Yutaka Matsuo
ICLR9
2025 Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words
abstract
Sparse autoencoders (SAEs) have gained a lot of attention as a promising tool to improve the interpretability of large language models (LLMs) by mapping the complex superposition of *polysemantic* neurons into *monosemantic* features and composing a sparse dictionary of words. However, traditional performance metrics like Mean Squared Error and $\mathrm{L}_{0}$ sparsity ignore the evaluation of the semantic representational power of SAEs - whether they can acquire interpretable monosemantic features while preserving the semantic relationship of words.For instance, it is not obvious whether a learned sparse feature could distinguish different meanings in one word. In this paper, we propose a suite of evaluations for SAEs to analyze the quality of monosemantic features by focusing on polysemous words. Our findings reveal that SAEs developed to improve the MSE-$\mathrm{L}_0$ Pareto frontier may confuse interpretability, which does not necessarily enhance the extraction of monosemantic features. The analysis of SAEs with polysemous words can also figure out the internal mechanism of LLMs; deeper layers and the Attention module contribute to distinguishing polysemy in a word. Our semantics-focused evaluation offers new insights into the polysemy and the existing SAE objective and contributes to the development of more practical SAEs.
Gouki Minegishi, Hiroki Furuta, Yusuke Iwasawa, Yutaka Matsuo
ICLR4
2025 Beyond Induction Heads: In-Context Meta Learning Induces Multi-Phase Circuit Emergence
abstract
Transformer-based language models exhibit In-Context Learning (ICL), where predictions are made adaptively based on context. While prior work links induction heads to ICL through a sudden jump in accuracy, this can only account for ICL when the answer is included within the context. However, an important property of practical ICL in large language models is the ability to meta-learn how to solve tasks from context, rather than just copying answers from context; how such an ability is obtained during training is largely unexplored. In this paper, we experimentally clarify how such meta-learning ability is acquired by analyzing the dynamics of the model’s circuit during training. Specifically, we extend the copy task from previous research into an In-Context Meta Learning setting, where models must infer a task from examples to answer queries. Interestingly, in this setting, we find that there are multiple phases in the process of acquiring such abilities, and that a unique circuit emerges in each phase, contrasting with the single-phases change in induction heads. The emergence of such circuits can be related to several phenomena known in large language models, and our analysis lead to a deeper understanding of the source of the transformer’s ICL ability.
Gouki Minegishi, Hiroki Furuta, Shohei Taniguchi, Yusuke Iwasawa, Yutaka Matsuo
ICML5
2025 A Comprehensive Survey on Physical Risk Control in the Era of Foundation Model-enabled Robotics
abstract
Recent Foundation Model-enabled robotics (FMRs) display greatly improved general-purpose skills, enabling more adaptable automation than conventional robotics. Their ability to handle diverse tasks thus creates new opportunities to replace human labor. However, unlike general foundation models, FMRs interact with the physical world, where their actions directly affect the safety of humans and surrounding objects, requiring careful deployment and control. Based on this proposition, our survey comprehensively summarizes robot control approaches to mitigate physical risks by covering all the lifespan of FMRs ranging from pre-deployment to post-accident stage. Specifically, we broadly divide the timeline into the following three phases: (1) pre-deployment phase, (2) pre-incident phase, and (3) post-incident phase. Throughout this survey, we find that there is much room to study (i) pre-incident risk mitigation strategies, (ii) research that assumes physical interaction with humans, and (iii) essential issues of foundation models themselves. We hope that this survey will be a milestone in providing a high-resolution analysis of the physical risks of FMRs and their control, contributing to the realization of a good human-robot relationship.
Takeshi Kojima, Yaonan Zhu, Yusuke Iwasawa, Toshinori Kitamura, Gang Yan 0003, Shu Morikuni, Ryosuke Takanami, Alfredo Solano, Tatsuya Matsushima, Akiko Murakami, Yutaka Matsuo
IJCAI11
2025 Language Models can Categorize System Inputs for Performance Analysis
abstract
Dominic Sobhani, Ruiqi Zhong, Edison Marrese-Taylor, Keisuke Sakaguchi, Yutaka Matsuo. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Dominic Sobhani, Ruiqi Zhong, Edison Marrese-Taylor, Keisuke Sakaguchi, Yutaka Matsuo
NAACL (Long Papers)5
2025 Provably Efficient RL under Episode-Wise Safety in Constrained MDPs with Linear Function Approximation
abstract
We study the reinforcement learning (RL) problem in a constrained Markov decision process (CMDP), where an agent explores the environment to maximize the expected cumulative reward while satisfying a single constraint on the expected total utility value in every episode. While this problem is well understood in the tabular setting, theoretical results for function approximation remain scarce. This paper closes the gap by proposing an RL algorithm for linear CMDPs that achieves $\widetilde{\mathcal{O}}(\sqrt{K})$ regret with an episode-wise zero-violation guarantee. Furthermore, our method is computationally efficient, scaling polynomially with problem-dependent parameters while remaining independent of the state space size. Our results significantly improve upon recent linear CMDP algorithms, which either violate the constraint or incur exponential computational costs.
Toshinori Kitamura, Arnob Ghosh, Tadashi Kozuno, Wataru Kumagai, Kazumi Kasaura, Kenta Hoshino, Yohei Hosoe, Yutaka Matsuo
NeurIPS8
2025 Topology of Reasoning: Understanding Large Reasoning Models through Reasoning Graph Properties
abstract
Recent large-scale reasoning models have achieved state-of-the-art performance on challenging mathematical benchmarks, yet the internal mechanisms underlying their success remain poorly understood. In this work, we introduce the notion of a reasoning graph, extracted by clustering hidden‐state representations at each reasoning step, and systematically analyze three key graph-theoretic properties: cyclicity, diameter, and small-world index, across multiple tasks (GSM8K, MATH500, AIME~2024). Our findings reveal that distilled reasoning models (e.g., DeepSeek-R1-Distill-Qwen-32B) exhibit significantly more recurrent cycles (about 5 per sample), substantially larger graph diameters, and pronounced small-world characteristics (about 6x) compared to their base counterparts. Notably, these structural advantages grow with task difficulty and model capacity, with cycle detection peaking at the 14B scale and exploration diameter maximized in the 32B variant, correlating positively with accuracy. Furthermore, we show that supervised fine-tuning on an improved dataset systematically expands reasoning graph diameters in tandem with performance gains, offering concrete guidelines for dataset design aimed at boosting reasoning capabilities. By bridging theoretical insights into reasoning graph structures with practical recommendations for data construction, our work advances both the interpretability and the efficacy of large reasoning models.
Gouki Minegishi, Hiroki Furuta, Takeshi Kojima, Yusuke Iwasawa, Yutaka Matsuo
NeurIPS5
2025 Inference-Time Text-to-Video Alignment with Diffusion Latent Beam Search
abstract
The remarkable progress in text-to-video diffusion models enables the generation of photorealistic videos, although the content of these generated videos often includes unnatural movement or deformation, reverse playback, and motionless scenes. Recently, an alignment problem has attracted huge attention, where we steer the output of diffusion models based on some measure of the content's goodness. Because there is a large room for improvement of perceptual quality along the frame direction, we should address which metrics we should optimize and how we can optimize them in the video generation. In this paper, we propose diffusion latent beam search with lookahead estimator, which can select a better diffusion latent to maximize a given alignment reward at inference time. We then point out that improving perceptual video quality with respect to alignment to prompts requires reward calibration by weighting existing metrics. This is because when humans or vision language models evaluate outputs, many previous metrics to quantify the naturalness of video do not always correlate with the evaluation. We demonstrate that our method improves the perceptual quality evaluated on the calibrated reward, VLMs, and human assessment, without model parameter update, and outputs the best generation compared to greedy search and best-of-N sampling under much more efficient computational cost. The experiments highlight that our method is beneficial to many capable generative models, and provide a practical guideline: we should prioritize the inference-time compute allocation into enabling the lookahead estimator and increasing the search budget, rather than expanding the denoising steps.
Yuta Oshima, Masahiro Suzuki 0001, Yutaka Matsuo, Hiroki Furuta
NeurIPS3
2025 Real-time data-efficient portrait stylization via geometric alignment
Zhuoru Li, Xuanyu Yin, Yusuke Iwasawa, Yutaka Matsuo, Jiaxian Guo
Neural Networks6
2025 Continual Pre-training on Character-level Noisy Texts Makes Decoder-based Language Models Robust Few-shot Learners
abstract
Abstract Recent decoder-based pre-trained language models (PLMs) generally use subword tokenizers. However, adding character-level perturbations drastically changes the delimitation of texts by the tokenizers, leading to the vulnerability of PLMs. This study proposes a method of continual pre-training to convert decoder-based PLMs with subword tokenizers into perturbation-robust few-shot in-context learners. Our method continually trains decoder-based PLMs to predict the next tokens conditioning on artificially created character-level noisy texts. Since decoder-based language models are auto-regressive, we skip noised words from the target optimization. In addition, to maintain the same word prediction performance under noisy text as clean text, our method employs word distribution matching between the original PLMs and training models. We conducted experiments on various subword-based PLMs, including GPT2, Pythia, Mistral, Gemma2, and Llama3, ranging from 1B to 8B parameters. The results demonstrate that our method consistently improves the performance of few-shot in-context learning on downstream tasks which contain actual typos or misspellings as well as artificial noise.1
Takeshi Kojima, Yutaka Matsuo, Yusuke Iwasawa
Trans. Assoc. Comput. Linguistics2
2024 Which Programming Language and What Features at Pre-training Stage Affect Downstream Logical Inference Performance?
abstract
Recent large language models (LLMs) have demonstrated remarkable generalization abilities in mathematics and logical reasoning tasks.Prior research indicates that LLMs pre-trained with programming language data exhibit high mathematical and reasoning abilities; however, this causal relationship has not been rigorously tested.Our research aims to verify which programming languages and features during pre-training affect logical inference performance.Specifically, we pre-trained decoderbased language models from scratch using datasets from ten programming languages (e.g., Python, C, Java) and three natural language datasets (Wikipedia, Fineweb, C4) under identical conditions.Thereafter, we evaluated the trained models in a few-shot in-context learning setting on logical reasoning tasks: FLD and bAbi, which do not require commonsense or world knowledge.The results demonstrate that nearly all models trained with programming languages consistently outperform those trained with natural languages, indicating that programming languages contain factors that elicit logic inference performance.In addition, we found that models trained with programming languages exhibit a better ability to follow instructions compared to those trained with natural languages.Further analysis reveals that the depth of Abstract Syntax Trees representing parsed results of programs also affects logical reasoning performance.These findings will offer insights into the essential elements of pretraining for acquiring the foundational abilities of LLMs. 1
Fumiya Uchiyama, Takeshi Kojima, Andrew Gambardella, Yusuke Iwasawa, Yutaka Matsuo
EMNLP6
2024 Paste and Harmonize via Denoising: Subject-Driven Image Editing with Frozen Pre-Trained Diffusion Model
abstract
Text-to-Image generative models have shown a remarkable ability to produce high-quality images. However, existing methods still face difficulties in exemplar-guided image editing without destroying the given objects’ identity in the exemplar image. To address this problem, we propose a new framework called Paste and Harmonize via Denoising, which leverages pre-trained diffusion models to facilitate the text-driven transfer of objects from an exemplar image to the edited image while preserving their appearance and characteristics. The framework consists of two main steps: paste and harmonize via denoising. In the paste step, an off-the-shelf text-driven model is utilized to localize the objects in the exemplar image. The editing task is naturally transformed into an image harmonization task by pasting the object patches into the edited image. In the harmonize via denoising step, we introduce an image harmonization module based on pre-trained diffusion models to blend the inserted object with the target image, producing a coherent and realistic image without compromising synthesis quality and preserving the text-driven style transfer editing ability. In the experiments, the qualitative comparisons with baselines demonstrate that our method achieves impressive performance in exemplar-based image editing on both training and in-the-wild images with high fidelity. More qualitative and quantitative results can be found at our website.
Jiaxian Guo, Paul Yoo, Yutaka Matsuo, Yusuke Iwasawa
ICASSP4
2024 Multimodal Web Navigation with Instruction-Finetuned Foundation Models
abstract
The progress of autonomous web navigation has been hindered by the dependence on billions of exploratory interactions via online reinforcement learning, and domain-specific model designs that make it difficult to leverage generalization from rich out-of-domain data. In this work, we study data-driven offline training for web agents with vision-language foundation models. We propose an instruction-following multimodal agent, WebGUM, that observes both webpage screenshots and HTML pages and outputs web navigation actions, such as click and type. WebGUM is trained by jointly finetuning an instruction-finetuned language model and a vision encoder with temporal and local perception on a large corpus of demonstrations. We empirically demonstrate this recipe improves the agent's ability of grounded multimodal perception, HTML comprehension, and multi-step reasoning, outperforming prior works by a significant margin. On the MiniWoB, we improve over the previous best offline methods by more than 45.8%, even outperforming online-finetuned SoTA, humans, and GPT-4-based agent. On the WebShop benchmark, our 3-billion-parameter model achieves superior performance to the existing SoTA, PaLM-540B. Furthermore, WebGUM exhibits strong positive transfer to the real-world planning tasks on the Mind2Web. We also collect 347K high-quality demonstrations using our trained models, 38 times larger than prior work, and make them available to promote future research in this direction.
Hiroki Furuta, Kuang-Huei Lee, Ofir Nachum, Yutaka Matsuo, Aleksandra Faust, Shixiang Gu, Izzeddin Gur
ICLR4
2024 A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis
abstract
Pre-trained large language models (LLMs) have recently achieved better generalization and sample efficiency in autonomous web automation. However, the performance on real-world websites has still suffered from (1) open domainness, (2) limited context length, and (3) lack of inductive bias on HTML. We introduce WebAgent, an LLM-driven agent that learns from self-experience to complete tasks on real websites following natural language instructions. WebAgent plans ahead by decomposing instructions into canonical sub-instructions, summarizes long HTML documents into task-relevant snippets, and acts on websites via Python programs generated from those. We design WebAgent with Flan-U-PaLM, for grounded code generation, and HTML-T5, new pre-trained LLMs for long HTML documents using local and global attention mechanisms and a mixture of long-span denoising objectives, for planning and summarization. We empirically demonstrate that our modular recipe improves the success on real websites by over 50%, and that HTML-T5 is the best model to solve various HTML understanding tasks; achieving 18.7% higher success rate than the prior method on MiniWoB web automation benchmark, and SoTA performance on Mind2Web, an offline task planning evaluation.
Izzeddin Gur, Hiroki Furuta, Austin V. Huang, Mustafa Safdari, Yutaka Matsuo, Douglas Eck, Aleksandra Faust
ICLR5
2024 GenDOM: Generalizable One-shot Deformable Object Manipulation with Parameter-Aware Policy
abstract
Due to the inherent uncertainty in their deformability during motion, previous methods in deformable object manipulation, such as rope and cloth, often required hundreds of real-world demonstrations to train a manipulation policy for each object, which hinders their applications in our ever-changing world. To address this issue, we introduce GenDOM, a framework that allows the manipulation policy to handle different deformable objects with only a single real-world demonstration. To achieve this, we augment the policy by conditioning it on deformable object parameters and training it with a diverse range of simulated deformable objects so that the policy can adjust actions based on different object parameters. At the time of inference, given a new object, GenDOM can estimate the deformable object parameters with only a single real-world demonstration by minimizing the disparity between the grid density of point clouds of real-world demonstrations and simulations in a differentiable physics simulator. Empirical validations on both simulated and real-world object manipulation setups clearly show that our method can manipulate different objects with a single demonstration and significantly outperforms the baseline in both environments (a 62% improvement for in-domain ropes and a 15% improvement for out-of-distribution ropes in simulation, as well as a 26% improvement for ropes and a 50% improvement for cloths in the real world), demonstrating the effectiveness of our approach in one-shot deformable object manipulation. https://sites.google.com/view/gendom/home.
So Kuroki, Jiaxian Guo, Tatsuya Matsushima, Takuya Okubo, Masato Kobayashi 0001, Yuya Ikeda, Ryosuke Takanami, Paul Yoo, Yutaka Matsuo, Yusuke Iwasawa
ICRA9
2024 Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration
abstract
Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for many applications. Can such a consolidation happen in robotics? Conventionally, robotic learning methods train a separate model for every application, every robot, and even every environment. Can we instead train "generalist" X-robot policy that can be adapted efficiently to new robots, tasks, and environments? In this paper, we provide datasets in standardized data formats and models to make it possible to explore this possibility in the context of robotic manipulation, alongside experimental results that provide an example of effective X-robot policies. We assemble a dataset from 22 different robots collected through a collaboration between 21 institutions, demonstrating 527 skills (160266 tasks). We show that a high-capacity model trained on this data, which we call RT-X, exhibits positive transfer and improves the capabilities of multiple robots by leveraging experience from other platforms. The project website is robotics-transformer-x.github.io.
Abigail O'Neill, Abhiram Maddukuri, Abhishek Gupta 0004, Abhishek Padalkar, Abraham Lee, Acorn Pooley, Agrim Gupta, Ajay Mandlekar, Ajinkya Jain, Albert Tung, Alex Bewley, Alex Irpan, Alexander Khazatsky, Anant Rai, Anchit Gupta, Andrew E. Wang, Anikait Singh, Animesh Garg, Aniruddha Kembhavi, Annie Xie, Anthony Brohan, Antonin Raffin, Archit Sharma, Arefeh Yavary, Arhan Jain, Ashwin Balakrishna, Ayzaan Wahid, Ben Burgess-Limerick, Bernhard Schölkopf, Blake Wulfe, Brian Ichter, Cewu Lu, Charles Xu 0003, Charlotte Le, Chelsea Finn, Chen Wang 0053, Chenfeng Xu, Cheng Chi 0001, Chenguang Huang, Christine Chan, Christopher Agia, Chuer Pan, Chuyuan Fu, Coline Devin, Danfei Xu, Daniel Morton, Danny Drieß, Daphne Chen, Deepak Pathak, Dhruv Shah, Dieter Büchler, Dinesh Jayaraman, Dmitry Kalashnikov, Dorsa Sadigh, Edward Johns, Ethan Paul Foster, Fangchen Liu, Federico Ceola, Fei Xia 0002, Feiyu Zhao, Freek Stulp, Gaoyue Zhou, Gaurav S. Sukhatme, Gautam Salhotra, Gilbert Feng, Giulio Schiavi, Glen Berseth, Gregory Kahn, Guanzhi Wang, Hao Su 0001, Haoshu Fang, Henghui Bao, Heni Ben Amor, Henrik I. Christensen, Hiroki Furuta, Homer Walke, Hongjie Fang, Huy Ha, Igor Mordatch, Ilija Radosavovic, Isabel Leal, Jacky Liang, Jad Abou-Chakra, Jaehyung Kim 0001, Jaimyn Drake, Jan Peters 0001, Jan Schneider 0007, Jasmine Hsu, Jeannette Bohg, Jeffrey T. Bingham, Jensen Gao, Jiaheng Hu, Jiajun Wu 0001, Jiankai Sun, Jianlan Luo, Jiayuan Gu, Jie Tan 0001, Jihoon Oh, Jimmy Wu, Jingpei Lu, Jitendra Malik, João Silvério, Joey Hejna, Jonathan Booher, Jonathan Tompson, Jonathan Yang, Jordi Salvador, Joseph J. Lim, Junhyek Han, Kanishka Rao, Karl Pertsch, Karol Hausman, Keegan Go, Keerthana Gopalakrishnan, Kenneth Y. Goldberg, Kendra Byrne, Kenneth Oslund, Kento Kawaharazuka, Kevin Black, Kevin Zhang 0002, Kiana Ehsani, Kiran Lekkala, Kirsty Ellis, Krishan Rana, Krishnan Srinivasan, Kuan Fang, Kunal Pratap Singh, Kuo-Hao Zeng, Kyle Hatch, Kyle Hsu, Laurent Itti, Yunliang Chen 0001, Lerrel Pinto, Li Fei-Fei 0001, Liam Tan, Linxi Fan, Lionel Ott, Lisa Lee, Luca Weihs, Magnum Chen, Marion Lepert, Marius Memmel, Masayoshi Tomizuka, Masha Itkina, Mateo Guaman Castro, Max Spero, Maximilian Du, Michael Ahn, Michael C. Yip, Mingtong Zhang 0003, Mingyu Ding, Minho Heo, Mohan Kumar Srirama, Mohit Sharma 0001, Moo Jin Kim, Naoaki Kanazawa, Nicklas Hansen 0001, Nicolas Heess, Nikhil J. Joshi, Niko Sünderhauf, Norman Di Palo, Nur Muhammad Shafiullah, Oier Mees, Oliver Kroemer, Osbert Bastani, Pannag R. Sanketi, Patrick Tree Miller, Patrick Yin, Paul Wohlhart, Peng Xu 0010, Peter David Fagan, Peter Mitrano, Pierre Sermanet, Pieter Abbeel, Priya Sundaresan, Qiuyu Chen, Rafael Rafailov, Ria Doshi, Roberto Martin Martin, Rohan Baijal, Rosario Scalise, Rose Hendrix, Roy Lin, Runjia Qian, Russell Mendonca, Rutav Shah, Ryan Hoque, Ryan Julian, Samuel Bustamante-Gomez, Sean Kirmani, Sergey Levine, Sherry Moore, Shikhar Bahl, Shivin Dass, Shubham D. Sonawani, Shuran Song, Sichun Xu, Siddhant Haldar, Siddharth Karamcheti, Simeon Adebola, Simon Guist, Soroush Nasiriany, Stefan Schaal, Stefan Welker, Stephen Tian, Subramanian Ramamoorthy, Sudeep Dasari, Suneel Belkhale, Sungjae Park, Suraj Nair 0003, Suvir Mirchandani, Takayuki Osa, Tanmay Gupta, Tatsuya Harada, Tatsuya Matsushima, Ted Xiao, Thomas Kollar, Tianhe Yu, Tianli Ding, Todor Davchev, Tony Z. Zhao, Travis Armstrong, Trevor Darrell, Trinity Chung, Vidhi Jain, Vincent Vanhoucke, Wolfram Burgard, Xiaolong Wang 0004, Xinghao Zhu, Xinyang Geng, Liangwei Xu, Yecheng Jason Ma 0001, Yejin Kim 0003, Yevgen Chebotar, Yilin Wu 0003, Yonatan Bisk, Yoonyoung Cho, Youngwoon Lee, Yuchen Cui, Yueh-Hua Wu, Yujin Tang, Yuke Zhu, Yunchu Zhang, Yunfan Jiang 0001, Yunshuang Li, Yunzhu Li, Yusuke Iwasawa, Yutaka Matsuo, Zehan Ma, Zichen Jeff Cui, Zichen Zhang 0016, Zipeng Lin
ICRA274
2024 Self-Recovery Prompting: Promptable General Purpose Service Robot System with Foundation Models and Self-Recovery
abstract
A general-purpose service robot (GPSR), which can execute diverse tasks in various environments, requires a system with high generalizability and adaptability to tasks and environments. In this paper, we first developed a top-level GPSR system for worldwide competition (RoboCup@Home2023) based on multiple foundation models. This system is both generalizable to variations and adaptive by prompting each model. Then, by analyzing the performance of the developed system, we found three types of failure in more realistic GPSR application settings: insufficient information, incorrect plan generation, and plan execution failure. We then propose the self-recovery prompting pipeline, which explores the necessary information and modifies its prompts to recover from failure. We experimentally confirm that the system with the self-recovery mechanism can accomplish tasks by resolving various failure cases. https://sites.google.com/view/srgpsr
Mimo Shirasaka, Tatsuya Matsushima, Soshi Tsunashima, Yuya Ikeda, Aoi Horo, So Ikoma, Chikaha Tsuji, Hikaru Wada, Tsunekazu Omija, Dai Komukai, Yutaka Matsuo, Yusuke Iwasawa
ICRA11
2024 On the Multilingual Ability of Decoder-based Pre-trained Language Models: Finding and Controlling Language-Specific Neurons
abstract
Takeshi Kojima, Itsuki Okimura, Yusuke Iwasawa, Hitomi Yanaka, Yutaka Matsuo. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Takeshi Kojima, Itsuki Okimura, Yusuke Iwasawa, Hitomi Yanaka, Yutaka Matsuo
NAACL-HLT5
2024 Media Bias Detection Across Families of Language Models
abstract
Iffat Maab, Edison Marrese-Taylor, Sebastian Padó, Yutaka Matsuo. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Iffat Maab, Edison Marrese-Taylor, Sebastian Padó, Yutaka Matsuo
NAACL-HLT4
2024 Geometric-Averaged Preference Optimization for Soft Preference Labels
abstract
Many algorithms for aligning LLMs with human preferences assume that human preferences are binary and deterministic. However, human preferences can vary across individuals, and therefore should be represented distributionally. In this work, we introduce the distributional soft preference labels and improve Direct Preference Optimization (DPO) with a weighted geometric average of the LLM output likelihood in the loss function. This approach adjusts the scale of learning loss based on the soft labels such that the loss would approach zero when the responses are closer to equally preferred. This simple modification can be easily applied to any DPO-based methods and mitigate over-optimization and objective mismatch, which prior works suffer from. Our experiments simulate the soft preference labels with AI feedback from LLMs and demonstrate that geometric averaging consistently improves performance on standard benchmarks for alignment research. In particular, we observe more preferable responses than binary labels and significant improvements where modestly-confident labels are in the majority.
Hiroki Furuta, Kuang-Huei Lee, Shixiang Gu, Yutaka Matsuo, Aleksandra Faust, Heiga Zen, Izzeddin Gur
NeurIPS4
2024 ADOPT: Modified Adam Can Converge with Any β2 with the Optimal Rate
Shohei Taniguchi, Keno Harada, Gouki Minegishi, Yuta Oshima, Seong Cheol Jeong, Go Nagahara, Tomoshi Iiyama, Masahiro Suzuki 0001, Yusuke Iwasawa, Yutaka Matsuo
NeurIPS10
2023 Unnatural Error Correction: GPT-4 Can Almost Perfectly Handle Unnatural Scrambled Text
abstract
While Large Language Models (LLMs) have achieved remarkable performance in many tasks, much about their inner workings remains unclear.In this study, we present novel experimental insights into the resilience of LLMs, particularly GPT-4, when subjected to extensive character-level permutations.To investigate this, we first propose the Scrambled Bench, a suite designed to measure the capacity of LLMs to handle scrambled input, in terms of both recovering scrambled sentences and answering questions given scrambled context.The experimental results indicate that multiple advanced LLMs demonstrate the capability akin to typoglycemia 1 , a phenomenon where humans can understand the meaning of words even when the letters within those words are scrambled, as long as the first and last letters remain in place.More surprisingly, we found that only GPT-4 nearly flawlessly processes inputs with unnatural errors, a task that poses significant challenges for other LLMs and often even for humans.Specifically, GPT-4 can almost perfectly reconstruct the original sentences from scrambled ones, decreasing the edit distance by 95%, even when all letters within each word are entirely scrambled.It is counter-intuitive that LLMs can exhibit such resilience despite severe disruption to input tokenization caused by scrambled text. 2
Takeshi Kojima, Yutaka Matsuo, Yusuke Iwasawa
EMNLP3
2023 A System for Morphology-Task Generalization via Unified Representation and Behavior Distillation
Hiroki Furuta, Yusuke Iwasawa, Yutaka Matsuo, Shixiang Gu
ICLR3
2023 Interaction-Based Disentanglement of Entities for Object-Centric World Models
Akihiro Nakano, Masahiro Suzuki 0001, Yutaka Matsuo
ICLR3
2023 Regularization and Variance-Weighted Regression Achieves Minimax Optimality in Linear MDPs: Theory and Practice
abstract
Mirror descent value iteration (MDVI), an abstraction of Kullback-Leibler (KL) and entropy-regularized reinforcement learning (RL), has served as the basis for recent high-performing practical RL algorithms. However, despite the use of function approximation in practice, the theoretical understanding of MDVI has been limited to tabular Markov decision processes (MDPs). We study MDVI with linear function approximation through its sample complexity required to identify an $\varepsilon$-optimal policy with probability $1-\delta$ under the settings of an infinite-horizon linear MDP, generative model, and G-optimal design. We demonstrate that least-squares regression weighted by the variance of an estimated optimal value function of the next state is crucial to achieving minimax optimality. Based on this observation, we present Variance-Weighted Least-Squares MDVI (VWLS-MDVI), the first theoretical algorithm that achieves nearly minimax optimal sample complexity for infinite-horizon linear MDPs. Furthermore, we propose a practical VWLS algorithm for value-based deep RL, Deep Variance Weighting (DVW). Our experiments demonstrate that DVW improves the performance of popular value-based deep RL algorithms on a set of MinAtar benchmarks.
Toshinori Kitamura, Tadashi Kozuno, Yunhao Tang, Nino Vieillard, Michal Valko, Jincheng Mei, Pierre Ménard, Mohammad Gheshlaghi Azar, Rémi Munos, Olivier Pietquin, Matthieu Geist, Csaba Szepesvári, Wataru Kumagai, Yutaka Matsuo
ICML15
2023 End-to-end Training of Deep Boltzmann Machines by Unbiased Contrastive Divergence with Local Mode Initialization
abstract
We address the problem of biased gradient estimation in deep Boltzmann machines (DBMs). The existing method to obtain an unbiased estimator uses a maximal coupling based on a Gibbs sampler, but when the state is high-dimensional, it takes a long time to converge. In this study, we propose to use a coupling based on the Metropolis-Hastings (MH) and to initialize the state around a local mode of the target distribution. Because of the propensity of MH to reject proposals, the coupling tends to converge in only one step with a high probability, leading to high efficiency. We find that our method allows DBMs to be trained in an end-to-end fashion without greedy pretraining. We also propose some practical techniques to further improve the performance of DBMs. We empirically demonstrate that our training algorithm enables DBMs to show comparable generative performance to other deep generative models, achieving the FID score of 10.33 for MNIST.
Shohei Taniguchi, Masahiro Suzuki 0001, Yusuke Iwasawa, Yutaka Matsuo
ICML4
2023 Target-Aware Contextual Political Bias Detection in News
abstract
Iffat Maab, Edison Marrese-Taylor, Yutaka Matsuo. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Iffat Maab, Edison Marrese-Taylor, Yutaka Matsuo
IJCNLP (1)3
2023 Improving the Robustness to Variations of Objects and Instructions with a Neuro-Symbolic Approach for Interactive Instruction Following
Kazutoshi Shinoda, Yuki Takezawa, Masahiro Suzuki 0001, Yusuke Iwasawa, Yutaka Matsuo
MMM (2)5
2023 DreamSparse: Escaping from Plato's Cave with 2D Diffusion Model Given Sparse Views
abstract
Synthesizing novel view images from a few views is a challenging but practical problem. Existing methods often struggle with producing high-quality results or necessitate per-object optimization in such few-view settings due to the insufficient information provided. In this work, we explore leveraging the strong 2D priors in pre-trained diffusion models for synthesizing novel view images. 2D diffusion models, nevertheless, lack 3D awareness, leading to distorted image synthesis and compromising the identity. To address these problems, we propose $\textit{DreamSparse}$, a framework that enables the frozen pre-trained diffusion model to generate geometry and identity-consistent novel view images. Specifically, DreamSparse incorporates a geometry module designed to capture features about spatial information from sparse views as a 3D prior. Subsequently, a spatial guidance model is introduced to convert rendered feature maps as spatial information for the generative process. This information is then used to guide the pre-trained diffusion model to encourage the synthesis of geometrically consistent images without further tuning. Leveraging the strong image priors in the pre-trained diffusion models, DreamSparse is capable of synthesizing high-quality novel views for both object and object-centric scene-level images and generalising to open-set images. Experimental results demonstrate that our framework can effectively synthesize novel view images from sparse views and outperforms baselines in both trained and open-set category images. More results can be found on our project page: https://sites.google.com/view/dreamsparse-webpage.
Paul Yoo, Jiaxian Guo, Yutaka Matsuo, Shixiang Gu
NeurIPS3
2023 Deep Learning Pipeline for Spotting Macro- and Micro-expressions in Long Video Sequences Based on Action Units and Optical Flow
Bo Yang 0057, Kazushi Ikeda, Gen Hattori, Masaru Sugano, Yusuke Iwasawa, Yutaka Matsuo
Pattern Recognit. Lett.7
2022 Generalized Decision Transformer for Offline Hindsight Information Matching
Hiroki Furuta, Yutaka Matsuo, Shixiang Gu
ICLR2
2022 Robustifying Vision Transformer without Retraining from Scratch by Test-Time Class-Conditional Feature Alignment
abstract
Vision Transformer (ViT) is becoming more popular in image processing. Specifically, we investigate the effectiveness of test-time adaptation (TTA) on ViT, a technique that has emerged to correct its prediction during test-time by itself. First, we benchmark various test-time adaptation approaches on ViT-B16 and ViT-L16. It is shown that the TTA is effective on ViT and the prior-convention (sensibly selecting modulation parameters) is not necessary when using proper loss function. Based on the observation, we propose a new test-time adaptation method called class-conditional feature alignment (CFA), which minimizes both the class-conditional distribution differences and the whole distribution differences of the hidden representation between the source and target in an online manner. Experiments of image classification tasks on common corruption (CIFAR-10-C, CIFAR-100-C, and ImageNet-C) and domain adaptation (digits datasets and ImageNet-Sketch) show that CFA stably outperforms the existing baselines on various datasets. We also verify that CFA is model agnostic by experimenting on ResNet, MLP-Mixer, and several ViT variants (ViT-AugReg, DeiT, and BeiT). Using BeiT backbone, CFA achieves 19.8% top-1 error rate on ImageNet-C, outperforming the existing test-time adaptation baseline 44.0%. This is a state-of-the-art result among TTA methods that do not need to alter training phase.
Takeshi Kojima, Yutaka Matsuo, Yusuke Iwasawa
IJCAI2
2022 Large Language Models are Zero-Shot Reasoners
abstract
Pretrained large language models (LLMs) are widely used in many sub-fields of natural language processing (NLP) and generally known as excellent few-shot learners with task-specific exemplars. Notably, chain of thought (CoT) prompting, a recent technique for eliciting complex multi-step reasoning through step-by-step answer examples, achieved the state-of-the-art performances in arithmetics and symbolic reasoning, difficult system-2 tasks that do not follow the standard scaling laws for LLMs. While these successes are often attributed to LLMs' ability for few-shot learning, we show that LLMs are decent zero-shot reasoners by simply adding ``Let's think step by step'' before each answer. Experimental results demonstrate that our Zero-shot-CoT, using the same single prompt template, significantly outperforms zero-shot LLM performances on diverse benchmark reasoning tasks including arithmetics (MultiArith, GSM8K, AQUA-RAT, SVAMP), symbolic reasoning (Last Letter, Coin Flip), and other logical reasoning tasks (Date Understanding, Tracking Shuffled Objects), without any hand-crafted few-shot examples, e.g. increasing the accuracy on MultiArith from 17.7% to 78.7% and GSM8K from 10.4% to 40.7% with large-scale InstructGPT model (text-davinci-002), as well as similar magnitudes of improvements with another off-the-shelf large model, 540B parameter PaLM. The versatility of this single prompt across very diverse reasoning tasks hints at untapped and understudied fundamental zero-shot capabilities of LLMs, suggesting high-level, multi-task broad cognitive capabilities may be extracted by simple prompting. We hope our work not only serves as the minimal strongest zero-shot baseline for the challenging reasoning benchmarks, but also highlights the importance of carefully exploring and analyzing the enormous zero-shot knowledge hidden inside LLMs before crafting finetuning datasets or few-shot exemplars.
Takeshi Kojima, Shixiang Gu, Machel Reid, Yutaka Matsuo, Yusuke Iwasawa
NeurIPS4
2022 Langevin Autoencoders for Learning Deep Latent Variable Models
abstract
Markov chain Monte Carlo (MCMC), such as Langevin dynamics, is valid for approximating intractable distributions. However, its usage is limited in the context of deep latent variable models owing to costly datapoint-wise sampling iterations and slow convergence. This paper proposes the amortized Langevin dynamics (ALD), wherein datapoint-wise MCMC iterations are entirely replaced with updates of an encoder that maps observations into latent variables. This amortization enables efficient posterior sampling without datapoint-wise iterations. Despite its efficiency, we prove that ALD is valid as an MCMC algorithm, whose Markov chain has the target posterior as a stationary distribution under mild assumptions. Based on the ALD, we also present a new deep latent variable model named the Langevin autoencoder (LAE). Interestingly, the LAE can be implemented by slightly modifying the traditional autoencoder. Using multiple synthetic datasets, we first validate that ALD can properly obtain samples from target posteriors. We also evaluate the LAE on the image generation task, and show that our LAE can outperform existing methods based on variational inference, such as the variational autoencoder, and other MCMC-based methods in terms of the test likelihood.
Shohei Taniguchi, Yusuke Iwasawa, Wataru Kumagai, Yutaka Matsuo
NeurIPS4
2022 Deep learning, reinforcement learning, and world models
abstract
Deep learning (DL) and reinforcement learning (RL) methods seem to be a part of indispensable factors to achieve human-level or super-human AI systems. On the other hand, both DL and RL have strong connections with our brain functions and with neuroscientific findings. In this review, we summarize talks and discussions in the "Deep Learning and Reinforcement Learning" session of the symposium, International Symposium on Artificial Intelligence and Brain Science. In this session, we discussed whether we can achieve comprehensive understanding of human intelligence based on the recent advances of deep learning and reinforcement learning algorithms. Speakers contributed to provide talks about their recent studies that can be key technologies to achieve human-level intelligence.
Yutaka Matsuo, Yann LeCun, Maneesh Sahani, Doina Precup, David Silver 0001, Masashi Sugiyama, Eiji Uchibe, Jun Morimoto
Neural Networks1
2022 Face-mask-aware Facial Expression Recognition based on Face Parsing and Vision Transformer
Bo Yang 0057, Kazushi Ikeda, Gen Hattori, Masaru Sugano, Yusuke Iwasawa, Yutaka Matsuo
Pattern Recognit. Lett.7
2021 Variational Inference for Learning Representations of Natural Language Edits
abstract
Document editing has become a pervasive component of production of information, with version control systems enabling edits to be efficiently stored and applied. In light of this, the task of learning distributed representations of edits has been recently proposed. With this in mind, we propose a novel approach that employs variational inference to learn a continuous latent space of vector representations to capture the underlying semantic information with regard to the document editing process. We achieve this by introducing a latent variable to explicitly model the aforementioned features. This latent variable is then combined with a document representation to guide the generation of an edited-version of this document. Additionally, to facilitate standardized automatic evaluation of edit representations, which has heavily relied on direct human input thus far, we also propose a suite of downstream tasks, PEER, specifically designed to measure the quality of edit representations in the context of natural language processing.
Edison Marrese-Taylor, Machel Reid, Yutaka Matsuo
AAAI3
2021 AfroMT: Pretraining Strategies and Reproducible Benchmarks for Translation of 8 African Languages
abstract
Reproducible benchmarks are crucial in driving progress of machine translation research.However, existing machine translation benchmarks have been mostly limited to highresource or well-represented languages.Despite an increasing interest in low-resource machine translation, there are no standardized reproducible benchmarks for many African languages, many of which are used by millions of speakers but have less digitized textual data.To tackle these challenges, we propose AFROMT, a standardized, clean, and reproducible machine translation benchmark for eight widely spoken African languages.We also develop a suite of analysis tools for system diagnosis taking into account unique properties of these languages.Furthermore, we explore the newly considered case of low-resource focused pretraining and develop two novel data augmentation-based strategies, leveraging word-level alignment information and pseudo-monolingual data for pretraining multilingual sequence-to-sequence models.We demonstrate significant improvements when pretraining on 11 languages, with gains of up to 2 BLEU points over strong baselines.We also show gains of up to 12 BLEU points over cross-lingual transfer baselines in data-constrained scenarios.All code and pretrained models will be released as further steps towards larger reproducible benchmarks for African languages.1
Machel Reid, Junjie Hu 0001, Graham Neubig, Yutaka Matsuo
EMNLP (1)4
2021 Group Equivariant Conditional Neural Processes
Makoto Kawano, Wataru Kumagai, Akiyoshi Sannai, Yusuke Iwasawa, Yutaka Matsuo
ICLR5
2021 Deployment-Efficient Reinforcement Learning via Model-Based Offline Optimization
Tatsuya Matsushima, Hiroki Furuta, Yutaka Matsuo, Ofir Nachum, Shixiang Gu
ICLR3
2021 Policy Information Capacity: Information-Theoretic Measure for Task Complexity in Deep Reinforcement Learning
abstract
Progress in deep reinforcement learning (RL) research is largely enabled by benchmark task environments. However, analyzing the nature of those environments is often overlooked. In particular, we still do not have agreeable ways to measure the difficulty or solvability of a task, given that each has fundamentally different actions, observations, dynamics, rewards, and can be tackled with diverse RL algorithms. In this work, we propose policy information capacity (PIC) – the mutual information between policy parameters and episodic return – and policy-optimal information capacity (POIC) – between policy parameters and episodic optimality – as two environment-agnostic, algorithm-agnostic quantitative metrics for task difficulty. Evaluating our metrics across toy environments as well as continuous control benchmark tasks from OpenAI Gym and DeepMind Control Suite, we empirically demonstrate that these information-theoretic metrics have higher correlations with normalized task solvability scores than a variety of alternatives. Lastly, we show that these metrics can also be used for fast and compute-efficient optimizations of key design parameters such as reward shaping, policy architectures, and MDP properties for better solvability by RL algorithms without ever running full RL experiments.
Hiroki Furuta, Tatsuya Matsushima, Tadashi Kozuno, Yutaka Matsuo, Sergey Levine, Ofir Nachum, Shixiang Gu
ICML4
2021 Co-Adaptation of Algorithmic and Implementational Innovations in Inference-based Deep Reinforcement Learning
abstract
Recently many algorithms were devised for reinforcement learning (RL) with function approximation. While they have clear algorithmic distinctions, they also have many implementation differences that are algorithm-independent and sometimes under-emphasized. Such mixing of algorithmic novelty and implementation craftsmanship makes rigorous analyses of the sources of performance improvements across algorithms difficult. In this work, we focus on a series of off-policy inference-based actor-critic algorithms -- MPO, AWR, and SAC -- to decouple their algorithmic innovations and implementation decisions. We present unified derivations through a single control-as-inference objective, where we can categorize each algorithm as based on either Expectation-Maximization (EM) or direct Kullback-Leibler (KL) divergence minimization and treat the rest of specifications as implementation details. We performed extensive ablation studies, and identified substantial performance drops whenever implementation details are mismatched for algorithmic choices. These results show which implementation or code details are co-adapted and co-evolved with algorithms, and which are transferable across algorithms: as examples, we identified that tanh Gaussian policy and network sizes are highly adapted to algorithmic types, while layer normalization and ELU are critical for MPO's performances but also transfer to noticeable gains in SAC. We hope our work can inspire future work to further demystify sources of performance improvements across multiple algorithms and allow researchers to build on one another's both algorithmic and implementational innovations.
Hiroki Furuta, Tadashi Kozuno, Tatsuya Matsushima, Yutaka Matsuo, Shixiang Gu
NeurIPS4
2021 Test-Time Classifier Adjustment Module for Model-Agnostic Domain Generalization
abstract
This paper presents a new algorithm for domain generalization (DG), \textit{test-time template adjuster (T3A)}, aiming to robustify a model to unknown distribution shift. Unlike existing methods that focus on \textit{training phase}, our method focuses \textit{test phase}, i.e., correcting its prediction by itself during test time. Specifically, T3A adjusts a trained linear classifier (the last layer of deep neural networks) with the following procedure: (1) compute a pseudo-prototype representation for each class using online unlabeled data augmented by the base classifier trained in the source domains, (2) and then classify each sample based on its distance to the pseudo-prototypes. T3A is back-propagation-free and modifies only the linear layer; therefore, the increase in computational cost during inference is negligible and avoids the catastrophic failure might caused by stochastic optimization. Despite its simplicity, T3A can leverage knowledge about the target domain by using off-the-shelf test-time data and improve performance. We tested our method on four domain generalization benchmarks, namely PACS, VLCS, OfficeHome, and TerraIncognita, along with various backbone networks including ResNet18, ResNet50, Big Transfer (BiT), Vision Transformers (ViT), and MLP-Mixer. The results show T3A stably improves performance on unseen domains across choices of backbone networks, and outperforms existing domain generalization methods.
Yusuke Iwasawa, Yutaka Matsuo
NeurIPS2
2021 Information-theoretic regularization for learning global features by sequential VAE
abstract
Abstract Sequential variational autoencoders (VAEs) with a global latent variable z have been studied for disentangling the global features of data, which is useful for several downstream tasks. To further assist the sequential VAEs in obtaining meaningful z, existing approaches introduce a regularization term that maximizes the mutual information (MI) between the observation and z. However, by analyzing the sequential VAEs from the information-theoretic perspective, we claim that simply maximizing the MI encourages the latent variable to have redundant information, thereby preventing the disentanglement of global features. Based on this analysis, we derive a novel regularization method that makes z informative while encouraging disentanglement. Specifically, the proposed method removes redundant information by minimizing the MI between z and the local features by using adversarial training. In the experiments, we trained two sequential VAEs, state-space and autoregressive model variants, using speech and image datasets. The results indicate that the proposed method improves the performance of downstream classification and data generation tasks, thereby supporting our information-theoretic perspective for the learning of global features.
Kei Akuzawa, Yusuke Iwasawa, Yutaka Matsuo
Mach. Learn.3
2021 Graph-based knowledge tracing: Modeling student proficiency using graph neural networks
abstract
Recent advancements in computer-assisted learning systems have caused an increase in the research in knowledge tracing, wherein student performance is predicted over time. Student coursework can potentially be structured as a graph. Incorporating this graph-structured nature into a knowledge tracing model as a relational inductive bias can improve its performance; however, previous methods, such as deep knowledge tracing, did not consider such a latent graph structure. Inspired by the recent successes of graph neural networks (GNNs), we herein propose a GNN-based knowledge tracing method, i.e., graph-based knowledge tracing. Casting the knowledge structure as a graph enabled us to reformulate the knowledge tracing task as a time-series node-level classification problem in the GNN. As the knowledge graph structure is not explicitly provided in most cases, we propose various implementations of the graph structure. Empirical validations on two open datasets indicated that our method could potentially improve the prediction of student performance and demonstrated more interpretable predictions compared to those of the previous methods, without the requirement of any additional information.
Hiromi Nakagawa, Yusuke Iwasawa, Yutaka Matsuo
Web Intell.3
2020 VCDM: Leveraging Variational Bi-encoding and Deep Contextualized Word Representations for Improved Definition Modeling
abstract
In this paper, we tackle the task of definition modeling, where the goal is to learn to generate definitions of words and phrases. Existing approaches for this task are discriminative, combining distributional and lexical semantics in an implicit rather than direct way. To tackle this issue we propose a generative model for the task, introducing a continuous latent variable to explicitly model the underlying relationship between a phrase used within a context and its definition. We rely on variational inference for estimation and leverage contextualized word embeddings for improved performance. Our approach is evaluated on four existing challenging benchmarks with the addition of two new datasets, “Cambridge” and the first non-English corpus “Robert”, which we release to complement our empirical study. Our Variational Contextual Definition Modeler (VCDM) achieves state-of-the-art performance in terms of automatic and human evaluation metrics, demonstrating the effectiveness of our approach.
Machel Reid, Edison Marrese-Taylor, Yutaka Matsuo
EMNLP (1)3
2020 Gravity of Location-Based Service: Analyzing the Effects for Mobility Pattern and Location Prediction
Keiichi Ochiai, Yusuke Fukazawa, Wataru Yamada, Hiroyuki Manabe, Yutaka Matsuo
ICWSM5
2020 Stabilizing Adversarial Invariance Induction from Divergence Minimization Perspective
abstract
Adversarial invariance induction (AII) is a generic and powerful framework for enforcing an invariance to nuisance attributes into neural network representations. However, its optimization is often unstable and little is known about its practical behavior. This paper presents an analysis of the reasons for the optimization difficulties and provides a better optimization procedure by rethinking AII from a divergence minimization perspective. Interestingly, this perspective indicates a cause of the optimization difficulties: it does not ensure proper divergence minimization, which is a requirement of the invariant representations. We then propose a simple variant of AII, called invariance induction by discriminator matching, which takes into account the divergence minimization interpretation of the invariant representations. Our method consistently achieves near-optimal invariance in toy datasets with various configurations in which the original AII is catastrophically unstable. Extentive experiments on four real-world datasets also support the superior performance of the proposed method, leading to improved user anonymization and domain generalization.
Yusuke Iwasawa, Kei Akuzawa, Yutaka Matsuo
IJCAI3
2020 Learning to Describe Editing Activities in Collaborative Environments: A Case Study on GitHub and Wikipedia
Edison Marrese-Taylor, Pablo Loyola, Jorge A. Balazs, Yutaka Matsuo
PACLIC4
2019 Adversarial Invariant Feature Learning with Accuracy Constraint for Domain Generalization
Kei Akuzawa, Yusuke Iwasawa, Yutaka Matsuo
ECML/PKDD (2)3
2019 Graph-based Knowledge Tracing: Modeling Student Proficiency Using Graph Neural Network
abstract
Recent advancements in computer-assisted learning systems have caused an increase in the research of knowledge tracing, wherein student performance on coursework exercises is predicted over time. From the viewpoint of data structure, the coursework can be potentially structured as a graph. Incorporating this graph-structured nature into the knowledge tracing model as a relational inductive bias can improve its performance; however, previous methods, such as deep knowledge tracing, did not consider such a latent graph structure. Inspired by the recent successes of the graph neural network (GNN), we herein propose a GNN-based knowledge tracing method, i.e., graph-based knowledge tracing. Casting the knowledge structure as a graph enabled us to reformulate the knowledge tracing task as a time-series node-level classification problem in the GNN. As the knowledge graph structure is not explicitly provided in most cases, we propose various implementations of the graph structure. Empirical validations on two open datasets indicated that our method could potentially improve the prediction of student performance and demonstrated more interpretable predictions compared to those of the previous methods, without the requirement of any additional information.
Hiromi Nakagawa, Yusuke Iwasawa, Yutaka Matsuo
WI3
2018 Content Aware Source Code Change Description Generation
abstract
We propose to study the generation of descriptions from source code changes by integrating the messages included on code commits and the intra-code documentation inside the source in the form of docstrings.Our hypothesis is that although both types of descriptions are not directly aligned in semantic terms -one explaining a change and the other the actual functionality of the code being modified-there could be certain common ground that is useful for the generation.To this end, we propose an architecture that uses the source codedocstring relationship to guide the description generation.We discuss the results of the approach comparing against a baseline based on a sequence-to-sequence model, using standard automatic natural language generation metrics as well as with a human study, thus offering a comprehensive view of the feasibility of the approach.
Pablo Loyola, Edison Marrese-Taylor, Jorge A. Balazs, Yutaka Matsuo, Fumiko Satoh
INLG4
2018 Expressive Speech Synthesis via Modeling Expressions with Variational Autoencoder
abstract
Recent advances in neural autoregressive models have improve the performance of speech synthesis (SS).However, as they lack the ability to model global characteristics of speech (such as speaker individualities or speaking styles), particularly when these characteristics have not been labeled, making neural autoregressive SS systems more expressive is still an open issue.In this paper, we propose to combine VoiceLoop, an autoregressive SS model, with Variational Autoencoder (VAE).This approach, unlike traditional autoregressive SS systems, uses VAE to model the global characteristics explicitly, enabling the expressiveness of the synthesized speech to be controlled in an unsupervised manner.Experiments using the VCTK and Bliz-zard2012 datasets show the VAE helps VoiceLoop to generate higher quality speech and to control the expressions in its synthesized speech by incorporating global characteristics into the speech generating process.
Kei Akuzawa, Yusuke Iwasawa, Yutaka Matsuo
INTERSPEECH3
2017 Extractive Summarization Using Multi-Task Learning with Document Classification
abstract
The need for automatic document summarization that can be used for practical applications is increasing rapidly.In this paper, we propose a general framework for summarization that extracts sentences from a document using externally related information.Our work is aimed at single document summarization using small amounts of reference summaries.In particular, we address document summarization in the framework of multitask learning using curriculum learning for sentence extraction and document classification.The proposed framework enables us to obtain better feature representations to extract sentences from documents.We evaluate our proposed summarization method on two datasets: financial report and news corpus.Experimental results demonstrate that our summarizers achieve performance that is comparable to stateof-the-art systems.
Masaru Isonuma, Toru Fujino, Junichiro Mori, Yutaka Matsuo, Ichiro Sakata
EMNLP4
2017 Privacy Issues Regarding the Application of DNNs to Activity-Recognition using Wearables and Its Countermeasures by Use of Adversarial Training
abstract
Deep neural networks have been successfully applied to activity recognition with wearables in terms of recognition performance. However, the black-box nature of neural networks could lead to privacy concerns. Namely, generally it is hard to expect what neural networks learn from data, and so they possibly learn features that highly discriminate user-information unintentionally, which increases the risk of information-disclosure. In this study, we analyzed the features learned by conventional deep neural networks when applied to data of wearables to confirm this phenomenon.Based on the results of our analysis, we propose the use of an adversarial training framework to suppress the risk of sensitive/unintended information disclosure. Our proposed model considers both an adversarial user classifier and a regular activity-classifier during training, which allows the model to learn representations that help the classifier to distinguish the activities but which, at the same time, prevents it from accessing user-discriminative information. This paper provides an empirical validation of the privacy issue and efficacy of the proposed method using three activity recognition tasks based on data of wearables. The empirical validation shows that our proposed method suppresses the concerns without any significant performance degradation, compared to conventional deep nets on all three tasks.
Yusuke Iwasawa, Kotaro Nakayama, Ikuko Eguchi Yairi, Yutaka Matsuo
IJCAI4
2017 Learning Feature Representations from Change Dependency Graphs for Defect Prediction
abstract
Given the heterogeneity of the data that can be extracted from the software development process, defect prediction techniques have focused on associating different sources of data with the introduction of faulty code, usually relying on handcrafted features. While these efforts have generated considerable progress over the years, little attention has been given to the fact that the performance of any predictive model depends heavily on the representation of the data used, and that different representations can lead to different results. We consider this a relevant problem, as it could be affecting directly the efforts towards generating safer software systems. Therefore, we propose to study the impact of the representation of the data in defect prediction models. To this end, we focus on the use of developer activity data, from which we structure dependency graphs. Then, instead of manually generating features, such as network metrics, we propose two models inspired by recent advances in representation learning which are able to automatically generate feature representations from graph data. These new representations are compared against manually crafted features for defect prediction in real world software projects. Our results show that automatically learned features are competitive, reaching increments in prediction performance up to 13%.
Pablo Loyola, Yutaka Matsuo
ISSRE2
2017 Inferring win-lose product network from user behavior
abstract
Various data mining techniques to extract product relations have been examined, especially in the context of building intelligent recommender systems. Most such techniques, however, specifically examine co-occurrences of browsed or purchased products on e-commerce websites, which provide little or no useful information related to the direct relation of superiority or the factor which forms that superiority. For marketers and product managers, understanding the competitive advantages of a given product is important to consolidate their product differentiation strategies.
Shuhei Iitsuka, Kazuya Kawakami, Seigen Hagiwara, Takayoshi Kawakami, Takayuki Hamada, Yutaka Matsuo
WI6
2016 Whole Brain Architecture Approach Is a Feasible Way Toward an Artificial General Intelligence
Hiroshi Yamakawa, Masahiko Osawa, Yutaka Matsuo
ICONIP (1)3
2015 Road Sensing: Personal Sensing and Machine Learning for Development of Large Scale Accessibility Map
abstract
This paper proposes a methodology for developing large scale accessibility map with personal sensing by using smart phone and machine learning technologies. The strength of the proposed method is its low cost data collection, which is a key to break through stagnations of accessibility map that currently applied to limited areas. This paper developed and evaluated a prototype system that estimates types of ground surfaces by applying supervised learning techniques to activity sensing data of wheelchair users recorded by a three-axis accelerometer, focusing on knowledge extraction and visualization. As a result of evaluation using nine wheelchair users' data with Support Vector Machine, three ground surface types, curb, tactile indicator, and slope, were detected with f-score (and accuracy) of 0.63 (0.92), 0.65 (0.85), and 0.54 (0.97) respectively.
Yusuke Iwasawa, Koya Nagamine, Yutaka Matsuo, Ikuko Eguchi Yairi
ASSETS3
2015 Website Optimization Problem and Its Solutions
abstract
Online controlled experiments are widely used to improve the performance of websites by comparison of user behavior related to different variations of the given website. Although such experiments might have an important effect on the key metrics to maximize, small-scale websites have difficulty applying this methodology because they have few users. Furthermore, the candidate variations increase exponentially with the number of elements that must be optimized. A testing method that finds a high-performing variation with a few samples must be devised to address these problems. As described herein, we formalize this problem as a website optimization problem and provide a basis to apply existing search algorithms to this problem. We further organize existing testing methods and extract devices to make the experiments more effective. By combining organized algorithms and devices, we propose a rapid testing method that detects high-performing variations with few users. We evaluated our proposed method using simulation experiments. Results show that it outperforms existing methods at any website scale. Moreover, we implemented our proposed method as an optimizer program and used it on an actual small-scale website. Results show that our proposed method can achieve 57% higher performance variation than that of the generally used A/B testing method. Therefore, our proposed method can optimize a website with fewer samples. The website optimization problem has broad application possibilities that are applicable not only to websites but also to manufactured goods.
Shuhei Iitsuka, Yutaka Matsuo
KDD2
2015 Understanding Rating Behaviour and Predicting Ratings by Identifying Representative Users
Rahul Kamath, Masanao Ochi, Yutaka Matsuo
PACLIC3
2014 MIGSOM: A SOM Algorithm for Large Scale Hyperlinked Documents Inspired by Neuronal Migration
Kotaro Nakayama, Yutaka Matsuo
DASFAA (1)2
2013 Identifying Customer Preferences about Tourism Products Using an Aspect-based Opinion Mining Approach
abstract
In this study we extend Bing Liu's aspect-based opinion mining technique to apply it to the tourism domain. Using this extension, we also offer an approach for considering a new alternative to discover consumer preferences about tourism products, particularly hotels and restaurants, using opinions available on the Web as reviews. An experiment is also conducted, using hotel and restaurant reviews obtained from TripAdvisor, to evaluate our proposals. Results showed that tourism product reviews available on web sites contain valuable information about customer preferences that can be extracted using an aspect-based opinion mining approach. The proposed approach proved to be very effective in determining the sentiment orientation of opinions, achieving a precision and recall of 90%. However, on average, the algorithms were only capable of extracting 35% of the explicit aspect expressions.
Edison Marrese-Taylor, Juan D. Velásquez 0001, Felipe Bravo-Marquez, Yutaka Matsuo
KES4
2013 Minimally Supervised Novel Relation Extraction Using a Latent Relational Mapping
abstract
The World Wide Web includes semantic relations of numerous types that exist among different entities. Extracting the relations that exist between two entities is an important step in various Web-related tasks such as information retrieval (IR), information extraction, and social network extraction. A supervised relation extraction system that is trained to extract a particular relation type (source relation) might not accurately extract a new type of a relation (target relation) for which it has not been trained. However, it is costly to create training data manually for every new relation type that one might want to extract. We propose a method to adapt an existing relation extraction system to extract new relation types with minimum supervision. Our proposed method comprises two stages: learning a lower dimensional projection between different relations, and learning a relational classifier for the target relation type with instance sampling. First, to represent a semantic relation that exists between two entities, we extract lexical and syntactic patterns from contexts in which those two entities co-occur. Then, we construct a bipartite graph between relation-specific (RS) and relation-independent (RI) patterns. Spectral clustering is performed on the bipartite graph to compute a lower dimensional projection. Second, we train a classifier for the target relation type using a small number of labeled instances. To account for the lack of target relation training instances, we present a one-sided under sampling method. We evaluate the proposed method using a data set that contains 2,000 instances for 20 different relation types. Our experimental results show that the proposed method achieves a statistically significant macroaverage F-score of 62.77. Moreover, the proposed method outperforms numerous baselines and a previously proposed weakly supervised relation extraction method.
Danushka Bollegala, Yutaka Matsuo, Mitsuru Ishizuka
IEEE Trans. Knowl. Data Eng.2
2013 Tweet Analysis for Real-Time Event Detection and Earthquake Reporting System Development
abstract
Twitter has received much attention recently. An important characteristic of Twitter is its real-time nature. We investigate the real-time interaction of events such as earthquakes in Twitter and propose an algorithm to monitor tweets and to detect a target event. To detect a target event, we devise a classifier of tweets based on features such as the keywords in a tweet, the number of words, and their context. Subsequently, we produce a probabilistic spatiotemporal model for the target event that can find the center of the event location. We regard each Twitter user as a sensor and apply particle filtering, which are widely used for location estimation. The particle filter works better than other comparable methods for estimating the locations of target events. As an application, we develop an earthquake reporting system for use in Japan. Because of the numerous earthquakes and the large number of Twitter users throughout the country, we can detect an earthquake with high probability (93 percent of earthquakes of Japan Meteorological Agency (JMA) seismic intensity scale 3 or more are detected) merely by monitoring tweets. Our system detects earthquakes promptly and notification is delivered much faster than JMA broadcast announcements.
Takeshi Sakaki, Makoto Okazaki, Yutaka Matsuo
IEEE Trans. Knowl. Data Eng.3
2012 Rating Prediction by Correcting User Rating Bias
abstract
We propose a novel method to improve the prediction accuracy on the rating prediction task by correcting the bias of user ratings. We demonstrate that the manner of user rating and review is biased and that it is necessary to correct this difference for more accurate prediction. Our proposed method comprises approaches based on the detection of each user value to ratings: The bias of the rating is detected using entropy of user rating and by updating word weights only when the words appear in the review, the problem of bias is reduced. We implement this idea by extending the Prank algorithm. We apply a review -- item matrix as a feature matrix instead of a user -- item matrix because of its volume of information. Our quantitative evaluation shows that our method improves the prediction accuracy (the Rank Loss measurement) significantly by 8.70 % compared with the normal Prank algorithm. Our proposed method helps users find out what they care about when buying something, and is applicable to newer variants of the Prank algorithm. Moreover, it is useful to most review sites because we use only rating and review data.
Masanao Ochi, Yutaka Matsuo, Makoto Okabe, Rikio Onai
Web Intelligence2
2012 Automatic Annotation of Ambiguous Personal Names on the Web
abstract
Personal name disambiguation is an important task in social network extraction, evaluation and integration of ontologies, information retrieval, cross‐document coreference resolution and word sense disambiguation. We propose an unsupervised method to automatically annotate people with ambiguous names on the Web using automatically extracted keywords. Given an ambiguous personal name, first, we download text snippets for the given name from a Web search engine. We then represent each instance of the ambiguous name by a term‐entity model (TEM), a model that we propose to represent the Web appearance of an individual. A TEM of a person captures named entities and attribute values that are useful to disambiguate that person from his or her namesakes (i.e., different people who share the same name). We then use group average agglomerative clustering to identify the instances of an ambiguous name that belong to the same person. Ideally, each cluster must represent a different namesake. However, in practice it is not possible to know the number of namesakes for a given ambiguous personal name in advance. To circumvent this problem, we propose a novel normalized cuts‐based cluster stopping criterion to determine the different people on the Web for a given ambiguous name. Finally, we annotate each person with an ambiguous name using keywords selected from the clusters. We evaluate the proposed method on a data set of over 2500 documents covering 200 different people for 20 ambiguous names. Experimental results show that the proposed method outperforms numerous baselines and previously proposed name disambiguation methods. Moreover, the extracted keywords reduce ambiguity of a name in an information retrieval task, which underscores the usefulness of the proposed method in real‐world scenarios.
Danushka Bollegala, Yutaka Matsuo, Mitsuru Ishizuka
Comput. Intell.2
2011 Using Graph Based Method to Improve Bootstrapping Relation Extraction
Haibo Li 0002, Danushka Bollegala, Yutaka Matsuo, Mitsuru Ishizuka
CICLing (2)3
2011 Collaborative exploratory search in real-world context
abstract
We propose Collaborative Exploratory Search (CES), which is an integration of dialog analysis and web search that involves multiparty collaboration to accomplish an exploratory information retrieval goal. Given a real-time dialog between users on a single topic; we define CES as the task of automatically detecting the topic of the dialog and retrieving task-relevant web pages to support the dialog. To recognize the task of the dialog, we apply the Author--Topic model as a topic model. Then, attribute extraction is applied to the dialog to obtain the attributes of the tasks. Finally, a specific search query is generated to identify the task-relevant information. We implement and evaluate the CES system for a commercial in-vehicle conversation. We also develop an iPad application that listens to conversations among users and continuously retrieves relevant web pages. Our experimental results reveal that the proposed method outperforms existing methods, which demonstrates the potential usefulness of collaborative exploratory search with practically usable accuracy levels.
Naoki Tani, Danushka Bollegala, Naiwala P. Chandrasiri, Keisuke Okamoto, Kazunari Nawa, Shuhei Iitsuka, Yutaka Matsuo
CIKM7
2011 Relation Adaptation: Learning to Extract Novel Relations with Minimum Supervision
Danushka Bollegala, Yutaka Matsuo, Mitsuru Ishizuka
IJCAI2
2011 Mining Longitudinal Network for Predicting Company Value
abstract
Real-world social networks are dynamic in nature. Companies continue to collaborate, align strategically, acquire, and merge over time, and receive positive/negative impact from other companies. Consequently, their performance changes with time. If one can understand what types of network changes affect a company's value, he/she can predict the future value of the company, grasp industry innovations, and make business more successful. However, it often requires continuous records of relational changes, which are often difficult to track for companies, and the models of mining longitudinal network are quite complicated. In this study, we developed algorithms and a system to infer large-scale evolutionary company networks from public news during 1981-2009. Then, based on how networks change over time, as well as the financial information of the companies, we predicted company profit growth. This is the first study of longitudinal network-mining-based company performance analysis in the literature
Yingzi Jin, Ching-Yung Lin, Yutaka Matsuo, Mitsuru Ishizuka
IJCAI3
2011 Automatic Discovery of Personal Name Aliases from the Web
abstract
An individual is typically referred by numerous name aliases on the web. Accurate identification of aliases of a given person name is useful in various web related tasks such as information retrieval, sentiment analysis, personal name disambiguation, and relation extraction. We propose a method to extract aliases of a given personal name from the web. Given a personal name, the proposed method first extracts a set of candidate aliases. Second, we rank the extracted candidates according to the likelihood of a candidate being a correct alias of the given name. We propose a novel, automatically extracted lexical pattern-based approach to efficiently extract a large set of candidate aliases from snippets retrieved from a web search engine. We define numerous ranking scores to evaluate candidate aliases using three approaches: lexical pattern frequency, word co-occurrences in an anchor text graph, and page counts on the web. To construct a robust alias detection system, we integrate the different ranking scores into a single ranking function using ranking support vector machines. We evaluate the proposed method on three data sets: an English personal names data set, an English place names data set, and a Japanese personal names data set. The proposed method outperforms numerous baselines and previously proposed name alias extraction methods, achieving a statistically significant mean reciprocal rank (MRR) of 0.67. Experiments carried out using location names and Japanese personal names suggest the possibility of extending the proposed method to extract aliases for different types of named entities, and for different languages. Moreover, the aliases extracted using the proposed method are successfully utilized in an information retrieval task and improve recall by 20 percent in a relation-detection task.
Danushka Bollegala, Yutaka Matsuo, Mitsuru Ishizuka
IEEE Trans. Knowl. Data Eng.2
2011 A Web Search Engine-Based Approach to Measure Semantic Similarity between Words
abstract
Measuring the semantic similarity between words is an important component in various tasks on the web such as relation extraction, community mining, document clustering, and automatic metadata extraction. Despite the usefulness of semantic similarity measures in these applications, accurately measuring semantic similarity between two words (or entities) remains a challenging task. We propose an empirical method to estimate semantic similarity using page counts and text snippets retrieved from a web search engine for two words. Specifically, we define various word co-occurrence measures using page counts and integrate those with lexical patterns extracted from text snippets. To identify the numerous semantic relations that exist between two given words, we propose a novel pattern extraction algorithm and a pattern clustering algorithm. The optimal combination of page counts-based co-occurrence measures and lexical pattern clusters is learned using support vector machines. The proposed method outperforms various baselines and previously proposed web-based semantic similarity measures on three benchmark data sets showing a high correlation with human ratings. Moreover, the proposed method significantly improves the accuracy in a community mining task.
Danushka Bollegala, Yutaka Matsuo, Mitsuru Ishizuka
IEEE Trans. Knowl. Data Eng.2
2010 Multi-View Clustering with Web and Linguistic Features for Relation Extraction
abstract
Binary semantic relation extraction is particularly useful for various NLP and Web applications. Currently Web-based methods and Linguistic-based methods are two types of leading methods for semantic relation extraction task. With a novel view on integrating linguistic analysis on local text with Web frequent information, we propose a multi-view co-clustering approach for semantic relation extraction. One is feature clustering by automatically learning clustering functions for Web features, linguistic features simultaneously based on a subset of entity pairs. The other is relation clustering, using the feature clustering functions to define learning function for relation extraction. Our experiments demonstrate the superiority of our clustering approach comparing with several state-of-the-art clustering methods.
Yulan Yan, Haibo Li 0002, Yutaka Matsuo, Zhenglu Yang, Mitsuru Ishizuka
APWeb3
2010 Multi-view Bootstrapping for Relation Extraction by Exploring Web Features and Linguistic Features
Yulan Yan, Haibo Li 0002, Yutaka Matsuo, Mitsuru Ishizuka
CICLing3
2010 How to Become Famous in the Microblog World
Takeshi Sakaki, Yutaka Matsuo
ICWSM2
2010 Relations Expansion: Extracting Relationship Instances from the Web
abstract
In this paper, we propose a Relation Expansion framework, which uses a few seed sentences marked up with two entities to expand a set of sentences containing target relations. During the expansion process, label propagation algorithm is used to select the most confident entity pairs and context patterns. The label propagation algorithm is a graph based semi-supervised learning method which models the entire data set as a weighted graph and the label score is propagated on this graph. We test the proposed framework with four relationships, the results show that the label propagation is quite competitive comparing with existing methods.
Haibo Li 0002, Yutaka Matsuo, Mitsuru Ishizuka
Web Intelligence2
2010 Relational duality: unsupervised extraction of semantic relations between entities on the web
abstract
Extracting semantic relations among entities is an important first step in various tasks in Web mining and natural language processing such as information extraction, relation detection, and social network mining. A relation can be expressed extensionally by stating all the instances of that relation or intensionally by defining all the paraphrases of that relation. For example, consider the ACQUISITION relation between two companies. An extensional definition of ACQUISITION contains all pairs of companies in which one company is acquired by another (e.g. (YouTube, Google) or (Powerset, Microsoft)). On the other hand we can intensionally define ACQUISITION as the relation described by lexical patterns such as X is acquired by Y, or Y purchased X, where X and Y denote two companies. We use this dual representation of semantic relations to propose a novel sequential co-clustering algorithm that can extract numerous relations efficiently from unlabeled data. We provide an efficient heuristic to find the parameters of the proposed coclustering algorithm. Using the clusters produced by the algorithm, we train an L1 regularized logistic regression model to identify the representative patterns that describe the relation expressed by each cluster. We evaluate the proposed method in three different tasks: measuring relational similarity between entity pairs, open information extraction (Open IE), and classifying relations in a social network system. Experiments conducted using a benchmark dataset show that the proposed method improves existing relational similarity measures. Moreover, the proposed method significantly outperforms the current state-of-the-art Open IE systems in terms of both precision and recall. The proposed method correctly classifies 53 relation types in an online social network containing 470; 671 nodes and 35; 652; 475 edges, thereby demonstrating its efficacy in real-world relation detection tasks.
Danushka Bollegala, Yutaka Matsuo, Mitsuru Ishizuka
WWW2
2010 Earthquake shakes Twitter users: real-time event detection by social sensors
abstract
Twitter, a popular microblogging service, has received much attention recently. An important characteristic of Twitter is its real-time nature. For example, when an earthquake occurs, people make many Twitter posts (tweets) related to the earthquake, which enables detection of earthquake occurrence promptly, simply by observing the tweets. As described in this paper, we investigate the real-time interaction of events such as earthquakes in Twitter and propose an algorithm to monitor tweets and to detect a target event. To detect a target event, we devise a classifier of tweets based on features such as the keywords in a tweet, the number of words, and their context. Subsequently, we produce a probabilistic spatiotemporal model for the target event that can find the center and the trajectory of the event location. We consider each Twitter user as a sensor and apply Kalman filtering and particle filtering, which are widely used for location estimation in ubiquitous/pervasive computing. The particle filter works better than other comparable methods for estimating the centers of earthquakes and the trajectories of typhoons. As an application, we construct an earthquake reporting system in Japan. Because of the numerous earthquakes and the large number of Twitter users throughout the country, we can detect an earthquake with high probability (96% of earthquakes of Japan Meteorological Agency (JMA) seismic intensity scale 3 or more are detected) merely by monitoring tweets. Our system detects earthquakes promptly and sends e-mails to registered users. Notification is delivered much faster than the announcements that are broadcast by the JMA.
Takeshi Sakaki, Makoto Okazaki, Yutaka Matsuo
WWW3
2009 Unsupervised Relation Extraction by Mining Wikipedia Texts Using Information from the Web
Yulan Yan, Naoaki Okazaki, Yutaka Matsuo, Zhenglu Yang, Mitsuru Ishizuka
ACL/IJCNLP3
2009 Ranking Learning on the Web by Integrating Network-Based Features
abstract
Many efforts are undertaken by people and companies to improve their popularity, growth, and power, the outcomes of which are all expressed as rankings (designated as target rankings). Are these rankings merely the results of its elements' own attributes? In the theory of social network analysis (SNA), the performance and power of actors are usually interpreted as relations and the relational structures they embedded. In this study, we propose an algorithm to generate and integrate network-based features systematically from a given social network that mined from the Web to learn a model for explaining target rankings. Experimental results for learning to rank researchers' productivity based on social networks confirm the effectiveness of our models. This paper specifically examines the application of a social network that provides an example of advanced utilization of social networks mined from the Web.
Yingzi Jin, Yutaka Matsuo, Mitsuru Ishizuka
ASONAM2
2009 A Relational Model of Semantic Similarity between Words using Automatically Extracted Lexical Pattern Clusters from the Web
Danushka Bollegala, Yutaka Matsuo, Mitsuru Ishizuka
EMNLP2
2009 Predicting Customer Models Using Behavior-Based Features in Shops
Junichiro Mori, Yutaka Matsuo, Hitoshi Koshiba, Kenro Aihara, Hideaki Takeda 0001
UMAP2
2009 Measuring the similarity between implicit semantic relations using web search engines
abstract
Measuring the similarity between implicit semantic relations is an important task in information retrieval and natural language processing. For example, consider the situation where you know an entity-pair (e.g. Google, YouTube), between which a particular relation holds (e.g. acquisition), and you are interested in retrieving other entity-pairs for which the same relation holds (e.g. Yahoo, Inktomi). Existing keyword-based search engines cannot be directly applied in this case because in keyword-based search, the goal is to retrieve documents that are relevant to the words used in the query -- not necessarily to the relations implied by a pair of words. Accurate measurement of relational similarity is an important step in numerous natural language processing tasks such as identification of word analogies, and classification of noun-modifier pairs. We propose a method that uses Web search engines to efficiently compute the relational similarity between two pairs of words. Our method consists of three components: representing the various semantic relations that exist between a pair of words using automatically extracted lexical patterns, clustering the extracted lexical patterns to identify the different semantic relations implied by them, and measuring the similarity between different semantic relations using an inter-cluster correlation matrix. We propose a pattern extraction algorithm to extract a large number of lexical patterns that express numerous semantic relations. We then present an efficient clustering algorithm to cluster the extracted lexical patterns. Finally, we measure the relational similarity between word-pairs using inter-cluster correlation. We evaluate the proposed method in a relation classification task. Experimental results on a dataset covering multiple relation types show a statistically significant improvement over the current state-of-the-art relational similarity measures.
Danushka Bollegala, Yutaka Matsuo, Mitsuru Ishizuka
WSDM2
2009 Measuring the similarity between implicit semantic relations from the web
abstract
Measuring the similarity between semantic relations that hold among entities is an important and necessary step in various Web related tasks such as relation extraction, information retrieval and analogy detection. For example, consider the case in which a person knows a pair of entities (e.g. Google, YouTube), between which a particular relation holds (e.g. acquisition). The person is interested in retrieving other such pairs with similar relations (e.g. Microsoft, Powerset). Existing keyword-based search engines cannot be applied directly in this case because, in keyword-based search, the goal is to retrieve documents that are relevant to the words used in a query -- not necessarily to the relations implied by a pair of words. We propose a relational similarity measure, using a Web search engine, to compute the similarity between semantic relations implied by two pairs of words. Our method has three components: representing the various semantic relations that exist between a pair of words using automatically extracted lexical patterns, clustering the extracted lexical patterns to identify the different patterns that express a particular semantic relation, and measuring the similarity between semantic relations using a metric learning approach. We evaluate the proposed method in two tasks: classifying semantic relations between named entities, and solving word-analogy questions. The proposed method outperforms all baselines in a relation classification task with a statistically significant average precision score of 0.74. Moreover, it reduces the time taken by Latent Relational Analysis to process 374 word-analogy questions from 9 days to less than 6 hours, with an SAT score of 51%.
Danushka Bollegala, Yutaka Matsuo, Mitsuru Ishizuka
WWW2
2009 Community gravity: measuring bidirectional effects by trust and rating on online social networks
abstract
Several attempts have been made to analyze customer behavior on online E-commerce sites. Some studies particularly emphasize the social networks of customers. Users' reviews and ratings of a product exert effects on other consumers' purchasing behavior. Whether a user refers to other users' ratings depends on the trust accorded by a user to the reviewer. On the other hand, the trust that is felt by a user for another user correlates with the similarity of two users' ratings. This bidirectional interaction that involves trust and rating is an important aspect of understanding consumer behavior in online communities because it suggests clustering of similar users and the evolution of strong communities. This paper presents a theoretical model along with analyses of an actual online E-commerce site. We analyzed a large community site in Japan: @cosme. The noteworthy characteristics of @cosme are that users can bookmark their trusted users; in addition, they can post their own ratings of products, which facilitates our analyses of the ratings' bidirectional effects on trust and ratings. We describe an overview of the data in @cosme, analyses of effects from trust to rating and vice versa, and our proposition of a measure of community gravity, which measures how strongly a user might be attracted to a community. Our study is based on the @cosme dataset in addition to the Epinions dataset. It elucidates important insights and proposes a potentially important measure for mining online social networks.
Yutaka Matsuo, Hikaru Yamamoto
WWW1
2008 Generating Useful Network-based Features for Analyzing Social Networks
Jun Karamon, Yutaka Matsuo, Mitsuru Ishizuka
AAAI2
2008 WWW sits the SAT: Measuring Relational Similarity on the Web
abstract
Measuring relational similarity between words is important in numerous natural language processing tasks such as solving analogy questions and classifying noun-modifier relations. We propose a method to measure the similarity between semantic relations that hold between two pairs of words using a web search engine. First, each pair of words is represented by a vector of automatically extracted lexical patterns. Then a Support Vector Machine is trained to recognize word pairs with similar semantic relations. We evaluate the proposed method on SAT multiple-choice word-analogy questions. The proposed method achieves a score of 40% which is comparable with relational similarity measures which use manually created resources such as WordNet. The proposed method significantly reduces the time taken by previously proposed computationally intensive methods, such as latent relational analysis, to process 374 analogy questions from 8 days to less than 6 hours.
Danushka Bollegala, Yutaka Matsuo, Mitsuru Ishizuka
ECAI2
2008 Extracting Topics and Innovators Using Topic Diffusion Process in Weblogs
Tadanobu Furukawa, Yutaka Matsuo, Ikki Ohmukai, Koki Uchiyama, Mitsuru Ishizuka
ICWSM2
2008 A Co-occurrence Graph-based Approach for Personal Name Alias Extraction from Anchor Texts
Danushka Bollegala, Yutaka Matsuo, Mitsuru Ishizuka
IJCNLP2
2008 Mining Scholarly Semantic Networks from the Web
abstract
With the increased usage of the Web and its availability of data, various scholarly information is now available on the Web. Extraction, aggregation, and visualization of such information is crucial for using our collective scholarly knowledge and expertise. Semantic network technology is prominent in structuring such knowledge. Our research project is designed to construct a scholarly semantic network from publicly available data that will be useful for researchers, business people in industry, and public servants in governmental organizations who are expected to make decisions appropriately. In this paper, we propose a practical architecture to construct a scholarly semantic network that integrates different scholarly entities such as researchers, papers, and keywords by taking into consideration ontological and web mining perspectives. We also provide an overview of our social network extraction system, which is used as an underpinning to build the architecture. The system, called POLYPHONET, employs several advanced web mining and semantic technologies to extract relations of researchers, to detect groups of researchers, and to obtain keywords for a researcher. The public installation of the system at several academic conferences provides evidence of the system's usability and potential to facilitate the discovery of scholarly knowledge.
Mizuki Oka, Yutaka Matsuo
IV2
2008 Relation Classification for Semantic Structure Annotation of Text
abstract
Confronting the challenges of annotating naturally occurring text into a semantically structured form to facilitate automatic information extraction, current semantic role labeling (SRL) systems have been specifically examining a semantic predicate-argument structure. Based on the concept description language for natural language (CDL.nl) which is intended to describe the concept structure of text using a set of pre-defined semantic relations, we develop a parser to add a new layer of semantic annotation of natural language sentences as an extension of SRL. With the assumption that all relation instances are detected, we present a relation classification approach facing the challenges of CDL.nl relation extraction. Preliminary evaluation on a manual dataset, using support vector machine, shows that CDL.nl relations can be classified with good performance.
Yulan Yan, Yutaka Matsuo, Mitsuru Ishizuka, Toshio Yokoi
Web Intelligence2
2008 Mining for personal name aliases on the web
abstract
We propose a novel approach to find aliases of a given name from the web. We exploit a set of known names and their aliases as training data and extract lexical patterns that convey information related to aliases of names from text snippets returned by a web search engine. The patterns are then used to find candidate aliases of a given name. We use anchor texts and hyperlinks to design a word co-occurrence model and define numerous ranking scores to evaluate the association between a name and its candidate aliases. The proposed method outperforms numerous baselines and previous work on alias extraction on a dataset of personal names, achieving a statistically significant mean reciprocal rank of 0.6718. Moreover, the aliases extracted using the proposed method improve recall by 20% in a relation-detection task.
Danushka Bollegala, Taiki Honma, Yutaka Matsuo, Mitsuru Ishizuka
WWW3
2007 Analyzing Reading Behavior by Blog Mining
Tadanobu Furukawa, Mitsuru Ishizuka, Yutaka Matsuo, Ikki Ohmukai, Koki Uchiyama
AAAI3
2007 Robust Estimation of Google Counts for Social Network Extraction
Yutaka Matsuo, Hironori Tomobe, Takuichi Nishimura
AAAI1
2007 Relation Extraction from Wikipedia Using Subtree Mining
Dat P. T. Nguyen, Yutaka Matsuo, Mitsuru Ishizuka
AAAI2
2007 A Self-Impact Analysis by Artificial Market Simulation
abstract
We constructed an evaluation system of the self-impact in a financial market using an artificial market and text-mining technology. Economic trends were first extracted from text data circulating in the real world. Then, the trends were inputted into the market simulation. Our simulation revealed that an operation by intervention could reduce over 70% of rate fluctuation in 1995. By the simulation results, the system was able to help for its user to find the exchange policy which can stabilize the yen-dollar rate
Kiyoshi Izumi, Hiroki Matsui, Yutaka Matsuo
CIDM3
2007 Extracting Social Networks Among Various Entities on the Web
Yingzi Jin, Yutaka Matsuo, Mitsuru Ishizuka
ESWC2
2007 Social Networks and Reading Bahavior in Blogosphere
Tadanobu Furukawa, Yutaka Matsuo, Ikki Ohmukai, Koki Uchiyama, Mitsuru Ishizuka
ICWSM2
2007 Diffusion of Recommendation through a Trust Network
Yutaka Matsuo, Hikaru Yamamoto
ICWSM1
2007 Inferring Long-term User Properties Based on Users' Location History
Yutaka Matsuo, Naoaki Okazaki, Kiyoshi Izumi, Yoshiyuki Nakamura, Takuichi Nishimura, Kôiti Hasida, Hideyuki Nakashima
IJCAI1
2007 Extracting Keyphrases to Represent Relations in Social Networks from Web
Junichiro Mori, Mitsuru Ishizuka, Yutaka Matsuo
IJCAI3
2007 An Integrated Approach to Measuring Semantic Similarity between Words Using Information Available on the Web
Danushka Bollegala, Yutaka Matsuo, Mitsuru Ishizuka
HLT-NAACL2
2007 Generating Social Network Features for Link-Based Classification
Jun Karamon, Yutaka Matsuo, Hikaru Yamamoto, Mitsuru Ishizuka
PKDD2
2007 An Augmented Tagging Scheme with Triple Tagging and Collective Filtering
abstract
Collaborative tagging is increasingly drawing attentions. However the keyword based tagging scheme has its limitations and it can be observed that tagging society are seeking and using new tagging patterns. This paper proposes a subject-predicate-object scheme for users to triple tag web resources. We first introduce the triple tag model and discuss its relations with existing tag schema and RDF model. Then a filter-based framework that supports the query of triple tags is proposed. The implementation and a case study in the comparison shopping domain are exhibited.
Jie Yang 0006, Yutaka Matsuo, Mitsuru Ishizuka
Web Intelligence2
2007 Measuring semantic similarity between words using web search engines
abstract
Article Share on Measuring semantic similarity between words using web search enginesWWW '07: Proceedings of the 16th international conference on World Wide WebMay 2007 Pages 757–766https://doi.org/10.1145/1242572.1242675Online:08 May 2007Publication History 123citation3,949DownloadsMetricsTotal Citations123Total Downloads3,949Last 12 Months94Last 6 weeks13 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Danushka Bollegala, Yutaka Matsuo, Mitsuru Ishizuka
WWW2
2007 Integration of artificial market simulation and text mining for market analysis
Kiyoshi Izumi, Hiroki Matsui, Yutaka Matsuo
Soft Comput.3
2007 POLYPHONET: An advanced social network extraction system from the Web
Yutaka Matsuo, Junichiro Mori, Masahiro Hamasaki, Takuichi Nishimura, Hideaki Takeda 0001, Kôiti Hasida, Mitsuru Ishizuka
J. Web Semant.1
2006 Spinning Multiple Social Networks for Semantic Web
Yutaka Matsuo, Masahiro Hamasaki, Yoshiyuki Nakamura, Takuichi Nishimura, Kôiti Hasida, Hideaki Takeda 0001, Junichiro Mori, Danushka Bollegala, Mitsuru Ishizuka
AAAI1
2006 Extracting Key Phrases to Disambiguate Personal Names on the Web
Danushka Bollegala, Yutaka Matsuo, Mitsuru Ishizuka
CICLing2
2006 Disambiguating Personal Names on the Web Using Automatically Extracted Key Phrases
Danushka Bollegala, Yutaka Matsuo, Mitsuru Ishizuka
ECAI2
2006 Graph-based Word Clustering using a Web Search Engine
Yutaka Matsuo, Takeshi Sakaki, Koki Uchiyama, Mitsuru Ishizuka
EMNLP1
2006 Doing Community: Co-construction of Meaning and Use with Interactive Information Kiosks
Tom Hope, Masahiro Hamasaki, Yutaka Matsuo, Yoshiyuki Nakamura, Noriyuki Fujimura, Takuichi Nishimura
UbiComp3
2006 Extracting Relations in Social Networks from the Web Using Similarity Between Collective Contexts
Junichiro Mori, Takumi Tsujishita, Yutaka Matsuo, Mitsuru Ishizuka
ISWC3
2006 Context-Aware Weblog to Enhance Communication among Participants in a Conference
Kosuke Numa, Hideaki Takeda 0001, Takuichi Nishimura, Yutaka Matsuo, Masahiro Hamasaki, Noriyuki Fujimura, Keisuke Ishida, Tom Hope, Yoshiyuki Nakamura, Satoshi Fujiyoshi, Kazuya Sakamoto, Hiroshi Nagata, Osamu Nakagawa, Eiji Shinbori
WEBIST (1)4
2006 POLYPHONET: an advanced social network extraction system from the web
abstract
Social networks play important roles in the Semantic Web: knowledge management, information retrieval, ubiquitous computing, and so on. We propose a social network extraction system called POLYPHONET, which employs several advanced techniques to extract relations of persons, detect groups of persons, and obtain keywords for a person. Search engines, especially Google, are used to measure co-occurrence of information and obtain Web documents.Several studies have used search engines to extract social networks from the Web, but our research advances the following points: First, we reduce the related methods into simple pseudocodes using Google so that we can build up integrated systems. Second, we develop several new algorithms for social networking mining such as those to classify relations into categories, to make extraction scalable, and to obtain and utilize person-to-word relations. Third, every module is implemented in POLYPHONET, which has been used at four academic conferences, each with more than 500 participants. We overview that system. Finally, a novel architecture called Super Social Network Mining is proposed; it utilizes simple modules using Google and is characterized by scalability and Relate-Identify processes: Identification of each entity and extraction of relations are repeated to obtain a more precise social network.
Yutaka Matsuo, Junichiro Mori, Masahiro Hamasaki, Keisuke Ishida, Takuichi Nishimura, Hideaki Takeda 0001, Kôiti Hasida, Mitsuru Ishizuka
WWW1
2005 Real-world oriented information sharing using social networks
abstract
While users disseminate various information in the open and widely distributed environment of the Semantic Web, determination of who shares access to particular information is at the center of looming privacy concerns. We propose a real-world-oriented information sharing system that uses social networks. The system automatically obtains users' social relationships by mining various external sources. It also enables users to analyze their social networks to provide awareness of the information dissemination process. Users can determine who has access to particular information based on the social relationships and network analysis.
Junichiro Mori, Tatsuhiko Sugiyama, Yutaka Matsuo
GROUP3
2005 Improving chronological ordering of sentences extracted from multiple newspaper articles
abstract
It is necessary to determine a proper arrangement of extracted sentences to generate a well-organized summary from multiple documents. This paper describes our Multi-Document Summarization (MDS) system for TSC-3. It specifically addresses an approach to coherent sentence ordering for MDS. An impediment to the use of chronological ordering, which is widely used by conventional summarization system, is that it arranges sentences without considering the presupposed information of each sentence. We propose a method to improve chronological ordering by resolving precedent information of arranging sentences. Combining the refinement algorithm with topical segmentation and chronological ordering, we address our experiments and metrics to test the effectiveness of MDS tasks. Results demonstrate that the proposed method significantly improves chronological sentence ordering. At the end of the paper, we also report an outline/evaluation of important sentence extraction and redundant clause elimination integrated in our MDS system.
Naoaki Okazaki, Yutaka Matsuo, Mitsuru Ishizuka
ACM Trans. Asian Lang. Inf. Process.2
2004 Improving Chronological Sentence Ordering by Precedence Relation
Naoaki Okazaki, Yutaka Matsuo, Mitsuru Ishizuka
COLING2
2004 Finding Social Network for Trust Calculation
Yutaka Matsuo, Hironori Tomobe, Kôiti Hasida, Mitsuru Ishizuka
ECAI1
2004 Smart navigation by lightweight components coordination in ubiquitous computing
abstract
We discuss a smart navigation system by route planning which makes active use of spatial semantics. The central concept of the system is coordination of lightweight components as services, such as business process modeling. Our proposed navigation system consists of three main components: a route planner, an object repository, and a space functionality retriever (SFR). The route planner is the component for inquisitions. The object repository deals with real-world information. Then, the SFR deals with knowledge about space represented by XML-based descriptions. Using these components, we realize a smart navigation system that enables a route planner to suggest alternate solutions to end-users. Our final target is to develop advanced location-based information services that surpass conventional ones, thereby realizing an intelligent space.
Shigeyoshi Hiratsuka, Yutaka Matsuo, Akio Sashima, Akira Takagi, Noriaki Izumi, Koichi Kurumatani
IROS2
2004 Spatial Function Representation and Retrieval
Yutaka Matsuo, Akira Takagi, Shigeyoshi Hiratsuka, Kôiti Hasida, Hideyuki Nakashima
PRICAI1
2004 Coherent Arrangement of Sentences Extracted from Multiple Newspaper Articles
Naoaki Okazaki, Yutaka Matsuo, Mitsuru Ishizuka
PRICAI2
2003 Mining Social Network of Conference Participants from the Web
abstract
In a ubiquitous computing environment, it is desirable to provide a user with information depending on a user's situation, such as time, location, user behavior, and social context. At conventions, such as academic conferences and exhibitions, where participants must register in advance, the social context of participants can be extracted from the Web using their names and affiliations without asking the participants many questions. Here, we attempt to extract the social network of participants from the Web, where a node represents a participant and an edge represents the relationship of two participants. Each edge is added using the number of pages retrieved by a search engine which include both participants names. Moreover, each edge has a label such as "coauthors" and "members of the same project" by applying classification rules to the page content. We show an example of the extracted network and make a preliminary evaluation. This network can be used in many information services, such as finding an appropriate introducer or negotiator, and who one should talk to in order to efficiently expand his/her network.
Yutaka Matsuo, Hironori Tomobe, Kôiti Hasida, Mitsuru Ishizuka
Web Intelligence1
2003 Average-Clicks: A New Measure of Distance on the World Wide Web
Yutaka Matsuo, Yukio Ohsawa, Mitsuru Ishizuka
J. Intell. Inf. Syst.1
2002 Two Transformations of Clauses into Constraints and Their Properties for Cost-Based Hypothetical Reasoning
Yutaka Matsuo, Mitsuru Ishizuka
PRICAI1
2002 Browsing support by highlighting keywords based on a user's browsing history
abstract
We develop a browsing support system which learns user's interests and highlights keywords based on a user's browsing history. Monitoring the user's access to the Web enable us to detect "familiar words" for the user. We extract keywords, which are relevant to the familiar words in the current page, and highlight them. The relevancy is measured by the biases of co-occurrence, called IRM (Interest Relevance Measure). Our system consists of three components; a proxy server which monitors access to the Web, a frequency server which stores frequency of words in the accessed Web pages, and a keyword extraction module. Preliminary reports are shown to evaluate the system.
Yutaka Matsuo, Hayato Fukuta, Mitsuru Ishizuka
SMC1
2002 Featuring web communities based on word co-occurrence structure of communications: 736
abstract
Textual communication in message boards is analyzed for classifying Web communities. We present a communication-content based generalization of an existing business-oriented classification of Web communities, using KeyGraph, a method for visualizing the co-occurrence relations between words and word clusters in text. Here, the text in a message board is analyzed with KeyGraph, and the structure obtained is shown to reflect the essence of the content-flow. The relation of this content-flow with participants' interests is then formalized. Three structure-features of relations between participants and words, determining the type of the community, are shown to be computed and visualized: (1) centralization (2) context coherence and (3) creative decisions. This helps in surveying the essence of a community, e.g. whether the community creates useful knowledge, how easy it is to join the community, and whether/why the community is good for making commercial advertisement.
Yukio Ohsawa, Hirotaka Soma, Yutaka Matsuo, Naohiro Matsumura, Masaki Usui
WWW3
2002 SL method for computing a near-optimal solution using linear and non-linear programming in cost-based hypothetical reasoning
Mitsuru Ishizuka, Yutaka Matsuo
Knowl. Based Syst.2
2001 KeyWorld: Extracting Keywords from a Document as a Small World
Yutaka Matsuo, Yukio Ohsawa, Mitsuru Ishizuka
Discovery Science1
2001 Average-Clicks: A New Measure of Distance on the World Wide Web
Yutaka Matsuo, Yukio Ohsawa, Mitsuru Ishizuka
Web Intelligence1
2000 Fast Hypothetical Reasoning by Parallel Processing
Yutaka Matsuo, Mitsuru Ishizuka
PRICAI1
1998 SL Method for Computing a Near-Optimal Solution Using Linear and Non-linear Programming in Cost-Based Hypothetical Reasoning
Mitsuru Ishizuka, Yutaka Matsuo
PRICAI2