EDBT 2026 Demo / reviewers in the wild / expert
Kotaro Kikuchi
dblp:191/1049
· DBLP profile ↗
12ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0003-1747-5945ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 6 · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Practical Evaluation of UI Element Detection for Automated Mobile Game TestingabstractDeveloping attractive mobile games requires extensive testing to ensure the quality and reliability of their rich graphics and user interface (UI). Frequent UI changes during development necessitate vision-based testing; however, detecting UI elements in mobile games has not been fully explored compared to web pages and mobile apps. In this study, we benchmark recent UI detection methods using our newly constructed dataset, which comprises three different types of mobile games. We compare the methods with several evaluation metrics, including three proposed metrics designed for automated testing. The experimental results show that existing methods still face difficulty in accurately detecting UI elements, yet they could be practically used for specific purposes, such as monkey testing. Nozomu Karai, Koya Ihara, Tatsuya Ishimoto, Daiki Kubo, Kotaro Kikuchi |
CoG | 5 |
| 2025 | ColorGPT: Leveraging Large Language Models for Multimodal Color Recommendation
Ding Xia, Naoto Inoue, Qianru Qiu, Kotaro Kikuchi |
ICDAR (5) | 4 |
| 2025 | Multimodal Markup Document Models for Graphic Design CompletionabstractWe introduce MarkupDM, a multimodal markup document model that represents graphic design as an interleaved multimodal document consisting of both markup language and images. Unlike existing holistic approaches that rely on an element-by-attribute grid representation, our representation accommodates variable-length elements, type-dependent attributes, and text content. Inspired by fill-in-the-middle training in code generation, we train the model to complete the missing part of a design document from its surrounding context, allowing it to treat various design tasks in a unified manner. Our model also supports image generation by predicting discrete image tokens through a specialized tokenizer with support for image transparency. We evaluate MarkupDM on three tasks, attribute value, image, and text completion, and demonstrate that it can produce plausible designs consistent with the given context. To further illustrate the flexibility of our approach, we evaluate our approach on a new instruction-guided design completion task where our instruction-tuned MarkupDM compares favorably to state-of-the-art image editing models, especially in textual completion. These findings suggest that multimodal language models with our document representation can serve as a versatile foundation for broad design automation. Kotaro Kikuchi, Ukyo Honda, Naoto Inoue, Mayu Otani, Edgar Simo-Serra, Kota Yamaguchi |
ACM Multimedia | 1 |
| 2024 | Retrieval-Augmented Layout Transformer for Content-Aware Layout GenerationabstractContent-aware graphic layout generation aims to automatically arrange visual elements along with a given content, such as an e-commerce product image. In this paper, we argue that the current layout generation approaches suffer from the limited training data for the high-dimensional layout structure. We show that a simple retrieval augmentation can significantly improve the generation quality. Our model, which is named Retrieval-Augmented Layout Transformer (RALF),retrieves nearest neighbor layout examples based on an input image and feeds these results into an autoregressive generator. Our model can apply retrieval augmentation to various controllable generation tasks and yield high-quality layouts within a unified architecture. Our extensive experiments show that RALF successfully generates content-aware layouts in both constrained and unconstrained settings and significantly outperforms the baselines.1 Daichi Horita, Naoto Inoue, Kotaro Kikuchi, Kota Yamaguchi, Kiyoharu Aizawa |
CVPR | 3 |
| 2024 | Fast Sprite Decomposition from Animated Graphics
Kotaro Kikuchi, Kota Yamaguchi |
ECCV (67) | 2 |
| 2023 | LayoutDM: Discrete Diffusion Model for Controllable Layout GenerationabstractControllable layout generation aims at synthesizing plausible arrangement of element bounding boxes with optional constraints, such as type or position of a specific element. In this work, we try to solve a broad range of layout generation tasks in a single model that is based on discrete state-space diffusion models. Our model, named Lay-outDM, naturally handles the structured layout data in the discrete representation and learns to progressively infer a noiseless layout from the initial input, where we model the layout corruption process by modality-wise discrete diffusion. For conditional generation, we propose to inject layout constraints in the form of masking or logit adjustment during inference. We show in the experiments that our Lay-outDM successfully generates high-quality layouts and outperforms both task-specific and task-agnostic baselines on several layout tasks.11Please find the code and models at: https://cyberagentailab.github.io/layout-drn. Naoto Inoue, Kotaro Kikuchi, Edgar Simo-Serra, Mayu Otani, Kota Yamaguchi |
CVPR | 2 |
| 2023 | Towards Flexible Multi-modal Document ModelsabstractCreative workflows for generating graphical documents involve complex inter-related tasks, such as aligning elements, choosing appropriate fonts, or employing aesthetically harmonious colors. In this work, we attempt at building a holistic model that can jointly solve many different design tasks. Our model, which we denote by FlexDM, treats vector graphic documents as a set of multi-modal elements, and learns to predict masked fields such as element type, position, styling attributes, image, or text, using a unified architecture. Through the use of explicit multitask learning and in-domain pre-training, our model can better capture the multi-modal relationships among the different document fields. Experimental results corroborate that our single FlexDM is able to successfully solve a multitude of different design tasks, while achieving performance that is competitive with task-specific and costly baselines.11Please find the code and models at: https://cyberagentailab.github.io/flex-dm Naoto Inoue, Kotaro Kikuchi, Edgar Simo-Serra, Mayu Otani, Kota Yamaguchi |
CVPR | 2 |
| 2023 | Generative Colorization of Structured Mobile Web PagesabstractColor is a critical design factor for web pages, affecting important factors such as viewer emotions and the overall trust and satisfaction of a website. Effective coloring requires design knowledge and expertise, but if this process could be automated through data-driven modeling, efficient exploration and alternative workflows would be possible. However, this direction remains underexplored due to the lack of a formalization of the web page colorization problem, datasets, and evaluation protocols. In this work, we propose a new dataset consisting of e-commerce mobile web pages in a tractable format, which are created by simplifying the pages and extracting canonical color styles with a common web browser. The web page colorization problem is then formalized as a task of estimating plausible color styles for a given web page content with a given hierarchical structure of the elements. We present several Transformer-based methods that are adapted to this task by prepending structural message passing to capture hierarchical relation-ships between elements. Experimental results, including a quantitative evaluation designed for this task, demonstrate the advantages of our methods over statistical and image colorization methods. The code is available at https://github.com/CyberAgentAILab/webcolor. Kotaro Kikuchi, Naoto Inoue, Mayu Otani, Edgar Simo-Serra, Kota Yamaguchi |
WACV | 1 |
| 2021 | Constrained Graphic Layout Generation via Latent OptimizationabstractIt is common in graphic design humans visually arrange various elements according to their design intent and semantics. For example, a title text almost always appears on top of other elements in a document. In this work, we generate graphic layouts that can flexibly incorporate such design semantics, either specified implicitly or explicitly by a user. We optimize using the latent space of an off-the-shelf layout generation model, allowing our approach to be complementary to and used with existing layout generation models. Our approach builds on a generative layout model based on a Transformer architecture, and formulates the layout generation as a constrained optimization problem where design constraints are used for element alignment, overlap avoidance, or any other user-specified relationship. We show in the experiments that our approach is capable of generating realistic layouts in both constrained and unconstrained generation tasks with a single model. The code is available at https://github.com/ktrk115/const_layout. Kotaro Kikuchi, Edgar Simo-Serra, Mayu Otani, Kota Yamaguchi |
ACM Multimedia | 1 |
| 2021 | Modeling Visual Containment for Web Page Layout OptimizationabstractAbstract Web pages have become fundamental in conveying information for companies and individuals, yet designing web page layouts remains a challenging task for inexperienced individuals despite web builders and templates. Visual containment, in which elements are grouped together and placed inside container elements, is an efficient design strategy for organizing elements in a limited display, and is widely implemented in most web page designs. Yet, visual containment has not been explicitly addressed in the research on generating layouts from scratch, which may be due to the lack of hierarchical structure. In this work, we represent such visual containment as a layout tree, and formulate the layout design task as a hierarchical optimization problem. We first estimate the layout tree from a given a set of elements, which is then used to compute tree‐aware energies corresponding to various desirable design properties such as alignment or spacing. Using an optimization approach also allows our method to naturally incorporate user intentions and create an interactive web design application. We obtain a dataset of diverse and popular real‐world web designs to optimize and evaluate various aspects of our method. Experimental results show that our method generates better quality layouts compared to the baseline method. Kotaro Kikuchi, Mayu Otani, Kota Yamaguchi, Edgar Simo-Serra |
Comput. Graph. Forum | 1 |
| 2018 | Fine-grained Video Retrieval using Query Phrases - Waseda_Meisei TRECVID 2017 AVS System -abstractIn this paper, a joint team from Waseda University and Meisei University (team name: Waseda_Meisei) report their efforts on the ad-hoc video search (AVS) task for the TRECVID benchmark, which is conducted annually by the National Institute of Standards and Technology (NIST). For the AVS task, a system is required to perform a fine-grained search of target videos from a large-scale video database using a query phrase including multiple keywords, such as objects, persons, scenes, and actions. The system we submitted has the following two characteristics. First, to improve the coverage rate of classes corresponding to keywords in query phrases, we prepared a large number of classifiers that can detect objects, persons, scenes, and actions, which were trained using various image and video datasets. Second, when choosing a concept classifier corresponding to a keyword, we introduced a mechanism that allows us to select additional concept classifiers by incorporating natural language processing techniques. We submitted multiple systems with these characteristics to the TRECVID 2017 AVS task and one of our systems ranked the highest among all the submitted systems from 22 teams. Kazuya Ueki, Koji Hirakawa, Kotaro Kikuchi, Tetsunori Kobayashi |
ICPR | 3 |
| 2018 | Analyzing Human Avoidance Behavior in Narrow PassageabstractTo ensure that humans and robots can safely coexist, the ability to recognize human behavior is a prerequisite for robots and a fundamental technical challenge for researchers. Current research can only recognize relatively simple cases of human behavior due to the lack of enough data and archetypally designed experiments. Our study elucidates human behavior in a systematic manner by observing the behavior of human subjects under more complex situations where they are surrounded by other people or objects to address the challenge. We focus on the following common situation that people pass each other through a narrow passage. We constructed a motion capture room with a narrow passage environment and measured the motion of human subjects performing different tasks. In addition, during the narrow passage experiment, we made subjects hold different daily necessaries (such as a backpack) to observe influences on human behavior. Our study found that passing and avoidance behavior exhibited by each of our subjects were significantly influenced by what kind of daily necessaries subjects carry. This research provides novel findings on human behavior in complex environments: in the case where subjects holding a handbag (a type of daily necessaries that stays next to one's body), they showed the tendency to be affected by the other subject and move more dynamically compared to the subject without anything or with other daily necessaries; in the case where subjects carrying a backpack (a type of daily necessaries on one's back) and looking at a smartphone, they also showed the tendency of being affected by others, but their motion is restricted compared to the subject without anything. Takayuki Nakatsuka, Tamon Miyake, Kotaro Kikuchi, Ayano Kobayashi, Yoshihiko Hayashi |
SMC | 3 |