EDBT 2026 Demo / reviewers in the wild / expert
Osamu Yoshie
dblp:14/5985
· DBLP profile ↗
49ranked-venue papers
0as first author
31since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 13 since 2021Databases, data management, data science and information retrieval · 11 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 11Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | E-ViC: Reasoning Beyond Text via Embodied Visual Chain for Spatial IntelligenceabstractJunbo Qi, Yi Zhang, Hanchu Ni, Che Liu, Zhimin Yao, Ruilin Yang, Xiancong Ren, Liangjian Wen, Wei Ge, Yuya Ieiri, Osamu Yoshie, Yong Dai, Xiaozhu Ju. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Junbo Qi, Yi Zhang 0001, Hanchu Ni, Che Liu 0002, Zhimin Yao, Ruilin Yang, Xiancong Ren, Liangjian Wen, Yuya Ieiri, Osamu Yoshie, Yong Dai 0001, Xiaozhu Ju |
ACL (1) | 11 |
| 2026 | DM3Net: Dual-Camera Super-Resolution via Domain Modulation and Multi-scale MatchingabstractDual-camera super-resolution is highly practical for smartphone photography that primarily super-resolve the wide-angle images using the telephoto image as a reference. In this paper, we propose DM3Net, a novel dual-camera super-resolution network based on Domain Modulation and Multi-scale Matching. To bridge the domain gap between the high-resolution domain and the degraded domain, we learn two compressed global representations from image pairs corresponding to the two domains. To enable reliable transfer of high-frequency structural details from the reference image, we design a multi-scale matching module that conducts patch-level feature matching and retrieval across multiple receptive fields to improve matching accuracy and robustness. Moreover, we also introduce Key Pruning to achieve a significant reduction in memory usage and inference time with little model performance sacrificed. Experimental results on three real-world datasets demonstrate that our DM3Net outperforms the state-of-the-art approaches. Cong Guan, Jiacheng Ying, Yuya Ieiri, Osamu Yoshie |
WACV | 4 |
| 2026 | Chart specification: Structural representations for incentivizing VLM reasoning in chart-to-code generationabstractVision-Language Models (VLMs) have shown promise in generating plotting code from chart images, yet achieving structural fidelity remains challenging. Existing approaches largely rely on supervised fine-tuning, encouraging surface-level token imitation rather than faithful modeling of chart structure, which often leads to hallucinated or semantically inconsistent outputs. We propose Chart Specification, a canonical structural representation that shifts training from mimicking training-code patterns to structure-grounded learning. By normalizing plotting code into structure-equivalent specifications, it enables (i) the construction of a structurally balanced training set, and (ii) a Spec-Align Reward that provides fine-grained, verifiable feedback on structural correctness for reinforcement learning. Under an explicit reasoning-to-code generation paradigm, this reward encourages structure-aware reasoning that produces constraint-consistent plotting code. Experiments on three public benchmarks show that our method consistently outperforms prior approaches. With only 3K training samples, we achieve strong data efficiency, surpassing leading baselines by up to 61.7% on complex benchmarks, and scaling to 4K samples establishes new state-of-the-art results across all evaluated metrics. Overall, our results demonstrate that precise structural supervision offers an efficient pathway to high-fidelity chart-to-code generation. Code and dataset are available at: https://github.com/Mighten/chart-specification-paper . Minggui He, Mingchen Dai, Yilun Liu 0001, Shimin Tao, Pufan Zeng, Osamu Yoshie, Yuya Ieiri |
Neurocomputing | 7 |
| 2026 | Active perception: Gaze-guided thinking for chart understandingabstractAnswering questions about charts presents a unique challenge for Vision-Language Models (VLMs). Unlike natural images, charts are structured artifacts governed by explicit visual grammar that demands pixel-level accuracy in visual perception. While recent VLMs demonstrate impressive reasoning abilities on chart tasks, a critical gap remains: their reasoning operates abstractly, disconnected from precise visual grounding. We introduce Active Perception, a framework that enables Gaze-Guided Thinking, a reasoning pattern that explicitly anchors abstract inference to concrete visual locations through coordinate-based operations (Locate, Trace, Extract, Compare). To instill this capability, we propose Skill Cultivation, a two-stage training strategy: Stage I injects coordinate-aware primitives via Supervised Fine-Tuning on ChartQAGaze-14K, our synthesized dataset of 14K coordinate-annotated reasoning chains; Stage II internalizes these skills into adaptive strategies via Reinforcement Learning with outcome-based rewards. Building upon Qwen2.5-VL-7B, Active Perception achieves state-of-the-art performance on ChartQA, improving overall accuracy from 78.96% to 82.44%, with particularly notable gains on the challenging Human split (75.76% to 81.28%). Qualitative analysis reveals emergent systematic chart-reading behaviors that mirror human visual strategies, demonstrating the effectiveness of spatially grounded reasoning for structured visual understanding. • We identify a critical gap in current VLMs for chart understanding: reasoning operates abstractly without precise visual grounding, limiting accurate data extraction from structured visualizations. • We propose Gaze-Guided Thinking , a reasoning pattern that anchors abstract inference to concrete visual locations through Coordinate Primitives ( Locate , Trace , Extract , Compare ), mimicking human chart scanning behavior. • We introduce Skill Cultivation , a two-stage training strategy combining SFT on ChartQAGaze-14K (14K coordinate-annotated reasoning chains) with outcome-based RL to inject and internalize spatially-grounded reasoning. • Active Perception achieves 82.44% overall accuracy on ChartQA (a 3.48% absolute improvement), with particularly strong performance on the Human split (81.28%, +5.52%), demonstrating that explicit visual grounding substantially enhances structured visual understanding. • Qualitative analysis reveals emergent human-like chart reading behaviors, where models systematically leverage coordinates for precise value extraction and spatial reasoning. Xin Huang 0027, Hongbing Li, Zejia Weng, Jia Wang 0025, Yeqing Shen, Haolong Yan, Kaijun Tan, Zheng Ge, Xiangyu Zhang 0005, Daxin Jiang, Osamu Yoshie |
Neurocomputing | 14 |
| 2026 | CLIP-driven rain perception: Adaptive deraining with pattern-aware network routing and mask-guided cross-attentionabstractExisting deraining models process all rainy images within a single network. However, different rain patterns have significant variations, which makes it challenging for a single network to handle diverse types of raindrops and streaks. To address this limitation, we propose a novel CLIP-driven rain perception network (CLIP-RPN) that leverages CLIP to automatically perceive rain patterns by computing visual-language matching scores and adaptively routing to sub-networks to handle different rain patterns, such as varying raindrop densities, streak orientations, and rainfall intensity. CLIP-RPN establishes semantic-aware rain pattern recognition through CLIP’s cross-modal visual-language alignment capabilities, enabling automatic identification of precipitation characteristics across different rain scenarios. This rain pattern awareness drives an adaptive subnetwork routing mechanism where specialized processing branches are dynamically activated based on the detected rain type, significantly enhancing the model’s capacity to handle diverse rainfall conditions. Furthermore, within sub-networks of CLIP-RPN, we introduce a mask-guided cross-attention mechanism (MGCA) that predicts precise rain masks at multi-scale to facilitate contextual interactions between rainy regions and clean background areas by cross-attention. We also introduces a dynamic loss scheduling mechanism (DLS) to adaptively adjust the gradients for the optimization process of CLIP-RPN. Compared with the commonly used l 1 or l 2 loss, DLS is more compatible with the inherent dynamics of the network training process, thus achieving enhanced outcomes. Our method achieves state-of-the-art performance across multiple datasets, particularly excelling in complex mixed datasets. Cong Guan, Osamu Yoshie |
Pattern Recognit. | 2 |
| 2026 | FluencyVE: Marrying Temporal-Aware Mamba with Bypass Attention for Video EditingabstractLarge-scale text-to-image diffusion models have achieved unprecedented success in image generation and editing. However, extending this success to video editing remains challenging. Recent video editing efforts have adapted pretrained text-to-image models by adding temporal attention mechanisms to handle video tasks. Unfortunately, these methods continue to suffer from temporal inconsistency issues and high computational overheads. In this study, we propose FluencyVE, which is a simple yet effective one-shot video editing approach. FluencyVE integrates the linear time-series module, Mamba, into a video editing model based on pretrained Stable Diffusion models, replacing the temporal attention layer. This enables global frame-level attention while reducing the computational costs. In addition, we employ low-rank approximation matrices to replace the query and key weight matrices in the causal attention, and use a weighted averaging technique during training to update the attention scores. This approach significantly preserves the generative power of the text-to-image model while effectively reducing the computational burden. Experiments and analyses demonstrate promising results in editing various attributes, subjects, and locations in real-world videos. Mingshu Cai, Osamu Yoshie, Yuya Ieiri |
IEEE Trans. Multim. | 3 |
| 2025 | Taming Text-to-Image Synthesis for Novices: User-centric Prompt Generation via Multi-turn GuidanceabstractYilun Liu, Minggui He, Feiyu Yao, Yuhe Ji, Shimin Tao, Jingzhou Du, Justin Li, Jian Gao, Zhang Li, Hao Yang, Boxing Chen, Osamu Yoshie. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yilun Liu 0001, Minggui He, Feiyu Yao, Yuhe Ji, Shimin Tao, Jingzhou Du, Justin Li, Hao Yang 0006, Boxing Chen, Osamu Yoshie |
EMNLP | 12 |
| 2025 | MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model EvaluationabstractWeihao Xuan, Rui Yang, Heli Qi, Qingcheng Zeng, Yunze Xiao, Aosong Feng, Dairui Liu, Yun Xing, Junjue Wang, Fan Gao, Jinghui Lu, Yuang Jiang, Huitao Li, Xin Li, Kunyu Yu, Ruihai Dong, Shangding Gu, Yuekang Li, Xiaofei Xie, Felix Juefei-Xu, Foutse Khomh, Osamu Yoshie, Qingyu Chen, Douglas Teodoro, Nan Liu, Randy Goebel, Lei Ma, Edison Marrese-Taylor, Shijian Lu, Yusuke Iwasawa, Yutaka Matsuo, Irene Li. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Weihao Xuan, Rui Yang 0016, Heli Qi, Qingcheng Zeng, Yunze Xiao, Aosong Feng, Dairui Liu, Yun Xing 0001, Jinghui Lu, Yuang Jiang, Huitao Li, Xin Li 0079, Kunyu Yu, Ruihai Dong, Shangding Gu, Yuekang Li, Xiaofei Xie, Felix Juefei-Xu, Foutse Khomh, Osamu Yoshie, Qingyu Chen 0001, Douglas Teodoro, Nan Liu 0003, Randy Goebel, Lei Ma 0003, Edison Marrese-Taylor, Shijian Lu, Yusuke Iwasawa, Yutaka Matsuo, Irene Li |
EMNLP | 22 |
| 2025 | ARM : nnU-Net with Arena Mechanism for Medical Image SegmentationabstractThe success of nnU-Net proves the significance of the rationality of workflow architecture and configuration settings in improving segmentation accuracy. However, since that, most efforts to improve U-Net have continued to address CNN inner limitations caused by architecture. These methods encountered challenges such as limited generalization, difficulty in managing data with varying distribution patterns. To tackle these issues, We designed a standardized processing workflow specifically tailored for the convolutional layers of U-Net : Arena Mechanism (ARM), inspired by game theory, which encompasses three stages: Recruit, Train and Fight. We augment the convolutional layers of U-Net with two additional processing branches (Recruit), serving as "challengers." These challengers are first optimized within a Hybrid Adaptive Weighting module to refine the feature representation of each internal channel (Train). During forward propagation, we design a cooperative loss and reward function to determine the optimal confidence distribution and fusion strategy for branch outputs (Fight). This standardized mechanism is designed to further nnU-Net’s core principle of "Automatic Adaptation." Experiments have demonstrated that this approach achieves state-of-the-art (SOTA) performance on both the ACDC, BraTS21 and KiTS datasets. Cong Guan, Tengfei Shao, Shenglei Li, Tomoji Kishi, Osamu Yoshie |
ICASSP | 6 |
| 2025 | MetaCert: Metabolic Attention Network Utilizing Uncertainty Estimation for Multimodal Aspect-Category-Sentiment Triple ExtractionabstractMultimodal Aspect-Category-Sentiment Triple Extraction (MACSTE) is a highly complex subtask within Multimodal Aspect-Based Sentiment Analysis (MABSA), requiring simultaneous attribute extraction and sentiment polarity prediction from image-text pairs. While existing research often emphasizes modality fusion and alignment, it frequently neglects the design of information flow pathways, leading to suboptimal utilization of complementary information. Additionally, modality-specific noise may compromise the robustness and accuracy of multimodal classification, with traditional filtering methods often degrading data quality. To overcome these challenges, we propose the Metabolic Attention Network Utilizing Uncertainty Estimation (MetaCert). MetaCert integrates two key components: the Metabolic Attention Mechanism (MAM), inspired by bio-chemical metabolic networks and enhanced by cross-attention for improved information exchange; and the Uncertainty Estimation Network (UEN), which optimizes the semantic contributions of each modality while preserving data integrity, thereby enhancing classification accuracy. Our approach achieves state-of-the-art (SOTA) results on the TWITTER-15 and TWITTER-17 datasets. Cong Guan, Tengfei Shao, Shenglei Li, Tomoji Kishi, Osamu Yoshie |
ICASSP | 6 |
| 2025 | Multi-Attribute guided Thermal Face Image Translation based on Latent Diffusion ModelabstractModern surveillance systems increasingly rely on multi-wavelength sensors and deep neural networks to recognize faces in infrared images captured at night. However, most facial recognition models are trained on visible light datasets, leading to substantial performance degradation on infrared inputs due to significant domain shifts. Early feature-based methods for infrared face recognition proved ineffective, prompting researchers to adopt generative approaches that convert infrared images into visible light images for improved recognition. This paradigm, known as Heterogeneous Face Recognition (HFR), faces challenges such as model and modality discrepancies, leading to distortion and feature loss in generated images. To address these limitations, this paper introduces a novel latent diffusion-based model designed to generate high-quality visible face images from thermal inputs while preserving critical identity features. A multi-attribute classifier is incorporated to extract key facial attributes from visible images, mitigating feature loss during infrared-to-visible image restoration. Additionally, we propose the Self-attn Mamba module, which enhances global modeling of cross-modal features and significantly improves inference speed. Experimental results on two benchmark datasets demonstrate the superiority of our approach, achieving state-of-the-art performance in both image quality and identity preservation. Mingshu Cai, Osamu Yoshie, Yuya Ieiri |
IJCB | 2 |
| 2025 | PADriver: Towards Personalized Autonomous DrivingabstractIn this paper, we propose PADriver, a novel closed-loop framework for personalized autonomous driving (PAD). Built upon Multi-modal Large Language Model (MLLM), PADriver takes streaming frames and personalized textual prompts as inputs. It autoaggressively performs scene understanding, danger level estimation and action decision. The predicted danger level reflects the risk of the potential action and provides an explicit reference for the final action, which corresponds to the preset personalized prompt. Moreover, we construct a closed-loop benchmark named PAD-Highway based on Highway-Env simulator to comprehensively evaluate the decision performance under traffic rules. The dataset contains 250 hours videos with high-quality annotation to facilitate the development of PAD behavior analysis. Experimental results on the constructed benchmark show that PADriver outperforms state-of-the-art approaches on different evaluation metrics, and enables various driving modes. Genghua Kou, Weixin Mao, Yingfei Liu, Osamu Yoshie, Tiancai Wang |
IJCNN | 7 |
| 2025 | GUI Exploration Lab: Enhancing Screen Navigation in Agents via Multi-Turn Reinforcement LearningabstractWith the rapid development of Large Vision Language Models, the focus of Graphical User Interface (GUI) agent tasks shifts from single-screen tasks to complex screen navigation challenges.
However, real-world GUI environments, such as PC software and mobile Apps, are often complex and proprietary, making it difficult to obtain the comprehensive environment information needed for agent training and evaluation. This limitation hinders systematic investigation and benchmarking of agent navigation capabilities.
To address this limitation, we introduce GUI Exploration Lab, a simulation environment engine for GUI agent navigation research that enables flexible definition and composition of screens, icons, and navigation graphs, while providing full access to environment information for comprehensive agent training and evaluation.
Through extensive experiments, we find that supervised fine-tuning enables effective memorization of fundamental knowledge, serving as a crucial foundation for subsequent training. Building on this, single-turn reinforcement learning further enhances generalization to unseen scenarios. Finally, multi-turn reinforcement learning encourages the development of exploration strategies through interactive trial and error, leading to further improvements in screen navigation performance.
We validate our methods on both static and interactive benchmarks, demonstrating that our findings generalize effectively to real-world scenarios.
These findings demonstrate the advantages of reinforcement learning approaches in GUI navigation and offer practical guidance for building more capable and generalizable GUI agents. Haolong Yan, Yeqing Shen, Xin Huang 0027, Jia Wang 0025, Kaijun Tan, Zhixuan Liang, Zheng Ge, Osamu Yoshie, Xiangyu Zhang 0005, Daxin Jiang |
NeurIPS | 9 |
| 2025 | MM-instruct: Generated visual instructions for large multimodal model alignmentabstractThis paper presents MM-Instruct, an automated pipeline for generating diverse and high-quality visual instruction data to better align large multimodal models (LMMs) with real-world use cases. While previous works have focused on question-answering data, their generated instruction datasets pose challenges for broader application scenarios. Additionally, manually collecting diverse instruction data at scale from users is prohibitively costly. To mitigate these issues, MM-Instruct leverages ChatGPT to automatically generate diverse instructions from a limited set of seed instructions through augmentation and summarization. It then uses an open-sourced large language model (LLM) to construct instruction-following answers and builds a large-scale visual instruction dataset. Evaluating the LLaVA-Instruct models trained with the generated data shows significant improvements in instruction-following capabilities compared to LLaVA-1.5 models. MM-Instruct releases its synthetic dataset containing diverse instructions and high-quality instruction-answer pairs to support training LMMs for real-world applications. Xin Huang 0027, Jihao Liu, Jinliang Zheng, Boxiao Liu, Jia Wang 0025, Yu Liu 0015, Hongsheng Li 0001, Osamu Yoshie |
Neurocomputing | 8 |
| 2025 | FastTalker: An unified framework for generating speech and conversational gestures from text
Zixin Guo, Minggui He, Osamu Yoshie |
Neurocomputing | 4 |
| 2025 | Conducting patch contrastive learning with mixture of experts on mixed datasets for medical image segmentationabstractAbstract Medical image segmentation is critical for accurate diagnosis, treatment planning, and surgical navigation. In recent years, large multitask segmentation models have often struggled due to the limited size of datasets and significant variability in target structures, image resolutions, and annotation standards. These variations can introduce task competitions during multitask model training, which hinder effective feature learning. To address these challenges, we propose PatchMoE, a unified framework designed to compensate for resolution discrepancies across datasets and feature conflicts arising in mixed-dataset training. PatchMoE is the first to introduce patch-based contrastive learning into medical image segmentation tasks, which divides images into equal-sized patches represented in 3D coordinate space. This novel approach ensures that mixed datasets with varying resolutions can be trained in a unified manner, preserving spatial relationships and enhancing contextual understanding. PatchMoE also incorporates a mixture of experts (MoE) mechanism into the decoder, which dynamically selects dataset-specific expert combinations. This design mitigates parameter conflicts through network sparsification, effectively resolving optimization conflicts in multitask datasets. The effectiveness of the proposed method was demonstrated in four independent segmentation tasks: retinal vessel (DRIVE), near-infrared blurred vessel (HVNIR), abdominal multiorgan (Synapse), and polyp segmentation (Kvasir-SEG). We compared performance using multiple metrics, including Dice score, Intersection over Union (IoU), and Hausdorff distance (HD). Compared with the state-of-the-art (SOTA) GCASCADE model, PatchMoE achieved an improvement of 3.04% in the mean Dice score across all tasks. The proposed method also achieved an average Dice score improvement of 0.88% compared to four independently trained SOTA models for each individual task. In summary, PatchMoE combines patch-based contrastive learning with dataset-informed expert gating to provide promising solutions for dataset conflicts in large transformer-based medical segmentation models. Jiazhe Wang, Osamu Yoshie, Yuya Ieiri |
Neural Comput. Appl. | 2 |
| 2024 | G-DGANet: Gated deep graph attention network with reinforcement learning for solving traveling salesman problem
Getu Fellek, Ahmed Farid, Shigeru Fujimura 0001, Osamu Yoshie, Goytom Gebreyesus |
Neurocomputing | 4 |
| 2024 | Learning hierarchical discrete prior for co-speech gesture generationabstractIn the context of Co-Speech Gesture Generation, Vector-Quantized Variational Autoencoder (VQ-VAE) based methods have shown promising results by separating the generation process into two stages: learning discrete gesture priors via pretraining for gesture reconstruction, which encodes gesture into a discrete codebook, followed by learning the mapping between speech audio and gesture codebook indices. This design leverages pretraining of motion VQVAE with the motion reconstruction task to improve the quality of generated gestures. However, the vanilla VQVAE’s codebook often fails to encode both low-level and high-level gesture features adequately, resulting in limited reconstruction quality and generation performance. To address this, we propose the Hierarchical Discrete Audio-to-Gesture (HD-A2G), which innovates (i) a two-stage hierarchical codebook structure for capturing high-level and low-level gesture priors, enabling the reconstruction of gesture details. (ii) it further integrates high-level and low-level feature using an AdaIn layer, effectively enhancing the learning of gesture’s rhythm and content. (iii) it explicitly maps text and audio onset features to the appropriate levels of the codebook, ensuring learning accurate hierarchical associations for the generation stage. Experimental results on the BEAT and Trinity datasets demonstrate that HD-A2G outperform the baseline method in both pretrained gesture reconstruction and audio-conditioned gesture generation with a clear margin, achieving the state-of-the-art performance qualitatively and quantitatively. Osamu Yoshie |
Neurocomputing | 2 |
| 2024 | Hierarchical graph fusion network and a new argumentative dataset for multiparty dialogue discourse parsingabstractDiscourse parsing in multi-party dialogue aims to extract the relationships between elementary discourse units (EDUs) such as arguments and utterances, and has numerous applications like chatbots or virtual assistants. Two significant challenges have been encountered in previous works: the obstacles in fusing various contexts in the modeling; and the lack of argumentative dialogue datasets in this field. To tackle context fusion challenges in the modeling, we introduce the Hierarchical Graph Fusion Network (HGFN). This method introduces sufficient contexts by hierarchically modeling the dialogue and minimizes context noise through a novel routing mechanism. It specifically: (1) Encodes multiple levels of contexts using hierarchical graph neural networks. During this stage, the router allows information exchange across different levels, expanding the model’s receptive field in dialogue. (2) Fuses matching signals from multiple levels of contexts with a fusion network. Here, the router restricts the information flows to the same context level, efficiently minimizing noise from irrelevant context. Furthermore, despite the significance of argumentative multi-party dialogue in real-world applications, this area remains largely unexplored due to dataset scarcity. To address this challenge, we develop two meticulously annotated datasets, MRDL and MRDR. Unlike the prevailing datasets that primarily focus on short and colloquial conversations, our datasets feature intricate argumentative dialogues and are publicly accessible at https://github.com/AI0Research/MRDL-and-MRDR. Our new datasets and the HGFN model could promote further advancements in this field. Extensive experiments are conducted and it is revealed that the HGFN model surpasses the state-of-the-art, particularly in complex, argumentative dialogues. Tiezheng Mao, Tianyong Hao, Jialing Fu, Osamu Yoshie |
Inf. Process. Manag. | 4 |
| 2024 | Seeing both sides: context-aware heterogeneous graph matching networks for extracting-related argumentsabstractAbstract Our research focuses on extracting exchanged views from dialogical documents through argument pair extraction (APE). The objective of this process is to facilitate comprehension of complex argumentative discourse by finding the related arguments. The APE comprises two stages: argument mining and argument matching. Researchers typically employ sequence labeling models for mining arguments and text matching models to calculate the relationships between them, thereby generating argument pairs. However, these approaches fail to capture long-distance contextual information and struggle to fully comprehend the complex structure of arguments. In our work, we propose the context-aware heterogeneous graph matching (HGMN) model for the APE task. First, we design a graph schema specifically tailored to argumentative texts, along with a heterogeneous graph attention network that effectively captures context information and structural information of arguments. Moreover, the text matching between arguments is converted into a graph matching paradigm and a multi-granularity graph matching model is proposed to handle the intricate relationships between arguments at various levels of granularity. In this way, the semantics of argument are modeled structurally and thus capture the complicated correlations between arguments. Extensive experiments are conducted to evaluate the HGMN model, including comparisons with existing methods and the GPT series of large language models (LLM). The results demonstrate that HGMN outperforms the state-of-the-art method. Tiezheng Mao, Osamu Yoshie, Jialing Fu, Weixin Mao |
Neural Comput. Appl. | 2 |
| 2023 | A Simple Framework for Text-Supervised Semantic SegmentationabstractText-supervised semantic segmentation is a novel research topic that allows semantic segments to emerge with image-text contrasting. However, pioneering methods could be subject to specifically designed network architectures. This paper shows that a vanilla contrastive language-image pretraining (CLIP) model is an effective text-supervised semantic segmentor by itself. First, we reveal that a vanilla CLIP is inferior to localization and segmentation due to its optimization being driven by densely aligning visual and language representations. Second, we propose the locality-driven alignment (LoDA) to address the problem, where CLIP optimization is driven by sparsely aligning local representations. Third, we propose a simple segmentation (SimSeg) framework. LoDA and SimSeg jointly amelio-rate a vanilla CLIP to produce impressive semantic segmentation results. Our method outperforms previous state-of-the-art methods on PASCAL VOC 2012, PASCAL Context and COCO datasets by large margins. Code and models are available at github.com/muyangyi/SimSeg. Muyang Yi, Quan Cui, Osamu Yoshie, Hongtao Lu 0001 |
CVPR | 5 |
| 2023 | Matching Intentions for Discourse Parsing in Multi-party Dialogues
Tiezheng Mao, Jialing Fu, Osamu Yoshie, Yimin Fu, Zhuyun Li |
IEA/AIE (2) | 3 |
| 2022 | Application of Multi-modal Fusion Attention Mechanism in Semantic Segmentation
Yunlong Liu 0008, Osamu Yoshie, Hiroshi Watanabe 0001 |
ACCV (7) | 2 |
| 2022 | Discriminability-Transferability Trade-Off: An Information-Theoretic Perspective
Quan Cui, Bingchen Zhao, Borui Zhao, Renjie Song, Boyan Zhou, Jiajun Liang, Osamu Yoshie |
ECCV (26) | 8 |
| 2022 | Contrastive Vision-Language Pre-training with Limited Resources
Quan Cui, Boyan Zhou, Weidong Yin, Osamu Yoshie, Yubo Chen 0004 |
ECCV (36) | 6 |
| 2022 | TRC-Unet: Transformer Connections for Near-infrared Blurred Image SegmentationabstractImaging blood vessel networks is useful in many biomedical applications, such as injection-assist, cancer detection, various surgery, and vein identification. In NIR (near-infrared) transillumination imaging, we can visualize the subcutaneous blood vessel network. However, such images are severely blurred by the strong scattering of body tissue, and it remains challenging for most models to accurately segment these blurred images. In addition, the convolution operation in the deep learning approach means that it extracts a mixture of blurred edges and clear centers, resulting in gradual distortion during upsampling. In this paper, we propose a novel and efficient deep learning model called TRC-Unet for segmenting blurred NIR images. The transformer connection (TRC) block extracts global spatial information from different scales by adaptively suppressing scattering and increasing the clarity of features. Our proposed transformer feature fusion (TFF) module closes the gap between the highly semantic feature maps of CNN and the adaptive fuzzy transformer output to enable a precise reconstruction of the segmentation. We evaluated TRC-Unet on both a simulated blurred DRIVE dataset and a NIR vessel dataset, and we achieved competitive results. (i.e., 83.86% Dice score on DRIVE and an average boost of 4.6% on simulated images at different depths). Jiazhe Wang, Osamu Yoshie, Koichi Shimizu |
ICPR | 2 |
| 2022 | Delving into the representation learning of deep hashingabstractSearching for the nearest neighbor is a fundamental problem in the computer vision field, and deep hashing has become one of the most representative and widely used methods, which learns to generate compact binary codes for visual data. In this paper, we first delve into the representation learning of deep hashing and surprisingly find that deep hashing could be a double-edged sword, i.e., deep hashing can accelerate the query speed and decrease the storage cost in the nearest neighbor search progress, but it greatly sacrifices the discriminability of deep representations especially with extremely short target code lengths. To solve this problem, we propose a two-step deep hashing learning framework. The first step focuses on learning deep discriminative representations with metric learning. Subsequently, the learning framework concentrates on simultaneously learning compact binary codes and preserving representations learned in the former step from being sacrificed. Extensive experiments on two general image datasets and four challenging image datasets validate the effectiveness of our proposed learning framework. Moreover, the side effect of deep hashing is successfully mitigated with our learning framework. Quan Cui, Osamu Yoshie |
Neurocomputing | 3 |
| 2022 | SST: Spatial and Semantic Transformers for Multi-Label Image RecognitionabstractMulti-label image recognition has attracted considerable research attention and achieved great success in recent years. Capturing label correlations is an effective manner to advance the performance of multi-label image recognition. Two types of label correlations were principally studied, i.e., the spatial and semantic correlations. However, in the literature, previous methods considered only either of them. In this work, inspired by the great success of Transformer, we propose a plug-and-play module, named the Spatial and Semantic Transformers (SST), to simultaneously capture spatial and semantic correlations in multi-label images. Our proposal is mainly comprised of two independent transformers, aiming to capture the spatial and semantic correlations respectively. Specifically, our Spatial Transformer is designed to model the correlations between features from different spatial positions, while the Semantic Transformer is leveraged to capture the co-existence of labels without manually defined rules. Other than methodological contributions, we also prove that spatial and semantic correlations complement each other and deserve to be leveraged simultaneously in multi-label image recognition. Benefitting from the Transformer's ability to capture long-range correlations, our method remarkably outperforms state-of-the-art methods on four popular multi-label benchmark datasets. In addition, extensive ablation studies and visualizations are provided to validate the essential components of our method. Quan Cui, Borui Zhao, Renjie Song, Xiaoqin Zhang 0002, Osamu Yoshie |
IEEE Trans. Image Process. | 6 |
| 2021 | OTA: Optimal Transport Assignment for Object DetectionabstractRecent advances in label assignment in object detection mainly seek to independently define positive/negative training samples for each ground-truth (gt) object. In this paper, we innovatively revisit the label assignment from a global perspective and propose to formulate the assigning procedure as an Optimal Transport (OT) problem – a well-studied topic in Optimization Theory. Concretely, we define the unit transportation cost between each demander (anchor) and supplier (gt) pair as the weighted summation of their classification and regression losses. After formulation, finding the best assignment solution is converted to solve the optimal transport plan at minimal transportation costs, which can be solved via Sinkhorn-Knopp Iteration. On COCO, a single FCOS-ResNet-50 detector equipped with Optimal Transport Assignment (OTA) can reach 40.7% mAP under 1× scheduler, outperforming all other existing assigning methods. Extensive experiments conducted on COCO and CrowdHuman further validate the effectiveness of our proposed OTA, especially its superiority in crowd scenarios. The code is available at https://github.com/Megvii-BaseDetection/OTA. Zheng Ge, Osamu Yoshie, Jian Sun 0001 |
CVPR | 4 |
| 2021 | Delving deep into the imbalance of positive proposals in two-stage object detectionabstractImbalance issue is a major yet unsolved bottleneck for the current object detection models. In this work, we observe two crucial yet never discussed imbalance issues. The first imbalance lies in the large number of low-quality RPN proposals, which makes the R-CNN module (i.e., post-classification layers) become highly biased towards the negative proposals in the early training stage. The second imbalance stems from the unbalanced ground-truth numbers across different testing images, resulting in the imbalance of the number of potentially existing positive proposals in testing phase. To tackle these two imbalance issues, we incorporates two innovations into Faster R-CNN: 1) an R-CNN Gradient Annealing (RGA) strategy to enhance the impact of positive proposals in the early training stage. 2) a set of Parallel R-CNN Modules (PRM) with different positive/negative sampling ratios during training on one same backbone. Our RGA and PRM can totally bring 2.0% improvements on AP on COCO minival. Experiments on CrowdHuman further validates the effectiveness of our innovations across various kinds of object detection tasks. Zheng Ge, Zequn Jie, Chengzheng Li, Osamu Yoshie |
Neurocomputing | 5 |
| 2021 | LLA: Loss-aware label assignment for dense pedestrian detectionabstractLabel assignment has been widely studied in general object detection because of its great impact on detectors’ performance. In the field of dense pedestrian detection, human bodies are often heavily entangled, making label assignment more important. However, none of the existing label assignment method focuses on crowd scenarios. Motivated by this, we propose Loss-aware Label Assignment (LLA) to boost the performance of pedestrian detectors in crowd scenarios. Concretely, LLA first calculates classification (cls) and regression (reg) losses between each anchor and ground-truth (GT) pair. A joint loss is then defined as the weighted summation of cls and reg losses as the assigning indicator. Finally, anchors with top K minimum joint losses for a certain GT box are assigned as its positive anchors. Anchors that are not assigned to any GT box are considered negative. LLA is simple but effective. Experiments on CrowdHuman and CityPersons show that such a simple label assigning strategy can boost MR by 9.53% and 5.47% on two famous one-stage detectors – RetinaNet and FCOS, becoming the first one-stage detector that surpasses Faster R-CNN in crowd scenarios. Zheng Ge, Osamu Yoshie |
Neurocomputing | 5 |
| 2020 | NMS by Representative Region: Towards Crowded Pedestrian Detection by Proposal PairingabstractAlthough significant progress has been made in pedestrian detection recently, pedestrian detection in crowded scenes is still challenging. The heavy occlusion between pedestrians imposes great challenges to the standard Non-Maximum Suppression (NMS). A relative low threshold of intersection over union (IoU) leads to missing highly overlapped pedestrians, while a higher one brings in plenty of false positives. To avoid such a dilemma, this paper proposes a novel Representative Region NMS (R2NMS) approach leveraging the less occluded visible parts, effectively removing the redundant boxes without bringing in many false positives. To acquire the visible parts, a novel Paired-Box Model (PBM) is proposed to simultaneously predict the full and visible boxes of a pedestrian. The full and visible boxes constitute a pair serving as the sample unit of the model, thus guaranteeing a strong correspondence between the two boxes throughout the detection pipeline. Moreover, convenient feature integration of the two boxes is allowed for the better performance on both full and visible pedestrian detection tasks. Experiments on the challenging CrowdHuman and CityPersons benchmarks sufficiently validate the effectiveness of the proposed approach on pedestrian detection in the crowded situation. Zheng Ge, Zequn Jie, Osamu Yoshie |
CVPR | 4 |
| 2020 | ExchNet: A Unified Hashing Network for Large-Scale Fine-Grained Image Retrieval
Quan Cui, Qing-Yuan Jiang, Xiu-Shen Wei, Wu-Jun Li, Osamu Yoshie |
ECCV (3) | 5 |
| 2020 | PS-RCNN: Detecting Secondary Human Instances in a Crowd via Primary Object SuppressionabstractDetecting human bodies in highly crowded scenes is a challenging problem. Two main reasons result in such a problem: 1). weak visual cues of heavily occluded instances can hardly provide sufficient information for accurate detection; 2). heavily occluded instances are easier to be suppressed by Non-Maximum-Suppression (NMS). To address these two issues, we introduce a variant of two-stage detectors called PS-RCNN. PS-RCNN first detects slightly/none occluded objects by an R-CNN [1] module (referred as P-RCNN), and then suppress the detected instances by human-shaped masks so that the features of heavily occluded instances can stand out. After that, PS-RCNN utilizes another R-CNN module specialized in heavily occluded human detection (referred as S-RCNN) to detect the rest missed objects by P-RCNN. Final results are the ensemble of the outputs from these two RCNNs. Moreover, we introduce a High Resolution RoI Align (HRRA) module to retain as much of fine-grained features of visible parts of the heavily occluded humans as possible. Our PS-RCNN significantly improves recall and AP by 4.49% and 2.92% respectively on CrowdHuman [2], compared to the baseline. Similar improvements on Widerperson [3] are also achieved by the PS-RCNN. Zheng Ge, Zequn Jie, Osamu Yoshie |
ICME | 5 |
| 2020 | DualBox: Generating BBox Pair with Strong Correspondence via Occlusion Pattern Clustering and Proposal RefinementabstractDespite the rapid development of pedestrian detection, the problem of dense pedestrian detection is still unsolved, especially the upper limit of Recall caused by Non-Maximum-Suppression (NMS). Out of this reason, R2NMS [1] is proposed to simultaneously detect full and visible body bounding boxes, by replacing the full body BBoxes with less occluded visible body BBoxes in the NMS algorithm, achieving a higher recall. However, the P-RPN and P-RCNN modules proposed in R2NMS for simultaneous high quality full and visible body prediction require non-trivial positive/negative assigning strategies for anchor BBoxes. To simplify the prerequisites and improve the utility of R2NMS, we incorporate clustering analysis into the learning of visible body proposals from full body proposals. Furthermore, to reduce the computation complexity caused by the large number of potential visible body proposals, we introduce a novel occlusion pattern prediction branch on top of the R-CNN module (i.e. F-RCNN) to select the best matched visible proposals for each full body proposals and then feed them into another R-CNN module (i.e. V-RCNN). Incorporated with R2NMS, our DualBox model can achieve competitive performance while only requires few hyper-parameters. We validate the effectiveness of the proposed approach on the CrowdHuman [2] and CityPersons [3] datasets. Experimental results show that our approach achieves promising performance for detecting both non-occluded and occluded pedestrians, especially heavily occluded ones. Zheng Ge, Chuyu Hu, Baiqiao Qiu, Osamu Yoshie |
ICPR | 5 |
| 2020 | Support software for Automatic Speech Recognition systems targeted for non-native speechabstractNowadays automatic speech recognition (ASR) systems can achieve higher and higher accuracy rates depending on the methodology applied and datasets used. The rate decreases significantly when the ASR system is being used with a non-native speaker of the language to be recognized. The main reason for this is specific pronunciation and accent features related to the mother tongue of that speaker, which influence the pronunciation. At the same time, an extremely limited volume of labeled non-native speech datasets makes it difficult to train, from the ground up, sufficiently accurate ASR systems for non-native speakers. Kacper Radzikowski, Osamu Yoshie, Robert M. Nowak |
iiWAS | 2 |
| 2019 | Accent neutralization for speech recognition of non-native speakersabstractThese days, automatic speech recognition (ASR) systems achieve higher and higher accuracy rates. The score drops significantly, in case when the ASR system is being used with a non-native speaker of the language to be recognized. The main reason is specific pronunciation and accent features. A limited volume of labeled non-native speech datasets makes it difficult to train new ASR systems for non-native speakers. Kacper Radzikowski, Mateusz Forc, Osamu Yoshie, Robert M. Nowak |
iiWAS | 4 |
| 2016 | A passive means based privacy protection method for the perceptual layer of IoTsabstractPrivacy protection in Internet of Things (IoTs) has long been the topic of extensive research in the last decade. The perceptual layer of IoTs suffers the most significant privacy disclosing because of the limitation of hardware resources. Data encryption and anonymization are the most common methods to protect private information for the perceptual layer of IoTs. However, these efforts are ineffective to avoid privacy disclosure if the communication environment exists unknown wireless nodes which could be malicious devices. Therefore, in this paper we derive an innovative and passive method called Horizontal Hierarchy Slicing (HHS) method to detect the existence of unknown wireless devices which could result negative means to the privacy. PAM algorithm is used to cluster the HHS curves and analyze whether unknown wireless devices exist in the communicating environment. Link Quality Indicator data are utilized as the network parameters in this paper. The simulation results show their effectiveness in privacy protection. Osamu Yoshie, Daoping Huang |
iiWAS | 2 |
| 2016 | Non-native English speakers' speech correction, based on domain focused documentabstractWith the increase in exchange programs, many international students worldwide can face communication problems. During lectures, usually English language is used for the communication between students and teachers. However both sides, not necessarily being native speakers of English, may misunderstand each other. In this paper we propose a method for correction of non-native English speakers' speech, based on the domain focused electronic document. The method relies on the results of speech recognition (SR) software, and uses them altogether with the document. Our approach consists of three steps. Firstly, document analysis in the preprocessing phase. Secondly, finding the document part corresponding to sentence from SR software, realised using the Hidden Markov Model (HMM) based method. Finally, the correction by calculating the score for each of candidate sentences, based on the result of SR software. The probability score combines keywords comparison, BM25F method and HMM based method scores. Highest score candidate is chosen as replacement. Kacper Radzikowski, Osamu Yoshie |
iiWAS | 3 |
| 2013 | Facial expression recognition by analyzing features of conceptual regionsabstractFacial expression recognition utilizes collection of information from characteristic actions to analyze emotions and mental states of a person. It has emerged as the pivotal research topics in areas such as human computer interaction, sentimental analysis and synthetic face animation over the last years. This paper proposes an approach for facial expression by discovering associations between visual feature and Local Binary Pattern (LBP). Unlike many previous studies, the proposed approach automatically tracks the facial area and segments face into meaningful areas based on description of Local Binary Pattern. And then it accumulates the probabilities throughout the frames from video data to capture the temporal characteristics of facial expressions by analyzing facial expressions. Through the proposed approach, the temporal variation of facial expression can be quantified in individual areas. Thus, the recognition process of facial expression tends to be more comprehensible without sacrificing results of recognition. The empirical evaluation results of the approach are realized using video data which is collected from 10 volunteers. The results demonstrated that the proposed approach can effectively segment face into specific area and recognize facial expression. Huiquan Zhang, Sha Luo, Osamu Yoshie |
ICIS | 3 |
| 2013 | Conversation Analysis Based on Interpersonal Relationship in Consensus BuildingabstractBuilding consensus is a process in which participants start with various opinions and reach an agreement as far as possible. During a discussion, participants exhibit different stances such as agree or disagree at other's utterances. In this paper, a stance analysis of participants by combining BBS tree structure and participant relation graph is proposed. The stance change is measured by an approach of information theory. Finally, simulations are provided to demonstrate the feasibility of the proposed method. Osamu Yoshie |
iiWAS | 2 |
| 2013 | Using Planning with Action Preference in Story GenerationabstractNowadays, plenty of researches focus on story generation which is widely used in computer games, education and training applications. It is highly desirable that the generated story should afford high user agency and at same time having capabilities to address user's interventions. In this paper, we apply planning, which is derived from artificial intelligence, to achieve this objective. With the use of planning, several solutions are produced, which contains a sequence of user's and system agents' actions. In addition, we propose the concept of Action Preference, which takes into account user's feedbacks, to evaluate all of the solutions after planning. Meanwhile a variant of hyperbolic tangent is utilized to calculate Action Preference. In order to evaluate its feasibility, an educational game was implemented on the basis of story generation. That result proves that planning with Action Preference is an effective approach in story generation. Samiullah Paracha, Osamu Yoshie |
MoMM | 4 |
| 2013 | A Fusion Approach for Facial Expression Using Local Binary Pattern and a Pseudo 3D Face ModelabstractThe discovery of association between local facial feature and regional feature of image for recognizing facial expression is one of the current challenges in the regard of facial expression recognition. In this paper, we propose an approach fusing different complementary methods that can facilitate recognition of facial expression by exploring the knowledge on the face. The approach incorporates the joint work of Local Binary Pattern to extract the holistic facial area for recognizing facial expression. To deal with regional feature of face, a pseudo 3Dface model is established for segmenting facial regions. Final, a propagation method is introduced to recognize the facial expression. Unlike many previous studies, the presented approach recognizes expressions from regional feature instead of holistic facial features or complete texture information. In order to validate our proposed approach, we have conducted experiments on the Extended Cohn-Kanade (CK+) facial expression database with an recognition rate of 91.7%. Huiquan Zhang, Osamu Yoshie |
SNPD | 2 |
| 2012 | Consensus building analysis using entropy in BBS treeabstractConsensus building is a process in which individuals collectively make a choice from the alternatives and contribute their best to make a common decision. During the discussion, there will appear a case that after a certain utterance, individuals will fix their positions or change their ideas. This certain utterance is called "clue". This paper describes a quantitative research method is put forward to measure the information entropy change during the discussion for tracing "clue" in BBS tree structure. Finally a simple simulation proves the theory. Yasushi Oda, Chikara Otani, Osamu Yoshie |
iiWAS | 4 |
| 2011 | Decision forest for multivariate time series analysisabstractNowadays with time series accounting for an increasingly large fraction of world's supply of data, there has been an explosion of interest in mining time series data. This paper proposes a multivariate time series classification model which is both effective in classifier's accuracy and comprehensibility. It is composed of two stages: a supervised clustering for pattern extraction and soft discretization decision forest. In supervised clustering, some real time series instances from the training dataset will be selected as class dedicated patterns. While in decision forest, the rule induction helps to improve the knowledge acquisition of the classifier. In addition, soft discretization would further improve the accuracy and comprehensibility of the classifier. Leyang Li, Osamu Yoshie |
iiWAS | 3 |
| 2010 | Query-biased summarization considering difference of paragraphsabstractMost conventional query-biased summarization methods generate the summary using extracted sentences based on similarity measure between all sentences in a document and the query. If there are plural sentences having high similarity to the query in the document, these methods cannot decide the sentence which the summary should be from. This paper proposes an algorithm adopting new indicator that shows the difference between one paragraph and the others. In a word space which is composed of all words in the target document, the algorithm determines the axis that maximizes the difference when a paragraph and the others are projected onto it. There are many combinations of a paragraph and a set of other paragraphs. For each combination, the above-mentioned axis that maximizes the difference and gives a conformity degree to the given query is calculated. With these conformities, the algorithm decides one paragraph for generating the summary. To obtain the axis, topic distinctiveness factor analysis is applied. The basic idea for making final summary is concatenating the sentences extracted from the paragraph. The resultant summary is evaluated from the following points of view: readability, understandability and the easiness to judge whether the link works well or not. Chikara Otani, Yasushi Oda, Osamu Yoshie |
iiWAS | 3 |
| 2006 | Low Cost Rendering Method for Virtual Factory Considering Interpolation of Occluded Objects
Hiroki Takahashi, Naoyuki Tamura, Toshihiko Furue, Osamu Yoshie |
MoMM | 4 |
| 2005 | Color Transformation Method for Universal Web View
Yayori Kasagi, Hiroki Takahashi, Osamu Yoshie |
iiWAS | 3 |
| 2004 | Agent Handling with VR Technology and Spatial Programming for Information Aggregation in Plants
Hiroki Takahashi, Osamu Yoshie |
iiWAS | 2 |