Xi Yang 0017

dblp:13/1520-17 · DBLP profile ↗
← Back
22ranked-venue papers
3as first author
19since 2021 · last 2026
0000-0001-5039-3680ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 3 first-author · 14 since 2021Human-computer interaction and ubiquitous computing · 10 · 10 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Learning Compact Latent Space for Representing Neural Signed Distance Functions with High-fidelity Geometry Details
abstract
Neural signed distance functions (SDFs) have been a vital representation to represent 3D shapes or scenes with neural networks. An SDF is an implicit function that can query signed distances at specific coordinates for recovering a 3D surface. Although implicit functions work well on a single shape or scene, they pose obstacles when analyzing multiple SDFs with high-fidelity geometry details, due to the limited information encoded in the latent space for SDFs and the loss of geometry details. To overcome these obstacles, we introduce a method to represent multiple SDFs in a common space, aiming to recover more high-fidelity geometry details with more compact latent representations. Our key idea is to take full advantage of the benefits of generalization-based and overfitting-based learning strategies, which manage to preserve high-fidelity geometry details with compact latent codes. Based on this framework, we also introduce a novel sampling strategy to sample training queries. The sampling can improve the training efficiency and eliminate artifacts caused by the influence of other SDFs. We report numerical and visual evaluations on widely used benchmarks to validate our designs and show advantages over the latest methods in terms of the representative ability and compactness.
Qiang Bai, Bojian Wu, Xi Yang 0017, Zhizhong Han
AAAI3
2026 Specializing Large Models for Oracle Bone Script Interpretation via Component-Grounded Multimodal Knowledge Augmentation
abstract
Jianing Zhang, Runan Li, Honglin Pang, Ding Xia, Zhou Zhu, Qian Zhang, Chuntao Li, Xi Yang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Runan Li, Honglin Pang, Ding Xia, Zhou Zhu, Xi Yang 0017
ACL (1)8
2025 LineArt: A Knowledge-guided Training-free High-quality Appearance Transfer for Design Drawing with Diffusion Model
abstract
Image rendering from line drawings is vital in design and image generation technologies reduce costs, yet professional line drawings demand preserving complex details. Text prompts struggle with accuracy, and image translation struggles with consistency and fine-grained control. We present LineArt, a framework that transfers complex appearance onto detailed design drawings, facilitating design and artistic creation. It generates high-fidelity appearance while preserving structural accuracy by simulating hierarchical visual cognition and integrating human artistic experience to guide the diffusion process. LineArt overcomes the limitations of current methods in terms of difficulty in fine-grained control and style degradation in design drawings. It requires no precise 3D modeling, physical property specifications, or network training, making it more convenient for design tasks. LineArt consists of two stages: a multi-frequency lines fusion module to supplement the input design drawing with detailed structural information and a two-part painting process for Base Layer Shaping and Surface Layer Coloring. We also present a new design drawing dataset, ProLines, for evaluation. The experiments show that LineArt performs better in accuracy, realism, and material precision compared to SOTAs. Project page: https://meaoxixi.github.io/LineArt/.
Hongzhen Li, Yichen Peng, Haoran Xie 0002, Xi Yang 0017
CVPR6
2025 JoruriPuppet: Learning Tempo-Changing Mechanisms Beyond the Beat for Music-to-Motion Generation with Expressive Metrics
abstract
In music-to-motion generation, the interplay between movements and music tempo variations significantly influences the emotional expressiveness and realism of performances. However, tempo-changing mechanisms remain underexplored in neural network-based music-to-motion tasks due to the scarcity of relevant datasets. Therefore, in this paper, we propose to use novel music features explicitly representing tempo variations, and introduce a dataset, JoruriPuppet, incorporating the Japanese traditional Jo-Ha-Kyu principle characterized by expressive tempo changes. Furthermore, we design three metrics to quantitatively evaluate the synchronization and expressiveness of generated motions. Experiments on our dataset highlight the limitations of SOTA methods in capturing fine-grained tempo changes. We demonstrate that integrating tempo-changing features into them improves neural network-based music-to-motion performance across existing datasets, validating the general effectiveness and applicability of our research. The dataset and source code are available at the project link: https://www.dr-lab.org/projects/joruripuppet/.
Ran Dong, Shaowen Ni, Xi Yang 0017
SIGGRAPH Asia3
2024 PairingNet: A Learning-Based Pair-Searching and -Matching Network for Image Fragments
Rixin Zhou, Ding Xia, Yi Zhang 0083, Honglin Pang, Xi Yang 0017
ECCV (59)5
2024 Speed Labeling: Non-stop Scrolling for Fast Image Labeling
abstract
This study presents “Speed Labeling”, an image-labeling technique to increase the efficiency of easy binary labeling tasks where an annotator can choose a label instantly. We first conduct a formative study to identify the factors affecting the efficiency of easy image labeling: image layout and image transition. Based on these results, we designed a novel labeling technique using non- stop scrolling. In conventional image labeling, the system moves to the next image only after the user assigns a label to the previous image. To maximize efficiency, our technique continuously scrolls images without waiting for the completion of labeling, assuming that the user gives labels at a mostly constant speed. The system dynamically adjusts the scrolling speed based on the labeling speed. Subsequently, we conduct a user study to compare the proposed “non-stop scrolling” technique to the conventional “stop-and-go scrolling” technique in an easy image-labeling task. The results showed that speed labeling requires less time (faster by 7%, 305 more images labeled per man-hour) to complete the labeling task than the conventional technique without a significant increase in errors. In addition, the results showed that speed labeling makes the labeling task more enjoyable for crowd workers and makes them feel more attentive during tasks.
Chia-Ming Chang 0003, Xi Yang 0017, Xiang 'Anthony' Chen, Takeo Igarashi
Graphics Interface3
2024 PDFChatAnnotator: A Human-LLM Collaborative Multi-Modal Data Annotation Tool for PDF-Format Catalogs
abstract
The document contains substantial unannotated data, necessitating extensive manual labeling efforts. To address this issue, we introduce PDFChatAnnotator, a human-LLM collaborative tool to collect multi-modal data from PDF catalogs. Initially, PDFChatAnnotator automatically employs our proposed multi-modal binding rules to link related data from different modalities and harnesses the information extraction capabilities of large language models (LLMs) to extract specific information from text descriptions. Furthermore, the tool empowers users to guide and refine the LLM’s annotations. During the annotation process, users can influence the LLM through multiple rounds of communication and example establishment via the provided interfaces. To assess the effectiveness of PDFChatAnnotator’s techniques, we conducted a technical evaluation using three catalogs with typical layouts as experimental data. The results showed that all accuracy rates for multi-modal binding exceeded 90%, and both the proposed "example establishment" and "interactive adjustment of requirements" contributed to enhanced accuracy rates.
Chia-Ming Chang 0003, Xi Yang 0017
IUI3
2024 SpaceEditing: A Latent Space Editing Interface for Integrating Human Knowledge into Deep Neural Networks
abstract
Human-centered AI aims to bridge the gap between machine decision-making and human understanding. However, even for classification tasks where deep neural networks have achieved superb performance, there are currently few methods that link humans and AI well, especially on domain-specific tasks. In this paper, we propose SpaceEditing, a 2D spatial layout tool that enables human users to interact with the latent space of deep neural networks. During the interaction process, the tool’s algorithm automatically processes user actions, providing feedback to the network and leveraging triplet loss to effectively learn from user-modified information. We evaluate SpaceEditing with three case studies: (1) an archaeology researcher uses a bronze dataset; (2) a deep learning researcher uses a garbage classification dataset; (3) six deep learning beginners use a head pose dataset. The experimental results demonstrate the effectiveness of our tool in integrating human knowledge and improving network performance.
Jiafu Wei, Ding Xia, Haoran Xie 0002, Chia-Ming Chang 0003, Xi Yang 0017
IUI6
2023 Parts2Words: Learning Joint Embedding of Point Clouds and Texts by Bidirectional Matching Between Parts and Words
abstract
Shape-Text matching is an important task of high-level shape understanding. Current methods mainly represent a 3D shape as multiple 2D rendered views, which obviously can not be understood well due to the structural ambiguity caused by self-occlusion in the limited number of views. To resolve this issue, we directly represent 3D shapes as point clouds, and propose to learn joint embedding of point clouds and texts by bidirectional matching between parts from shapes and words from texts. Specifically, we first segment the point clouds into parts, and then leverage optimal transport method to match parts and words in an optimized feature space, where each part is represented by aggregating features of all points within it and each word is abstracted by its contextual information. We optimize the feature space in order to enlarge the similarities between the paired training samples, while simultaneously maximizing the margin between the unpaired ones. Experiments demonstrate that our method achieves a significant improvement in accuracy over the SOTAs on multi-modal retrieval tasks under the Text2Shape dataset. Codes are available at here.
Chuan Tang, Xi Yang 0017, Bojian Wu, Zhizhong Han
CVPR2
2023 Multi-Granularity Archaeological Dating of Chinese Bronze Dings Based on a Knowledge-Guided Relation Graph
abstract
The archaeological dating of bronze dings has played a critical role in the study of ancient Chinese history. Current archaeology depends on trained experts to carry out bronze dating, which is time-consuming and labor-intensive. For such dating, in this study, we propose a learning-based approach to integrate advanced deep learning techniques and archaeological knowledge. To achieve this, we first collect a large-scale image dataset of bronze dings, which contains richer attribute information than other existing fine-grained datasets. Second, we introduce a multihead classifier and a knowledge-guided relation graph to mine the relationship between attributes and the ding era. Third, we conduct comparison experiments with various existing methods, the results of which show that our dating method achieves a state-of-the-art performance. We hope that our data and applied networks will enrich fine-grained classification re-search relevant to other interdisciplinary areas of expertise. The dataset and source code used are included in our supplementary materials, and will be open after submission owing to the anonymity policy. Source codes and data are available at: https://github.com/zhourixin/bronze-Ding
Rixin Zhou, Jiafu Wei, Ruihua Qi, Xi Yang 0017
CVPR5
2023 Efficient Human-in-the-loop System for Guiding DNNs Attention
abstract
Attention guidance is used to address dataset bias in deep learning, where the model relies on incorrect features to make decisions. Focusing on image classification tasks, we propose an efficient human-in-the-loop system to interactively direct the attention of classifiers to regions specified by users, thereby reducing the effect of co-occurrence bias and improving the transferability and interpretability of a deep neural network (DNN). Previous approaches for attention guidance require the preparation of pixel-level annotations and are not designed as interactive systems. We herein present a new interactive method that allows users to annotate images via simple clicks. Additionally, we identify a novel active learning strategy that can significantly reduce the number of annotations. We conduct both numerical evaluations and a user study to evaluate the proposed system using multiple datasets. Compared with the existing non-active-learning approach, which typically relies on considerable amounts of polygon-based segmentation masks to fine-tune or train the DNNs, our system can obtain fine-tuned networks on biased datasets in a more time- and cost-efficient manner and offers a more user-friendly experience. Our experimental results show that the proposed system is efficient, reasonable, and reliable. Our code is publicly available at https://github.com/ultratykis/Guiding-DNNs-Attention.
Xi Yang 0017, Chia-Ming Chang 0003, Haoran Xie 0002, Takeo Igarashi
IUI2
2023 A two-step surface-based 3D deep learning pipeline for segmentation of intracranial aneurysms
abstract
The exact shape of intracranial aneurysms is critical in medical diagnosis and surgical planning. While voxel-based deep learning frameworks have been proposed for this segmentation task, their performance remains limited. In this study, we offer a two-step surface-based deep learning pipeline that achieves significantly better results. Our proposed model takes a surface model of an entire set of principal brain arteries containing aneurysms as input and returns aneurysm surfaces as output. A user first generates a surface model by manually specifying multiple thresholds for time-of-flight magnetic resonance angiography images. The system then samples small surface fragments from the entire set of brain arteries and classifies the surface fragments according to whether aneurysms are present using a point-based deep learning network (PointNet++). Finally, the system applies surface segmentation (SO-Net) to surface fragments containing aneurysms. We conduct a direct comparison of the segmentation performance of our proposed surface-based framework and an existing voxel-based method by counting voxels: our framework achieves a much higher Dice similarity (72%) than the prior approach (46%).
Xi Yang 0017, Ding Xia, Taichi Kin, Takeo Igarashi
Comput. Vis. Media1
2023 SyncLabeling: A Synchronized Audio Segmentation Interface for Mobile Devices
abstract
Manual audio segmentation is a time-consuming process, especially when there is more than one sound playing simultaneously that needs to be segmented and annotated (e.g., target and background sounds). In conventional audio annotation interfaces, users need to repeatedly pause and replay the audio to complete an overlap segmentation task, which is very inefficient. In this paper, we propose "SyncLabeling," a synchronized audio segmentation interface for smartphones that allows users to segment and annotate two overlapping sounds in a single audio stream at a time using a game-like labeling interface on mobile devices. We conducted a user study to compare the proposed SyncLabeling interface with a conventional audio annotation interface on four types of audio segmentation tasks. The results showed that the proposed interface is much more efficient than the conventional interface (2.4× faster) under comparable annotation accuracy in most tasks. In addition, more than half of the participants enjoyed using the proposed SyncLabeling interface and showed willingness to use it.
Chia-Ming Chang 0003, Xi Yang 0017, Takeo Igarashi
Proc. ACM Hum. Comput. Interact.3
2022 A Drawing Support System for Sketching Aging Anime Faces
abstract
Drawing the facial features of anime characters at different ages is challenging in the creation process. Since characters’ facial features at different ages have obvious differences, it is difficult, especially for novices to accurately illustrate the age features of anime characters. Conventional data-driven drawing interfaces for anime characters focus on the visual features of hair, emotion, and coloring but fail to provide age-specific drawing guidance. To address this gap, we propose AgeFace, an interactive drawing interface that assists users in creating facial features with age-specific features based on user input strokes. We evaluated the usability of AgeFace by a user experience experiment and a comparison experiment with baseline approaches. The results verified that AgeFace could achieve better performance in usability and better support in the creative process than baseline systems.
Sicheng Li 0002, Haoran Xie 0002, Xi Yang 0017, Chia-Ming Chang 0003, Kazunori Miyata
CW3
2022 Interactive Drawing Interface for Editing Scene Graph
abstract
Scene graphs have been widely used in various visual applications, such as image retrieval and generation. However, it is time consuming and challenging to draw scene graphs, especially for images with complex scenes. To resolve this issue, we propose an interactive editing interface for scene graph representation in which users can draw scene graphs simply and conveniently. We provide alternatives for frequently used objects, attributes, and relationships in graph drawing. The proposed function design can greatly reduce the drawing time cost and improve drawing quality. We conducted a user study and confirmed that the proposed interface can help users draw the desired scene graph, as well as providing a good user experience.
Xusheng Du, Chia-Ming Chang 0003, Xi Yang 0017, Haoran Xie 0002
CW4
2022 An Empirical Study on the Effect of Quick and Careful Labeling Styles in Image Annotation
Chia-Ming Chang 0003, Xi Yang 0017, Takeo Igarashi
Graphics Interface2
2022 DualLabel: Secondary Labels for Challenging Image Annotation
Chia-Ming Chang 0003, Xi Yang 0017, Haoran Xie 0002, Takeo Igarashi
Graphics Interface3
2022 Fine-tuning Deep Neural Networks by Interactively Refining the 2D Latent Space of Ambiguous Images
abstract
Deep neural networks (DNNs) have achieved excellent results currently in classification, while they may still suffer from ambiguous images which are similar across classes. By contrast, humans have a relatively good ability to distinguish these categories of images. Therefore, we propose a human-in-the-loop solution to assist the network to better classify the images by leveraging human knowledge. To achieve this, we project the high-dimensional latent space trained by the network onto a two-dimensional workspace. The users can interactively modify the projected coordinates of inputs on the workspace using our designed tools, then the modified information will be fed back to the network to fine-tune it, which in turn affects the network's classification results, thereby improving the accuracy of network classification.
Jiafu Wei, Haoran Xie 0002, Chia-Ming Chang 0003, Xi Yang 0017
IJCAI4
2022 Data-Driven Multi-modal Partial Medical Image Preregistration by Template Space Patch Mapping
Ding Xia, Xi Yang 0017, Oliver van Kaick, Taichi Kin, Takeo Igarashi
MICCAI (6)2
2020 IntrA: 3D Intracranial Aneurysm Dataset for Deep Learning
abstract
Medicine is an important application area for deep learning models. Research in this field is a combination of medical expertise and data science knowledge. In this paper, instead of 2D medical images, we introduce an open-access 3D intracranial aneurysm dataset, IntrA, that makes the application of points-based and mesh-based classification and segmentation models available. Our dataset can be used to diagnose intracranial aneurysms and to extract the neck for a clipping operation in medicine and other areas of deep learning, such as normal estimation and surface reconstruction. We provide a large-scale benchmark of classification and part segmentation by testing state-of-the-art networks. We also discuss the performance of each method and demonstrate the challenges of our dataset. The published dataset can be accessed here: https://github.com/intra2d2019/IntrA.
Xi Yang 0017, Ding Xia, Taichi Kin, Takeo Igarashi
CVPR1
2020 G2MF-WA: Geometric multi-model fitting with weakly annotated data
abstract
In this paper we address the problem of geometric multi-model fitting using a few weakly annotated data points, which has been little studied so far. In weak annotating (WA), most manual annotations are supposed to be correct yet inevitably mixed with incorrect ones. Such WA data can naturally arise through interaction in various tasks. For example, in the case of homography estimation, one can easily annotate points on the same plane or object with a single label by observing the image. Motivated by this, we propose a novel method to make full use of WA data to boost multi-model fitting performance. Specifically, a graph for model proposal sampling is first constructed using the WA data, given the prior that WA data annotated with the same weak label has a high probability of belonging to the same model. By incorporating this prior knowledge into the calculation of edge probabilities, vertices (i.e., data points) lying on or near the latent model are likely to be associated and further form a subset or cluster for effective proposal generation. Having generated proposals, o-expansion is used for labeling, and our method in return updates the proposals. This procedure works in an iterative way. Extensive experiments validate our method and show that it produces noticeably better results than state-of-the-art techniques in most cases.
Chao Zhang 0030, Xuequan Lu, Katsuya Hotta, Xi Yang 0017
Comput. Vis. Media4
2017 Interactive visualization of assembly instruction for stone tools restoration
abstract
3D exploded views have been widely used for assembly instructions in many fields but have seldom been used in archaeology. For studying stone tools, relics are repeatedly assembled using indistinct traditional illustrations. We apply this powerful presentation technique on the stone tool models, and study algorithms for point cloud data. In addition to presenting principles for the restoration of stone tools, we designed our system based on archaeological rules. In this paper, we propose an interactive visualization method of assembly instruction using 3D visualization technology to assist in the efficient restoration of stone tools.
Xi Yang 0017, Katsutsugu Matsuyama, Kouichi Konno
PacificVis1