VLDB 2026 Research / reviewers in the wild / expert
Xiaoyang Mao
dblp:55/3814
· DBLP profile ↗
130ranked-venue papers
9as first author
58since 2021 · last 2026
0000-0001-9531-3197ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 106 · 7 first-author · 46 since 2021Human-computer interaction and ubiquitous computing · 54 · 4 first-author · 23 since 2021Artificial intelligence and machine learning · 11 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Neighborhood constrained attention for lightweight image super-resolutionabstractIn recent years, to improve image super-resolution performance, several studies have explored integrating convolutional modules with vision transformers (ViTs) to enhance the local feature modeling of ViTs. However, these hybrid approaches often introduce inconsistencies in feature representation, redundant information, and an increased number of parameters, ultimately limiting both performance and computational efficiency. To overcome these challenges, we propose a novel neighborhood constrained attention (NCA) mechanism that enables transformers to effectively capture both local and global features without requiring additional convolutional modules. Specifically, we first divide the window into a set of grids, treating them as local features, and then explore both intra- and inter-relationships within and across these local features, using them as constraints to refine window attention. Furthermore, instead of relying on averaging or other heuristic schemes for assigning labels to local features, we combine them through a linear transformation, ensuring label accuracy and uniqueness. Extensive experiments demonstrate that the proposed NCA not only outperforms other state-of-the-art lightweight approaches on public benchmark datasets but also excels in engineering image datasets, such as automated defect detection and product quality inspection, while requiring fewer parameters and lower computational costs. Notably, compared to 4 SwinIR-light (SwinIR: Image Restoration Using Swin Transformer), NCA achieves an average performance gain of 0.28 dB across five public test sets while reducing network parameters by 27% and computational complexity (floating point operations, FLOPs) by 30%. Code and models are obtainable at https://github.com/hms-source/NCA . Zhenyang Zhu, Xiaoyang Mao |
Eng. Appl. Artif. Intell. | 3 |
| 2026 | A plug-and-play intra-class variance suppression framework for industrial anomaly detectionabstractAnomaly detection methods leveraging unsupervised learning are expected to find broad application across diverse sectors, especially in inspecting defects of industrial products. This potential is largely due to their resilience against the unpredictability of anomaly types and the imbalances of learning data across classes. Central to these methods is the premise that feature extractors or image reconstructors, when trained solely on normal data, are incapable of fully replicating the features or inputs of anomalous data. As a result, anomalies could be detected by thresholding the deviations in the extracted features or the reconstructed outputs. However, finding an optimal threshold that effectively separates anomalous from normal data remains a substantial challenge in real-world scenarios. The inherent variability within normal data itself is a significant factor contributing to this challenge. In this study, we introduce a simple yet powerful intra-class variance suppression framework that enables anomaly detection models to suppress intra-class variability by learning compact representations of normal data. We evaluate the proposed framework on three established unsupervised anomaly detection paradigms, namely generative adversarial learning, knowledge distillation, and reverse distillation. Experiments are conducted on multiple benchmark datasets, including handwritten digit images, natural object images, industrial anomaly detection benchmarks, and two additional real-world industrial datasets. The results demonstrate that the proposed framework consistently improves anomaly detection and localization performance, particularly in practical industrial quality inspection scenarios. Yixuan Ju, Prawit Buayai, Gangyong Jia, Xiaoyang Mao |
Eng. Appl. Artif. Intell. | 5 |
| 2026 | Memory-efficient divide-and-conquer attention for lightweight image super-resolutionabstractRecently, transformer-based methods have achieved significant progress in lightweight image super-resolution (SR). However, most of these approaches primarily aim to improve either inference speed or reconstruction quality, while overlooking memory consumption, thereby limiting their practicality on resource-constrained devices. In this paper, we propose a memory-efficient divide-and-conquer attention (MEDCA) for SR, which substantially reduces memory usage while achieving notable improvements in reconstruction performance and competitive inference speed. To address the high memory and space complexity of standard window-based self-attention (WSA), MEDCA adopts a divide-and-conquer strategy. Specifically, the input features are first split into multiple subspaces along the channel dimension. Each subspace is further partitioned into multiple windows, which are then evenly divided into two parts using distinct asymmetric strategies. Self-attention is independently applied to each part, and the outputs are aggregated to form the final representation. Compared with traditional WSA methods such as SwinIR, MEDCA reduces the space complexity within each subspace by half. Furthermore, we design multiple asymmetric partitioning strategies that allow the model to extract features from a broader spatial context, thereby enabling it to capture richer spatial information and enhance its representation capacity. Extensive experiments demonstrate that MEDCA significantly reduces memory consumption while outperforming existing lightweight state-of-the-art methods and maintaining competitive inference speed across multiple public benchmark datasets. In particular, compared with the state-of-the-art method HiT-SRF, MEDCA improves the average performance by 0.12dB across five public test sets, while maintaining comparable inference time and requiring only 29.7% of the memory used by HiT-SRF. The code and models are provided at https://github.com/hms-source/MEDCA. Zhenyang Zhu, Xiaoyang Mao |
Neural Networks | 3 |
| 2026 | A ViT-like unified multi-scale learning network for efficient image super-resolutionabstractRecently, vision transformer (ViT)-based super-resolution (SR) models have achieved strong reconstruction performance but suffer from substantial computational and memory overhead due to self-attention operations, limiting their applicability in resource-constrained scenarios. Although several CNN-based alternatives attempt to replicate the global modeling capability of ViTs, their limited receptive fields often struggle to achieve competitive performance. To address this challenge, we propose a ViT-like unified multi-scale learning network (UMLN) that achieves a favorable trade-off between reconstruction quality and computational efficiency. Instead of relying on memory-intensive self-attention, our ViT-like design adopts a unified multi-scale feature aggregation mechanism that generates a global weight matrix through efficient convolutional operations. This unified module enables cross-scale interaction and long-range dependency modeling while significantly reducing parameter redundancy and memory consumption compared to attention-based approaches. In addition, we design an efficient feature extraction module CGDConv that effectively captures both local textures and non-local contextual information. Extensive experiments demonstrate that our UMLN outperforms existing efficient SR methods on public benchmark datasets, matching the performance of state-of-the-art lightweight ViT-based approaches while reducing significantly computational cost. Notably, compared to the × 4 SRFormer-light, our UMLN-L achieves an average PSNR improvement of 0.14 dB across five public benchmark datasets, while requiring only 68% of the computational complexity (e.g., FLOPs). Code and models are obtainable at https://github.com/hms-source/UMLN . • We propose a unified multi-scale feature extraction strategy that shares a single processing module across different scales, effectively reducing parameters and improving efficiency. • We propose a ViT-like weighted fusion mechanism to aggregate multi-scale features, enabling more effective global context modeling and overcoming the performance limitations of existing methods. • We design an effective feature extraction module, CGDConv, to capture both high- and low-frequency information within the network. • Extensive experiment results on public benchmark datasets demonstrate that the proposed UMLN achieves competitive performance in terms of model complexity, computational efficiency, and reconstruction quality compared to state-of-the-art approaches. Zhenyang Zhu, Xiaoyang Mao |
Pattern Recognit. | 3 |
| 2026 | The Effects of Visual-Olfactory Interactions With Moving Particles on EEG-Based Emotional Classification in AR EnvironmentsabstractCross-modal perception, the integration of information from multiple senses, plays a critical role in shaping emotional experiences. This study examines the interactions between visual and olfactory stimuli and their effects on emotional responses, a topic rarely addressed in prior research. Experiments employed five distinct visual stimulation methods that were combined with olfactory stimuli. Participants' emotional responses were assessed via surveys and electroencephalography (EEG) signal analysis. The study varied the color and movement direction of augmented particles to investigate their impact on EEG signals and emotional states. The findings demonstrated significant differences in emotional state classification under the influence of visual-olfactory interactions. Specifically, with backward-moving particles with matching colors (M4), classification accuracy was comparable to that of unimodal olfactory conditions (M1). Other visual stimuli generally caused confusion in classifying emotional responses. The increased valence ratings for pleasant aromas across all visual conditions did not consistently align with EEG-based classification results, suggesting that visual stimuli may introduce complexities into neural signals. These results highlight the intricate dynamics of multisensory interactions, emphasizing the role of visual stimuli in modulating emotional responses. The findings also suggest the potential of visual-olfactory interactions in developing augmented reality (AR) systems. By aligning visual and olfactory cues, AR environments can enhance the user experience and create immersive emotional landscapes, leading to applications for mood modulation and stress relief. This study underscores the relevance of multisensory integration in advancing emotion analysis and affective computing. Ye-Ji Jin, Xiaoyang Mao, Masaki Omata, Won-Du Chang |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2026 | A Swin-style shifted pooling cross-aggregation network for efficient image super-resolutionabstractAbstract Swin transformer-based methods have achieved impressive performance in image super-resolution (SR) due to their ability to effectively model long-range spatial dependencies. However, the core component, window-based self-attention (WSA), introduces considerable computational overhead, which limits their applicability on resource-constrained devices. To address these issues, we propose a Swin-style shifted pooling cross-aggregation network (SPCAN) for image SR, which achieves high computational efficiency while maintaining excellent reconstruction quality. Specifically, we adopt max pooling-based downsampling as a lightweight alternative to WSA for extracting low-frequency features and introduce a shifted pooling mechanism that emulates the shifted window strategy of Swin transformers within a convolutional neural network (CNN) framework. This mechanism is embedded within a cross-aggregation module to facilitate efficient inter-region feature interaction. Moreover, we generalize the pooling operation from square to rectangular regions to enhance the model’s ability to capture spatial dependencies across different orientations. Extensive experiments on public SR benchmarks demonstrate that the proposed method achieves competitive reconstruction accuracy while offering significantly better efficiency compared with existing state-of-the-art methods. The source code and pretrained models are available at: https://github.com/hms-source/SPCAN . Zhenyang Zhu, Xiaoyang Mao |
Vis. Comput. | 3 |
| 2025 | Latent Interpretation for Multi-view Face Synthesis Across GAN and Diffusion via Conditional Reconstruction
Yixuan Ju, Zhenyang Zhu, Xiaoyang Mao |
CGI (3) | 5 |
| 2025 | Semantic Compensation and Localization for Color Vision DeficiencyabstractColor Vision Deficiency (CVD) affects approximately 8 % of men and 0.5 % of women worldwide, creating barriers in color-coded visual interactions. Existing assistive solutions, such as image recoloring and pattern overlays, often fail to provide accurate color identification or precise location in complex scenes. This paper presents a novel framework that integrates vision-language models (VLMs) with open-vocabulary object detection to enhance color understanding for individuals with CVD. Our approach employs physiologically grounded CVD simulation algorithms to generate CVD-simulated images, then computes pixel-level difference maps between original and CVDsimulated images to identify perceptually challenging regions. This difference information guides the vision-language model's attention toward color-critical objects. After training on a specially constructed CVD dataset, the model generates precise color-aware descriptions targeting challenging areas. Our openvocabulary object detection component then provides accurate localization of key objects with color attributes. Experimental results demonstrate effective identification of visually confusing regions, accurate color-aware descriptions, and precise spatial location. This work provides a practical approach for developing inclusive visual assistance systems for the CVD community. Zhencheng Chen, Zhenyang Zhu, Xiaoyang Mao |
CW | 3 |
| 2025 | Cross-Domain Personal Identification Framework Based on Intraoral PhotographsabstractIn disaster scenarios, forensic experts typically rely on manual comparison between clinical intraoral photos and onsite photos to identify remains. However, this may introduce subjective bias and reduce overall efficiency. Moreover, on-site photos are generally more complex due to variations in illumination, longitudinal changes, and imaging angles. These challenges are further exacerbated by differences in capture devices, shooting conditions, and the inconsistent use of intraoral retractor. To figure out these challenges, this study proposes a deep learningbased intraoral photographs used identification approach which uses image domain translation. To validate the effectiveness of the proposed method, we constructed a multi-domain intraoral photo dataset comprising over$\mathbf{4, 0 0 0}$images from more than 100 individuals, covering both clinical and on-site domains. Through the image domain transfer technology, intraoral photographs in the clinic environment are simulated as images in the on-site environment, thereby improving the model's recognition ability for on-site images. Experimental results demonstrate that domain adaptation significantly enhances cross-domain identification performance. The results confirm the potential of the proposed approach to achieve accurate and efficient identity verification in complex, real-world forensic scenarios. Zhean Ma, Zhenfei Wang, Zhenyang Zhu, Kotaro Kubo, Xiaoyang Mao |
CW | 6 |
| 2025 | ClothMotion: Pose-Guided Temporal Consistent Garment Mask GenerationabstractWith the advancement of diffusion models and Transformer-based architectures, pose-guided human image animation has achieved remarkable progress. However, existing methods often treat garments as static textures attached to the body, ignoring the non-rigid nature of clothing deformation during motion. This modeling limitation leads to unrealistic artifacts such as skirt misalignment and sleeve discontinuity, especially in the presence of loose or dynamic garments. To address this issue, ClothMotion is introduced as an appearance flow-based approach that explicitly models the evolution of garment contours driven by pose changes. The method employs a multi-scale flow estimation framework to predict dense non-rigid displacement fields, which deform a reference garment mask according to the target pose sequence. The resulting warped mask serves as a strong geometric prior for a pose-aware decoder, enabling temporally consistent and spatially coherent segmentation across frames. To support training, a dynamic garment segmentation dataset is constructed. Pseudo-labels are generated on TikTok and UBC-Fashion videos using a DeepFashion2-pretrained detector. A global majority voting strategy and manual correction are applied to improve annotation consistency and accuracy. Experimental results show that ClothMotion significantly improves mask accuracy and garment stability, achieving 0.85 mIoU on real-world sequences. When integrated into existing animation frameworks such as DisCo, the method further enhances visual quality and structural consistency. These results highlight the importance of explicit contour modeling in pose-guided image animation and demonstrate improved controllability over dynamic garment behavior, offering a generalizable path toward realistic clothing synthesis. Huiyun Qiu, Zhenyang Zhu, Xiaoyang Mao |
CW | 4 |
| 2025 | Swin-WNet: Boundary-Aware Semantic Segmentation for Oral Squamous Cell CarcinomaabstractOral squamous cell carcinoma (OSCC) poses a significant threat to public health due to its severity and the laborintensive process of image analysis required by physicians. Intercellular bridges are bridge-like structures that connect adjacent cells and indicate the differentiation level of a cancer. Although intercellular bridges are known to disappear as differentiation decreases, pathologists and clinicians evaluate the presence of intercellular bridges to assess the degree of differentiation of cancer. While state-of-the-art (SOTA) deep learning methods, such as U-Net and its variants, perform well on uniform and clearly delineated objects (e.g., cell, lung, etc.), accurately segmenting intricate objects (e.g., intercellular bridge, retinal vessel) remains challenging due to their complex topologies, fine branches, and irregular morphological changes. This paper aims to propose a method that effectively utilize boundary information, particularly targeting intricate objects. Our approach, inspired by Swin-UNet, employs the Swin Transformer, comprising a feature encoder, two decoders (semantic decoder and boundary decoder) and an attention-guided fusion module to enhance the model's ability to segment intricate objects. By applying the constraints of the boundary decoder, the feature encoder's ability to encode structural information is enhanced without compromising semantic information extraction and representation. Furthermore, fusing the outputs of the boundary decoder and semantic decoder further strengthens the detail of structural information. To validate the generalizability of our method, we conducted comparative experiments on one private and one public dataset. The results demonstrate that our method outperforms SOTA methods. Zhenfei Wang, Zhenyang Zhu, Kunio Yoshizawa, Masahiro Toyoura, Naoki Oishi, Xiaoyang Mao |
CW | 6 |
| 2025 | PG-VTON: Front-And-Back Garment Guided Panoramic Gaussian Virtual Try-On With Diffusion ModelingabstractABSTRACT Virtual try‐on (VTON) technology enables the rapid creation of realistic try‐on experiences, which makes it highly valuable for the metaverse and e‐commerce. However, 2D VTON methods struggle to convey depth and immersion, while existing 3D methods require multi‐view garment images and face challenges in generating high‐fidelity garment textures. To address the aforementioned limitations, this paper proposes a panoramic Gaussian VTON framework guided solely by front‐and‐back garment information, named PG‐VTON, which uses an adapted local controllable diffusion model for generating virtual dressing effects in specific regions. Specifically, PG‐VTON adopts a coarse‐to‐fine architecture consisting of two stages. The coarse editing stage employs a local controllable diffusion model with a score distillation sampling (SDS) loss to generate coarse garment geometries with high‐level semantics. Meanwhile, the refinement stage applies the same diffusion model with a photometric loss not only to enhance garment details and reduce artifacts but also to correct unwanted noise and distortions introduced during the coarse stage, thereby effectively enhancing realism. To improve training efficiency, we further introduce a dynamic noise scheduling (DNS) strategy, which ensures stable training and high‐fidelity results. Experimental results demonstrate the superiority of our method, which achieves geometrically consistent and highly realistic 3D virtual try‐on generation. Shengwei Sang, Guojun Dai, Xiaoyang Mao, Wenhui Zhou 0001 |
Comput. Animat. Virtual Worlds | 5 |
| 2025 | Sufficient learning: mining denser high-quality pixel-level labels for edge detection
Wenya Yang, Wen Wu 0008, Xiuting Tao, Xiaoyang Mao |
Neural Comput. Appl. | 5 |
| 2025 | Boosting lightweight single image super-resolution via global prior featureabstractRecently, lightweight vision transformer (ViT)-based single image super-resolution (SISR) has gained significant attention. However, many existing lightweight methods struggle to achieve satisfactory performance due to the aggressive reduction in the number of parameters. Therefore, to improve the performance of lightweight networks, we propose a novel global feature prior self-attention network. First, conventional window-based self-attention methods typically apply attention mechanisms indiscriminately to all pixels within a window. This can lead to artifacts and texture blurring. To mitigate this issue, we leverage prior knowledge to identify texture-related pixels within the window and perform self-attention operations specifically on these pixels. Second, to enhance the network's ability to capture critical information and structural details, we introduce an efficient global feature extraction method. Finally, while transformers excel at capturing global features and low-frequency information, they often struggle with extracting local features and high-frequency information. Therefore, we integrate a local complementary module into the shift window attention to compensate for the transformer's shortcomings in extracting local and high-frequency features. Extensive experiments demonstrate that the proposed method outperforms all other state-of-the-art lightweight approaches. Code and models are obtainable at https://github.com/hms-source/GFPSAN. Zhenyang Zhu, Xiaoyang Mao |
Neural Networks | 3 |
| 2025 | Non-invasive estimation of Shine Muscat grape color and sensory evaluation from standard camera imagesabstractAbstract This study proposes a non-invasive method to estimate both color and sensory attributes of Shine Muscat grapes from standard camera images. First, we focus on color estimation by integrating a Vision Transformer (ViT) feature extractor with interquartile range (IQR)-based outlier removal. Experimental results show that our approach achieves 97.2% accuracy, significantly outperforming Convolutional Neural Network (CNN) models. This improvement underscores the importance of capturing global contextual information to differentiate subtle color variations in grape ripeness. Second, we address human sensory evaluation by collecting questionnaire responses on 13 attributes (e.g., “Sweetness,” “Overall taste rating”), each rated on a five-point scale. Because these ratings tend to cluster around midrange values (labels “2,” “3,” and “4”), we initially limit the dataset to the extreme labels “1” (“lowest grade”) and “5” (“highest grade”) for binary classification. Three attributes—“Overall color,” “Sweetness,” and “Overall taste rating”—exhibit relatively high classification accuracies of 79.9%, 75.1%, and 75.7%, respectively. By contrast, the other 10 attributes reach only 50%–66%, suggesting that subjective variations and limited visual cues pose significant challenges. Overall, the proposed approach demonstrates the feasibility of an image-based system that integrates color estimation and sensory evaluation to support more objective, data-driven harvest timing decisions for Shine Muscat grapes. Ryosuke Shimazu, Chee Siang Leow, Prawit Buayai, Xiaoyang Mao, Wan-Young Chung, Hiromitsu Nishizaki |
Vis. Comput. | 4 |
| 2025 | High similarity controllable face anonymization based on dynamic identity perception
Jiayi Xu 0002, Yixuan Ju, Xiaoyang Mao, Shanqing Zhang |
Vis. Comput. | 4 |
| 2025 | A content-aware image editing method for perceptual size restoration based on seam carvingabstractAbstract To alleviate the gap between the perceptual size of Thing of Interest (ToI) in a real scene and its appearance in a photograph, ToI region enlargement methods based on seam carving have been proposed in previous studies. However, state-of-the-art methods may suffer from issues such as a lack of automation and distortion caused by seam concentrations. In this study, we propose an image editing method for ToI enlargement based on seam carving. The proposed method incorporates the segment anything model to automatically generate ToI mask. To address the issue of seam concentrations, a wide-area energy strategy is introduced. Additionally, to improve the image quality of the enlarged ToI region, a super-resolution-based post-processing mechanism is put forward. Subjective experiments were conducted to evaluate the effectiveness of the proposed method. The experimental results suggest that the proposed method effectively enlarges the ToI region while suppressing distortion caused by seam concentrations. Zhenyang Zhu, Naohiko Ishikawa, Xiaoyang Mao |
Vis. Comput. | 3 |
| 2024 | A Semi-automatic Quality Assessment System for Capturing High-quality Fundus ImageabstractThe precision of medical diagnoses based on images is inextricably linked to the quality and clarity of the images. Poor image quality can impede accurate diagnosis and pose challenges for both physicians and machine learning algorithms in interpreting images. By automating the process of fundus image quality assessment during capture, we can ensure that only high-quality images are used, improving diagnostic accuracy. Accordingly, we propose a novel approach for automatically assessing the quality of fundus images using deep learning techniques. Our method incorporates retinal vessel segmentation into RGB images to create four-channel images and then trains a deep learning model on these images to identify the image focus and quality of the images. It can evaluate the quality of fundus images and request image recapture with the adjustment of specific parameters that are classified as poor quality. Our proposed method has the potential to improve the diagnostic accuracy and efficiency of retinal disease diagnosis, particularly in telemedicine settings. By automating the process of fundus image quality assessment, we can ensure that only high-quality images are used for diagnosis, thus improving diagnostic precision. It can serve as an efficient screening tool in the initial stage of acquiring high-quality fundus images. Asif Mohammed Arfi, Masahiro Toyoura, Kenji Kashiwagi, Satoshi Nishiguchi, Kentaro Go, Zhenyang Zhu, Xiaoyang Mao |
CW | 7 |
| 2024 | Tilt-Invariant Lemon Size Estimation Using RGB-D Camera ImagesabstractThe size of fruits significantly impacts their market value. For lemons, the size grade is determined by the cross-sectional diameter, necessitating the harvest of lemons that meet specific size criteria. Current manual measurement methods, involving metal rings, may degrade lemon quality and are labor-intensive. This study proposes a non-contact method for estimating lemon diameter using RGB-D images and deep learning. Our approach detects lemons and their tips, utilizing depth information and the position of the tip relative to the boundary of detected lemon mask to estimate size irrespective of the lemon’s tilt. With 2,038 images of indoor and outdoor green lemons, our method achieved a Mean Absolute Error (MAE) of 2.94 mm for fully visible lemons, though accuracy decreased with occlusions. These findings suggest that accurate, tilt-independent lemon size measurement is feasible in field conditions, providing a valuable tool for harvest support. Ayuna Dohi, Prawit Buayai, Ki-Ryong Kwon, Xiaoyang Mao |
CW | 4 |
| 2024 | Seam Carving-based Image Partial Enlargement Method for Perceptual Impression ReflectionabstractIn addressing the issue that the Thing of Interest (ToI) in a photograph appears significantly smaller than that being directly perceived from that in the physical world, image editing techniques have been developed to enlarge ToI while maintaining appearance of other areas. However, these methods can suffer from issues such as loss of spatial perception, distortions of objects, the lack of automation, etc. To address these issues, we propose a novel ToI enlarging method based on Seam Carving method. The proposed method is composed of seam insertion and deletion processing, and introduces wide area seam energy mechanism, which considers impact of inserted or deleted seams on adjacent areas, in addition to preservation to salient objects, image structure, gradient and mask energies. Moreover, the proposed method introduces the Segment Anything Model to automatically generate masks representing the ToI. To validate the effectiveness of the proposed method, we conducted two subjective evaluation experiments in this study. The experimental results demonstrate that the proposed method can produce resulting images that reflect perceptual impressions, and can preserve area other than ToI. Naohiko Ishikawa, Zhenyang Zhu, Jong-Nam Kim, Xiaoyang Mao |
CW | 4 |
| 2024 | Customer Activity Detection Using YOLOv8 and Status Order AlgorithmabstractThis paper presents a method to analyze table utilization and customer movements in fast food courts through video surveillance data using methods from machine learning. The system that has been proposed YOLOv8 for object detection, DeepSORT for object tracking and a custom status order algorithm to classify the table activities. A real-world implementation on existing infrastructure in SA such as the Tapah Southbound RSA (Rest and Service Area) Food Court were undertaken to address challenges posed by non-optimal camera angles, occlusion issues etc. Based on people (people) and objects (bowls, plates, bottles or cups), the system classifies table statuses to eat, drink, eat_drink sit-down empty. To improve the classification accuracy and to address this unstable detection issue, a status order hierarchy was introduce. This paper summarizes the system performance of occupancy rates, activity durations and customer patterns analysis while highlighting its limitations by potential embedding in started a business intelligence for food court operations. Consistent with our confirmation, we find high table utilization rates and clear behavioural patterns that can provide insights for service optimization or layout improvement. It also points to avenues for further improvement in data collection methods that cater to improved detection of nuanced activity under challenging visual settings. Ammar Zakaria, Ahmad Shakaff Ali Yeon, Syed Muhammad Mamduh, Latifah Kamarudin, Muhammad Reza Zainal Abidin, Retnam Visvanathan, Xiaoyang Mao, Norbazlan Mohd Yusof |
CW | 8 |
| 2024 | Versatile and Easy-to-Operate Grading System for High-Grade Table Grapes: Leveraging Deep Learning, Computer Vision, and IoTabstractGrapes, ranking among the top fruits globally, undergo essential grading processes to ensure quality and market readiness. However, conventional grading methods rely heavily on subjective human expertise, leading to inconsistencies and inefficiencies. To address this, we present a novel grape grading system integrating computer vision, artificial intelligence (AI), and IoT technologies. Our system utilizes deep neural networks (DNNs) based prediction model to accurately grade the grapes, and a sensing station equipped with cameras and weight sensors to capture images and weight data of grape bunches. Our research focuses on Shine Muscat grapes, a prominent variety in Japan. The system implements a multimodal grading approach that combines image and weight information. Our system offers portability, affordability, and operational versatility that prioritizes the safety of the grape bunch. Experimentation with different DNN architectures and input data configurations reveals the superiority of ResNet-18 for grading classification using multiple images from different angles, achieving $85.71 \%$ accuracy. On the other hand, when the weight information is included with the images, ResNet-50 performs the best with $82.86 \%$ accuracy. Additionally, we analyze mispredictions across grade classes, highlighting the model’s challenges in discerning subtle differences between closely ranked grades. Muhammad Faris Bin Kamarudzaman, Prawit Buayai, Yin Suan Tan, Latifah Kamarudin, Xiaoyang Mao |
CW | 5 |
| 2024 | Wearable Ear EEG Device for Emotion Recognition in Human-Robot InteractionabstractWearable electroencephalography (EEG) devices are emerging as crucial tools in human-robot interaction (HRI), enabling intuitive and effective communication between humans and robots. These devices non-invasively measure brain activity, providing real-time insights into a user’s mental state, intentions, and cognitive load. This paper explores the advancements and applications of wearable ear EEG technology in HRI, with a focus on detecting human intent, monitoring emotional and cognitive states, and delivering real-time feedback for adaptive robot behavior. The benefits of wearable EEG over traditional scalp EEG, such as enhanced user comfort, reduced setup time, and improved long-term wearability, are thoroughly examined. Additionally, the paper covers advanced applications, including emotion-aware adaptive interactions, neurofeedback training, and direct robot control via brain-computer interfaces (BCIs). A review of the latest advancements in device miniaturization, integration with other wearable sensors, and applications in virtual reality, gaming, and healthcare is provided. Future directions focus on expanding applications, enhancing human-AI collaboration, and improving accuracy and reliability in HRI. The studies discussed in this paper represent our state-of-the-art research across various fields and applications of wearable EEG systems. This study highlights the potential of wearable Ear-EEG devices to revolutionize human-robot interactions by enhancing customization, responsiveness, and overall efficacy. Ngoc-Dau Mai, Kentaro Go, Xiaoyang Mao, Wan-Young Chung |
CW | 3 |
| 2024 | Passive Fatigue Assessment in Augmented Reality Workspaces: Behavioral Cues Indicators for Workers with Intellectual DisabilitiesabstractThis study identifies challenges in user interface design for applying Augmented Reality (AR) technology to support work for individuals with intellectual disabilities. The main focus is on difficulty in self-assessing fatigue levels. The research methodology involved conducting experiments using HoloLens 2, analyzing changes in biometric information during fatigue, and examining the relationship between information display position and head orientation. Results indicate that changes in head height and orientation could potentially serve as fatigue indicators. In conclusion, while AR technology is effective in supporting work for individuals with intellectual disabilities, special considerations are necessary. Fatigue detection using biometric information and optimization of information display positions are crucial, and these findings may lead to safer and less burdensome use of AR. Kaishi Naito, Daisuke Inoue 0004, Prawit Buayai, Wan-Young Chung, Xiaoyang Mao |
CW | 5 |
| 2024 | Application of Super-Resolution (SR) for Thrips Detection and ClassificationabstractThis study presents an advanced automatic thrips counting and classification system, leveraging a novel SuperResolution (SR) technique, named KSVD_DR to enhance image analysis accuracy. We developed and validated detection models using a diverse dataset that included high-resolution scanned images and smartphone-captured images of blue and yellow traps, both with and without plastic wrap. This approach ensured robust performance across various real-world agricultural settings. The application of SR improved the detection accuracy from $66.5 \%$ to $\mathbf{8 9. 7} \%$, as measured by the mean Average Precision at $\mathbf{5 0 \%}$ Intersection over Union (mAP50). The overall testing accuracy achieved was $81.2 \%$, with specific accuracies of $80.3 \%$ for images with plastic wrap and $83.4 \%$ for those without confirming the system’s effectiveness in both laboratory and field conditions. Additionally, SR processing enhanced thrips classification accuracy from $58.5 \%$ to $\mathbf{6 5. 3 \%}$ across six distinct thrips classes, demonstrating its potential to refine species-specific identification. Future developments will focus on expanding outdoor data collection to validate and enhance system performance under varying environmental conditions and to improve detection accuracy at higher confidence levels. The study also aims to refine the classification model by incorporating more diverse data inputs and exploring advanced machine learning techniques, enhancing the ability to differentiate between thrips species effectively. Suit Mun Ng, Prawit Buayai, Latifah Kamarudin, Haniza Yazid, Xiaoyang Mao |
CW | 5 |
| 2024 | High Quality Color Estimation of Shine Muscat Grape Using Vision TransformerabstractCurrently, skilled farmers judge the ripeness of the Shine Muscat grape variety by looking at the color on the surface of the grapes. However, the color of Shine Muscat grapes does not change much as they grow, and there are individual differences in the way the color is perceived. Furthermore, the same color can look very different depending on the exposure to sunlight and shadows. Therefore, there is a need for a system that can quantitatively determine the color of Shine Muscat grapes to pass on the harvesting techniques of experienced farmers to amateurs and inexperienced farmers. This research aims to improve the accuracy of the color estimation of Shine Muscat grapes using deep learning. We propose a method to estimate the color of individual grapes using a color estimation model with a self-attention mechanism, from which the color of the whole bunch is estimated. A Vision Transformer model with a self-attention mechanism was found to improve the color estimation accuracy to $96.9 \%$. Furthermore, by eliminating outliers using the interquartile range, a color estimation accuracy of $97.2 \%$ could be achieved, demonstrating the effectiveness of the new color estimation model. Ryosuke Shimazu, Chee Siang Leow, Prawit Buayai, Koji Makino, Xiaoyang Mao, Hiromitsu Nishizaki |
CW | 5 |
| 2024 | AR Grape Thinning SupportabstractRecent advancements in deep neural networks (DNN) and augmented reality (AR) have improved the efficiency and automation of agriculture. This study proposes an AR grape thinning support system to assist in grape thinning operations. The proposed system uses DNN to predict grape berries that need to be thinned and uses the optical see-through Head-Mounted Display (HMD) HoloLens 2 to superimpose contour information over the real target berry, making it easy for users to identify the berry to be thinned. Additionally, the hand-tracking function of HoloLens 2 is utilized to monitor the thinning operation in real-time and provide voice instructions to improve work efficiency and user experience. The evaluation experiment compared three interfaces: “image only”, “image with contour overlay”, and “image with contour overlay and voice instructions”, evaluating using the metric of time taken to thin one grape cluster and the usability and user experience. The results showed that the image with contour overlay and voice instructions could significantly improve usability. Shun Tamura, Prawit Buayai, Won-Du Chang, Xiaoyang Mao |
CW | 4 |
| 2024 | Self-verificated UAED model for edge detectionabstractThis study enhances the UAED (Universal Adversarial Edge Detection) model to improve edge detection accuracy. We introduce cross-validated adaptive weights and a “majority voting” mechanism for historical predictions to optimize the model’s performance. The BSDS dataset is preprocessed by resizing images to $320 \times 480$ pixels and splitting them into $160 \times 240$ pixel sub-images. The “majority voting” mechanism generates more reliable labels, reducing noise. We also enhance the model by generating a mask to evaluate pixel edge reliability and dynamically adjusting loss weights based on edge prediction consistency. A key contribution is our alternating loss function strategy, which serves as a regularizer, preventing overfitting and improving generalization. Results show significant improvements in OIS and ODS scores, and a slight improvement in AP scores, validating the effectiveness of our method. Yizhe Yao, Jiaxuan Xie, Xiaoyang Mao |
CW | 4 |
| 2024 | Visual Coherence Face Anonymization Algorithm Based on Dynamic Identity PerceptionabstractIn the era of the meta-universe and the proliferation of personalized social networks, interactive behaviors like sharing personal and family photos pose an escalating risk of privacy breaches and identity exposure. A potential remedy lies in substituting real images with anonymized face images in public contexts. While existing face anonymization methods often replace substantial portions of face images, the resultant faces lack sufficient similarity to the originals. To address this, we propose a anonymization model leveraging saliency analysis to detect identity relevant facial region, preserving visual coherence and avoiding recognition by face recognition systems. Our model comprises two integral networks: the Dynamic Identity Perception Network (DIPNet) and the improved PSPNet. DIPNet, in particular, encompasses two vital sub-modules: the dynamic region perception module detects identity relevant region; the anonymization region control module governs the size of region through thresholding, thereby dominating the preservation of identity independent features and the degree of anonymization. The improved PSPNet produces high-quality identity anonymized faces. Experimental results demonstrate that our method yields realistic anonymized faces, retaining original features and deceiving face recognition systems, safeguarding privacy in the modern digital landscape. Shanqing Zhang, Yixuan Ju, Xiaoyang Mao, Jiayi Xu 0002 |
FG | 4 |
| 2024 | DAS-COD: Depth-Aware Camouflaged Object Detection via Swin TransformerabstractThe successful integration of depth information into salient object detection tasks has catalyzed research interest in depth-enhanced camouflaged object detection (COD) tasks. However, the challenges associated with acquiring depth information pose significant hurdles to this task, especially given the lack of RGB-D datasets tailored specifically for COD. Consequently, employing depth estimation techniques to generate pseudo-depth information emerges as a viable solution in the realm of depth-enhanced COD tasks. In this study, we propose an architecture, DAS-COD (Depth-Aware Swin Transformer COD), that integrates Swin Transformer model with depth estimation techniques for the purpose of camouflaged object detection. In particular, we use a dual-stream Swin Transformer backbone to extract feature maps from different modalities. These maps are then enhanced with a multi-modal feature enhancement module. Additionally, to address the inherent discrepancies between pseudo-depth maps and actual depth information, we incorporate an edge-aware module to significantly improve the accuracy of boundary delineation in the predicted outcomes. We tested our proposed method on three different COD datasets. Our results show that the model achieves state-of-the-art performance across these camouflaged object detection datasets. Chenye Lu, Min Tan 0005, Xiaoyang Mao, Zilin Xia |
SMC | 4 |
| 2024 | Image recoloring for color vision deficiency compensation using Swin transformerabstractAbstract People with color vision deficiency (CVD) have difficulty in distinguishing differences between colors. To compensate for the loss of color contrast experienced by CVD individuals, a lot of image recoloring approaches have been proposed. However, the state-of-the-art methods suffer from the failures of simultaneously enhancing color contrast and preserving naturalness of colors [without reducing the Quality of Vision (QOV)], high computational cost, etc. In this paper, we propose an image recoloring method using deep neural network, whose loss function takes into consideration the naturalness and contrast, and the network is trained in an unsupervised manner. Moreover, Swin transformer layer, which has long-range dependency mechanism, is adopted in the proposed method. At the same time, a dataset, which contains confusing color pairs to CVD individuals, is newly collected in this study. To evaluate the performance of the proposed method, quantitative and subjective experiments have been conducted. The experimental results showed that the proposed method is competitive to the state-of-the-art methods in contrast enhancement and naturalness preservation and has a real-time advantage. The code and model will be made available at https://github.com/Ligeng-c/CVD_swin . Ligeng Chen, Zhenyang Zhu, Wangkang Huang, Kentaro Go, Xiaoyang Mao |
Neural Comput. Appl. | 6 |
| 2024 | Boosting Deep Unsupervised Edge Detection via Segment Anything ModelabstractSegment anything model (SAM), a vision foundation network trained on a massive segmentation corpus, exhibits a superior boundary localization capability for nature images. This work aims to leverage such strengths to develop a deep unsupervised edge detection (UED) framework for alleviating the high reliance on dense labeling. However, applying vanilla SAM to edge detection fails to identify the salient edge cues but only the semantic boundary. This article introduces a lightweight adapter-tuning scheme to learn detailed edge information for filling the gap between boundary and edge, enabling a well-fitting even with limited training data. Moreover, considering the low-quality pseudo labels used in our UED framework, we propose two training strategies, adaptive progressive learning and gradient-guided pseudo label updating, to alleviate the impact of noisy labels from traditional UED methods. Extensive experiments demonstrate that our method achieves comparable results to state-of-the-art fully supervised edge detectors. Wenya Yang, Wen Wu 0008, Hongshuai Qin, Kangming Yan, Xiaoyang Mao |
IEEE Trans. Ind. Informatics | 6 |
| 2024 | Annotate less but perform better: weakly supervised shadow detection via label augmentation
Wen Wu 0008, Wenya Yang, Xiaoyang Mao |
Vis. Comput. | 5 |
| 2024 | SLIM: A transparent structurized self-learning interpolation method for super-resolution images
Xiaoyang Mao |
Vis. Comput. | 3 |
| 2024 | Personalized hairstyle and hair color editing based on multi-feature fusionabstractAbstract In the metaverse era, virtual design of hairstyle becomes very popular for personalized aesthetics. As hair design tasks can be decomposed into hair attribute editing and generation, the development of generative adversarial networks (GANs) has significantly prompted its development. The majority of the existing algorithms focus on transferring the overall hair region from one face to another, which ignore fine control over the color and geometric features. Furthermore, these algorithms may result in unnatural generation results. In this paper, we propose a hair modification framework that learns hairstyle information from a reference face mask and color information from a guidance face image. Firstly, the features of the input face image and reference images are extracted through a group of encoders, and then divided into feature vectors of coarse, medium, and fine levels. Secondly, multi-level feature vectors are fused in the latent space using attention-based modulation modules. Finally, the fused feature vector is passed through a StyleGAN generator to generate face images with specified hairstyle and hair color. Experimental results show that the proposed method can finely simulate the hairstyle transition between long and short hair under the constraint of the reference mask, and can produce realistic fusion effects in the hair-covered regions, such as ears, neck, and forehead. Various hair dyeing effects that adapt to personalized characteristics are demonstrated, as facial features including skin color and hair texture are preserved when transferring the hair color. Jiayi Xu 0002, Chenming Zhang, Weikang Zhu, Li Li 0014, Xiaoyang Mao |
Vis. Comput. | 6 |
| 2024 | Edge detection using multi-scale closest neighbor operator and grid partition
Wenya Yang, Xiaoyang Mao |
Vis. Comput. | 4 |
| 2024 | Fast image recoloring for red-green anomalous trichromacy with contrast enhancement and naturalness preservationabstractAbstract Color vision deficiency (CVD) is an eye disease caused by genetics that reduces the ability to distinguish colors, affecting approximately 200 million people worldwide. In response, image recoloring approaches have been proposed in existing studies for CVD compensation, and a state-of-the-art recoloring algorithm has even been adapted to offer personalized CVD compensation; however, it is built on a color space that is lacking perceptual uniformity, and its low computation efficiency hinders its usage in daily life by individuals with CVD. In this paper, we propose a fast and personalized degree-adaptive image-recoloring algorithm for CVD compensation that considers naturalness preservation and contrast enhancement. Moreover, we transferred the simulated color gamut of the varying degrees of CVD in RGB color space to CIE L*a*b* color space, which offers perceptual uniformity. To verify the effectiveness of our method, we conducted quantitative and subject evaluation experiments, demonstrating that our method achieved the best scores for contrast enhancement and naturalness preservation. Haiqiang Zhou, Wangkang Huang, Zhenyang Zhu, Kentaro Go, Xiaoyang Mao |
Vis. Comput. | 6 |
| 2023 | Seamless Image Editing for Perceptual Size Restoration Based on Seam Carving
Naohiko Ishikawa, Zhenyang Zhu, Jong-Nam Kim, Wan-Young Chung, Kentaro Go, Xiaoyang Mao |
CGI (1) | 6 |
| 2023 | Metamorphopsia Insepction System Based on Relevance FeedbackabstractPeople with metamorphopsia suffer from perceiving things in a distorted way. Various methods for examining metamorphopsia have been suggested in the current literature, with the most advanced techniques demonstrating the ability to yield quantitative measurements. However, these cutting-edge methods necessitate extended examination durations and impose challenging manipulations on patients. In this study, our objective is to enhance the time efficiency of the inspection process and alleviate the burden placed on the user. We propose a novel user-friendly quantitative inspection system which utilizes interactive reinforcement learning. Instead of having users directly operate the system, we ask them to evaluate the stimuli generated by the system. Based on their evaluations, the system gradually refines the deformation map representing the distortion perceived by the user. The reinforcement learning scheme is implemented using relevance feedback approach based on optimum-path forest classifier. To evaluate the effectiveness of the proposed system, subjective evaluation experiments involving simulated and real metamorphopsia participants were conducted in this study. The experimental findings reveal that, when compared to the state-of-the-art method, our proposed system yields comparable inspection out-comes while significantly reducing both the inspection duration and the mental workload. Zhenyang Zhu, Katsuhito Moritake, Kenji Kashiwagi, Masahiro Toyoura, Kentaro Go, Issei Fujishiro, Xiaoyang Mao |
SMC | 7 |
| 2023 | Augmented Aroma: The Influence of Augmented Particles' Movement and Color on Emotion during Olfactory PerceptionabstractThis study investigates the impact of visual augmentation on the olfactory system by analyzing users’ emotional responses. Augmented particles were presented using HoloLens through five methods, involving adjustment in color and movement, alongside six odors. Through the experiments with 30 participants, we discovered that augmented particles could intensify or reduce emotional reactions based on their colors and movement directions. Ye-Ji Jin, Masaki Omata, Won-Du Chang, Xiaoyang Mao |
VRST | 4 |
| 2023 | How to use extra training data for better edge detection?
Wenya Yang, Wen Wu 0008, Xiuting Tao, Xiaoyang Mao |
Appl. Intell. | 5 |
| 2023 | Make Segment Anything Model Perfect on Shadow DetectionabstractCompared to models pre-trained on ImageNet, the segment anything model (SAM) has been trained on a massive segmentation corpus, excelling in both generalization ability and boundary localization. However, these strengths are still insufficient to enhance shadow detection without additional training, and it raises the question: do we still need precise manual annotations to fine-tune SAM for high detection accuracy? This paper proposes an annotation-free framework for deep unsupervised shadow detection (USD) by leveraging SAM’s capabilities. The key lies in how to exploit the abilities acquired from a large-scale corpus and utilize them to improve downstream tasks. Instead of directly fine-tuning SAM, we propose a prompt-like tuning method to inject task-specific cues into SAM in a light-weight manner, namely ShadowSAM. This adaptation manner can ensure a good fitting when training data is limited. Moreover, considering that the pseudo labels used in our framework are generated by traditional USD approaches and may contain severe label noises, we propose an illumination and texture-guided updating strategy to selectively boost the quality of pseudo masks. To further improve the model’s robustness, we design a mask diversity index to establish easy-to-hard subsets for incremental curriculum learning. Extensive experiments on benchmark datasets (i.e., SBU, UCF, ISTD, and CUHK-Shadow) demonstrate that our unsupervised solution can achieve comparable performance to state-of-the-art (SOTA) fully supervised methods. Our code is available at this repository. Wen Wu 0008, Wenya Yang, Hongshuai Qin, Xiantao Wu, Xiaoyang Mao |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | A Pilot Study on the AR Interface Design for People with Intellectual DisabilitiesabstractWe have identified two problems for people with intellectual disabilities when using Augmented Reality (AR) devices. One is not noticing the presented information and the other is the difficulty in pressing buttons in the AR space. To solve the first problem, we propose two methods to attract a user's attention, one is to place the information window near the user's gaze point, and dynamically adapt the window based on eye-tracking data. The other is to blink the information window. As a solution to the second problem, we propose a method that detects the user's button-pressing gesture, and when the gesture is detected, the system presses the button on behalf of the user. The effectiveness of the proposed methods was validated through the user studies with 7 participants of varying degrees of intellectual disability involved. Kaishi Naito, Daisuke Inoue 0004, Prawit Buayai, Xiaoyang Mao |
CW | 4 |
| 2022 | Personalized Image Recoloring for Color Vision Deficiency CompensationabstractSeveral image recoloring methods have been proposed to compensate for the loss of contrast caused by color vision deficiency (CVD). However, these methods only work for dichromacy (a case in which one of the three types of cone cells loses its function completely), while the majority of CVD is anomalous trichromacy (another case in which one of the three types of cone cells partially loses its function). In this paper, a novel degree-adaptable recoloring algorithm is presented, which recolors images by minimizing an objective function constrained by contrast enhancement and naturalness preservation. To assess the effectiveness of the proposed method, a quantitative evaluation using common metrics and subjective studies involving 14 volunteers with varying degrees of CVD are conducted. The results of the evaluation experiment show that the proposed personalized recoloring method outperforms the state-of-the-art methods, achieving desirable contrast enhancement adapted to different degrees of CVD while preserving naturalness as much as possible. Zhenyang Zhu, Masahiro Toyoura, Kentaro Go, Kenji Kashiwagi, Issei Fujishiro, Tien-Tsin Wong, Xiaoyang Mao |
IEEE Trans. Multim. | 7 |
| 2022 | Appropriate grape color estimation based on metric learning for judging harvest timingabstractAbstract The color of a bunch of grapes is a very important factor when determining the appropriate time for harvesting. However, judging whether the color of the bunch is appropriate for harvesting requires experience and the result can vary by individuals. In this paper, we describe a system to support grape harvesting based on color estimation using deep learning. To estimate the color of a bunch of grapes, bunch detection, grain detection, removal of pest grains, and color estimation are required, for which deep learning-based approaches are adopted. In this study, YOLOv5, an object detection model that considers both accuracy and processing speed, is adopted for bunch detection and grain detection. For the detection of diseased grains, an autoencoder-based anomaly detection model is also employed. Since color is strongly affected by brightness, a color estimation model that is less affected by this factor is required. Accordingly, we propose multitask learning that uses metric learning. The color estimation model in this study is based on AlexNet. Metric learning was applied to train this model. Brightness is an important factor affecting the perception of color. In a practical experiment using actual grapes, we empirically selected the best three image channels from RGB and CIELAB (L*a*b*) color spaces and we found that the color estimation accuracy of the proposed multi-task model, the combination with “L” channel from L*a*b color space and “GB” from RGB color space for the grape image (represented as “LGB” color space), was 72.1%, compared to 21.1% for the model which used the normal RGB image. In addition, it was found that the proposed system was able to determine the suitability of grapes for harvesting with an accuracy of 81.6%, demonstrating the effectiveness of the proposed system. Tatsuyoshi Amemiya, Chee Siang Leow, Prawit Buayai, Koji Makino, Xiaoyang Mao, Hiromitsu Nishizaki |
Vis. Comput. | 5 |
| 2022 | Image recoloring for Red-Green dichromats with compensation range-based naturalness preservation and refined dichromacy gamut
Wangkang Huang, Zhenyang Zhu, Ligeng Chen, Kentaro Go, Xiaoyang Mao |
Vis. Comput. | 6 |
| 2022 | Enhancing edge indicator for visual field loss compensation for homonymous hemianopia patients
Keisuke Ichinose, Issei Fujishiro, Masahiro Toyoura, Kenji Kashiwagi, Kentaro Go, Xiaoyang Mao |
Vis. Comput. | 7 |
| 2022 | LineM: assessing metamorphopsia symptom using line manipulation task
Zhenyang Zhu, Masahiro Toyoura, Issei Fujishiro, Kentaro Go, Kenji Kashiwagi, Xiaoyang Mao |
Vis. Comput. | 6 |
| 2021 | Development of a Support System for Judging the Appropriate Timing for Grape HarvestingabstractThe color of grape bunches is a significant factor when harvesting grapes at the appropriate timing. Judging the suitable color for shipment requires experience and varies from one person to another. We herein describe a support system for grape harvesting based on color estimation. To estimate the color of a bunch of grapes, bunch detection, grain detection, removal of diseased grains, and color estimation should be performed. Models based on deep learning are employed for this series of processes. Since color is strongly affected by sunlight, we propose a multitask model that considers sunlight exposure to achieve a robust color estimation model that exhibits decreased sensitivity to sunlight. Our results show that the color estimation accuracy of the model is 76% when sunlight exposure is not considered and 81% when sunlight exposure is considered. In addition, we performed a practical field test of the developed harvest support system in an actual grape field. The results show that our support system can determine the appropriateness of grape harvest with an accuracy of 90%, demonstrating the effectiveness of the system. Tatsuyoshi Amemiya, Kodai Akiyama, Chee Siang Leow, Prawit Buayai, Koji Makino, Xiaoyang Mao, Hiromitsu Nishizaki |
CW | 6 |
| 2021 | End-to-End Inflorescence Measurement for Supporting Table Grape Trimming with Augmented RealityabstractInflorescence trimming is a crucial process to produce high-quality table grapes. It can eliminate nutrient competition in a bunch and makes it less vulnerable to disease development. After trimming, the remaining part of the inflorescence should have a target length decided by the grape variety. This is challenging for novice farmers because of the time constraint. The farmer needs to finish trimming the inflorescence before the berries develop. This paper proposes a novel end-to-end inflorescence length measurement method for supporting a trimming process with augmented reality technology. The proposed technique makes use of the state-of-the-art deep neural network model for detecting the inflorescence area, as well as the scissors from the images captured with a camera installed on an optical see-through head-mounted display. A new algorithm is designed to estimate the length of the remaining inflorescence with the screw of the scissors loop as the calibrator. The estimated length is then visualized on the head-mounted display to support the farmer in performing the trimming correctly and efficiently. The experiment, conducted with real inflorescence trimming tasks, shows that the mean absolute error of the length estimation is only 0.19 cm, which is small enough for use in real applications. Prawit Buayai, Kabin Yok-In, Daisuke Inoue 0004, Chee Siang Leow, Hiromitsu Nishizaki, Koji Makino, Xiaoyang Mao |
CW | 7 |
| 2021 | Supporting Vine Vegetation Status Observation Using ARabstractAugmented reality (AR) is a technology that expands information by superimposing digital information, such as virtual objects, on the real world using smartphones, smart glasses, and head-mounted displays. It is used in a variety of situations. In this paper, we propose a system that allows vine farmers to investigate effectively the vegetation condition of trellising-style vineyards using a head mounted display and AR technology. The experiment results show that by using a hybrid navigation approach include showing the whole vineyard in a small window and showing the details only, when necessary, the proposed system enable the farmers to move to a location with concern accurately. Daisuke Inoue 0004, Prawit Buayai, Hiromitsu Nishizaki, Koji Makino, Xiaoyang Mao |
CW | 5 |
| 2021 | Eye-Tracker-Free Compensation for MetamorphopsiaabstractMetamorphopsia is a symptom caused by abnormalities in the retina, and people with metamorphopsia experience distortions in their field of view. To compensate for this disorder, computer-based distortion assessment and compensation methods have been proposed. In the state-of-the-art method, compensation is carried out by dynamically deforming the image with a manipulation map, which is bound to the symptom of an individual user with metamorphopsia, according to the user's gaze data captured by an eye tracker. However, this method suffers from the instability and inaccuracy of the eye tracker, which leads to a poor compensation effect. In this paper, we propose a novel method for metamorphopsia compensation without using an eye tracker. The proposed method generates a compensation image by simultaneously applying multiple manipulation maps to the input image. To evaluate the effectiveness of the proposed method, a preliminary subjective experiment using a reading task and involving 10 participants with normal vision was conducted. The evaluation results show that the proposed method can compensate for visual distortion during reading tasks. Katsuhito Moritake, Zhenyang Zhu, Masahiro Toyoura, Kentaro Go, Kenji Kashiwagi, Issei Fujishiro, Xiaoyang Mao |
CW | 7 |
| 2021 | Augmented Reality based Support System using Three-Dimensional-Printed Model with Accelerometer at Medical Specimen MuseumsabstractMedical and nursing students often use medical specimens to enrich their knowledge of anatomy and physiology. Most medical specimens exhibited at the medical specimen museum are in a sliced form; therefore, visitors can only view the planar structure. However, it is important to understand the three-dimensional (3D) structure of the human body in medical learning. In this study, we proposed an augmented reality (AR)-based system that presents a virtual object and realized a virtual hands-on exhibit using a 3D-printed model of the human anatomy by displaying the positional relationship between the 3D-printed model and the specimen, using an accelerometer embedded in the 3D-printed model. The system successfully estimated the posture of the 3D-printed model of the human brain from the positional values measured by the sensor and matched it with the posture of the virtual object. The applicability of the proposed system was evaluated by conducting a questionnaire survey. As per the feedback from the visitors, they were able to operate the AR information by moving the 3D-printed model in different directions, making it easier to understand the actual 3D structure of the human body. Atsushi Sugiura, Toshihiro Kitama, Xiaoyang Mao |
CW | 3 |
| 2021 | Fast contrast and naturalness preserving image recolouring for dichromats
Zhenyang Zhu, Kentaro Go, Masahiro Toyoura, Xiaoyang Mao |
Comput. Graph. | 6 |
| 2021 | Deep-based Self-refined Face-top CoordinationabstractFace-top coordination, which exists in most clothes-fitting scenarios, is challenging due to varieties of attributes, implicit correlations, and tradeoffs between general preferences and individual preferences. We present a Deep-Based Self-Refined (DBSR) system to simulate face-top coordination based on intuition evaluation. To this end, we first establish a well-coordinated face-top (WCFT) dataset from fashion databases and communities. Then, we use a jointly trained CNN Deep Canonical Correlation Analysis (DCCA) method to bridge the semantic face-top gap based on the WCFT dataset to deal with general preferences. Subsequently, an irrelevance-based Optimum-path Forest (OPF) method is developed to adapt the results to individual preferences iteratively. Experimental results and user study demonstrate the effectiveness of our method. Xiaoyang Mao, Mengdi Xu, Xiaogang Jin 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2021 | Adaptive semantic attribute decoupling for precise face image editing
Yixuan Ju, Xiaoyang Mao, Jiayi Xu 0002 |
Vis. Comput. | 3 |
| 2021 | Matching a composite sketch to a photographed face using fused HOG and deep feature models
Jiayi Xu 0002, Xinying Xue, Yitiao Wu, Xiaoyang Mao |
Vis. Comput. | 4 |
| 2021 | Image recoloring for color vision deficiency compensation: a surveyabstractAbstract People with color vision deficiency (CVD) have a reduced capability to discriminate different colors. This impairment can cause inconveniences in the individuals’ daily lives and may even expose them to dangerous situations, such as failing to read traffic signals. CVD affects approximately 200 million people worldwide. In order to compensate for CVD, a significant number of image recoloring studies have been proposed. In this survey, we briefly review the representative existing recoloring methods and categorize them according to their methodological characteristics. Concurrently, we summarize the evaluation metrics, both subjective and quantitative, introduced in the existing studies and compare the state-of-the-art studies using the experimental evaluation results with the quantitative metrics. Zhenyang Zhu, Xiaoyang Mao |
Vis. Comput. | 2 |
| 2020 | Using an Eye Tracking Device to Discriminate Different Symptoms in GlaucomaabstractWorldwide, 79.6 million people experience glaucoma, which can cause visual field loss to the individuals. Visual field examination plays an important role in the detection of glaucoma. However, current visual field examination approaches have disadvantages, such as requirements for expensive equipment, long testing time, location restrictions, and more. In the present study, we propose an assessment based on saccadic reaction time (SRT) to overcome the issues in the existing approaches. To confirm the effectiveness of our method, we simulated different stages of glaucoma using a visual field defect simulation system. The result of the visual experiment showed that the discrimination method using SRT can distinguish different symptoms with less testing time. Lina Chen, Kentaro Go, Yuichiro Kinoshita, Kenji Kashiwagi, Masahiro Toyoura, Issei Fujishiro, Xiaoyang Mao |
CW | 7 |
| 2020 | Visual Field Loss Compensation for Homonymous Hemianopia Patients Using Edge IndicatorabstractHomonymous hemianopia (HH) is one kind of visual field defect that the right or left half of the visual fields of both eyes are missing. HH is caused by damage to the neural pathways of the brain and a complete recovery is usually difficult. This study proposes a new information compensation method for HH patients using optical-see-through head mounted display. To avoid causing the occlusion, the proposed technique uses indicators placed at the boundary between the lost and remaining sides of visual field to notify patients about the changes in the lost side. Experiment involving simulated HH participants was conducted to verify the effectiveness of the proposed method and compare with the existing method using overlaid overview window. Objective and subjective evaluation results show that the proposed method can partially alleviate the occlusion problem of overlaid overview window approach and has lower physical and mental demand. Keisuke Ichinose, Issei Fujishiro, Kenji Kashiwagi, Xiaoyang Mao, Masahiro Toyoura, Kentaro Go |
CW | 4 |
| 2020 | Different Eye Movement Patterns on Simulated Visual Field Defects in a Video-watching TaskabstractVisual field defects (VFD) can be caused by a variety of conditions. Checking and tracking the progression of VFD is an important part of an eye assessment. Although the use of standard automatic perimetry (SAP) is very popular for VFD diagnosis, it limits the population because of its high requirement for patients. We used a video-watching task as a replacement modality, which precludes the long period of fixation and uses the on-screen gaze to replace the button response. We developed a simulation system to mimic the different types of VFD in people with a normal pattern.We hypothesize that patients with VFD need more eye movement to compensate for the unseen area. We proposed a metric that indicates the gross eye movements toward a specific direction and found a significant difference between the VFD and normal pattern. Furthermore, we found videos that show the unique eye movement pattern in different eye conditions. Changtong Mao, Kentaro Go, Yuichiro Kinoshita, Kenji Kashiwagi, Masahiro Toyoura, Issei Fujishiro, Jianjun Li 0001, Xiaoyang Mao |
CW | 8 |
| 2020 | Evaluation of Color Vision Compensation Algorithms for People with Varying Degrees of Color Vision DeficiencyabstractPeople with color vision deficiency (CVD) may have difficulty in discriminating colors. To improve their color perception, several compensation methods have been proposed which considered naturalness maintenance and contrast emphasis. All these methods are based on the simulation model of severe CVD and hence it is not clear whether are also effective for people with light CVD. In this paper, we conduct subjective study to evaluate the effectiveness of these methods for people with varying degrees of CVD. Zhenyang Zhu, Masahiro Toyoura, Xiaoyang Mao |
CW | 5 |
| 2020 | AffectI: A Game for Diverse, Reliable, and Efficient Affective Image AnnotationabstractAn important application of affective image annotation is affective image content analysis, which aims to automatically understand the emotion being brought to viewers by image contents. The so-called subjective perception issue, i.e., different viewers may have different emotional responses to the same image, makes it difficult to link image features with the expected perceived emotion. Due to the ability to learn features, recent deep learning technologies have opened a new window on affective image content analysis, which has led to a growing demand for affective image annotation technologies to build large reliable training datasets. This paper proposes a novel affective image annotation technique, AffectI, for efficiently collecting diverse and reliable emotional labels with the estimate emotion distribution for images based on the concept of Game With a Purpose (GWAP). AffectI features three novel mechanisms: a selection mechanism for ensuring all emotion words being fairly evaluated for collecting diverse and reliable labels; an estimation mechanism for estimating the emotion distribution by aggregating partial pairwise comparisons of the emotion words for collecting the labels effectively and efficiently; an incentive mechanism shows the comparison between current player and her opponents as well as all past players to promote the interest of players and also contributes the reliability and diversity. Our experimental results demonstrate that AffectI is superior to existing methods in terms of being able to collect more diverse and reliable labels. The advantage of using GWAP for reducing the frustration of evaluators was also confirmed through subjective evaluation. Xingkun Zuo, Jiyi Li, Qili Zhou, Jianjun Li 0001, Xiaoyang Mao |
ACM Multimedia | 5 |
| 2020 | Enhancing visual performance of hemianopia patients using overview window
Issei Fujishiro, Kentaro Go, Masahiro Toyoura, Kenji Kashiwagi, Xiaoyang Mao |
Comput. Graph. | 6 |
| 2019 | Eyes-Free Text Entry with EdgeWrite Alphabets for Round-Face SmartwatchesabstractSmall wearable information devices such as smart-watches are available everywhere. A user can input text in these devices while doing other tasks. However, eyes-free typing, that is, without looking at the input space, on these devices remains difficult to execute. This paper describes our project on developing eyes-free text input systems. Not only an auditory feedback feature was developed for the EdgeWrite text input system for round-face smartwatches, but also a simple template-based gesture recognition engine was implemented. The experimental results show that the developed system maintains reasonable text input speed and error rate for eyes-free under standing and walking conditions. Kentaro Go, Mei Kikawa, Yuichiro Kinoshita, Xiaoyang Mao |
CW | 4 |
| 2019 | Visual Assessment of Distorted View for Metamorphopsia Patient by Interactive Line ManipulationabstractThe number of individuals with Age-related Macular Degeneration (AMD) is rapidly increasing. One of the main symptoms of AMD is "metamorphopsia," or distorted vision, which not only makes it difficult for individuals with AMD to do detailed-oriented tasks but also makes sufferers more vulnerable to certain risks in day-to-day life. Traditional clinical approaches to assess metamorphopsia have lacked mechanisms for quantifying the degree of distortion in space, making it impossible to know exactly how individuals with the condition see things. This paper proposes a new method for quantifying distortion in space and visualizing AMD patients' distorted views via line manipulation. By visualizing the distorted views stemming from metamorphopsia, the method gives doctors and others an intuitive picture of how patients see the world and thereby enables a broad range of options for treatment and support. Hiromichi Ichige, Masahiro Toyoura, Kentaro Go, Kenji Kashiwagi, Issei Fujishiro, Xiaoyang Mao |
CW | 6 |
| 2019 | Composite Sketch Recognition Using Multi-scale Hog Features and Semantic AttributesabstractComposite sketch recognition belongs to heterogeneous face recognition research, which is of great important in the field of criminal investigation. Because composite face sketch and photo belong to different modalities, robust representation of face feature cross different modalities is the key to recognition. Considering that composite sketch lacks texture details in some area, using texture features only may result in low recognition accuracy, this paper proposes a composite sketch recognition algorithm based on multi-scale Hog features and semantic attributes. Firstly, the global Hog features of the face and the local Hog features of each face component are extracted to represent the contour and detail features. Then the global and detail features are fused according to their importance at score level. Finally, semantic attributes are employed to reorder the matching results. The proposed algorithm is validated on PRIP-VSGC database and UoM-SGFS database, and achieves rank 10 identification accuracy of 88.6% and 96.7% respectively, which demonstrates that the proposed method outperforms other state-of-the-art methods. Xinying Xue, Jiayi Xu 0002, Xiaoyang Mao |
CW | 3 |
| 2019 | Computational Alleviation of Homonymous Visual Field Defect with OST-HMD: The Effect of Size and Position of Overlaid Overview WindowabstractVisual field defect (VFD) refers to a symptom in which a patient loses part of his/her field of view (FoV). Medical therapy can halt the progression of VFD, but complete recovery is impossible. In this paper, we propose a computational method for alleviating the restricted FoV with an optical see-through head-mounted display (OST-HMD), where an overview scene captured by the installed camera is overlaid on the persisting FoV. Since the overview window occludes with the real world scene, there is a trade-off between the augmented contextual information and the local unscreened information. We hypothesized that such a trade-off can be resolved by taking into consideration the size of the overview window and its displacement from the center of the unimpaired FoV. We therefore conducted an empirical evaluation through a Whac-A-Mole type of task with ten VFD-imitative subjects, where three sizes of an overview window with a fixed aspect ratio and seven positions in terms of elevation and azimuth were used combinatorially on an OST-HMD to find the best size and position of the overview window. It was statistically proven that for left-sided homonymous VFD-imitative subjects, the performance of the task was better when the medium-sized overview window was placed in the lower right position. The obtained result can legitimate default settings for the proposed VFD alleviation method. Kentaro Go, Kenji Kashiwagi, Masahiro Toyoura, Xiaoyang Mao, Issei Fujishiro |
CW | 5 |
| 2019 | Pencil Drawing Video Rendering Using Convolutional NetworksabstractAbstract Traditional pencil drawing rendering algorithms when applied to video may suffer from temporal inconsistency and shower‐door effect due to the stochastic noise models employed. This paper attempts to resolve these problems with deep learning. Recently, many research endeavors have demonstrated that feed‐forward Convolutional Neural Networks (CNNs) are capable of using a reference image to stylize a whole video sequence while removing the shower‐door effect in video style transfer applications. Compared with video style transfer, pencil drawing video is more sensitive to the inconsistency of texture and requires a stronger expression of pencil hatching. Thus, in this paper we develop an approach by combining a latest Line Integral Convolution (LIC) based method, specializing in realistically simulating pencil drawing images, with a new feed‐forward CNN that can eliminate the shower‐door effect successfully. Taking advantage of optical flow, we adopt a feature‐map‐level temporal loss function and propose a new framework to avoid the temporal inconsistency between consecutive frames, enhancing the visual impression of pencil strokes and tone. Experimental comparisons with the existing feed‐forward CNNs have demonstrated that our method can generate temporally more stable and visually more pleasant pencil drawing video results in a faster manner. Dingkun Yan, Yun Sheng, Xiaoyang Mao |
Comput. Graph. Forum | 3 |
| 2019 | Naturalness- and information-preserving image recoloring for red-green dichromatsabstractMore than 100 million individuals around the world suffer from color vision deficiency (CVD). Image recoloring algorithms have been proposed to compensate for CVD. This study has proposed a new recoloring algorithm to make up shortages of contrast enhancement and naturalness preservation of the state-of-the-art methods. The recoloring task is formulated as an optimization problem that is solved by using the colors in a simulated CVD color space to maximize contrast and to preserve the original color as much as possible. In addition, the dominant colors are extracted for recoloring. They are then propagated to the whole image so that the optimization problem could be solved at a reasonable cost independent of the image size. In the quantitative evaluation, the results of the proposed method are competitive with those of the best existing method. The evaluation involving subjects with CVD demonstrates that the proposed method outperforms the state-of-the-art method in preserving both the information and the naturalness of images. Zhenyang Zhu, Masahiro Toyoura, Kentaro Go, Issei Fujishiro, Kenji Kashiwagi, Xiaoyang Mao |
Signal Process. Image Commun. | 6 |
| 2019 | Generating Jacquard Fabric Pattern With Visual ImpressionsabstractWith jacquard fabric, designers can create complex patterns by freely defining the over-under relationships between warp yarns and weft yarns at each grid point or intersections in the fabric. Binary images are one way of representing the over-under relationships of warp and weft yarns at the grid points in a fabric pattern; an image requires an optimal number of warp-weft intersections-not too many, not too few-to produce a weave with both aesthetic and functional merits. This study proposes a method for generating jacquard fabric patterns that reproduce the visual impressions of given input images on jacquard fabric. Our method makes it possible to assign specific fabric dither masks to individual regions. These fabric dither masks can preserve the overall tone of the input image in the dithered image while appropriately controlling the number of warp-weft intersections in the pattern. As the fabric dither masks give users a considerable degree of freedom in defining texture frequencies and directional properties, the proposed method captures the visual impression of a given input image by enabling users to apply fabric dither masks to match textural features-either automatically or interactively. The former approach involves assigning the mask with the closest textural resemblance to each target region, ultimately producing a fabric pattern with a tone and texture that matches the input image. The interactive approach, meanwhile, provides the designer with an interactive interface for assigning masks. This allows designers to experiment with different expressions and track their progress visually, emphasizing or subduing specific areas at their own discretion. Masahiro Toyoura, Tetsuya Igarashi, Xiaoyang Mao |
IEEE Trans. Ind. Informatics | 3 |
| 2019 | Caricature synthesis with feature deviation matching under example-based framework
Masahiro Toyoura, Xiaoyang Mao |
Vis. Comput. | 3 |
| 2019 | Processing images for red-green dichromats compensation via naturalness and information-preservation considered recoloringabstractColor vision deficiency (CVD) is caused by anomalies in the cone cells of the human retina. It affects approximately 200 million individuals throughout the world. Although previous studies have proposed compensation methods, contrast and naturalness preservation have not been adequately and simultaneously addressed in the state-of-the-art studies. This paper focuses on red–green dichromats’ compensation and proposes a recoloring algorithm that combines contrast enhancement and naturalness preservation in a unified optimization model. In this implementation, representative color extraction and edit propagation methods are introduced to maintain global and local information in the recolored image. The quantitative evaluation results showed that the proposed method is competitive with state-of-the-art methods. A subjective experiment was also conducted and the evaluation results revealed that the proposed method obtained the best scores in preserving both naturalness and information for individuals with severe red–green CVD. Zhenyang Zhu, Masahiro Toyoura, Kentaro Go, Issei Fujishiro, Kenji Kashiwagi, Xiaoyang Mao |
Vis. Comput. | 6 |
| 2018 | Suggesting the Appropriate Number of Observers for Predicting Video Saliency with Eye-Tracking DataabstractAccurately predicting video saliency is important for applications such as video quality assessment, summary, compression, and retargeting. As the automatic saliency models for videos suffer from problems of inaccuracy, determining video saliency from data on the human gaze is a promising approach. Due to differences in individual observers, however, eye-tracking data of a certain number of observers are usually required to compute a visual attention map close to the ground truth. Although it has become cheaper to acquire human eye-tracking data thanks to the lower price of equipment, it is still not easy to carry out studies with a large number of observers. To keep the balance between accuracy and expense, this paper proposes a new method for suggesting the appropriate number of observers needed in eye-tracking experiments for a given video. Through carefully analyzing eye-tracking data of various video clips, we found videos can be classified into four types based on the number of observers required to approach the ground truth. A new support vector machine (SVM) classifier was trained to automatically classify videos into one of the four typical types. Chuancai Li, Jiayi Xu 0002, Jianjun Li 0001, Xiaoyang Mao |
CGI | 4 |
| 2018 | Parallel and efficient approximate nearest patch matching for image editing applications
Hanli Zhao, Heyang Guo, Xiaogang Jin 0001, Jianbing Shen, Xiaoyang Mao, Junru Liu |
Neurocomputing | 5 |
| 2018 | A Cascaded Algorithm for Image Quality Assessment and Image Denoising Based on CNN for Image Security and AuthorizationabstractWith the rapid development of Internet technology, images on the Internet are used in various aspects of people’s lives. The security and authorization of images are strongly dependent on image quality. Some potential problems have also emerged, among which the quality assessment and denoising of images are particularly evident. This paper proposes a novel NR-IQA method based on the dual convolutional neural network structure, which combines saliency detection with the human visual system (HSV), used as a weighting function to reflect the important distortion caused by the local area. The model is trained using gray and color features in the HSV space. It is applied to the parameter selection of an image denoising algorithm. The experiment proves that our proposed method can accurately evaluate image quality in the process of denoising. It provides great help in parameter optimization iteration and improves the performance of the algorithm. Through experiments, we obtain both improved image quality and a reasonable result of subject assessment when the cascaded algorithm is applied in image security and authorization. Jianjun Li 0001, Jie Yu 0007, Lanlan Xu, Xinying Xue, Chin-Chen Chang 0001, Xiaoyang Mao |
Secur. Commun. Networks | 6 |
| 2018 | Visual attention prediction for images with leading line structure
Issei Mochizuki, Masahiro Toyoura, Xiaoyang Mao |
Vis. Comput. | 3 |
| 2017 | Auto-framing based on user camera movementabstractWe propose a novel approach to assisting users with searching the optimal composition of a photograph. In existing studies, the process of detecting the object in a given photo occurred via only image processing, however the result does not always include the object of user's interest. A major technique contribution of our approach is to exploit the user's motion to understand the user's subjective interest in a scene. User's subjective interest and objective structure information of the scene are combined to estimate the best composition based on aesthetic measures. We named this system Auto-Framing. The evaluation result shows that estimated optimal composition closes to the ground-truth. We will embed our technique in an actual camera to enable both automatic detection of compositions and real-time guidance functionality. Tomoya Sawada, Masahiro Toyoura, Xiaoyang Mao |
CGI | 3 |
| 2017 | Synthesis of Facial Images Based on Relevance FeedbackabstractWe propose a dialogic system based on a relevance feedback strategy that allows for the semiautomatic synthesis of a facial image that only exists in a user's mind. The user is presented with several facial images and judges whether each one resembles the face that he or she is imagining. Based on the feedback from the user, a set of sample facial images are used to train an Optimum-Path Forest classifying the relevance of facial images. An interpolation method is then employed to synthesize new facial images that closely resemble the imagined face. A series of experiments are conducted to evaluate and verify the effectiveness and efficiency of the proposed technique. Caie Xu, Shota Fushimi, Masahiro Toyoura, Jiayi Xu 0002, Xiaoyang Mao |
CW | 6 |
| 2017 | Marbling-based creative modelling
Shufang Lu, Xiaogang Jin 0001, Aubrey Jaffer, Craig S. Kaplan, Xiaoyang Mao |
Vis. Comput. | 6 |
| 2017 | CGI 2017 Editorial (TVCJ)
Xiaoyang Mao, Daniel Thalmann, Marina L. Gavrilova |
Vis. Comput. | 1 |
| 2016 | Painterly Image Generation Using Scene-Aware Style TransferringabstractIn this paper, we propose a method for painterly image generation that uses an example painting with a similar scene to reflect the style of an original work in great detail. The styles of specific painters and methods often employ different colors and brushwork for each individual subject. Likewise, the connections between various subjects in a work also affect the colors and brushwork used. Our method takes input images, searches an example database for paintings with similar scenes, i.e., paintings in which the subjects have similar positional relationships and connections, and transfers the color and brushwork of the paintings to the corresponding subjects of the target images to generate painterly images that reflect specific styles in great detail. In order to ensure close linkage between various elements and to reproduce styles faithfully, our method applies the GIST approach proposed by Oliva et al. to the process of searching for paintings with similar scenes before performing style transfers. Masahiro Toyoura, Noriyuki Abe, Xiaoyang Mao |
CW | 3 |
| 2016 | Visualizing the lesson process in active learning classesabstractActive learning classes, which aim at increasing student participation in class, demand more management skills from the instructor than a conventional lecture class does. However, the instructor rarely recognizes how his/her lessons are different from those of others. The instructor cannot know exactly how one of his/her lessons is different from his/her previous week's lesson. This class-to-class comparison is effective in improving classes. This paper proposes a method for automatically visualizing the process and content of classes. Although there are ways to visualize the contents of classes manually, these approaches involve considerable investments of time and money. Machine learning techniques can automate the visualization. Our method estimated content with an average accuracy of 72.4%. Through our visualization, we confirmed that individual instructors use time differently from others and use their own time differently from lesson to lesson. Masahiro Toyoura, Mayato Sakaguchi, Xiaoyang Mao, Masanori Hanawa, Masayuki Murakami |
FIE | 3 |
| 2016 | Retrieval of clothing images based on relevance feedback with focus on collar designs
Masahiro Toyoura, Kazumi Shimizu, Xiaoyang Mao |
Vis. Comput. | 5 |
| 2016 | Example-based caricature generation with exaggeration control
Masahiro Toyoura, Jiayi Xu 0002, Fumio Ohnuma, Xiaoyang Mao |
Vis. Comput. | 5 |
| 2015 | Relevance Feedback Based Retrieval of Cloth Image with Focus on Collar DesignabstractAlthough many online shops allow users to search for clothing items by categories or keywords, it is usually a time consuming task to find the item of preferred design. Also users are usually not allowed to specify the details of design. This paper presents a new technology for extracting the feature vectors capturing the details of collar design. A prototype system based on relevance feedback is also developed allowing users search for cloth images with preferred collar design. The effectiveness of the proposed technique has been validated through a subject study. Kazumi Shimizu, Masahiro Toyoura, Xiaoyang Mao |
CW | 4 |
| 2015 | Robust Match Fusion Using OptimizationabstractIn this paper, we present a novel patch-based match and fusion algorithm by taking account of moving scene in a multiple exposure image sequence using optimization. A uniform iterative approach is developed to match and find the corresponding patches in different exposure images, which are then fused in each iteration. Our approach does not need to align the input multiple exposure images before the fusion process. Considering that the pixel values are affected by various exposure time, we design a new patch-based energy function that will be optimized to improve the matching accuracy. An efficient patch-based exposure fusion approach using the random walker algorithm is developed to preserve the moving objects from the input multiple exposure images. To the best of our knowledge, our algorithm is the first patch-based exposure fusion work to preserve the moving objects of dynamic scenes that does not need the registration process of different exposure images. Experimental results of moving scenes demonstrate that our algorithm achieves visually pleasing fusion results without ghosting artifacts, while the results produced by the state-of-the-art exposure fusion and tone mapping algorithms exhibit different levels of ghosting artifacts. Xiameng Qin, Jianbing Shen, Xiaoyang Mao, Xuelong Li 0001, Yunde Jia |
IEEE Trans. Cybern. | 3 |
| 2015 | Structured-Patch Optimization for Dense CorrespondenceabstractThis paper presents a new method to compute the dense correspondences between two images by using the energy optimization and the structured patches. In terms of the property of the sparse feature and the principle that nearest sub-scenes and neighbors are much more similar, we design a new energy optimization to guide the dense matching process and find the reliable correspondences. The sparse features are also employed to design a new structure to describe the patches. Both transformation and deformation with the structured patches are considered and incorporated into an energy optimization framework. Thus, our algorithm can match the objects robustly in complicated scenes. Finally, a local refinement technique is proposed to solve the perturbation of the matched patches. Experimental results demonstrate that our method outperforms the state-of-the-art matching algorithms. Xiameng Qin, Jianbing Shen, Xiaoyang Mao, Xuelong Li 0001, Yunde Jia |
IEEE Trans. Multim. | 3 |
| 2015 | Hidden message in a deformation-based texture
Jiayi Xu 0002, Xiaoyang Mao, Xiaogang Jin 0001, Aubrey Jaffer, Shufang Lu, Li Li 0014, Masahiro Toyoura |
Vis. Comput. | 2 |
| 2014 | A Study on Perceived Similarity between Photograph and Shape Exaggerated CaricatureabstractThis paper investigates the relationship between the extent of exaggeration in a caricature and its face identification ability. As face recognition is largely influenced by facial deformations, we focused on finding the borderline between likeness and unlikeness by applying gradual alterations to the face shape of the subject being studied. Suggestions on manipulating the degree of similarity when generating a caricature will be given. The experimental environment in this research can be used as a user-friendly caricature generation system based on Exaggerating the Difference From the Mean face, which allows a user to freely control each generation step and design his or her own unique caricature portrait. Jiayi Xu 0002, Xiaoyang Mao, Masahiro Toyoura, Xiaogang Jin 0001 |
CW | 3 |
| 2014 | Example-Based Automatic Caricature GenerationabstractCaricature is a popular artistic media widely used for effective communications. The fascination of caricature lies in its expressive depiction of a person's prominent features, which is usually realized through the so called exaggeration technique. This paper proposes a new example based automatic caricature generation system supporting the exaggeration of visual appearance features. The system comprises the construction of a learning database and the generation of caricatures. The construction of the learning database links the pairs of facial images and corresponding caricatures. Given an input face, the system automatically compute the feature vectors of facial parts and hairstyle, and search the learning database for the exaggerated parts by using the most prominent features. Experimental results show that our system can achieve the control over the degree of exaggeration and the exaggerated results can better represent the features of the subjects. Kouki Tajima, Jiayi Xu 0002, Masahiro Toyoura, Xiaoyang Mao |
CW | 5 |
| 2014 | A natural click interface for AR systems with a single camera
Atsushi Sugiura, Masahiro Toyoura, Xiaoyang Mao |
Graphics Interface | 3 |
| 2014 | Mono-spectrum marker: an AR marker robust to image blur and defocus
Masahiro Toyoura, Haruhito Aruga, Matthew Turk 0001, Xiaoyang Mao |
Vis. Comput. | 4 |
| 2013 | Automatic pencil drawing generation using saliency mapabstractAn artist usually does not draw all the areas in a picture homogeneously, but tries to make the work more expressive by emphasizing what is important while eliminating irrelevant details. We present a technique for automatically converting an input image into a pencil drawing with such effect of emphasis and elimination [HATA, et al. 2012]. The technique combines Saliency Map [ITTI, et al. 1998] and Line Integral Convolution(LIC) based pencil drawing filter [MAO, et al. 2001]. Saliency map is used to predict the focus of attention in the input image. Multi-resolution pyramid is used to locally adapt the density and appearance of pencil strokes to the degree of attention defined with saliency map Michitaka Hata, Masahiro Toyoura, Xiaoyang Mao |
SAP | 3 |
| 2013 | Stego-Marbling-TextureabstractWe present stego-marbling-texture, a new and unique texture design method which allows users to deliver personalized messages with beautiful marbling textures. Our approach is inspired by the success of the recent work on modeling traditional marbling operations as mathematical functions. The encrypter transforms an input image or a text message into an intricate marbling pattern using marbling operations defined as reversible functions, and the decrypter recovers the input image or message through reversing the process of marbling operations. When applying marbling operations, the parameters of operations are automatically recorded, encrypted, and then invisibly embedded into the marbling pattern to create a stego-marbling-texture. In this way, the decrypter can be implemented as a stand along software, enabling the receiver to extract the hidden message from the stego-marbling-texture without requiring any extra information from the sender. To ensure that the message is unnoticeably and beautifully covered by the marbling texture, we propose a new technique for automatically creating a background which is harmonious with the input message based on a set of visual perception cues. Jiayi Xu 0002, Xiaoyang Mao, Xiaogang Jin 0001, Aubrey Jaffer, Shufang Lu, Li Li 0014, Masahiro Toyoura |
CAD/Graphics | 2 |
| 2013 | Emotion Estimation from Biological Signals and Its Application to an Emotional Painting ToolabstractThis paper describes a technique for estimating the emotion of a user from the biological signals of user's central nervous system, such as cerebral blood flow and brain wave. The proposed technique uses multiple regression analysis in providing a high resolution measure to the emotional valence, which could not be realized with the existing methods based on peripheral nervous system. To demonstrate the effectiveness of the proposed emotion estimation technique in emotion based interaction, we also implemented an emotional painting tool that dynamically adapts the colors of brush and the outline of canvas to the estimated emotion of the user. The tool allows users to create original images that reflect their emotion. Masaki Omata, Daisuke Kanuka, Xiaoyang Mao |
CW | 3 |
| 2013 | Clikable Virtual Button in Real SpaceabstractSummary form only given. Clicking a virtual object is the most fundamental and important interaction in Augmented Reality (AR). This paper presents a new natural click interface for AR systems. Through a primary study, we found the acceleration of fingertips provides cues for detecting click gesture and succeeded in use it for recognizing natural click gestures with a single camera. The proposed technique was evaluated through a virtual calculator application. Atsushi Sugiura, Masahiro Toyoura, Xiaoyang Mao |
CW | 3 |
| 2013 | Detecting Markers in Blurred and Defocused ImagesabstractPlanar markers enable an augmented reality (AR) system to estimate the pose of objects from images containing them. However, conventional markers are difficult to detect in blurred or defocused images. We propose a new marker and a new detection and identification method that is designed to work under such conditions. The problem of conventional markers is that their patterns consist of high-frequency components such as sharp edges which are attenuated in blurred or defocused images. Our marker consists of a single low-frequency component. We call it a mono-spectrum marker. The mono-spectrum marker can be detected in real time with a GPU. In experiments, we confirm that the mono-spectrum marker can be accurately detected in blurred and defocused images in real time. Using these markers can increase the performance and robustness of AR systems and other vision applications that require detection or tracking of defined markers. Masahiro Toyoura, Haruhito Aruga, Matthew Turk 0001, Xiaoyang Mao |
CW | 4 |
| 2013 | ActVis: Activity Visualization in VideosabstractWe present ActVis, which is a computer-aided video surveillance system for detecting and visualizing the activation levels of multiple objects in a video. ActVis indicates "something is happening" in a video. A user arranges panels indicating the regions of focusing objects on the video screen. Temporal differential as an activation level in a panel is detected by the system, and a corresponding seek bar representing the level is generated. In general, high-level features, such as body posture or facial direction/expression, cannot be extracted when the target object is partially occluded in video, or it is not human. By employing the temporal differential as a low-level feature and the metaphor of a level meter, our system can notify a user "when something happens." The user can explore high-level features of the moment. Potential applications of ActVis include the analysis of student activation levels in classroom for professional development of faculty, and observations of wild animals for ecological investigation. Masahiro Toyoura, Satoshi Nishiguchi, Xiaoyang Mao, Masayuki Murakami |
CW | 3 |
| 2013 | Film Comic Generation with Eye Tracking
Tomoya Sawada, Masahiro Toyoura, Xiaoyang Mao |
MMM (1) | 3 |
| 2013 | Real-time directional stylization of images and videos
Hanli Zhao, Xiaogang Jin 0001, Xiaoyang Mao |
Multim. Tools Appl. | 3 |
| 2012 | Affective Doodle: a painting tool reflecting user emotionabstractWe present Affective Doodle as a novel concept of interactive painting which senses the emotion of its user and adapts the drawing parameters to it in real time. As depicted in Figure 1, Affective Doodle is realized by looping through the following 3 Steps: Masaki Omata, Daisuke Kanuka, Xiaoyang Mao, Atsumi Imamiya |
SAP | 3 |
| 2012 | Using eye-tracking data for automatic film comic creationabstractA film comic is a kind of art work representing a movie story as a comic. It uses the images of the movie as panels. Verbal information such as dialogue and narrations is represented in word balloons. A key issue in creating film comics is how to select images which are significant in conveying the story of the movie. Such significance of images is inherently semantic and context-dependent and hence, technologies purely based on image analysis usually fail to produce good results. On the other hand, the word balloon arrangement requires understanding not only the semantic of images but also the verbal information, which is difficult except for the case the script of the movie is available. This paper describes a new attempt to use eye-tracking data for the automatic creation of a film comic from a movie. Patterns of eye movement are analyzed for detecting the change of scenes and gaze information is used for automatically finding the location for inserting and directing the word balloons. Our experiments showed that the proposed technique can largely improve the selection of significant images compared with the method using image features only and realize the automatic balloon arrangement. Masahiro Toyoura, Tomoya Sawada, Mamoru Kunihiro, Xiaoyang Mao |
ETRA | 4 |
| 2012 | Film Comic Reflecting Camera-Works
Masahiro Toyoura, Mamoru Kunihiro, Xiaoyang Mao |
MMM | 3 |
| 2012 | Hairstyle Suggestion Using Statistical Learning
Masahiro Toyoura, Xiaoyang Mao |
MMM | 3 |
| 2012 | Be-code: information embedding for logo imagesabstract2D image codes are widely used for inputting information to mobile devices with cameras. QR code is a typical example of such codes. Conventional image codes do not have semantics in images themselves. Such codes may spoil the quality of design when attached on commercial products. We propose a novel code, Be-code, for embedding data into logo images. Be-code is generated from a logo image by modifying some of the edge pixels according to the bits to be embedded, but avoiding causing obvious change to the appearance of the original logo image. Algorithm for extracting information from the camera captured Be-code is also presented. The original logo image is not required in decoding. In experiments, embedded data could be decoded at a high success rate. Kazuha Yamada, Yasuna Yamashita, Masahiro Toyoura, Xiaoyang Mao, Satoshi Takatsu, Chihiro Sugawara |
MoMM | 4 |
| 2012 | Digital Camouflage Images Using Two-scale DecompositionabstractAbstract We present an alternative approach to create digital camouflage images which follows human's perception intuition and complies with the physical creation procedure of artists. Our method is based on a two‐scale decomposition scheme of the input images. We modify the large‐scale layer of the background image by considering structural importance based on energy optimization and the detail layer by controlling its spatial variation. A gradient correction is presented to prevent halo artifacts. Users can control the difficulty level of perceiving the camouflage effect through a few parameters. Our camouflage images are natural and have less long coherent edges in the hidden region. Experimental results show that our algorithm yields visually pleasing camouflage images. Xiaogang Jin 0001, Xiaoyang Mao |
Comput. Graph. Forum | 3 |
| 2012 | Automatic generation of accentuated pencil drawing with saliency map and LIC
Michitaka Hata, Masahiro Toyoura, Xiaoyang Mao |
Vis. Comput. | 3 |
| 2011 | Color-Mood-Aware Clothing Re-texturingabstractIn this paper, we present a novel color-mood-aware technique to re-texture clothing in a photograph. An efficient classification algorithm is developed to classify clothing textures using color mood scheme. To re-texture the clothing, our approach first computes the gradient maps for the cloth region to be replaced and then calculates the texture distortion coordinates on the projected cloth region according to the gradient maps. After the user selects a target clothing texture from the classified clothing texture database, the lighting and shading effects on the original photograph is transferred using the HSV color space. Experimental results show that the proposed approach successfully re-textures the clothes in photographs while preserving the geometry and lighting features. Jianbing Shen, Hanqiu Sun, Xiaoyang Mao, Yanwen Guo 0001, Xiaogang Jin 0001 |
CAD/Graphics | 3 |
| 2010 | BioMetal gloveabstractWe propose a new haptic device for rendering contact sensation of virtual objects in camera-based Augmented Reality (AR) environments. Haptic feedback can help a user to intuitively sense virtual objects. For vision-impaired users, it also means a transfer from optical information observed in the cameras to haptic information. In our system, the contact between the virtual objects and the user's hand is detected with cameras. Therefore, when presenting the contact sensation, optical markers on the hand should not be occluded from the cameras so as to avoid disturbing the estimation of 3D position and posture of the hand. To fulfill such a requirement, we used BioMetal, a promising and versatile light and thin material that shrinks when electric current is applied, which provides the haptic feedback. Strings of BioMetal were stitched onto our proposed BioMetal glove. Because BioMetal does not shrink instantly when energized, a major challenge is how to deal with the time lag. We address this problem by setting buffering regions for pre-heating the BioMetal strings. Masahiro Toyoura, Tatsuya Shono, Xiaoyang Mao |
VRST | 3 |
| 2009 | AtelierM++: a fast and accurate marbling system
Hanli Zhao, Xiaogang Jin 0001, Shufang Lu, Xiaoyang Mao, Jianbing Shen |
Multim. Tools Appl. | 4 |
| 2009 | Real-time saliency-aware video abstraction
Hanli Zhao, Xiaoyang Mao, Xiaogang Jin 0001, Jianbing Shen, Jieqing Feng |
Vis. Comput. | 2 |
| 2008 | Using multiple data sources to get closer insights into user cost and task performanceabstractThis pilot study explores the use of combining multiple data sources (subjective, physical, physiological, and eye tracking) in understanding user cost and behavior. Specifically, we show the efficacy of such objective measurements as heart rate variability (HRV), and pupillary response in evaluating user cost in game environments, along with subjective techniques, and investigate eye and hand behavior at various levels of user cost. In addition, a method for evaluating task performance at the micro-level is developed by combining eye and hand data. Four findings indicate the great potential value of combining multiple data sources to evaluate interaction: first, spectral analysis of HRV in the low frequency band shows significant sensitivity to changes in user cost, modulated by game difficulty—the result is consistent with subjective ratings, but pupillary response fails to accord with user cost in this game environment; second, eye saccades seem to be more sensitive to user cost changes than eye fixation number and duration, or scanpath length; third, a composite index based on eye and hand movements is developed, and it shows more sensitivity to user cost changes than a single eye or hand measurement; finally, timeline analysis of the ratio of eye fixations to mouse clicks demonstrates task performance changes and learning effects over time. We conclude that combining multiple data sources has a valuable role in human–computer interaction (HCI) evaluation and design. Tao Lin 0006, Atsumi Imamiya, Xiaoyang Mao |
Interact. Comput. | 3 |
| 2008 | Real-time feature-aware video abstraction
Hanli Zhao, Xiaogang Jin 0001, Jianbing Shen, Xiaoyang Mao, Jieqing Feng |
Vis. Comput. | 4 |
| 2007 | Deformation-based interactive texture design using energy optimization
Jianbing Shen, Xiaogang Jin 0001, Xiaoyang Mao, Jieqing Feng |
Vis. Comput. | 3 |
| 2006 | Completion-based texture design using deformation
Jianbing Shen, Xiaogang Jin 0001, Xiaoyang Mao, Jieqing Feng |
Vis. Comput. | 3 |
| 2005 | Sketchy hairstylesabstractWe present an intuitive, interactive modeling and rendering system for creating non-photorealistic hairstyle images. The main feature of our system lies in its user-friendly sketch interface that allows the user to generate his/her desired hairstyle simply by drawing a few free-form strokes on the 3D model of a scalp. Hairstyles are modeled with a polygon based technique called cluster polygons, and are rendered expressively to obtain non-photorealistic images similar to those found in hand-drawn animations and cartoons. Xiaoyang Mao, Shiho Isobe, Ken Anjyo, Atsumi Imamiya |
Computer Graphics International | 1 |
| 2004 | Robust object-identification from inaccurate recognition-based inputsabstractEyesight and speech are two channels that humans naturally use to communicate with each other. However both the eye tracking and the speech recognition technique existing are still far from perfect. This work explored how to integrate two (or more) error-prone sources of information on users' selection of objects in a visual interface. The implemented system integrated a commercial speech recognition system with gaze tracking in order to improve recognition results. In addition, we employed a new measure of the rate of mutual disambiguation for the multimodal system and conducted an experimental evaluation. Qiaohui Zhang, Kentaro Go, Atsumi Imamiya, Xiaoyang Mao |
AVI | 4 |
| 2004 | Sketch Interface Based Expressive Hairstyle Modelling and RenderingabstractWe present a new system for interactively modeling and rendering nonphotorealistic hairstyles. The system is featured with a user-friendly sketch interface allowing a user generate his/her desired hairstyle simply by drawing a few free-form strokes. Hairstyles are modeled with a new polygon based technique called cluster polygon, and can be rendered expressively to obtain hair images similar to those found in cell animation and cartoons. Xiaoyang Mao, Hiroki Kato, Atsumi Imamiya, Ken Anjyo |
Computer Graphics International | 1 |
| 2004 | Resolving ambiguities of a gaze and speech interfaceabstractThe recognition ambiguity of a recognition-based user interface is inevitable. Multimodal architecture should be an effective means to reduce the ambiguity, and contribute to error avoidance and recovery, compared with a unimodal one. But does the multimodal architecture always perform better than the unimode at any time? If not, when does it perform better than unimode, and when is it the optimum? Furthermore, how can modalities best be combined to gain the advantage of synergy? Little is known about these issues in the literature available. In this paper we try to give the answer through analyzing integration strategies for gaze and speech modalities, together with an evaluation experiment verifying these analyses. The approach involves studying the mutual correction cases and investigating when the mutual correction phenomena will occur. The goal of this study is to gain insights into integration strategies, and develop an optimum system to make error-prone recognition technologies perform at a more stable and robust level within a multimodal architecture. Qiaohui Zhang, Atsumi Imamiya, Kentaro Go, Xiaoyang Mao |
ETRA | 4 |
| 2004 | Overriding errors in a speech and gaze multimodal architectureabstractThis work explores how to use the gaze and the speech command simultaneously to select an object on the screen. Multimodal systems have long been a key mean to reduce the recognition errors of individual components. But the multimodal system generates errors as well. This present study tries to classify the multimodal errors, analyze the reasons causing these errors, and propose the solutions for eliminating them. The goal of this study is to gain insight into multimodal integration errors, and to develop an error self-recoverable multimodal architecture so as to make the error-prone recognition technologies perform at a more stable and robust level within multimodal architecture. Qiaohui Zhang, Atsumi Imamiya, Kentaro Go, Xiaoyang Mao |
IUI | 4 |
| 2004 | Colored Pencil Filter with Custom ColorsabstractThis paper presents a new technique for automatically converting digital images into colored pencil drawings. In real colored pencil drawing, artists use different colors for different regions, and add pure colors directly onto paper to build the target color through optical blending. To support such effect, our technique extends the existing technique for reproducing color image with custom inks to automatically select the best color set for individual regions in a source image. Then layers of stroke image for each color are generated and superimposed with the Kubelka-Munk optical compositing model. We also allow users to specify regions and to customize the color set for a specified region interactively. The proposed technique can be easily extended to simulate other artistic media featured with optical color blending, such as pastel and wax crayons. Shigefumi Yamamoto, Xiaoyang Mao, Atsumi Imamiya |
PG | 2 |
| 2000 | Automatic Generation of Hair Texture with Line Integral ConvolutionabstractSynthesis of hair images is one of the most important and challenging computer graphics problems. We propose a new technique for automatically generating realistic human hair texture on 3D models of human characters. The idea is inspired by the similarity between the texture of human hair and the texture generated by the LIC algorithm. The proposed technique generates the texture of human hair using the vector field defining the directions of hair strands as the input to the 3D LIC algorithm. Xiaoyang Mao, Makoto Kikukawa, Kouichi Kashio, Atsumi Imamiya |
IV | 1 |
| 1998 | Multi-Granularity Noise for Curvilinear Grid LIC
Xiaoyang Mao, Lichan Hong, Arie E. Kaufman, Noboru Fujita, Makoto Kikukawa, Atsumi Imamiya |
Graphics Interface | 1 |
| 1998 | Image-guided streamline placement on curvilinear grid surfacesabstractThe success of using a streamline technique for visualizing a vector field usually depends largely on the choice of adequate seed points. G. Turk and D. Banks (1996) developed an elegant technique for automatically placing seed points to achieve a uniform distribution of streamlines on a 2D vector field. Their method uses an energy function calculated from the low-pass filtered streamline image to guide the optimization process of the streamline distribution. This paper proposes a new technique for creating evenly distributed streamlines on 3D parametric surfaces found in curvilinear grids. We make use of Turk and Banks's 2D algorithm by first mapping the vectors on a 3D surface into the computational space of the curvilinear grid. To take into the consideration the mapping distortion caused by the uneven grid density in a curvilinear grid, a new energy function is designed and used for guiding the placement of streamlines in the computational space with desired local densities. Xiaoyang Mao, Yuji Hatanaka, Hidenori Higashida, Atsumi Imamiya |
IEEE Visualization | 1 |
| 1996 | Splatting of Non Rectilinear Volumes Through Stochastic ResamplingabstractThe paper extends the conventional splatting algorithm for volume rendering non rectilinear grids. A stochastic sampling technique called Poisson sphere/ellipsoid is employed to adaptively resample a non rectilinear grid with a set of randomly distributed points whose energy support extents are well approximated by spheres or ellipsoids. Then volume rendered images can be generated by splatting the scalar values at the new sample points with filter kernels corresponding to these spheres and ellipsoids. Experiments have been carried out to investigate the image quality as well as the time/space efficiency of the new approach, and the results suggest that our approach can be regarded as an alternative for existing fast volume rendering techniques of non rectilinear grids. Xiaoyang Mao |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 1995 | Interactive Visualization of Mixed Scalar and Vector FieldsabstractThis paper describes an approach for interactive visualization of mixed scalar and vector fields, in which vector icons are generated from pre-voxelized icon templates and volume-rendered together with the volumetric scalar data. This approach displays simultaneously the global structure of the scalar field and the detailed features of the vector field. Interactive visualization is achieved with incremental image update, by re-rendering only a small portion of the image wherever and whenever a change occurs. This technique supports a set of interactive visualization tools, including change of vector field visualization parameters, real-time animation of vector icons advected within the scalar field, a zooming lens, and a local probe. Lichan Hong, Xiaoyang Mao, Arie E. Kaufman |
IEEE Visualization | 2 |
| 1995 | Splatting of Curvilinear VolumesabstractThe paper presents a splatting algorithm for volume rendering of curvilinear grids. A stochastic sampling technique called Poisson sphere/ellipsoid sampling is employed to adaptively resample a curvilinear grid with a set of randomly distributed points whose energy support extents are well approximated by spheres and ellipsoids. Filter kernels corresponding to these spheres and ellipsoids are used to generate the volume rendered image of the curvilinear grid with a conventional footprint evaluation algorithm. Experimental results show that our approach can be regarded as an alternative to existing fast volume rendering techniques of curvilinear grids. Xiaoyang Mao, Lichan Hong, Arie E. Kaufman |
IEEE Visualization | 1 |
| 1990 | A translucent display algorithm for G-octree represented grey-scale imagesabstractAbstract A new translucent display algorithm is proposed for grey‐scale images represented by G‐octrees. The algorithm generates a special triangle quadtree as the display image for the isometric projection of a G‐octree, while traversing the G‐octree from front to back, with respect to the viewpoint. A method of using the translucent effect to analyse the interiors of three‐dimensional objects is also described, together with its implementation results. Xiaoyang Mao, Issei Fujishiro, Tosiyasu L. Kunii |
Comput. Animat. Virtual Worlds | 1 |
| 1986 | G-quadtree: A hierarchical representation of gray-scale digital images
Tosiyasu L. Kunii, Issei Fujishiro, Xiaoyang Mao |
Vis. Comput. | 3 |