VLDB 2026 Research / reviewers in the wild / expert
Zhenyang Zhu
dblp:214/3695
· DBLP profile ↗
28ranked-venue papers
8as first author
24since 2021 · last 2026
0000-0003-1023-3193ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 7 first-author · 18 since 2021Human-computer interaction and ubiquitous computing · 9 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Neighborhood constrained attention for lightweight image super-resolutionabstractIn recent years, to improve image super-resolution performance, several studies have explored integrating convolutional modules with vision transformers (ViTs) to enhance the local feature modeling of ViTs. However, these hybrid approaches often introduce inconsistencies in feature representation, redundant information, and an increased number of parameters, ultimately limiting both performance and computational efficiency. To overcome these challenges, we propose a novel neighborhood constrained attention (NCA) mechanism that enables transformers to effectively capture both local and global features without requiring additional convolutional modules. Specifically, we first divide the window into a set of grids, treating them as local features, and then explore both intra- and inter-relationships within and across these local features, using them as constraints to refine window attention. Furthermore, instead of relying on averaging or other heuristic schemes for assigning labels to local features, we combine them through a linear transformation, ensuring label accuracy and uniqueness. Extensive experiments demonstrate that the proposed NCA not only outperforms other state-of-the-art lightweight approaches on public benchmark datasets but also excels in engineering image datasets, such as automated defect detection and product quality inspection, while requiring fewer parameters and lower computational costs. Notably, compared to 4 SwinIR-light (SwinIR: Image Restoration Using Swin Transformer), NCA achieves an average performance gain of 0.28 dB across five public test sets while reducing network parameters by 27% and computational complexity (floating point operations, FLOPs) by 30%. Code and models are obtainable at https://github.com/hms-source/NCA . Zhenyang Zhu, Xiaoyang Mao |
Eng. Appl. Artif. Intell. | 2 |
| 2026 | Memory-efficient divide-and-conquer attention for lightweight image super-resolutionabstractRecently, transformer-based methods have achieved significant progress in lightweight image super-resolution (SR). However, most of these approaches primarily aim to improve either inference speed or reconstruction quality, while overlooking memory consumption, thereby limiting their practicality on resource-constrained devices. In this paper, we propose a memory-efficient divide-and-conquer attention (MEDCA) for SR, which substantially reduces memory usage while achieving notable improvements in reconstruction performance and competitive inference speed. To address the high memory and space complexity of standard window-based self-attention (WSA), MEDCA adopts a divide-and-conquer strategy. Specifically, the input features are first split into multiple subspaces along the channel dimension. Each subspace is further partitioned into multiple windows, which are then evenly divided into two parts using distinct asymmetric strategies. Self-attention is independently applied to each part, and the outputs are aggregated to form the final representation. Compared with traditional WSA methods such as SwinIR, MEDCA reduces the space complexity within each subspace by half. Furthermore, we design multiple asymmetric partitioning strategies that allow the model to extract features from a broader spatial context, thereby enabling it to capture richer spatial information and enhance its representation capacity. Extensive experiments demonstrate that MEDCA significantly reduces memory consumption while outperforming existing lightweight state-of-the-art methods and maintaining competitive inference speed across multiple public benchmark datasets. In particular, compared with the state-of-the-art method HiT-SRF, MEDCA improves the average performance by 0.12dB across five public test sets, while maintaining comparable inference time and requiring only 29.7% of the memory used by HiT-SRF. The code and models are provided at https://github.com/hms-source/MEDCA. Zhenyang Zhu, Xiaoyang Mao |
Neural Networks | 2 |
| 2026 | A ViT-like unified multi-scale learning network for efficient image super-resolutionabstractRecently, vision transformer (ViT)-based super-resolution (SR) models have achieved strong reconstruction performance but suffer from substantial computational and memory overhead due to self-attention operations, limiting their applicability in resource-constrained scenarios. Although several CNN-based alternatives attempt to replicate the global modeling capability of ViTs, their limited receptive fields often struggle to achieve competitive performance. To address this challenge, we propose a ViT-like unified multi-scale learning network (UMLN) that achieves a favorable trade-off between reconstruction quality and computational efficiency. Instead of relying on memory-intensive self-attention, our ViT-like design adopts a unified multi-scale feature aggregation mechanism that generates a global weight matrix through efficient convolutional operations. This unified module enables cross-scale interaction and long-range dependency modeling while significantly reducing parameter redundancy and memory consumption compared to attention-based approaches. In addition, we design an efficient feature extraction module CGDConv that effectively captures both local textures and non-local contextual information. Extensive experiments demonstrate that our UMLN outperforms existing efficient SR methods on public benchmark datasets, matching the performance of state-of-the-art lightweight ViT-based approaches while reducing significantly computational cost. Notably, compared to the × 4 SRFormer-light, our UMLN-L achieves an average PSNR improvement of 0.14 dB across five public benchmark datasets, while requiring only 68% of the computational complexity (e.g., FLOPs). Code and models are obtainable at https://github.com/hms-source/UMLN . • We propose a unified multi-scale feature extraction strategy that shares a single processing module across different scales, effectively reducing parameters and improving efficiency. • We propose a ViT-like weighted fusion mechanism to aggregate multi-scale features, enabling more effective global context modeling and overcoming the performance limitations of existing methods. • We design an effective feature extraction module, CGDConv, to capture both high- and low-frequency information within the network. • Extensive experiment results on public benchmark datasets demonstrate that the proposed UMLN achieves competitive performance in terms of model complexity, computational efficiency, and reconstruction quality compared to state-of-the-art approaches. Zhenyang Zhu, Xiaoyang Mao |
Pattern Recognit. | 2 |
| 2026 | A Swin-style shifted pooling cross-aggregation network for efficient image super-resolutionabstractAbstract Swin transformer-based methods have achieved impressive performance in image super-resolution (SR) due to their ability to effectively model long-range spatial dependencies. However, the core component, window-based self-attention (WSA), introduces considerable computational overhead, which limits their applicability on resource-constrained devices. To address these issues, we propose a Swin-style shifted pooling cross-aggregation network (SPCAN) for image SR, which achieves high computational efficiency while maintaining excellent reconstruction quality. Specifically, we adopt max pooling-based downsampling as a lightweight alternative to WSA for extracting low-frequency features and introduce a shifted pooling mechanism that emulates the shifted window strategy of Swin transformers within a convolutional neural network (CNN) framework. This mechanism is embedded within a cross-aggregation module to facilitate efficient inter-region feature interaction. Moreover, we generalize the pooling operation from square to rectangular regions to enhance the model’s ability to capture spatial dependencies across different orientations. Extensive experiments on public SR benchmarks demonstrate that the proposed method achieves competitive reconstruction accuracy while offering significantly better efficiency compared with existing state-of-the-art methods. The source code and pretrained models are available at: https://github.com/hms-source/SPCAN . Zhenyang Zhu, Xiaoyang Mao |
Vis. Comput. | 2 |
| 2025 | Latent Interpretation for Multi-view Face Synthesis Across GAN and Diffusion via Conditional Reconstruction
Yixuan Ju, Zhenyang Zhu, Xiaoyang Mao |
CGI (3) | 4 |
| 2025 | Semantic Compensation and Localization for Color Vision DeficiencyabstractColor Vision Deficiency (CVD) affects approximately 8 % of men and 0.5 % of women worldwide, creating barriers in color-coded visual interactions. Existing assistive solutions, such as image recoloring and pattern overlays, often fail to provide accurate color identification or precise location in complex scenes. This paper presents a novel framework that integrates vision-language models (VLMs) with open-vocabulary object detection to enhance color understanding for individuals with CVD. Our approach employs physiologically grounded CVD simulation algorithms to generate CVD-simulated images, then computes pixel-level difference maps between original and CVDsimulated images to identify perceptually challenging regions. This difference information guides the vision-language model's attention toward color-critical objects. After training on a specially constructed CVD dataset, the model generates precise color-aware descriptions targeting challenging areas. Our openvocabulary object detection component then provides accurate localization of key objects with color attributes. Experimental results demonstrate effective identification of visually confusing regions, accurate color-aware descriptions, and precise spatial location. This work provides a practical approach for developing inclusive visual assistance systems for the CVD community. Zhencheng Chen, Zhenyang Zhu, Xiaoyang Mao |
CW | 2 |
| 2025 | Cross-Domain Personal Identification Framework Based on Intraoral PhotographsabstractIn disaster scenarios, forensic experts typically rely on manual comparison between clinical intraoral photos and onsite photos to identify remains. However, this may introduce subjective bias and reduce overall efficiency. Moreover, on-site photos are generally more complex due to variations in illumination, longitudinal changes, and imaging angles. These challenges are further exacerbated by differences in capture devices, shooting conditions, and the inconsistent use of intraoral retractor. To figure out these challenges, this study proposes a deep learningbased intraoral photographs used identification approach which uses image domain translation. To validate the effectiveness of the proposed method, we constructed a multi-domain intraoral photo dataset comprising over$\mathbf{4, 0 0 0}$images from more than 100 individuals, covering both clinical and on-site domains. Through the image domain transfer technology, intraoral photographs in the clinic environment are simulated as images in the on-site environment, thereby improving the model's recognition ability for on-site images. Experimental results demonstrate that domain adaptation significantly enhances cross-domain identification performance. The results confirm the potential of the proposed approach to achieve accurate and efficient identity verification in complex, real-world forensic scenarios. Zhean Ma, Zhenfei Wang, Zhenyang Zhu, Kotaro Kubo, Xiaoyang Mao |
CW | 3 |
| 2025 | ClothMotion: Pose-Guided Temporal Consistent Garment Mask GenerationabstractWith the advancement of diffusion models and Transformer-based architectures, pose-guided human image animation has achieved remarkable progress. However, existing methods often treat garments as static textures attached to the body, ignoring the non-rigid nature of clothing deformation during motion. This modeling limitation leads to unrealistic artifacts such as skirt misalignment and sleeve discontinuity, especially in the presence of loose or dynamic garments. To address this issue, ClothMotion is introduced as an appearance flow-based approach that explicitly models the evolution of garment contours driven by pose changes. The method employs a multi-scale flow estimation framework to predict dense non-rigid displacement fields, which deform a reference garment mask according to the target pose sequence. The resulting warped mask serves as a strong geometric prior for a pose-aware decoder, enabling temporally consistent and spatially coherent segmentation across frames. To support training, a dynamic garment segmentation dataset is constructed. Pseudo-labels are generated on TikTok and UBC-Fashion videos using a DeepFashion2-pretrained detector. A global majority voting strategy and manual correction are applied to improve annotation consistency and accuracy. Experimental results show that ClothMotion significantly improves mask accuracy and garment stability, achieving 0.85 mIoU on real-world sequences. When integrated into existing animation frameworks such as DisCo, the method further enhances visual quality and structural consistency. These results highlight the importance of explicit contour modeling in pose-guided image animation and demonstrate improved controllability over dynamic garment behavior, offering a generalizable path toward realistic clothing synthesis. Huiyun Qiu, Zhenyang Zhu, Xiaoyang Mao |
CW | 2 |
| 2025 | Swin-WNet: Boundary-Aware Semantic Segmentation for Oral Squamous Cell CarcinomaabstractOral squamous cell carcinoma (OSCC) poses a significant threat to public health due to its severity and the laborintensive process of image analysis required by physicians. Intercellular bridges are bridge-like structures that connect adjacent cells and indicate the differentiation level of a cancer. Although intercellular bridges are known to disappear as differentiation decreases, pathologists and clinicians evaluate the presence of intercellular bridges to assess the degree of differentiation of cancer. While state-of-the-art (SOTA) deep learning methods, such as U-Net and its variants, perform well on uniform and clearly delineated objects (e.g., cell, lung, etc.), accurately segmenting intricate objects (e.g., intercellular bridge, retinal vessel) remains challenging due to their complex topologies, fine branches, and irregular morphological changes. This paper aims to propose a method that effectively utilize boundary information, particularly targeting intricate objects. Our approach, inspired by Swin-UNet, employs the Swin Transformer, comprising a feature encoder, two decoders (semantic decoder and boundary decoder) and an attention-guided fusion module to enhance the model's ability to segment intricate objects. By applying the constraints of the boundary decoder, the feature encoder's ability to encode structural information is enhanced without compromising semantic information extraction and representation. Furthermore, fusing the outputs of the boundary decoder and semantic decoder further strengthens the detail of structural information. To validate the generalizability of our method, we conducted comparative experiments on one private and one public dataset. The results demonstrate that our method outperforms SOTA methods. Zhenfei Wang, Zhenyang Zhu, Kunio Yoshizawa, Masahiro Toyoura, Naoki Oishi, Xiaoyang Mao |
CW | 2 |
| 2025 | Boosting lightweight single image super-resolution via global prior featureabstractRecently, lightweight vision transformer (ViT)-based single image super-resolution (SISR) has gained significant attention. However, many existing lightweight methods struggle to achieve satisfactory performance due to the aggressive reduction in the number of parameters. Therefore, to improve the performance of lightweight networks, we propose a novel global feature prior self-attention network. First, conventional window-based self-attention methods typically apply attention mechanisms indiscriminately to all pixels within a window. This can lead to artifacts and texture blurring. To mitigate this issue, we leverage prior knowledge to identify texture-related pixels within the window and perform self-attention operations specifically on these pixels. Second, to enhance the network's ability to capture critical information and structural details, we introduce an efficient global feature extraction method. Finally, while transformers excel at capturing global features and low-frequency information, they often struggle with extracting local features and high-frequency information. Therefore, we integrate a local complementary module into the shift window attention to compensate for the transformer's shortcomings in extracting local and high-frequency features. Extensive experiments demonstrate that the proposed method outperforms all other state-of-the-art lightweight approaches. Code and models are obtainable at https://github.com/hms-source/GFPSAN. Zhenyang Zhu, Xiaoyang Mao |
Neural Networks | 2 |
| 2025 | A content-aware image editing method for perceptual size restoration based on seam carvingabstractAbstract To alleviate the gap between the perceptual size of Thing of Interest (ToI) in a real scene and its appearance in a photograph, ToI region enlargement methods based on seam carving have been proposed in previous studies. However, state-of-the-art methods may suffer from issues such as a lack of automation and distortion caused by seam concentrations. In this study, we propose an image editing method for ToI enlargement based on seam carving. The proposed method incorporates the segment anything model to automatically generate ToI mask. To address the issue of seam concentrations, a wide-area energy strategy is introduced. Additionally, to improve the image quality of the enlarged ToI region, a super-resolution-based post-processing mechanism is put forward. Subjective experiments were conducted to evaluate the effectiveness of the proposed method. The experimental results suggest that the proposed method effectively enlarges the ToI region while suppressing distortion caused by seam concentrations. Zhenyang Zhu, Naohiko Ishikawa, Xiaoyang Mao |
Vis. Comput. | 1 |
| 2025 | Enhanced fine-grained visual classification through lightweight Transformer integration and auxiliary information fusion
Zhenyang Zhu, Ketai He |
Vis. Comput. | 1 |
| 2024 | A Semi-automatic Quality Assessment System for Capturing High-quality Fundus ImageabstractThe precision of medical diagnoses based on images is inextricably linked to the quality and clarity of the images. Poor image quality can impede accurate diagnosis and pose challenges for both physicians and machine learning algorithms in interpreting images. By automating the process of fundus image quality assessment during capture, we can ensure that only high-quality images are used, improving diagnostic accuracy. Accordingly, we propose a novel approach for automatically assessing the quality of fundus images using deep learning techniques. Our method incorporates retinal vessel segmentation into RGB images to create four-channel images and then trains a deep learning model on these images to identify the image focus and quality of the images. It can evaluate the quality of fundus images and request image recapture with the adjustment of specific parameters that are classified as poor quality. Our proposed method has the potential to improve the diagnostic accuracy and efficiency of retinal disease diagnosis, particularly in telemedicine settings. By automating the process of fundus image quality assessment, we can ensure that only high-quality images are used for diagnosis, thus improving diagnostic precision. It can serve as an efficient screening tool in the initial stage of acquiring high-quality fundus images. Asif Mohammed Arfi, Masahiro Toyoura, Kenji Kashiwagi, Satoshi Nishiguchi, Kentaro Go, Zhenyang Zhu, Xiaoyang Mao |
CW | 6 |
| 2024 | Seam Carving-based Image Partial Enlargement Method for Perceptual Impression ReflectionabstractIn addressing the issue that the Thing of Interest (ToI) in a photograph appears significantly smaller than that being directly perceived from that in the physical world, image editing techniques have been developed to enlarge ToI while maintaining appearance of other areas. However, these methods can suffer from issues such as loss of spatial perception, distortions of objects, the lack of automation, etc. To address these issues, we propose a novel ToI enlarging method based on Seam Carving method. The proposed method is composed of seam insertion and deletion processing, and introduces wide area seam energy mechanism, which considers impact of inserted or deleted seams on adjacent areas, in addition to preservation to salient objects, image structure, gradient and mask energies. Moreover, the proposed method introduces the Segment Anything Model to automatically generate masks representing the ToI. To validate the effectiveness of the proposed method, we conducted two subjective evaluation experiments in this study. The experimental results demonstrate that the proposed method can produce resulting images that reflect perceptual impressions, and can preserve area other than ToI. Naohiko Ishikawa, Zhenyang Zhu, Jong-Nam Kim, Xiaoyang Mao |
CW | 2 |
| 2024 | Image recoloring for color vision deficiency compensation using Swin transformerabstractAbstract People with color vision deficiency (CVD) have difficulty in distinguishing differences between colors. To compensate for the loss of color contrast experienced by CVD individuals, a lot of image recoloring approaches have been proposed. However, the state-of-the-art methods suffer from the failures of simultaneously enhancing color contrast and preserving naturalness of colors [without reducing the Quality of Vision (QOV)], high computational cost, etc. In this paper, we propose an image recoloring method using deep neural network, whose loss function takes into consideration the naturalness and contrast, and the network is trained in an unsupervised manner. Moreover, Swin transformer layer, which has long-range dependency mechanism, is adopted in the proposed method. At the same time, a dataset, which contains confusing color pairs to CVD individuals, is newly collected in this study. To evaluate the performance of the proposed method, quantitative and subjective experiments have been conducted. The experimental results showed that the proposed method is competitive to the state-of-the-art methods in contrast enhancement and naturalness preservation and has a real-time advantage. The code and model will be made available at https://github.com/Ligeng-c/CVD_swin . Ligeng Chen, Zhenyang Zhu, Wangkang Huang, Kentaro Go, Xiaoyang Mao |
Neural Comput. Appl. | 2 |
| 2024 | Fast image recoloring for red-green anomalous trichromacy with contrast enhancement and naturalness preservationabstractAbstract Color vision deficiency (CVD) is an eye disease caused by genetics that reduces the ability to distinguish colors, affecting approximately 200 million people worldwide. In response, image recoloring approaches have been proposed in existing studies for CVD compensation, and a state-of-the-art recoloring algorithm has even been adapted to offer personalized CVD compensation; however, it is built on a color space that is lacking perceptual uniformity, and its low computation efficiency hinders its usage in daily life by individuals with CVD. In this paper, we propose a fast and personalized degree-adaptive image-recoloring algorithm for CVD compensation that considers naturalness preservation and contrast enhancement. Moreover, we transferred the simulated color gamut of the varying degrees of CVD in RGB color space to CIE L*a*b* color space, which offers perceptual uniformity. To verify the effectiveness of our method, we conducted quantitative and subject evaluation experiments, demonstrating that our method achieved the best scores for contrast enhancement and naturalness preservation. Haiqiang Zhou, Wangkang Huang, Zhenyang Zhu, Kentaro Go, Xiaoyang Mao |
Vis. Comput. | 3 |
| 2023 | Seamless Image Editing for Perceptual Size Restoration Based on Seam Carving
Naohiko Ishikawa, Zhenyang Zhu, Jong-Nam Kim, Wan-Young Chung, Kentaro Go, Xiaoyang Mao |
CGI (1) | 2 |
| 2023 | Metamorphopsia Insepction System Based on Relevance FeedbackabstractPeople with metamorphopsia suffer from perceiving things in a distorted way. Various methods for examining metamorphopsia have been suggested in the current literature, with the most advanced techniques demonstrating the ability to yield quantitative measurements. However, these cutting-edge methods necessitate extended examination durations and impose challenging manipulations on patients. In this study, our objective is to enhance the time efficiency of the inspection process and alleviate the burden placed on the user. We propose a novel user-friendly quantitative inspection system which utilizes interactive reinforcement learning. Instead of having users directly operate the system, we ask them to evaluate the stimuli generated by the system. Based on their evaluations, the system gradually refines the deformation map representing the distortion perceived by the user. The reinforcement learning scheme is implemented using relevance feedback approach based on optimum-path forest classifier. To evaluate the effectiveness of the proposed system, subjective evaluation experiments involving simulated and real metamorphopsia participants were conducted in this study. The experimental findings reveal that, when compared to the state-of-the-art method, our proposed system yields comparable inspection out-comes while significantly reducing both the inspection duration and the mental workload. Zhenyang Zhu, Katsuhito Moritake, Kenji Kashiwagi, Masahiro Toyoura, Kentaro Go, Issei Fujishiro, Xiaoyang Mao |
SMC | 1 |
| 2022 | Personalized Image Recoloring for Color Vision Deficiency CompensationabstractSeveral image recoloring methods have been proposed to compensate for the loss of contrast caused by color vision deficiency (CVD). However, these methods only work for dichromacy (a case in which one of the three types of cone cells loses its function completely), while the majority of CVD is anomalous trichromacy (another case in which one of the three types of cone cells partially loses its function). In this paper, a novel degree-adaptable recoloring algorithm is presented, which recolors images by minimizing an objective function constrained by contrast enhancement and naturalness preservation. To assess the effectiveness of the proposed method, a quantitative evaluation using common metrics and subjective studies involving 14 volunteers with varying degrees of CVD are conducted. The results of the evaluation experiment show that the proposed personalized recoloring method outperforms the state-of-the-art methods, achieving desirable contrast enhancement adapted to different degrees of CVD while preserving naturalness as much as possible. Zhenyang Zhu, Masahiro Toyoura, Kentaro Go, Kenji Kashiwagi, Issei Fujishiro, Tien-Tsin Wong, Xiaoyang Mao |
IEEE Trans. Multim. | 1 |
| 2022 | Image recoloring for Red-Green dichromats with compensation range-based naturalness preservation and refined dichromacy gamut
Wangkang Huang, Zhenyang Zhu, Ligeng Chen, Kentaro Go, Xiaoyang Mao |
Vis. Comput. | 2 |
| 2022 | LineM: assessing metamorphopsia symptom using line manipulation task
Zhenyang Zhu, Masahiro Toyoura, Issei Fujishiro, Kentaro Go, Kenji Kashiwagi, Xiaoyang Mao |
Vis. Comput. | 1 |
| 2021 | Eye-Tracker-Free Compensation for MetamorphopsiaabstractMetamorphopsia is a symptom caused by abnormalities in the retina, and people with metamorphopsia experience distortions in their field of view. To compensate for this disorder, computer-based distortion assessment and compensation methods have been proposed. In the state-of-the-art method, compensation is carried out by dynamically deforming the image with a manipulation map, which is bound to the symptom of an individual user with metamorphopsia, according to the user's gaze data captured by an eye tracker. However, this method suffers from the instability and inaccuracy of the eye tracker, which leads to a poor compensation effect. In this paper, we propose a novel method for metamorphopsia compensation without using an eye tracker. The proposed method generates a compensation image by simultaneously applying multiple manipulation maps to the input image. To evaluate the effectiveness of the proposed method, a preliminary subjective experiment using a reading task and involving 10 participants with normal vision was conducted. The evaluation results show that the proposed method can compensate for visual distortion during reading tasks. Katsuhito Moritake, Zhenyang Zhu, Masahiro Toyoura, Kentaro Go, Kenji Kashiwagi, Issei Fujishiro, Xiaoyang Mao |
CW | 2 |
| 2021 | Fast contrast and naturalness preserving image recolouring for dichromats
Zhenyang Zhu, Kentaro Go, Masahiro Toyoura, Xiaoyang Mao |
Comput. Graph. | 2 |
| 2021 | Image recoloring for color vision deficiency compensation: a surveyabstractAbstract People with color vision deficiency (CVD) have a reduced capability to discriminate different colors. This impairment can cause inconveniences in the individuals’ daily lives and may even expose them to dangerous situations, such as failing to read traffic signals. CVD affects approximately 200 million people worldwide. In order to compensate for CVD, a significant number of image recoloring studies have been proposed. In this survey, we briefly review the representative existing recoloring methods and categorize them according to their methodological characteristics. Concurrently, we summarize the evaluation metrics, both subjective and quantitative, introduced in the existing studies and compare the state-of-the-art studies using the experimental evaluation results with the quantitative metrics. Zhenyang Zhu, Xiaoyang Mao |
Vis. Comput. | 1 |
| 2020 | Evaluation of Color Vision Compensation Algorithms for People with Varying Degrees of Color Vision DeficiencyabstractPeople with color vision deficiency (CVD) may have difficulty in discriminating colors. To improve their color perception, several compensation methods have been proposed which considered naturalness maintenance and contrast emphasis. All these methods are based on the simulation model of severe CVD and hence it is not clear whether are also effective for people with light CVD. In this paper, we conduct subjective study to evaluate the effectiveness of these methods for people with varying degrees of CVD. Zhenyang Zhu, Masahiro Toyoura, Xiaoyang Mao |
CW | 2 |
| 2019 | Naturalness- and information-preserving image recoloring for red-green dichromatsabstractMore than 100 million individuals around the world suffer from color vision deficiency (CVD). Image recoloring algorithms have been proposed to compensate for CVD. This study has proposed a new recoloring algorithm to make up shortages of contrast enhancement and naturalness preservation of the state-of-the-art methods. The recoloring task is formulated as an optimization problem that is solved by using the colors in a simulated CVD color space to maximize contrast and to preserve the original color as much as possible. In addition, the dominant colors are extracted for recoloring. They are then propagated to the whole image so that the optimization problem could be solved at a reasonable cost independent of the image size. In the quantitative evaluation, the results of the proposed method are competitive with those of the best existing method. The evaluation involving subjects with CVD demonstrates that the proposed method outperforms the state-of-the-art method in preserving both the information and the naturalness of images. Zhenyang Zhu, Masahiro Toyoura, Kentaro Go, Issei Fujishiro, Kenji Kashiwagi, Xiaoyang Mao |
Signal Process. Image Commun. | 1 |
| 2019 | Processing images for red-green dichromats compensation via naturalness and information-preservation considered recoloringabstractColor vision deficiency (CVD) is caused by anomalies in the cone cells of the human retina. It affects approximately 200 million individuals throughout the world. Although previous studies have proposed compensation methods, contrast and naturalness preservation have not been adequately and simultaneously addressed in the state-of-the-art studies. This paper focuses on red–green dichromats’ compensation and proposes a recoloring algorithm that combines contrast enhancement and naturalness preservation in a unified optimization model. In this implementation, representative color extraction and edit propagation methods are introduced to maintain global and local information in the recolored image. The quantitative evaluation results showed that the proposed method is competitive with state-of-the-art methods. A subjective experiment was also conducted and the evaluation results revealed that the proposed method obtained the best scores in preserving both naturalness and information for individuals with severe red–green CVD. Zhenyang Zhu, Masahiro Toyoura, Kentaro Go, Issei Fujishiro, Kenji Kashiwagi, Xiaoyang Mao |
Vis. Comput. | 1 |
| 2018 | QGLG Automatic Energy Gear-Shifting Mechanism with Flexible QoS Constraint in Cyber-Physical Systems: Designing, Analysis, and EvaluationabstractThis article describes how with the continuous expansion on the volume of data produced by sensors in Cyber Physical Systems, the scale of the cloud storage system has become larger. This will lead to the problems of a high energy consumption rate and a low utilization becoming a serious issue. In order to enhance the effective energy consumption, reduce the invalid energy consumption, and supply more flexible QoS for users in CPS, this article proposes an automatic energy gear-shifting mechanism with flexible QoS constraints (QGLG). The QGLG predicts system load of the follow-up period through a support vector machine model. According to the current system load, the predicted load, and the flexible QoS, QGLG automatically up-shifts and down-shifts among nodes. Substantive results from the simulation experiments done on GridSim show that the QGLG can achieve energy consumption reduction while satisfying the user's flexible QoS requirements. Compared with a similar energy-reducing mechanism, QGLG has its obvious advantage when considering the requirements of user with energy saved notwithstanding. Xindong You, Yeli Li, Zhenyang Zhu, Lifeng Yu, Dawei Sun 0001 |
J. Database Manag. | 3 |