Zhiying Zhou

dblp:46/2826 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 6 · 5 first-authorGraphics, computer vision, multimedia, augmented reality and games · 5 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Software engineering, systems software and programming languages · 2Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Generative modeling · 56% Segmentation and scene understanding · 28% Deep learning architectures and training · 8%
Computer graphics and multimedia
2 papers
Virtual and augmented reality · 100%

Topics — the 11 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
0.912025
Noise-Consistent Siamese-Diffusion for Medical Image Synthesis and Segmentation · CVPR 2025
Computer vision › Segmentation and scene understanding
medical image segmentation
0.912025
Noise-Consistent Siamese-Diffusion for Medical Image Synthesis and Segmentation · CVPR 2025
Machine learning › Generative modeling › image generation
medical image synthesis
0.912025
Noise-Consistent Siamese-Diffusion for Medical Image Synthesis and Segmentation · CVPR 2025
Machine learning › Deep learning architectures and training
data augmentation
0.312025
Noise-Consistent Siamese-Diffusion for Medical Image Synthesis and Segmentation · CVPR 2025
Computer vision › 3D vision › 3d reconstruction
shape from silhouette
0.122005
Real-Time 3D Human Capture System for Mixed-Reality Art and Entertainment · IEEE Trans. Vis. Comput. Graph. 2005
Live three-dimensional content for augmented reality · IEEE Trans. Multim. 2005
Computer vision › 3D vision
3d reconstruction
0.112005
Live three-dimensional content for augmented reality · IEEE Trans. Multim. 2005
Computer vision › 3D vision › 3d human reconstruction
human volumetric capture
0.112005
Real-Time 3D Human Capture System for Mixed-Reality Art and Entertainment · IEEE Trans. Vis. Comput. Graph. 2005
Virtual and augmented reality
augmented reality
0.112005
Live three-dimensional content for augmented reality · IEEE Trans. Multim. 2005
Virtual and augmented reality › tracking
marker-based tracking
0.012005
Live three-dimensional content for augmented reality · IEEE Trans. Multim. 2005
Virtual and augmented reality
mixed reality
0.012005
Real-Time 3D Human Capture System for Mixed-Reality Art and Entertainment · IEEE Trans. Vis. Comput. Graph. 2005
Personal fabrication and tangible interfaces
tangible interaction
0.012005
Real-Time 3D Human Capture System for Mixed-Reality Art and Entertainment · IEEE Trans. Vis. Comput. Graph. 2005

Methods — techniques the papers use, named apart from their topics

siamese diffusion · 0.9noise consistency loss · 0.9mask diffusion · 0.9image diffusion · 0.9multi-camera capture · 0.2view synthesis · 0.1shape-from-silhouette · 0.1
YearPublicationVenuePosition
2025 Noise-Consistent Siamese-Diffusion for Medical Image Synthesis and Segmentation
abstract
Deep learning has revolutionized medical image segmentation, yet its full potential remains constrained by the paucity of annotated datasets. While diffusion models have emerged as a promising approach for generating synthetic image-mask pairs to augment these datasets, they paradoxically suffer from the same data scarcity challenges they aim to mitigate. Traditional mask-only models frequently yield low-fidelity images due to their inability to adequately capture morphological intricacies, which can critically compromise the robustness and reliability of segmentation models. To alleviate this limitation, we introduce Siamese-Diffusion, a novel dual-component model comprising Mask-Diffusion and Image-Diffusion. During training, a Noise Consistency Loss is introduced between these components to enhance the morphological fidelity of Mask-Diffusion in the parameter space. During sampling, only Mask-Diffusion is used, ensuring diversity and scalability. Comprehensive experiments demonstrate the superiority of our method. Siamese-Diffusion boosts SANet’s mDice and mIoU by 3.6% and 4.4% on the Polyps, while UNet improves by 1.52% and 1.64% on the ISIC2018.
Kunpeng Qiu, Zhiying Zhou, Mingjie Sun, Yongxin Guo 0002
CVPR3
2025 Adaptively Distilled ControlNet: Accelerated Training and Superior Sampling for Medical Image Synthesis
Kunpeng Qiu, Zhiying Zhou, Yongxin Guo 0002
MICCAI (10)2
2024 Learn From Zoom: Decoupled Supervised Contrastive Learning For WCE Image Classification
abstract
Accurate lesion classification in Wireless Capsule Endoscopy (WCE) images is vital for early diagnosis and treatment of gastrointestinal (GI) cancers. However, this task is confronted with challenges like tiny lesions and background interference. Additionally, WCE images exhibit higher intra-class variance and inter-class similarities, adding complexity. To tackle these challenges, we propose Decoupled Supervised Contrastive Learning for WCE image classification, learning robust representations from zoomed-in WCE images generated by Saliency Augmentor. Specifically, We use uniformly down-sampled WCE images as anchors and WCE images from the same class, especially their zoomed-in images, as positives. This approach empowers the Feature Extractor to capture rich representations from various views of the same image, facilitated by Decoupled Supervised Contrastive Learning. Training a linear Classifier on these representations within 10 epochs yields an impressive 92.01% overall accuracy, surpassing the prior state-of-the-art (SOTA) by 0.72% on a blend of two publicly accessible WCE datasets. Code is available at: https://github.com/Qiukunpeng/DSCL.
Kunpeng Qiu, Zhiying Zhou, Yongxin Guo 0002
ICASSP2
2018 Proximal robust factorization for piecewise planar reconstruction
Wen Kou, Loong Fah Cheong, Zhiying Zhou
Comput. Vis. Image Underst.3
2007 The Role of 3-D Sound in Human Reaction and Performance in Augmented Reality Environments
abstract
Three-dimensional sound's effectiveness in virtual reality (VR) environments has been widely studied. However, due to the big differences between VR and augmented reality (AR) systems in registration, calibration, perceptual difference of immersiveness, navigation, and localization, it is important to develop new approaches to seamlessly register virtual 3-D sound in AR environments and conduct studies on 3-D sound's effectiveness in AR context. In this paper, we design two experimental AR environments to study the effectiveness of 3-D sound both quantitatively and qualitatively. Two different tracking methods are applied to retrieve the 3-D position of virtual sound sources in each experiment. We examine the impacts of 3-D sound on improving depth perception and shortening task completion time. We also investigate its impacts on immersive and realistic perception, different spatial objects identification, and subjective feeling of "human presence and collaboration". Our studies show that applying 3-D sound is an effective way to complement visual AR environments. It helps depth perception and task performance, and facilitates collaborations between users. Moreover, it enables a more realistic environment and more immersive feeling of being inside the AR environment by both visual and auditory means. In order to make full use of the intensity cues provided by 3-D sound, a process to scale the intensity difference of 3-D sound at different depths is designed to cater small AR environments. The user study results show that the scaled 3-D sound significantly increases the accuracy of depth judgments and shortens the searching task completion time. This method provides a necessary foundation for implementing 3-D sound in small AR environments. Our user study results also show that this process does not degrade the intuitiveness and realism of an augmented audio reality environment
Zhiying Zhou, Adrian David Cheok, Xubo Yang
IEEE Trans. Syst. Man Cybern. Part A1
2006 Identifying use cases in source code
Lu Zhang 0023, Zhiying Zhou, Dan Hao 0001, Jiasu Sun
J. Syst. Softw.3
2005 Live three-dimensional content for augmented reality
abstract
We describe an augmented reality system for superimposing three-dimensional (3-D) live content onto two-dimensional fiducial markers in the scene. In each frame, the Euclidean transformation between the marker and the camera is estimated. The equivalent virtual view of the live model is then generated and rendered into the scene at interactive speeds. The 3-D structure of the model is calculated using a fast shape-from-silhouette algorithm based on the outputs of 15 cameras surrounding the subject. The novel view is generated by projecting rays through each pixel of the desired image and intersecting them with the 3-D structure. Pixel color is estimated by taking a weighted sum of the colors of the projections of this 3-D point in nearby real camera images. Using this system, we capture live human models and present them via the augmented reality interface at a remote location. We can generate 384/spl times/288 pixel images of the models at 25 fps, with a latency of <100 ms. The result gives the strong impression that the model is a real 3-D part of the scene.
Farzam Farbiz, Adrian David Cheok, Wei Liu 0009, Zhiying Zhou, Ke Xu 0004, Simon Prince, Mark Billinghurst, Hirokazu Kato 0001
IEEE Trans. Multim.4
2005 Real-Time 3D Human Capture System for Mixed-Reality Art and Entertainment
abstract
A real-time system for capturing humans in 3D and placing them into a mixed reality environment is presented in this paper. The subject is captured by nine cameras surrounding her. Looking through a head-mounted-display with a camera in front pointing at a marker, the user can see the 3D image of this subject overlaid onto a mixed reality scene. The 3D images of the subject viewed from this viewpoint are constructed using a robust and fast shape-from-silhouette algorithm. The paper also presents several techniques to produce good quality and speed up the whole system. The frame rate of our system is around 25 fps using only standard Intel processor-based personal computers. Besides a remote live 3D conferencing and collaborating system, we also describe an application of the system in art and entertainment, named Magic Land, which is a mixed reality environment where captured avatars of human and 3D computer generated virtual animations can form an interactive story and play with each other. This system demonstrates many technologies in human computer interaction: mixed reality, tangible interaction, and 3D communication. The result of the user study not only emphasizes the benefits, but also addresses some issues of these technologies.
Ta Huynh Duy Nguyen, Tran Cong Thien Qui, Ke Xu 0004, Adrian David Cheok, Sze Lee Teo, Zhiying Zhou, Asitha Mallawaarachchi, Shang Ping Lee, Wei Liu 0009, Hui Siang Teo, Le Nam Thang, Yu Li 0024, Hirokazu Kato 0001
IEEE Trans. Vis. Comput. Graph.6
2004 An interactive 3D exploration narrative interface for storytelling
abstract
Storytelling is an effective and important educational means for children. With the augmented reality (AR) technology, storytelling becomes more and more interactive and intuitive in the sense of human computer interaction. Although AR technology is not new, it's potential in education is just beginning to be explored. In this paper, we present a 3D mixed media story cube which uses a foldable cube as the tangible and interactive storytelling interface. Here, we embed both the concept of AR and the concept of tangible interaction. Multiple modalities including speech, 3D audio, 3D graphics and touch are used to provide the user (especially children) with multi-sensory experiences in the process of storytelling. Our research explores a new interface for children education.
Zhiying Zhou, Adrian David Cheok, Jiun Horng Pan, Yu Li 0024
IDC1
2004 An experimental study on the role of software synthesized 3D sound in augmented reality environments
abstract
Investigation of augmented reality (AR) environments has become a popular research topic for engineers, computer and cognitive scientists. Although application oriented studies focused on audio AR environments have been published, little work has been done to vigorously study and evaluate the important research questions of the effectiveness of 3D sound in the AR context, and to what extent the addition of 3D sound would contribute to the AR experience. Thus, we have developed two AR environments and performed vigorous experiments with human subjects to study the effects of 3D sound in the AR context. The study concerns two scenarios. In the first scenario, one participant must use vision only and vision with 3D sound to judge the relative depth of augmented virtual objects. In the second scenario, two participants must co-operate to perform a joint task in a game-based AR environment. Hence, the goals of this study are (1) to access the impact of 3D sound on depth perception in a single-camera AR environment, (2) to study the impact of 3D sound on task performance and the feeling of ‘human presence and collaboration’, (3) to better understand the role of 3D sound in human-computer and human–human interactions, (4) to investigate if gender can affect the impact of 3D sound in AR environments. The outcomes of this research can have a useful impact on the development of audio AR systems which provide more immersive, realistic and entertaining experiences by introducing 3D sound. Our results suggest that 3D sound in AR environment significantly improves the accuracy of depth judgment and improves task performance. Our results also suggest that 3D sound contributes significantly to the feeling of ‘human presence and collaboration’ and helps the subjects to ‘identify spatial objects’.
Zhiying Zhou, Adrian David Cheok, Xubo Yang
Interact. Comput.1
2004 An experimental study on the role of 3D sound in augmented reality environment
abstract
Investigation of augmented reality (AR) environments has become a popular research topic for engineers, computer and cognitive scientists. Although application oriented studies focused on audio AR environments have been published, little work has been done to vigorously study and evaluate the important research questions of the effectiveness of three-dimensional (3D) sound in the AR context, and to what extent the addition of 3D sound would contribute to the AR experience. Thus, we have developed two AR environments and performed vigorous experiments with human subjects to study the effects of 3D sound in the AR context. The study concerns two scenarios. In the first scenario, one participant must use vision only and vision with 3D sound to judge the relative depth of augmented virtual objects. In the second scenario, two participants must cooperate to perform a joint task in a game-based AR environment. Hence, the goals of this study are (1) to access the impact of 3D sound on depth perception in a single-camera AR environment, (2) to study the impact of 3D sound on task performance and the feeling of ‘human presence and collaboration’, (3) to better understand the role of 3D sound in human–computer and human–human interactions, (4) to investigate if gender can affect the impact of 3D sound in AR environments. The outcomes of this research can have a useful impact on the development of audio AR systems, which provide more immersive, realistic and entertaining experiences by introducing 3D sound. Our results suggest that 3D sound in AR environment significantly improves the accuracy of depth judgment and improves task performance. Our results also suggest that 3D sound contributes significantly to the feeling of human presence and collaboration and helps the subjects to ‘identify spatial objects’.
Zhiying Zhou, Adrian David Cheok, Xubo Yang
Interact. Comput.1
2004 3D story cube: an interactive tangible user interface for storytelling with 3D graphics and audio
Zhiying Zhou, Adrian David Cheok, Jiun Horng Pan
Pers. Ubiquitous Comput.1
2003 Discovering Use Cases from Source Code using the Branch-Reserving Call Graph
abstract
Understanding the behavior of a software system is an important problem in program comprehension. Use cases have been accepted as an effective means for describing behavioral requirements for a software system. We propose a novel approach for obtaining use cases from source code. The central idea of our approach is to use the branch-reserving call graph (BRCG) as the intermediate representation of a software program. We also provide strategies for pruning the BRCG to avoid generating too many fine-grained use cases. Use cases, which may just undergo some minor modifications from human experts, can be generated through traversing the pruned BRCG. The contributions of our approach are three-fold, i) This method represents a compromised approach, which differs from both the static and dynamic approaches for use case discovery, ii) This method takes into consideration the fact that it is the branch statements that separate one use case from another in source code. iii) This method can avoid intensive human involvement in determining the final set of use cases. We have also performed a case study for this method on a GNU system.
Lu Zhang 0023, Zhiying Zhou, Dan Hao 0001, Jiasu Sun
APSEC3
2002 Touch-Space: Mixed Reality Game Space Based on Ubiquitous, Tangible, and Social Computing
Adrian David Cheok, Xubo Yang, Zhiying Zhou, Mark Billinghurst, Hirokazu Kato 0001
Pers. Ubiquitous Comput.3