Yi Xiao 0004

dblp:61/6921-4 · DBLP profile ↗
← Back
46ranked-venue papers
15as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 22 · 7 first-author · 7 since 2021Artificial intelligence and machine learning · 21 · 9 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 InspirationGraph for Progressive Design Space Exploration
abstract
Text-to-image (T2I) models demonstrate strong generative capabilities and are increasingly used in design. However, their support for early exploratory ideation remains limited. Their linear, one-shot interaction paradigm aligns more closely with convergent, refinement-oriented stages of design. To address this gap, we present an interaction paradigm supporting early-stage ideation with T2I models, with a particular focus on novice designers. It introduces a dimension–attribute dictionary to guide prompt construction progressively and employs a dynamic, editable tree structure to help users organize and navigate their design space. Based on this paradigm, we developed a prototyping tool named InspirationGraph, focusing on the product design field. The results from a user study involving 24 participants highlight how this structured exploration approach supports divergent thinking and reduces cognitive load. We also uncover varying ideation patterns among designers and offer actionable insights into how T2I systems can be reimagined to better support the early-stage design.
Suxiang Ling, Yi Xiao 0004, Ruoxuan Ma, Guangpeng Wei, Andrew Chi-Sing Leung
CHI2
2026 RPGAgent: Driving Coherent Story-to-Play Generation with an LLM-Based Multi-Agent System
abstract
Recent advances in LLMs have enabled new possibilities for creative content generation, yet their use in game design is often limited by poor integration across creative components, particularly for novice designers aiming to rapidly prototype playable concepts. Guided by the Elemental Tetrad framework, we present RPGAgent, an LLM-driven multi-agent system specifically designed to assist novice game creators in transforming a short story outline into a playable game. Specialized agents exchange structured data to generate coherent narrative, scene, and gameplay mechanics, ensuring structural correctness and consistency between story and world. By combining LLM-based generation with procedural content creation, the system offers a controllable and interpretable workflow. In a within-subjects study with 18 participants, RPGAgent outperformed a GPT-assisted baseline in both user experience and creative satisfaction during the prototyping of playable RPGs. These results demonstrate the potential of collaborative multi-agent frameworks for structured, AI-assisted game design.
Shunan Zhang, Yi Xiao 0004, Ruoxuan Ma, Andrew Chi-Sing Leung
CHI2
2026 CoNode: Visualizing Workflows for Knowledge Reuse and Recombination in Team-AI Collaborative Design
abstract
In early-stage industrial design, teams generate essential but fragile process knowledge—semantic tags, sketches, exploration paths—that is rarely captured or reused but which may be useful at latter design stages, and AI could be used for this purpose. Yet existing AI creativity tools remain outcome-oriented, offering limited support for preserving, tracing, or recombining underlying reasoning. Our formative study (N=6) revealed persistent challenges in team–AI ideation across sessions and collaborators, including semantic–visual fragmentation, context loss, and cross-tool disruption. These insights inspired CoNode, a two-layer system that embeds AI nodes within a shared whiteboard through triplet workflows and augments them with workflow-level consolidation, reuse, and recombination via the CoSense module. We conducted a two-stage evaluation: User Study I (N=12) validates CoNode's foundational interaction paradigm layer, and User Study II (N=30) evaluates its process-oriented knowledge layer. Results show that CoNode significantly improves knowledge consolidation, reuse, and recombination, effectively facilitating the collaborative processes and demonstrating how generative AI can evolve process knowledge across collaborative rounds.
Yi Xiao 0004, Guangpeng Wei, Suxiang Ling, Ruoxuan Ma, Andrew Chi-Sing Leung
CHI2
2026 FGFDL: Frame grouping and feature dissimilarity learning for reverse video recognition
Xianyi Zhu, Yi Xiao 0004, Yan Zheng 0003
Pattern Recognit. Lett.2
2025 Med-SER: Enhancing Reasoning Interpretability in Medical Visual Question Answering via Structured Chain-of-Thought
abstract
Existing Medical Visual Question Answering (Med-VQA) methods typically rely on either direct answer generation or Chain-of-Thought (CoT) reasoning, both of which suffer from limited interpretability or logical inconsistency. To overcome these challenges, we propose Med-SER, a novel framework featuring Structured Chain-of-Thought (SCoT), which decomposes reasoning into four clinically grounded stages: Summary, Caption, Reasoning, and Conclusion. To facilitate training, we construct VQA-RAD-SCoT, the first Med-VQA dataset annotated with structured reasoning chains. Med-SER further introduces a Dual-Channel Visual Projection (DCVP) module to extract both holistic and stage-specific visual features, and a Dual Dynamic Supervision (DDS) mechanism combining adaptive stage-aware weighting and logical consistency loss. Experiments on VQA-RAD-SCoT demonstrate that Med-SER demonstrates the potential of Med-SER to establish a new interpretable and trustworthy paradigm for Med-VQA.
Jinhao Qiao, Yi Xiao 0004, Hongshan Yu, Yan Zheng 0003
BIBM5
2024 VAG: Voxel Attenuation Grid For Sparse-View CBCT Reconstruction
abstract
Sparse view CBCT reconstruction has become one of the important research fields to reduce the radiation impact of CT scanning. However, the reconstruction of high-quality 3D CT volumes from sparse and noisy CBCT data still faces challenges such as slow convergence, long computation time, and increased noise. In light of these issues, we propose a voxel attenuation grid representation to explicitly model the attenuation field of the 3D CT volume. Since this representation does not involve the implementation of neural networks, our method for reconstruction is extremely fast. Furthermore, trim regularization and total variation regularization terms are introduced on top of the mean square error loss to optimize the voxel attenuation grid and significantly reduce the noise in the reconstructed 3D CT volume. Experiments on the NSCLC dataset demonstrate the superiority of our method and its potential in clinical applications. Our code will be available at: https://github.com/qiaodongxing/VAG.
Jinhao Qiao, Yi Xiao 0004, Hongshan Yu, Yan Zheng 0003
ICIP4
2024 Temporal vectorized visibility for direct illumination of animated models
abstract
Direct illumination rendering is an important technique in computer graphics. Precomputed radiance transfer algorithms can provide high quality rendering results in real time, but they can only support rigid models. On the other hand, ray tracing algorithms are flexible and can gracefully handle animated models. With NVIDIA RTX and the AI denoiser, we can use ray tracing algorithms to render visually appealing results in real time. Visually appealing though, they can deviate from the actual one considerably. We propose a visibility-boundary edge oriented infinite triangle bounding volume hierarchy (BVH) traversal algorithm to dynamically generate visibility in vector form. Our algorithm utilizes the properties of visibility-boundary edges and infinite triangle BVH traversal to maximize the efficiency of the vector form visibility generation. A novel data structure, temporal vectorized visibility, is proposed, which allows visibility in vector form to be shared across time and further increases the generation efficiency. Our algorithm can efficiently render close-to-reference direct illumination results. With the similar processing time, it provides a visual quality improvement around 10 dB in terms of peak signal-to-noise ratio (PSNR) w.r.t. the ray tracing algorithm reservoir-based spatiotemporal importance resampling (ReSTIR).
Zhenni Wang, Tze-Yui Ho, Yi Xiao 0004, Andrew Chi-Sing Leung
Comput. Vis. Media3
2023 Semantic-Aware Gated Fusion Network For Interactive Colorization
abstract
Deep neural networks boost many successful colorization methods, including automatic, interactive, and exemplar-based methods. Among them, interactive methods with global and/or local inputs are probably the most flexible to accurately add colors to a gray image. However, due to the sparseness of input-semantic correspondences, existing methods encounter difficulties in distributing inputs into correct regions. Moreover, they simply add or concatenate the features of different inputs to the network before color reconstruction, which cannot balance the influences of different inputs. To this end, we propose a novel interactive colorization network, which explicitly builds input-semantic correspondences with an attention mechanism and proposes a gated feature fusion module to balance the influences of global and local inputs. We further apply a differentiable histogram loss to impose a smooth impact of the global inputs. Extensive experiments demonstrate that our method can flexibly control the results and outperforms other state-of-the-art interactive methods.
Yi Xiao 0004, Yan Zheng 0003, Zhenni Wang, Andrew Chi-Sing Leung
ICASSP2
2023 A Local Correspondence-Aware Hybrid CNN-GCN Model for Single-Image Human Body Reconstruction
abstract
Reconstructing a 3D human body mesh from a monocular image is a challenging inverse problem because of occlusion and complicated human articulations. Recent deep learning-based methods have made significant progress in single-image human reconstruction. Most of these works are either model-based methods or model-free methods. However, model-based methods always suffer detail losses due to the limited parameter space, and model-free methods are hard to directly recover satisfactory results from images due to the use of a shared global feature for all vertices and the domain gap between 2D regular images and 3D irregular meshes. To resolve these issues, we propose a hybrid model, which combines the advantages of both model based approach and model-free approach to estimate a 3D human mesh in a coarse-to fine manner. Initially, we utilize a convolutional neural network (CNN) to estimate the parameters of a Skinned Multi-Person Linear Model (SMPL), which allows us to generate a coarse human mesh. After that, the vertex coordinates of the coarse human mesh are further refined by a graph convolutional neural network (GCN). Unlike previous GCN-based methods, whose vertex coordinates are recovered from a shared global feature, we propose a LOcal CorRespondence-Aware (LOCRA) module to extract local special features for each vertex. To make the local features related to the human pose, we also add a keypoint-related loss to supervise the training process of the LOCRA module. Experiments demonstrate that our hybrid model with the LOCRA module outperforms existing methods on multiple public benchmarks.
Qingping Sun, Yi Xiao 0004, Shizhe Zhou, Andrew Chi-Sing Leung, Xin Su 0004
IEEE Trans. Multim.2
2022 Histogram-Guided Semantic-Aware Colorization
abstract
User-guided colorization can predict the colors of a grayscale image according to user inputs, including exemplar images, local inputs and global inputs. Global inputs-based methods are probably the easiest ones to use, but are hard to distribute the input colors into correct regions, due to the lack of color-semantic correspondences. In this paper, we propose a novel histogram-guided semantic-aware colorization method, which explicitly builds the correspondences between global colors and local features with an attention mechanism and uses a differentiable histogram loss to impose the histogram of the results. Our method starts with a semantic-aware subnetwork to build the color-semantic correspondences, followed by a colorization subnetwork to reconstruct the color channels. Experiments demonstrate that our method can effectively control the results with the input histogram. Extensive visual, numerical and user study comparisons show that our method outperforms other global input-based state-of-the-art methods in color naturalness and consistency.
Yi Xiao 0004, Qingping Sun, Fangqiang Xu, Andrew Chi-Sing Leung
ICASSP2
2022 DNN-kWTA With Bounded Random Offset Voltage Drifts in Threshold Logic Units
abstract
The dual neural network-based$k$-winner-take-all (DNN-$k$WTA) is an analog neural model that is used to identify the$k$largest numbers from$n$inputs. Since threshold logic units (TLUs) are key elements in the model, offset voltage drifts in TLUs may affect the operational correctness of a DNN-$k$WTA network. Previous studies assume that drifts in TLUs follow some particular distributions. This brief considers that only the drift range, given by$[-\Delta, \Delta]$, is available. We consider two drift cases: time-invariant and time-varying. For the time-invariant case, we show that the state of a DNN-$k$WTA network converges. The sufficient condition to make a network with the correct operation is given. Furthermore, for uniformly distributed inputs, we prove that the probability that a DNN-$k$WTA network operates properly is greater than$(1-2\Delta)^{n}$. The aforementioned results are generalized for the time-varying case. In addition, for the time-invariant case, we derive a method to compute the exact convergence time for a given data set. For uniformly distributed inputs, we further derive the mean and variance of the convergence time. The convergence time results give us an idea about the operational speed of the DNN-$k$WTA model. Finally, simulation experiments have been conducted to validate those theoretical results.
Wenhao Lu, Andrew Chi-Sing Leung, John Sum, Yi Xiao 0004
IEEE Trans. Neural Networks Learn. Syst.4
2022 Interactive Deep Colorization and its Application for Image Compression
abstract
Recent methods based on deep learning have shown promise in converting grayscale images to colored ones. However, most of them only allow limited user inputs (no inputs, only global inputs, or only local inputs), to control the output colorful images. The possible difficulty lies in how to differentiate the influences of different inputs. To solve this problem, we propose a two-stage deep colorization method allowing users to control the results by flexibly setting global inputs and local inputs. The key steps include enabling color themes as global inputs by extracting K mean colors and generating K-color maps to define a global theme loss, and designing a loss function to differentiate the influences of different inputs without causing artifacts. We also propose a color theme recommendation method to help users choose color themes. Based on the colorization model, we further propose an image compression scheme, which supports variable compression ratios in a single network. Experiments on colorization show that our method can flexibly control the colorized results with only a few inputs and generate state-of-the-art results. Experiments on compression show that our method achieves much higher image quality at the same compression ratio when compared to the state-of-the-art methods.
Yi Xiao 0004, Peiyao Zhou, Yan Zheng 0003, Andrew Chi-Sing Leung, Ladislav Kavan
IEEE Trans. Vis. Comput. Graph.1
2021 Edge-Aware Multi-Scale Progressive Colorization
abstract
Image colorization recovers a colorful image from a grayscale one. Trained by large-scale datasets, recent deep neural networks based methods can produce impressive colorful images. However, they usually directly train a single network using training images of fixed resolution. It is hard for such a single network to learn the features of different scales for colorization. Moreover, they are prone to generate color bleedings and blurry details around objects boundaries. To address these problems, we propose a novel edge-aware multi-scale progressive network (EMSPN). The key idea is to train a series of multi-scale networks in a progressive manner, so that the network in finer scales can leverage the outputs of its previous scale. In addition, we also propose an edge-map loss to effectively prevent bleedings and blurs around the image edges. Experimental results show that our work outperforms existing methods and achieves state-of-the-art results.
Guanghua Tan, Yi Xiao 0004, Fangqiang Xu, Andrew Chi-Sing Leung
ICASSP3
2021 Blind image super-resolution based on prior correction network
Yihao Luo, Yi Xiao 0004, Xianyi Zhu, Tianjiang Wang, Qi Feng 0003, Zehan Tan
Neurocomputing3
2021 History-based attention in Seq2Seq model for multi-label text classification
Yaoqiang Xiao, Jin Yuan 0002, Songrui Guo, Yi Xiao 0004, Zhiyong Li 0001
Knowl. Based Syst.5
2020 Sketchppnet: A Joint Pixel and Point Convolutional Neural Network For Low Resolution Sketch Image Recognition
abstract
Sketch recognition using deep neural networks have become a recent trend. However, traditional pixel (image) based convolutional neural networks show poor recognizing performance on low resolution (LR) sketch image due to the loss of image details. To solve this problem, we propose a joint pixel and point convolutional neural network for LR sketch image recognition. The network, equipped with both image convolution and point convolution, can simultaneously handle both the image and point representation of sketches. Furthermore, we propose a hybrid classifier, a corresponding loss function, and a training scheme to better extract features for recognition. Experimental results show that our method outperforms state-of-art deep neural networks.
Xianyi Zhu, Yi Xiao 0004, Yan Zheng 0003, Guanghua Tan, Shizhe Zhou
ICASSP2
2020 2D freehand sketch labeling using CNN and CRF
Xianyi Zhu, Yi Xiao 0004, Yan Zheng 0003
Multim. Tools Appl.2
2020 Stroke classification for sketch segmentation by fine-tuning a developmental VGGNet16
Xianyi Zhu, Jin Yuan 0002, Yi Xiao 0004, Yan Zheng 0003, Zheng Qin 0001
Multim. Tools Appl.3
2020 Gated CNN: Integrating multi-scale feature layers for object detection
Jin Yuan 0002, Heng-Chang Xiong, Yi Xiao 0004, Weili Guan, Meng Wang 0001, Richang Hong, Zhiyong Li 0001
Pattern Recognit.3
2020 Image Captioning with a Joint Attention Mechanism by Visual Concept Samples
abstract
The attention mechanism has been established as an effective method for generating caption words in image captioning; it explores one noticed subregion in an image to predict a related caption word. However, even though the attention mechanism could offer accurate subregions to train a model, the learned captioner may predict wrong, especially for visual concept words, which are the most important parts to understand an image. To tackle the preceding problem, in this article we propose Visual Concept Enhanced Captioner, which employs a joint attention mechanism with visual concept samples to strengthen prediction abilities for visual concepts in image captioning. Different from traditional attention approaches that adopt one LSTM to explore one noticed subregion each time, Visual Concept Enhanced Captioner introduces multiple virtual LSTMs in parallel to simultaneously receive multiple subregions from visual concept samples. Then, the model could update parameters by jointly exploring these subregions according to a composite loss function. Technically, this joint learning is helpful in finding the common characters of a visual concept, and thus it enhances the prediction accuracy for visual concepts. Moreover, by integrating diverse visual concept samples from different domains, our model can be extended to bridge visual bias in cross-domain learning for image captioning, which saves the cost for labeling captions. Extensive experiments have been conducted on two image datasets (MSCOCO and Flickr30K), and superior results are reported when comparing to state-of-the-art approaches. It is impressive that our approach could significantly increase BLUE-1 and F1 scores, which demonstrates an accuracy improvement for visual concepts in image captioning.
Jin Yuan 0002, Songrui Guo, Yi Xiao 0004, Zhiyong Li 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2019 PA-RetinaNet: Path Augmented RetinaNet for Dense Object Detection
Guanghua Tan, Yi Xiao 0004
ICANN (2)3
2019 Action Recognition Based on Divide-and-Conquer
Guanghua Tan, Yi Xiao 0004
ICANN (3)3
2019 Interactive Deep Colorization Using Simultaneous Global and Local Inputs
abstract
Colorization methods using deep neural networks have become a recent trend. However, most of them do not allow user inputs, or only allow limited user inputs (only global inputs or only local inputs), to control the output colorful images. The possible reason is that it's difficult to differentiate the influence of different kind of user inputs in network training. To solve this problem, we propose a novel deep colorization method allowing inputting global and local inputs simultaneously or individually, which is not supported in previous deep colorization methods. The key steps include designing a neural network model that can appropriately combine the different inputs, and designing an appropriate loss function that can differentiate the influence of different inputs. Experimental results show that our method can magnificently control the colorized results and generate state-of-art results.
Yi Xiao 0004, Peiyao Zhou, Yan Zheng 0003, Andrew Chi-Sing Leung
ICASSP1
2019 Joint residual pyramid for joint image super-resolution
Yan Zheng 0003, Yi Xiao 0004, Xianyi Zhu, Jin Yuan 0002
J. Vis. Commun. Image Represent.3
2018 Gradient-Guided DCNN for Inverse Halftoning and Image Expanding
Yi Xiao 0004, Chao Pan 0009, Yan Zheng 0003, Xianyi Zhu, Zheng Qin 0001, Jin Yuan 0002
ACCV (4)1
2018 Part-Level Sketch Segmentation and Labeling Using Dual-CNN
Xianyi Zhu, Yi Xiao 0004, Yan Zheng 0003
ICONIP (1)2
2018 Joint Residual Pyramid for Depth Map Super-Resolution
Yi Xiao 0004, Yan Zheng 0003, Xianyi Zhu
PRICAI (1)1
2018 PatchSwapper: A novel real-time single-image editing technique by region-swapping
Shizhe Zhou, Chengfeng Zhou, Yi Xiao 0004, Guanghua Tan
Comput. Graph.3
2018 Scribble-based gradient mesh recoloring
Yi Xiao 0004, Ning Dou, Andrew Chi-Sing Leung, Yukun Lai
Multim. Tools Appl.2
2018 Summed Area Tables for Cube Maps
abstract
The original Summed Area Table (SAT) structure is designed for handling 2D rectangular data. Due to the nature of spherical functions, the SAT structure cannot handle cube maps directly. This paper proposes a new SAT structure for cube maps and develops the corresponding lookup algorithm. Our formulation starts by considering a cube map as part of an auxiliary 3D function defined in a 3D rectangular space. We interpret the 2D integration process over the cube map surface as a 3D integration over the auxiliary 3D function. One may suggest that we can create a 3D SAT for this auxiliary function, and then use the 3D SAT to achieve the 3D integration. However, it is not practical to generate or store this special 3D SAT directly. This 3D SAT has some nice properties that allow us to store it in a storage-friendly data structure, namely Summed Area Cube Map (SACM). A SACM can be stored in a standard cube map texture. The lookup algorithm of our SACM structure can be implemented efficiently on current graphics hardware. In addition, the SACM structure inherits the favorable properties of the original SAT structure.
Yi Xiao 0004, Tze-Yui Ho, Andrew Chi-Sing Leung
IEEE Trans. Vis. Comput. Graph.1
2017 Efficient image colorization based on seed pixel selection
Bo Ou, Yi Xiao 0004
Multim. Tools Appl.3
2016 Objective Function and Learning Algorithm for the General Node Fault Situation
abstract
Fault tolerance is one interesting property of artificial neural networks. However, the existing fault models are able to describe limited node fault situations only, such as stuck-at-zero and stuck-at-one. There is no general model that is able to describe a large class of node fault situations. This paper studies the performance of faulty radial basis function (RBF) networks for the general node fault situation. We first propose a general node fault model that is able to describe a large class of node fault situations, such as stuck-at-zero, stuck-at-one, and the stuck-at level being with arbitrary distribution. Afterward, we derive an expression to describe the performance of faulty RBF networks. An objective function is then identified from the formula. With the objective function, a training algorithm for the general node situation is developed. Finally, a mean prediction error (MPE) formula that is able to estimate the test set error of faulty networks is derived. The application of the MPE formula in the selection of basis width is elucidated. Simulation experiments are then performed to demonstrate the effectiveness of the proposed method.
Yi Xiao 0004, Ruibin Feng, Andrew Chi-Sing Leung, John Sum
IEEE Trans. Neural Networks Learn. Syst.1
2015 Optimization-Based Gradient Mesh Colour Transfer
abstract
Abstract In vector graphics, gradient meshes represent an image object by one or more regularly connected grids. Every grid point has attributes as the position, colour and gradients of these quantities specified. Editing the attributes of an existing gradient mesh (such as the colour gradients) is not only non‐intuitive but also time‐consuming. To facilitate user‐friendly colour editing, we develop an optimization‐based colour transfer method for gradient meshes. The key idea is built on the fact that we can approximate a colour transfer operation on gradient meshes with a linear transfer function. In this paper, we formulate the approximation as an optimization problem, which aims to minimize the colour distribution of the example image and the transferred gradient mesh. By adding proper constraints, i.e. image gradients, to the optimization problem, the details of the gradient meshes can be better preserved. With the linear transfer function, we are able to edit the colours and colour gradients of the mesh points automatically, while preserving the structure of the gradient mesh. The experimental results show that our method can generate pleasing recoloured gradient meshes.
Yi Xiao 0004, Andrew Chi-Sing Leung, Yukun Lai, Tien-Tsin Wong
Comput. Graph. Forum1
2015 GPU Accelerated Self-Organizing Map for High Dimensional Data
Yi Xiao 0004, Ruibin Feng, Zi-Fa Han, Andrew Chi-Sing Leung
Neural Process. Lett.1
2015 Online Training for Open Faulty RBF Networks
Yi Xiao 0004, Ruibin Feng, Andrew Chi-Sing Leung, John Sum
Neural Process. Lett.1
2015 Properties and Performance of Imperfect Dual Neural Network-Based k WTA Networks
abstract
The dual neural network (DNN)-based k -winner-take-all ( k WTA) model is an effective approach for finding the k largest inputs from n inputs. Its major assumption is that the threshold logic units (TLUs) can be implemented in a perfect way. However, when differential bipolar pairs are used for implementing TLUs, the transfer function of TLUs is a logistic function. This brief studies the properties of the DNN- kWTA model under this imperfect situation. We prove that, given any initial state, the network settles down at the unique equilibrium point. Besides, the energy function of the model is revealed. Based on the energy function, we propose an efficient method to study the model performance when the inputs are with continuous distribution functions. Furthermore, for uniformly distributed inputs, we derive a formula to estimate the probability that the model produces the correct outputs. Finally, for the case that the minimum separation ∆min of the inputs is given, we prove that if the gain of the activation function is greater than 1/4∆min max(ln 2n, 2 ln 1 - ϵ/ϵ ), then the network can produce the correct outputs with winner outputs greater than 1-ϵ and loser outputs less than ϵ, where ϵ is the threshold less than 0.5.
Ruibin Feng, Andrew Chi-Sing Leung, John Sum, Yi Xiao 0004
IEEE Trans. Neural Networks Learn. Syst.4
2015 All-Frequency Direct Illumination with Vectorized Visibility
abstract
Many existing pre-computed radiance transfer (PRT) approaches for all-frequency lighting store the information of a 3D object in the pre-vertex manner. To preserve the fidelity of high frequency effects, the 3D object must be tessellated densely. Otherwise, rendering artifacts due to interpolation may appear. This paper presents an all-frequency lighting algorithm for direct illumination based on a new visibility representation which approximates a visibility function using a sequence of 3D vectors. The algorithm is able to construct the visibility function of an on-screen pixel on-the-fly. Hence even though the 3D object is not tessellated densely, the rendering artifacts can be suppressed greatly. Besides, a summed area table based rendering algorithm, which is able to handle the integration over a non-axis aligned polygon, is developed. Using our approach, we can rotate lighting environment, change view point, and adjust the shininess of the 3D object in a real-time manner. Experimental results show that our approach can render plausible all-frequency lighting effects for direct illumination in real-time, especially for specular shadows, which are difficult for other methods to obtain.
Tze-Yui Ho, Yi Xiao 0004, Ruibin Feng, Andrew Chi-Sing Leung, Tien-Tsin Wong
IEEE Trans. Vis. Comput. Graph.2
2014 Tensor Completion Based on Structural Information
Zi-Fa Han, Ruibin Feng, Longting Huang, Yi Xiao 0004, Andrew Chi-Sing Leung, Hing-Cheung So
ICONIP (2)4
2013 GPU Accelerated Spherical K-Means Training
Yi Xiao 0004, Ruibin Feng, Andrew Chi-Sing Leung, John Sum
ICONIP (2)1
2013 Concentric Spherical Representation for Omnidirectional Soft Shadow
abstract
Abstract Soft shadows play an important role in photo‐realistic rendering. Although there are many efficient soft shadow algorithms, most of them focus on the one‐side light source situation, where a planar light source is on the outside of the scene. In fact, in many situations, such as games, light sources are omnidirectional. They may be surrounded by a number of 3D objects. This paper proposes a soft shadow algorithm for the omnidirectional situation. We develop a concentric spherical representation to model the behaviour of omnidirectional light sources. To provide better rendering results, a novel summed area table based filtering scheme for spherical functions is proposed. In addition, we utilize unicube mapping, which samples the spherical space more uniformly, to further improve the filtering quality.
Yi Xiao 0004, Andrew Chi-Sing Leung, Tze-Yui Ho, Tien-Tsin Wong
Comput. Graph. Forum1
2013 Example-Based Color Transfer for Gradient Meshes
abstract
Editing a photo-realistic gradient mesh is a tough task. Even only editing the colors of an existing gradient mesh can be exhaustive and time-consuming. To facilitate user-friendly color editing, we develop an example-based color transfer method for gradient meshes, which borrows the color characteristics of an example image to a gradient mesh. We start by exploiting the constraints of the gradient mesh, and accordingly propose a linear-operator-based color transfer framework. Our framework operates only on colors and color gradients of the mesh points and preserves the topological structure of the gradient mesh. Bearing the framework in mind, we build our approach on PCA-based color transfer. After relieving the color range problem, we incorporate a fusion-based optimization scheme to improve color similarity between the reference image and the recolored gradient mesh. Finally, a multi-swatch transfer scheme is provided to enable more user control. Our approach is simple, effective, and much faster than color transferring the rastered gradient mesh directly. The experimental results also show that our method can generate pleasing recolored gradient meshes.
Yi Xiao 0004, Andrew Chi-Sing Leung, Yukun Lai, Tien-Tsin Wong
IEEE Trans. Multim.1
2012 Decouple implementation of weight decay for recursive least square
Andrew Chi-Sing Leung, Yi Xiao 0004, Kwok-Wo Wong
Neural Comput. Appl.2
2012 Self-organizing map-based color palette for high-dynamic range texture compression
Yi Xiao 0004, Andrew Chi-Sing Leung, Ping-Man Lam, Tze-Yui Ho
Neural Comput. Appl.1
2012 Analysis on the Convergence Time of Dual Neural Network-Based WTA
abstract
A k-winner-take-all (kWTA) network is able to find out the k largest numbers from n inputs. Recently, a dual neural network (DNN) approach was proposed to implement the kWTA process. Compared to the conventional approach, the DNN approach has much less number of interconnections. A rough upper bound on the convergence time of the DNN-kWTA model, which is expressed in terms of input variables, was given. This brief derives the exact convergence time of the DNN-kWTA model. With our result, we can study the convergence time without spending excessive time to simulate the network dynamics. We also theoretically study the statistical properties of the convergence time when the inputs are uniformly distributed. Since a nonuniform distribution can be converted into a uniform one and the conversion preserves the ordering of the inputs, our theoretical result is also valid for nonuniformly distributed inputs.
Yi Xiao 0004, Yuxin Liu 0008, Andrew Chi-Sing Leung, John Sum, Kevin I.-J. Ho
IEEE Trans. Neural Networks Learn. Syst.1
2011 Comparison between the Applications of Fragment-Based and Vertex-Based GPU Approaches in K-Means Clustering of Time Series Gene Expression Data
Yau-King Lam, Wuchao Situ, Peter Wai-Ming Tsang, Andrew Chi-Sing Leung, Yi Xiao 0004
ICONIP (1)5
2011 A GPU implementation for LBG and SOM training
Yi Xiao 0004, Andrew Chi-Sing Leung, Tze-Yui Ho, Ping-Man Lam
Neural Comput. Appl.1