Frank Guan

dblp:202/0204 · also Frank Yunqing Guan, Yunqing Guan 0001 · DBLP profile ↗
← Back
12ranked-venue papers
0as first author
8since 2021 · last 2026
0000-0003-0549-142XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 7 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multi-process thermodynamic graph learning for 2D fluid simulation
abstract
Solving partial differential equations (PDEs) for fluid simulation is computationally expensive, especially when dealing with complex geometries and high-resolution meshes. Recent advances in physics-informed graph neural networks (PIGNNs) have demonstrated potential in approximating such simulations more efficiently. Particularly, thermodynamic informed graph neural networks (TIGNNs) offer a promising data-driven alternative to traditional PDE solvers for fluid simulations. However, existing TIGNN implementations suffer from significant training inefficiencies, requiring prolonged runtimes and high memory consumption due to the need to maintain large parameter matrices in GPU memory. Inspired by multi-processor strategies in deformable solid simulations, we propose a novel multi-processor thermodynamic-informed graph neural network (MP-TIGNN) architecture to significantly accelerate training without compromising accuracy. Our approach enables faster convergence and reduces memory usage with mixed-precision training by leveraging fully sharded data parallelism (FSDP) across multiple GPUs. Experimental results show that our approach reduces training time by approximately 70% compared to the original setup while maintaining similar prediction accuracy.
Aik Beng Ng, Simon See, Zhengkui Wang, Frank Guan
Virtual Real. Intell. Hardw.5
2025 A Lightweight Method for Generating Precise Number of 3D Object Instances from Text Prompts
abstract
Generating the precise number of 3D object instances from a user prompt is critical for building virtual worlds and augmenting 3D datasets. However, current generative models often struggle to honor explicit quantity instructions, even in straightforward prompts involving a single object type (e.g. “six chairs”), mainly because they process content generation holistically, prioritizing visual realism and global coherence, which often leads to the fuzziness in quantity fidelity of the results. We propose BatchGen, a new framework that enhances the precision of generating the exact number of individual 3D object instances from text, consistently delivering accurate results. There are two stages to our design: 1) A joint intent-slot model processes and interprets the user's utterance, and 2) the slot outputs serve as the text prompts to a tailored Text-to-3D model, which generates the correct number of specified 3D object instances. Experiments demonstrate that compared to existing methods, BatchGen excels at this task. Conceptually, our approach introduces a semantic enforcement mechanism towards ensuring that the generated output aligns exactly with the user's intent. This paves the way for exact, semantically accurate generation. In this paper, we focus specifically on ensuring the correct quantity of objects is generated.
Jin Qi Yeo, Aik Beng Ng, Simon See, Frank Guan
CW5
2025 Amodal3R: Amodal 3D Reconstruction from Occluded 2D Images
Tianhao Wu 0015, Chuanxia Zheng, Frank Guan, Andrea Vedaldi, Tat-Jen Cham
ICCV3
2025 Dynamic Eyebox Steering for Improved Pinlight AR Near-Eye Displays
abstract
An optical-see-through near-eye display (NED) for augmented reality (AR) allows the user to perceive virtual and real imagery simultaneously. Existing technologies for optical-see-through AR NEDs involve trade-offs between key metrics such as field of view (FOV), eyebox size, form factor, etc. We have enhanced an existing compact wide-FOV pinlight AR NED design with real-time 3D pupil localization in order to dynamically steer and thus effectively enlarge the usable eyebox. This is achieved with a dual-camera rig that captures stereoscopic views of the pupils. The 3D pupil location is used to dynamically calculate a display pattern that spatio-temporally modulates the light entering the wearer's eyes. We have built a demonstrable compact prototype and have conducted a user study that indicates the effectiveness of our eyebox steering method (e.g., without eyebox steering, in 10.5% of our tests, users were unable to perceive the test pattern correctly before experiment timeout; with eyebox steering, that fraction decreased dramatically to 1.25%). This is a small yet crucial step in making simple wide-FOV pinlight NEDs usable for human users and not just as demonstration prototypes filmed with a precisely positioned camera standing in for the user's eye. Further contributions of this paper include a detailed description of display design, calibration technique, and user study design, all of which may benefit other NED research.
Xinxing Xia, Zheye Yu, Dongyu Qiu, Andrei State, Tat-Jen Cham, Frank Guan, Henry Fuchs
IEEE Trans. Vis. Comput. Graph.6
2024 Efficient and Accurate Semi-Automatic Neuron Tracing with Extended Reality
abstract
Neuron tracing, alternately referred to as neuron reconstruction, is the procedure for extracting the digital representation of the three-dimensional neuronal morphology from stacks of microscopic images. Achieving accurate neuron tracing is critical for profiling the neuroanatomical structure at single-cell level and analyzing the neuronal circuits and projections at whole-brain scale. However, the process often demands substantial human involvement and represents a nontrivial task. Conventional solutions towards neuron tracing often contend with challenges such as non-intuitive user interactions, suboptimal data generation throughput, and ambiguous visualization. In this paper, we introduce a novel method that leverages the power of extended reality (XR) for intuitive and progressive semi-automatic neuron tracing in real time. In our method, we have defined a set of interactors for controllable and efficient interactions for neuron tracing in an immersive environment. We have also developed a GPU-accelerated automatic tracing algorithm that can generate updated neuron reconstruction in real time. In addition, we have built a visualizer for fast and improved visual experience, particularly when working with both volumetric images and 3D objects. Our method has been successfully implemented with one virtual reality (VR) headset and one augmented reality (AR) headset with satisfying results achieved. We also conducted two user studies and proved the effectiveness of the interactors and the efficiency of our method in comparison with other approaches for neuron tracing.
Zexin Yuan, Jiaqi Xi, Ziqin Gao, Ying Li 0028, Xiaoqiang Zhu, Yun Stone Shi, Frank Guan
IEEE Trans. Vis. Comput. Graph.8
2023 Integration of Virtual Reality with Intelligent Tutoring for High Fidelity Air Traffic Control Training
abstract
Air traffic control plays a significant role and service by ground-based air traffic controllers (ATC) in providing specific and clear advisory guidance to pilots at every stage of a flight. Specifically, the air traffic control purpose is to ensure safety procedures and protocols are adhered to avoid collisions, and to ensure organized and systematic flow of air traffic on the ground and in the air. In this paper, we present the design and implementation of an immersive and collaborative Virtual Reality (VR) training platform that is scalable and cost-effective compared to traditional method of training ATCs in a physical mock-up of a 360-degree air control tower simulator. The use of immersive VR technology through Head Mounted Display (HMD) would not only solve the space constraints but also immerse users in their tasks while supporting better management and analysis of the complex data produced during training. Through the integration of intelligent tutoring that actively tracks the training progress of the trainee, the system facilitates personalized training that has been shown to significantly improve the learning experience.
Alvin Toong Shoon Chan, Frank Guan, Saw Han Soo, Haris Lim Hao Li
CSEDU (2)3
2023 Message from the ISMAR 2023 General Chairs
abstract
It is our great pleasure to welcome you to the 22nd IEEE International Symposium on Mixed and Augmented Reality (ISMAR), held from 16 to 20 October 2023 in Sydney, Australia. ISMAR stands as the foremost international academic conference in the fields of Augmented Reality (AR), Mixed Reality (MR) and Virtual Reality (VR). Over the years, it has continually expanded its horizon, delving into the latest developments in AR, MR, and VR within both commercial and research domains. The conference is organized and supported by IEEE, IEEE Computer Society, and IEEE VGTC.
Barrett Ens, Denis Kalkofen, Gelareh Mohammadi, Frank Guan
ISMAR4
2022 An AI-empowered Cloud Solution towards End-to-End 2D-to-3D Image Conversion for Autostereoscopic 3D Display
abstract
Autostereoscopic displays allow the users to view the 3D content on electronic displays without wearing any glasses. However, the content for glass-free 3D displays needs to be in 3D format such that novel views could be synthesized. Unfortunately, nowadays images/videos are still normally captured in 2D which cannot be directly utilized for glass-free 3D displays. In this paper, we introduce an AI-empowered cloud solution towards end-to-end 2D-to-3D image conversion for autostereoscopic 3D displays, or “CONVAS (3D)” in short. Taking a single 2D image as the input, CONVAS (3D) is able to automatically convert the input 2D image and generate an image suitable for a target autostereoscopic 3D display. It is implemented on a web-based server such that it can allow the users to submit the conversion task and to retrieve the results without geographical constraints.
Jun Wei Lim, Jin Qi Yeo, Xinxing Xia, Frank Guan
VRST4
2020 Towards Eyeglass-style Holographic Near-eye Displays with Statically
abstract
Holography is perhaps the only method demonstrated so far that can achieve a wide field of view (FOV) and a compact eyeglass-style form factor for augmented reality (AR) near-eye displays (NEDs). Unfortunately, the eyebox of such NEDs is impractically small (~ <; 1mm). In this paper, we introduce and demonstrate a design for holographic NEDs with a practical, wide eyebox of ~ 10mm and without any moving parts, based on holographic lenslets. In our design, a holographic optical element (HOE) based on a lenslet array was fabricated as the image combiner with expanded eyebox. A phase spatial light modulator (SLM) alters the phase of the incident laser light projected onto the HOE combiner such that the virtual image can be perceived at different focus distances, which can reduce the vergence-accommodation conflict (VAC). We have successfully implemented a bench-top prototype following the proposed design. The experimental results show effective eyebox expansion to a size of ~ 10mm. With further work, we hope that these design concepts can be incorporated into eyeglass-size NEDs.
Xinxing Xia, Frank Guan, Andrei State, Praneeth Chakravarthula, Tat-Jen Cham, Henry Fuchs
ISMAR2
2020 VEGO: A novel design towards customizable and adjustable head-mounted display for VR
abstract
Virtual Reality (VR) technologies have advanced fast and have been applied to a wide spectrum of sectors in the past few years. VR can provide an immersive experience to users by generating virtual images and displaying the virtual images to the user with a head-mounted display (HMD) which is a primary component of VR. Normally, an HMD contains a list of hardware components, e.g., housing pack, micro LCD display, microcontroller, optical lens, etc. Settings of VR HMD to accommodate the user's inter-pupil distance (IPD) and the user's eye focus power are important for the user's experience with VR. Although various methods have been developed towards IPD and focus adjustments for VR HMD, the increased cost and complexity impede the possibility for users who wish to assemble their own VR HMD for various purposes, e.g., DIY teaching, etc. In our paper, we present a novel design towards building a customizable and adjustable HMD for VR in a cost-effective manner. Modular design methodology is adopted, and the VR HMD can be easily printed with 3D printers. The design also features adjustable IPD and variable distance between the optical lens and the display. It can help to mitigate the vergence and accommodation conflict issue. A prototype of the customizable and adjustable VR HMD has been successfully built up with off-the-shelf components. A VR software program running on Raspberry Pi board has been developed and can be utilized to show the VR effects. A user study with 20 participants is conducted with positive feedback on our novel design. Modular design can be successfully applied for building up VR HMD with 3D printing. It helps to promote the wide application of VR at affordable costs while featuring flexibility and adjustability.
Jia Ming Lee, Xinxing Xia, Clemen Yun Da Ow, Felix Chua, Frank Guan
Virtual Real. Intell. Hardw.5
2019 Towards a Switchable AR/VR Near-eye Display with Accommodation-Vergence and Eyeglass Prescription Support
abstract
In this paper, we present our novel design for switchable AR/VR near-eye displays which can help solve the vergence-accommodation-conflict issue. The principal idea is to time-multiplex virtual imagery and real-world imagery and use a tunable lens to adjust focus for the virtual display and the see-through scene separately. With this novel design, prescription eyeglasses for near- and far-sighted users become unnecessary. This is achieved by integrating the wearer's corrective optical prescription into the tunable lens for both virtual display and see-through environment. We built a prototype based on the design, comprised of micro-display, optical systems, a tunable lens, and active shutters. The experimental results confirm that the proposed near-eye display design can switch between AR and VR and can provide correct accommodation for both.
Xinxing Xia, Frank Guan, Andrei State, Praneeth Chakravarthula, Kishore Rathinavel, Tat-Jen Cham, Henry Fuchs
IEEE Trans. Vis. Comput. Graph.2
2018 Towards Efficient 3D Calibration for Different Types of Multi-view Autostereoscopic 3D Displays
abstract
A novel and efficient 3D calibration method for different types of autostereoscopic multi-view 3D displays is presented in this paper. In our method, a camera is placed at different locations within the viewing volume of a 3D display to capture a series of images that relate to the subset of light rays emitted by the 3D display and arriving at each of the camera positions. Gray code patterns modulate the images shown on the 3D display, helping to significantly reduce the number of images captured by the camera and thereby accelerate the process of calculating the correspondence relationship between the pixels on the 3D display and the locations of the capturing camera. The proposed calibration method has been successfully tested on two different types of multi-view 3D displays and can be easily generalized for calibrating other types of such displays. The experimental results show that this novel 3D calibration method can also be used to improve the image quality by reducing the frequently observed crosstalk that typically exists when multiple users are simultaneously viewing multi-view 3D displays from a range of viewing positions.
Xinxing Xia, Frank Guan, Andrei State, Tat-Jen Cham, Henry Fuchs
CGI2