VLDB 2026 Research / reviewers in the wild / expert
Qi Guo 0009
dblp:67/398-9
· DBLP profile ↗
15ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0002-8329-7668ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CLAP: Cross-Layer Adaptive Pipelining Inference Scheduling for Resource-Efficient Edge-Cloud Vision SystemsabstractWith the rapid growth of video-based applications, edge-cloud collaboration has become a mainstream paradigm for large-scale visual inference. However, existing edge-cloud systems primarily emphasize task offloading and static resource allocation, often overlooking the dynamic and heterogeneous nature of real-world scenarios. The significant variability in scene complexity across tasks leads to inefficient system performance. In this article, we propose CLAP, a cross-layer adaptive pipelining inference scheduling framework for edge-cloud vision systems. First, CLAP introduces a lightweight multiscale scene-aware module that accurately characterizes the visual complexity of incoming tasks at different granularities with minimal overhead. Based on this complexity profile, we design an adaptive multi-stage pipeline scheduling strategy, which dynamically adjusts processing granularity and selectively activates stages across edge and cloud nodes. Furthermore, we formulate the resource allocation as a multi-agent decision-making problem and employ cross-layer reinforcement learning to optimize task distribution under complex objectives, efficiently balancing accuracy, delay, and energy consumption. Extensive evaluations on public datasets demonstrate that CLAP can improve the throughput by more than 2.1x compared to traditional cloud-only and edge-only solutions while meeting accuracy requirements. Compared to state-of-the-art edge-cloud methods, CLAP achieves a 3% improvement in inference accuracy while simultaneously reducing end-to-end resource overhead, delay, and energy consumption by over 35%, proving its effectiveness in dynamic, large-scale vision applications. Zheming Yang, Wen Ji 0003, Qi Guo 0009, Jian Zhao 0006, Xingzhou Zhang, Yangyu Zhang, Yang You 0001 |
ACM Trans. Archit. Code Optim. | 3 |
| 2025 | Focal Split: Untethered Snapshot Depth from Differential DefocusabstractWe introduce Focal Split, a handheld, snapshot depth camera with fully onboard power and computing based on depth-from-differential-defocus (DfDD). Focal Split is passive, avoiding power consumption of light sources. Its achromatic optical system simultaneously forms two differentially defocused images of the scene, which can be independently captured using two photosensors in a snapshot. The data processing is based on the DfDD theory, which efficiently computes a depth and a confidence value for each pixel with only 500 floating point operations (FLOPs) per pixel from the camera measurements. We demonstrate a Focal Split prototype, which comprises a handheld custom camera system connected to a Raspberry Pi 5 for real-time data processing. The system consumes 4.9 W and is powered on a 5 V, 10,000 mAh battery. The prototype can measure objects with distances from 0.4 m to 1.2 m, outputting 480×360 sparse depth maps at 2.1 frames per second (FPS) using unoptimized Python scripts. Focal Split is DIY friendly. A comprehensive guide to building your own Focal Split depth camera, code, and additional data can be found at https://focal-split.qiguo.org. Junjie Luo 0009, John Mamish, Alan Fu, Thomas Concannon, Josiah D. Hester, Emma Alexander, Qi Guo 0009 |
CVPR | 7 |
| 2025 | Blurry-Edges: Photon-Limited Depth Estimation from Defocused BoundariesabstractExtracting depth information from photon-limited, defocused images is challenging because depth from defocus (DfD) relies on accurate estimation of defocus blur, which is fundamentally sensitive to image noise. We present a novel approach to robustly measure object depths from photon-limited images along the defocused boundaries. It is based on a new image patch representation, Blurry-Edges, that explicitly stores and visualizes a rich set of low-level patch information, including boundaries, color, and smoothness. We develop a deep neural network architecture that predicts the Blurry-Edges representation from a pair of differently defocused images, from which depth can be calculated using a closed-form DfD relation we derive. The experimental results on synthetic and real data show that our method achieves the highest depth estimation accuracy on photon-limited images compared to a broad range of state-of-the-art DfD methods. Charles James Wagner, Junjie Luo 0009, Qi Guo 0009 |
CVPR | 4 |
| 2025 | Depth from Coupled Optical DifferentiationabstractAbstract We propose depth from coupled optical differentiation, a low-computation passive-lighting 3D sensing mechanism. It is based on our discovery that per-pixel object distance can be rigorously determined by a coupled pair of optical derivatives of a defocused image using a simple, closed-form relationship. Unlike previous depth-from-defocus (DfD) methods that leverage higher-order spatial derivatives of the image to estimate scene depths, the proposed mechanism’s use of only first-order optical derivatives makes it significantly more robust to noise. Furthermore, unlike many previous DfD algorithms with requirements on aperture code, this relationship is proved to be universal to a broad range of aperture codes. We build the first 3D sensor based on depth from coupled optical differentiation. Its optical assembly includes a deformable lens and a motorized iris, which enables dynamic adjustments to the optical power and aperture radius. The sensor captures two pairs of images: one pair with a differential change of optical power and the other with a differential change of aperture scale. From the four images, a depth and confidence map can be generated with only 36 floating point operations per output pixel (FLOPOP), more than ten times lower than the previous lowest passive-lighting depth sensing solution to our knowledge. Additionally, the depth map generated by the proposed sensor demonstrates more than twice the working range of previous DfD methods while using significantly lower computation. Junjie Luo 0009, Emma Alexander, Qi Guo 0009 |
Int. J. Comput. Vis. | 4 |
| 2024 | Generative Quanta Color ImagingabstractThe astonishing development of single-photon cameras has created an unprecedented opportunity for scientific and industrial imaging. However, the high data throughput generated by these 1-bit sensors creates a significant bottleneck for low-power applications. In this paper, we explore the possibility of generating a color image from a single binary frame of a single-photon camera. We evidently find this problem being particularly difficult to standard colorization approaches due to the substantial degree of exposure variation. The core innovation of our paper is an exposure synthesis model framed under a neural ordinary differential equation (Neural ODE) that allows us to generate a contin-uum of exposures from a single observation. This innovation ensures consistent exposure in binary images that col-orizers take on, resulting in notably enhanced colorization. We demonstrate applications of the method in single-image and burst colorization and show superior generative performance over baselines. Project website can be found at https://vishal-s-p.github.io/projects/2023/generative_quanta_color.html Vishal Purohit, Junjie Luo 0009, Yiheng Chi, Qi Guo 0009, Stanley H. Chan, Qiang Qiu 0001 |
CVPR | 4 |
| 2024 | CT-Bound: Robust Boundary Detection from Noisy Images Via Hybrid Convolution and Transformer Neural NetworksabstractWe present CT-Bound, a robust and fast boundary detection method for very noisy images using a hybrid Convolution and Transformer neural network. The proposed architecture decomposes boundary estimation into two tasks: local detection and global regularization. During the local detection, the model uses a convolutional architecture to predict the boundary structure of each image patch in the form of a predefined local boundary representation, the field-of-junctions (FoJ) [9]. Then, it uses a feed-forward transformer architecture to globally refine the boundary structures of each patch to generate an edge map and a smoothed color map simultaneously. Our quantitative analysis shows that CT-Bound outperforms the previous best algorithms in edge detection on very noisy images. It also increases the edge detection accuracy of FoJ-based methods while having a 3-time speed improvement. Finally, we demonstrate that CT-Bound can produce boundary and color maps on real captured images without extra fine-tuning and real-time boundary map and color map videos at ten frames per second. Junjie Luo 0009, Qi Guo 0009 |
MMSP | 3 |
| 2023 | Polarization Multi-Image Synthesis with Birefringent MetasurfacesabstractOptical metasurfaces composed of precisely engineered nanostructures have gained significant attention for their ability to manipulate light and implement distinct functionalities based on the properties of the incident field. Computational imaging systems have started harnessing this capability to produce sets of coded measurements that benefit certain tasks when paired with digital post-processing. Inspired by these works, we introduce a new system that uses a birefringent metasurface with a polarizer-mosaicked photosensor to capture four optically-coded measurements in a single exposure. We apply this system to the task of incoherent opto-electronic filtering, where digital spatial-filtering operations are replaced by simpler, per-pixel sums across the four polarization channels, independent of the spatial filter size. In contrast to previous work on incoherent opto-electronic filtering that can realize only one spatial filter, our approach can realize a continuous family of filters from a single capture, with filters being selected from the family by adjusting the post-capture digital summation weights. To find a metasurface that can realize a set of user-specified spatial filters, we introduce a form of gradient descent with a novel regularizer that encourages light efficiency and a high signal-to-noise ratio. We demonstrate several examples in simulation and with fabricated prototypes, including some with spatial filters that have prescribed variations with respect to depth and wavelength. Dean Hazineh, Soon Wei Daniel Lim, Qi Guo 0009, Federico Capasso, Todd E. Zickler |
ICCP | 3 |
| 2023 | JAVP: Joint-Aware Video Processing with Edge-Cloud Collaboration for DNN InferenceabstractCurrently, massive video inference tasks are processed through edge-cloud collaboration. However, the diverse scenarios make it difficult to allocate the inference tasks efficiently, resulting in many wasted resources. In this paper, we propose a joint-aware video processing (JAVP) architecture for edge-cloud collaboration. First, we develop a multiscale complexity-aware model for predicting task complexity and determining its suitability for edge or cloud servers. The task is subsequently efficiently scheduled to the appropriate servers by integrating complexity with an adaptive resource-aware optimization algorithm. For input tasks, JAVP can dynamically and intelligently select the most appropriate server. The evaluation results on public datasets show that JAVP can improve the through-put by more than 70% compared to traditional cloud-only solutions while meeting accuracy requirements. And JAVP can improve the accuracy by 3%-5% and reduce delay and energy consumption by 16%-50% compared to state-of-the-art edge-cloud solutions. Zheming Yang, Wen Ji 0003, Qi Guo 0009, Zhi Wang 0001 |
ACM Multimedia | 3 |
| 2022 | Design of Personalization Warehouse Management Platform Based on SaaS ModelabstractSaaS is widely used in the fields of operation management, business process outsourcing, data analysis, and information security. However, with the improvement of the degree of information, the generalized SaaS platform is unable to satisfy the requirements of enterprise personalization. SaaS products face great challenges in personalized customization technology due to the application of multi-tenant architecture. The challenges include multi-tenant customization in SaaS model, data isolation during customization, and mapping of multi-tenant virtual warehouse locations to actual locations. Therefore, we decompose personalization technology into metadata driver, cloud data placement and mapping mechanism for research. We decompose the personalized technologies into metadata drive, cloud data placement, and mapping mechanism. In order to solve the problem that traditional personalization customization is unable to be applied in the SaaS field, we propose a multi-tenant personalized warehousing mode architecture, designed a personalization warehouse management platform based on SaaS model, and realized SaaS-based data isolation, interface customization, rapid positioning, and on-demand customization services. Experimental results show that the proposed multi-tenant personalized warehouse model can achieve data isolation and customization, and reflect the advantages of highly automated warehousing, shared storage resources, and on-demand customization in terms of warehouse management, data security, and user experience. Qi Guo 0009, Hongbo Sun 0004, Wen Ji 0003 |
CSCWD | 1 |
| 2020 | Raycast Calibration for Augmented Reality HMDs with Off-Axis Reflective CombinersabstractAugmented reality overlays virtual objects on the real world. To do so, the head mounted display (HMD) needs to be calibrated to establish a mapping between 3D points in the real world with 2D pixels on display panels. This distortion is a high-dimensional function that also depends on pupil position and varifocal settings. We present Raycast calibration, an efficient approach to geometrically calibrate AR displays with off-axis reflective combiners. Our approach requires a small amount of data to estimate a compact, physics-based, and ray-traceable model of the HMD optics. We apply this technique to automatically calibrate an AR prototype with display, SLAM and eye-tracker, without user in the loop. Qi Guo 0009, Huixuan Tang, Aaron Schmitz, Yang Lou, Alexander Fix, Steven Lovegrove, Hauke Strasdat |
ICCP | 1 |
| 2018 | Tackling 3D ToF Artifacts Through Learning and the FLAT Dataset
Qi Guo 0009, Iuri Frosio, Orazio Gallo, Todd E. Zickler, Jan Kautz |
ECCV (1) | 1 |
| 2018 | Focal Flow: Velocity and Depth from Differential Defocus Through Motion
Emma Alexander, Qi Guo 0009, Sanjeev J. Koppal, Steven J. Gortler, Todd E. Zickler |
Int. J. Comput. Vis. | 2 |
| 2018 | Multi-Perspective Tracking for Intelligent VehicleabstractThe multi-camera array has drawn attention of researchers in recent years, and has been configured and deployed on intelligent vehicle to capture the panoramic views. Understanding surroundings is crucial for the ego-vehicle. This paper presents a Multi-perspective Tracking (MPT) framework for intelligent vehicle. An iterative search procedure is proposed to associate detections and tracklets in different perspectives. This procedure iteratively assigns determined states and estimates non-determined states for the detections and tracklets. An inherent determined and non-determined graph is utilized to reinforce this procedure. For more reliable associations between perspectives, a Siamese convolutional neural network is employed to learn feature representation. The supervised classification and verification signals are added to train the network. The features in different conventional stages are integrated together as the discriminative appearance model. The experiments are conducted on a MPT data set with five perspectives. The proposed framework is tested in each pair of adjacent perspectives for the ability to associate target objects between perspectives. Xiangyang Ji, Guanwen Zhang, Qi Guo 0009 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2017 | Focal Track: Depth and Accommodation with Oscillating Lens DeformationabstractThe focal track sensor is a monocular and computationally efficient depth sensor that is based on defocus controlled by a liquid membrane lens. It synchronizes small lens oscillations with a photosensor to produce real-time depth maps by means of differential defocus, and it couples these oscillations with bigger lens deformations that adapt the defocus working range to track objects over large axial distances. To create the focal track sensor, we derive a texture-invariant family of equations that relate image derivatives to scene depth when a lens changes its focal length differentially. Based on these equations, we design a feed-forward sequence of computations that: robustly incorporates image derivatives at multiple scales; produces confidence maps along with depth; and can be trained endto- end to mitigate against noise, aberrations, and other non-idealities. Our prototype with 1-inch optics produces depth and confidence maps at 100 frames per second over an axial range of more than 75cm. Qi Guo 0009, Emma Alexander, Todd E. Zickler |
ICCV | 1 |
| 2016 | Focal Flow: Measuring Distance and Velocity with Defocus and Differential Motion
Emma Alexander, Qi Guo 0009, Sanjeev J. Koppal, Steven J. Gortler, Todd E. Zickler |
ECCV (3) | 2 |