VLDB 2026 Research / reviewers in the wild / expert
Qiqi Hou
dblp:168/2419
· DBLP profile ↗
13ranked-venue papers
8as first author
6since 2021 · last 2025
0009-0009-3472-6401ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 6 first-author · 5 since 2021Artificial intelligence and machine learning · 8 · 5 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Sort-free Gaussian Splatting via Weighted Sum RenderingabstractRecently, 3D Gaussian Splatting (3DGS) has emerged as a significant advancement in 3D scene reconstruction, attracting considerable attention due to its ability to recover high-fidelity details while maintaining low complexity. Despite the promising results achieved by 3DGS, its rendering performance is constrained by its dependence on costly non-commutative alpha-blending operations. These operations mandate complex view dependent sorting operations that introduce computational overhead, especially on the resource-constrained platforms such as mobile phones. In this paper, we propose Weighted Sum Rendering, which approximates alpha blending with weighted sums, thereby removing the need for sorting. This simplifies implementation, delivers superior performance, and eliminates the ``popping'' artifacts caused by sorting. Experimental results show that optimizing a generalized Gaussian splatting formulation to the new differentiable rendering yields competitive image quality. The method was implemented and tested in a mobile device GPU, achieving on average $1.23\times$ faster rendering. Qiqi Hou, Randall Rauwendaal, Hoang Le, Farzad Farhadzadeh, Fatih Porikli, Alex Bourd, Amir Said |
ICLR | 1 |
| 2024 | Low-Latency Neural Stereo StreamingabstractThe rise of new video modalities like virtual reality or autonomous driving has increased the demand for efficient multi-view video compression methods, both in terms of rate-distortion (R-D) performance and in terms of delay and runtime. While most recent stereo video compression approaches have shown promising performance, they compress left and right views sequentially, leading to poor parallelization and runtime performance. This work presents Low-Latency neural codec for Stereo video Streaming (LLSS), a novel parallel stereo video coding method designed for fast and efficient low-latency stereo video streaming. Instead of using a sequential cross-view motion compensation like existing methods, LLSS introduces a bidirectional feature shifting module to directly exploit mutual information among views and encode them effectively with a joint cross-view prior model for entropy coding. Thanks to this design, LLSS processes left and right views in parallel, minimizing latency; all while substantially improving R-D performance compared to both existing neural and conventional codecs. Qiqi Hou, Farzad Farhadzadeh, Amir Said, Guillaume Sautière, Hoang Le |
CVPR | 1 |
| 2024 | Neural Graphics Texture Compression Supporting Random Access
Farzad Farhadzadeh, Qiqi Hou, Hoang Le, Amir Said, Randall Rauwendaal, Alex Bourd, Fatih Porikli |
ECCV (37) | 2 |
| 2024 | Auxiliary Features-Guided Super Resolution for Monte Carlo RenderingabstractAbstract This paper investigates super‐resolution to reduce the number of pixels to render and thus speed up Monte Carlo rendering algorithms. While great progress has been made to super‐resolution technologies, it is essentially an ill‐posed problem and cannot recover high‐frequency details in renderings. To address this problem, we exploit high‐resolution auxiliary features to guide super‐resolution of low‐resolution renderings. These high‐resolution auxiliary features can be quickly rendered by a rendering engine and at the same time provide valuable high‐frequency details to assist super‐resolution. To this end, we develop a cross‐modality transformer network that consists of an auxiliary feature branch and a low‐resolution rendering branch. These two branches are designed to fuse high‐resolution auxiliary features with the corresponding low‐resolution rendering. Furthermore, we design Residual Densely Connected Swin Transformer groups to learn to extract representative features to enable high‐quality super‐resolution. Our experiments show that our auxiliary features‐guided super‐resolution method outperforms both super‐resolution methods and Monte Carlo denoising methods in producing high‐quality renderings. Qiqi Hou, Feng Liu 0015 |
Comput. Graph. Forum | 1 |
| 2022 | A Perceptual Quality Metric for Video Frame Interpolation
Qiqi Hou, Abhijay Ghildyal, Feng Liu 0015 |
ECCV (15) | 1 |
| 2021 | Fast Monte Carlo Rendering via Multi-Resolution Sampling
Qiqi Hou, Carl S. Marshall, Selvakumar Panneer, Feng Liu 0015 |
Graphics Interface | 1 |
| 2019 | Context-Aware Image Matting for Simultaneous Foreground and Alpha EstimationabstractNatural image matting is an important problem in computer vision and graphics. It is an ill-posed problem when only an input image is available without any external information. While the recent deep learning approaches have shown promising results, they only estimate the alpha matte. This paper presents a context-aware natural image matting method for simultaneous foreground and alpha matte estimation. Our method employs two encoder networks to extract essential information for matting. Particularly, we use a matting encoder to learn local features and a context encoder to obtain more global context information. We concatenate the outputs from these two encoders and feed them into decoder networks to simultaneously estimate the foreground and alpha matte. To train this whole deep neural network, we employ both the standard Laplacian loss and the feature loss: the former helps to achieve high numerical performance while the latter leads to more perceptually plausible results. We also report several data augmentation strategies that greatly improve the network's generalization performance. Our qualitative and quantitative experiments show that our method enables high-quality matting for a single natural image. Qiqi Hou, Feng Liu 0015 |
ICCV | 1 |
| 2018 | Face alignment recurrent network
Qiqi Hou, Jinjun Wang, Ruibin Bai, Sanping Zhou, Yihong Gong |
Pattern Recognit. | 1 |
| 2018 | Deep ranking model by large adaptive margin learning for person re-identification
Sanping Zhou, Jinjun Wang, Qiqi Hou |
Pattern Recognit. | 4 |
| 2018 | Large Margin Learning in Set-to-Set Similarity Comparison for Person ReidentificationabstractPerson reidentification aims at matching images of the same person across disjoint camera views, which is a challenging problem in multimedia analysis, multimedia editing, and content-based media retrieval communities. The major challenge lies in how to preserve similarity of the same person across video footages with large appearance variations, while discriminating different individuals. To address this problem, conventional methods usually consider the pairwise similarity between persons by only measuring the point-to-point distance. In this paper, we propose using a deep learning technique to model a novel set-to-set (S2S) distance, in which the underline objective focuses on preserving the compactness of intraclass samples for each camera view, while maximizing the margin between the intraclass set and interclass set. The S2S distance metric consists of three terms, namely, the class-identity term, the relative distance term, and the regularization term. The class-identity term keeps the intraclass samples within each camera view gathering together, the relative distance term maximizes the distance between the intraclass class set and interclass set across different camera views, and the regularization term smoothes the parameters of the deep convolutional neural network. As a result, the final learned deep model can effectively find out the matched target to the probe object among various candidates in the video gallery by learning discriminative and stable feature representations. Using the CUHK01, CUHK03, PRID2011, and Market1501 benchmark datasets, we extensively conducted comparative evaluations to demonstrate the advantages of our method over the state-of-the-art approaches. Sanping Zhou, Jinjun Wang, Qiqi Hou, Yihong Gong, Nanning Zheng 0001 |
IEEE Trans. Multim. | 4 |
| 2017 | Part-aware trajectories association across non-overlapping uncalibrated cameras
De Cheng, Yihong Gong, Jinjun Wang, Qiqi Hou, Nanning Zheng 0001 |
Neurocomputing | 4 |
| 2015 | Facial landmark detection via cascade multi-channel convolutional neural networkabstractThis paper presents a novel cascade multi-channel convolutional neural networks(CMC-CNN) approach for face alignment. Several CNN are jointly used for the finally output. In our method, each stage CNN takes the local region around the landmarks as input, and each local patches does convolution separately, which can lead network to learn local high-level features. Then a fully connected layer is put to learn global information from these local features. Our methods has achieves the state-of-the-art results when tested on the 300 Face in-the-Wild(300-W) dataset. Qiqi Hou, Jinjun Wang, Lele Cheng, Yihong Gong |
ICIP | 1 |
| 2015 | Robust Deep Auto-encoder for Occluded Face RecognitionabstractOcclusions by sunglasses, scarf, hats, beard, shadow etc, can significantly reduce the performance of face recognition systems. Although there exists a rich literature of researches focusing on face recognition with illuminations, poses and facial expression variations, there is very limited work reported for occlusion robust face recognition. In this paper, we present a method to restore occluded facial regions using deep learning technique to improve face recognition performance. Inspired by SSDA for facial occlusion removal with known occlusion type and explicit occlusion location detection from a preprocessing step, this paper further introduces Double Channel SSDA (DC-SSDA) which requires no prior knowledge of the types and the locations of occlusions. Experimental results based on CMU-PIE face database have showed that, the proposed method is robust to a variety of occlusion types and locations, and the restored faces could yield significant recognition performance improvements over occluded ones. Lele Cheng, Jinjun Wang, Yihong Gong, Qiqi Hou |
ACM Multimedia | 4 |