EDBT 2026 Demo / reviewers in the wild / expert
Qiang Qu 0004
dblp:92/5150-4
· DBLP profile ↗
13ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0002-6648-5050ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Ultralight Polarity-Split Neuromorphic SNN for Event-Stream Super-ResolutionabstractEvent cameras offer unparalleled advantages such as high temporal resolution, low latency, and high dynamic range. However, their limited spatial resolution poses challenges for fine-grained perception tasks. In this work, we propose an ultra-lightweight, stream-based event-to-event super-resolution method based on Spiking Neural Networks (SNNs), designed for real-time deployment on resource-constrained devices. To further reduce model size, we introduce a novel Dual-Forward Polarity-Split Event Encoding strategy that decouples positive and negative events into separate forward paths through a shared SNN. Furthermore, we propose a Learnable Spatio-temporal Polarity-aware Loss (LearnSTPLoss) that adaptively balances temporal, spatial, and polarity consistency using learnable uncertainty-based weights. Experimental results demonstrate that our method achieves competitive super-resolution performance on multiple datasets while significantly reducing model size and inference time. The lightweight design enables embedding the module into event cameras or using it as an efficient front-end preprocessing for downstream vision tasks. Chuanzhi Xu, Haoxian Zhou, Langyi Chen, Vera Chung, Qiang Qu 0004 |
AAAI | 5 |
| 2026 | Dynolayout: robust layout estimation from event-stream for extended reality under dynamic scenarios
Xucheng Guo, Qiang Qu 0004, Guangrong Zhao, Yuanfeng Zhou, Yiran Shen 0001 |
CCF Trans. Pervasive Comput. Interact. | 4 |
| 2026 | NVS-SQA: Exploring Self-Supervised Quality Representation Learning for Neurally Synthesized Scenes Without ReferencesabstractNeural View Synthesis (NVS), such as NeRF and 3D Gaussian Splatting, effectively creates photorealistic scenes from sparse viewpoints, typically evaluated by quality assessment methods like PSNR, SSIM, and LPIPS. However, these full-reference methods, which compare synthesized views to reference views, may not fully capture the perceptual quality of neurally synthesized scenes (NSS), particularly due to the limited availability of dense reference views. Furthermore, the challenges in acquiring human perceptual labels hinder the creation of extensive labeled datasets, risking model overfitting and reduced generalizability. To address these issues, we propose NVS-SQA, a NSS quality assessment method to learn no-reference quality representations through self-supervision without reliance on human labels. Traditional self-supervised learning predominantly relies on the "same instance, similar representation" assumption and extensive datasets. However, given that these conditions do not apply in NSS quality assessment, we employ heuristic cues and quality scores as learning objectives, along with a specialized contrastive pair preparation process to improve the effectiveness and efficiency of learning. The results show that NVS-SQA outperforms 17 no-reference methods by a large margin (i.e., on average 109.5% in SRCC, 98.6% in PLCC, and 91.5% in KRCC over the second best) and even exceeds 16 full-reference methods across all evaluation metrics (i.e., 22.9% in SRCC, 19.1% in PLCC, and 18.6% in KRCC over the second best). Qiang Qu 0004, Yiran Shen 0001, Xiaoming Chen 0006, Vera Chung, Tom Weidong Cai, Tongliang Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | FinCast: A Foundation Model for Financial Time-Series ForecastingabstractFinancial time-series forecasting is critical for maintaining economic stability, guiding informed policymaking, and promoting sustainable investment practices. However, it remains challenging due to various underlying pattern shifts. These shifts arise primarily from three sources: temporal non-stationarity (distribution changes over time), multi-domain diversity (distinct patterns across financial domains such as stocks, commodities, and futures), and varying temporal resolutions (patterns differing across per-second, hourly, daily, or weekly indicators). While recent deep learning methods attempt to address these complexities, they frequently suffer from overfitting and typically require extensive domain-specific fine-tuning. To overcome these limitations, we introduce FinCast, the first foundation model specifically designed for financial time-series forecasting, trained on large-scale financial datasets. Remarkably, FinCast exhibits robust zero-shot performance, effectively capturing diverse patterns without domain-specific fine-tuning. Comprehensive empirical and qualitative evaluations demonstrate that FinCast surpasses existing state-of-the-art methods, highlighting its strong generalization capabilities. Zhuohang Zhu, Qiang Qu 0004, Vera Chung |
CIKM | 3 |
| 2025 | Can Large Language Models Grasp Event Signals? Exploring Pure Zero-Shot Event-based RecognitionabstractRecent advancements in event-based zero-shot object recognition have demonstrated promising results. However, these methods heavily depend on extensive training and are inherently constrained by the characteristics of CLIP. To the best of our knowledge, this research is the first study to explore the understanding capabilities of large language models (LLMs) for event-based visual content. We demonstrate that LLMs can achieve event-based object recognition without additional training or fine-tuning in conjunction with CLIP, effectively enabling pure zero-shot event-based recognition. Particularly, we evaluate the ability of GPT-4o / 4turbo and two other open-source LLMs to directly recognize event-based visual content. Extensive experiments are conducted across three benchmark datasets, systematically assessing the recognition accuracy of these models. The results show that LLMs, especially when enhanced with well-designed prompts, significantly improve event-based zero-shot recognition performance. Notably, GPT-4o outperforms the compared models and exceeds the recognition accuracy of state-of-the-art event-based zero-shot methods on N-ImageNet by five orders of magnitude. Zongyou Yu, Qiang Qu 0004, Xiaoming Chen 0006, Chen Wang 0043 |
ICASSP | 2 |
| 2025 | EEG2Gaussian: Decoding and Visualizing Visual-Evoked EEG for VR Scenes Using 3D Gaussian SplattingabstractDecoding and visualizing brain activity evoked by visual stimuli is critical for both understanding neural mechanisms and advancing brain-computer interfaces (BCIs). However, non-invasive signals such as Electroencephalogram (EEG) present significant challenges due to their inherently low signal-to-noise ratios. Although recent deep learning methods have resolved this task, most approaches are confined to 2D visualizations that fail to capture the complexities of real-world 3D perception. In this research, we investigate the relationship between EEG signals and 3D visual stimuli presented in virtual reality (VR) scenes, aiming to extract taskrelevant semantics from the EEG responses elicited by these stimuli. We introduce EEG2Gaussian, a novel framework for decoding and visualizing visual-evoked EEG signals by reconstructing immersive VR scenes using 3D Gaussian Splatting. The framework consists of three stages. The preprocessing stage removes noise and artifacts from raw EEG signals to provide cleaner input for subsequent processing. In the encoding stage, we propose a Neural Temporal-Frequency Encoder (NTF-Encoder) to extract temporal and frequency features using fused channel and band attention mechanisms, and disentangles them into high-level and low-level semantic representations. In the decoding stage, a 3D EEG Decoder takes these multi-level features through separate pathways as conditional inputs to guide the reconstruction of semantically consistent VR scenes. Furthermore, we construct a VR-EEG dataset that pairs real-time EEG recordings with VR scenes, and analyze how different types of scenes affect EEG responses across frequency bands. Our experimental results show that EEG2Gaussian can reconstruct VR scenes that are semantically aligned with the visual stimuli. Ablation studies verify the effectiveness of channel and band attention in EEG feature encoding, and demonstrate that combining high-level and low-level semantic features enhances the consistency and interpretability of the reconstructed scenes. Qiang Qu 0004, Xiaoming Chen 0006, Longfei Han, Yiran Shen 0001 |
ISMAR | 2 |
| 2025 | MirrorPose: Enabling Full-Body Gestures Interaction for Head-Mounted Devices with a Full-Length MirrorabstractHuman-computer interaction based on full-body gestures has been successfully adopted in various applications, such as motionsensing games. Typically, full-body gestures are captured using vision-based pose estimation or multiple inertial measurement units (IMUs) attached to the limbs. Gesture-based interactions in virtual and augmented reality environments allow for seamless and intuitive engagement across virtual and real domains. However, due to the design of head-mounted devices, only partial body tracking—such as hand tracking—is typically available for interactions. Capturing full-body pose using head-mounted sensors is inherently challenging due to device placement constraints. Furthermore, the limited computational resources of AR devices (e.g., constrained processing power and memory bandwidth) present significant challenges for the real-time deployment of sophisticated 3D human pose estimation architectures. To address these challenges, we propose MirrorPose, a lightweight framework that integrates a 3D pose estimation network (PoseARNet), optimized for the resource constraints of AR headsets and the dynamic viewpoint changes inherent in mirror-mediated spatial perception. This design enables practical, full-body gesture interaction on AR devices. To demonstrate its practicality, we developed a 3D virtual teaching application on Microsoft HoloLens 2, to enhance students' understanding of human poses. Extensive experiments and evaluations confirm that our system provides users with accurate and timely feedback. The codes and dataset are available at https://github.com/zhchlong/mirror_pose. Xingwang Xue, Xiyu Sheng, Qiang Qu 0004, Yiran Shen 0001 |
ISMAR | 4 |
| 2025 | VF-Lens: Enhancing Visual Perception of Visually Impaired Users in VR via Adversarial Learning with Visual Field AttentionabstractThis research aims to enhance the image perception of visually impaired users in VR environments. We propose VF-Lens, a model that adaptively compensates for light sensitivity based on the user’s visual field impairment, acting as a virtual lens between the visually impaired users and the VR world. VF-Lens is designed as a tailored generative adversarial learning model with a generator and discriminator, offering applicability to various types of visual impairments while bypassing engineering complexities. The generator creates a "hyperimage" tailored to the user’s visual field impairment, which then undergoes a particular regression process to predict and replicate the real perception of the visually impaired user. The discriminator then evaluates the similarity between the replicated perception and the original image. Through adversarial training, the generator can produce hyperimages that adapt to the user’s visual field parameters, enabling them to perceive the image more similarly to normal-vision users. We further improve VF-Lens by proposing new "visual field attention" mechanisms that prioritize and refine visual information in the user’s visual field. Extensive evaluation, encompassing both visually impaired participants and simulations, has been conducted to demonstrate the effectiveness of VF-Lens in improving visual perception for visually impaired users. Moreover, we establish a standardized evaluation process involving tailored metrics as well as objective and subjective evaluations to promote reusability and comparability for future research in this field. Xiaoming Chen 0006, Dehao Han, Qiang Qu 0004, Yiran Shen 0001 |
VR | 3 |
| 2024 | E2HQV: High-Quality Video Generation from Event Camera via Theory-Inspired Model-Aided Deep LearningabstractThe bio-inspired event cameras or dynamic vision sensors are capable of asynchronously capturing per-pixel brightness changes (called event-streams) in high temporal resolution and high dynamic range. However, the non-structural spatial-temporal event-streams make it challenging for providing intuitive visualization with rich semantic information for human vision. It calls for events-to-video (E2V) solutions which take event-streams as input and generate high quality video frames for intuitive visualization. However, current solutions are predominantly data-driven without considering the prior knowledge of the underlying statistics relating event-streams and video frames. It highly relies on the non-linearity and generalization capability of the deep neural networks, thus, is struggling on reconstructing detailed textures when the scenes are complex. In this work, we propose E2HQV, a novel E2V paradigm designed to produce high-quality video frames from events. This approach leverages a model-aided deep learning framework, underpinned by a theory-inspired E2V model, which is meticulously derived from the fundamental imaging principles of event cameras. To deal with the issue of state-reset in the recurrent components of E2HQV, we also design a temporal shift embedding module to further improve the quality of the video frames. Comprehensive evaluations on the real world event camera datasets validate our approach, with E2HQV, notably outperforming state-of-the-art approaches, e.g., surpassing the second best by over 40% for some evaluation metrics. Qiang Qu 0004, Yiran Shen 0001, Xiaoming Chen 0006, Vera Chung, Tongliang Liu |
AAAI | 1 |
| 2024 | EvRepSL: Event-Stream Representation via Self-Supervised Learning for Event-Based VisionabstractEvent-stream representation is the first step for many computer vision tasks using event cameras. It converts the asynchronous event-streams into a formatted structure so that conventional machine learning models can be applied easily. However, most of the state-of-the-art event-stream representations are manually designed and the quality of these representations cannot be guaranteed due to the noisy nature of event-streams. In this paper, we introduce a data-driven approach aiming at enhancing the quality of event-stream representations. Our approach commences with the introduction of a new event-stream representation based on spatial-temporal statistics, denoted as EvRep. Subsequently, we theoretically derive the intrinsic relationship between asynchronous event-streams and synchronous video frames. Building upon this theoretical relationship, we train a representation generator, RepGen, in a self-supervised learning manner accepting EvRep as input. Finally, the event-streams are converted to high-quality representations, termed as EvRepSL, by going through the learned RepGen (without the need of fine-tuning or retraining). Our methodology is rigorously validated through extensive evaluations on a variety of mainstream event-based classification and optical flow datasets (captured with various types of event cameras). The experimental results highlight not only our approach's superior performance over existing event-stream representations but also its versatility, being agnostic to different event cameras and tasks. Qiang Qu 0004, Xiaoming Chen 0006, Vera Chung, Yiran Shen 0001 |
IEEE Trans. Image Process. | 1 |
| 2024 | NeRF-NQA: No-Reference Quality Assessment for Scenes Generated by NeRF and Neural View Synthesis MethodsabstractNeural View Synthesis (NVS) has demonstrated efficacy in generating high-fidelity dense viewpoint videos using a image set with sparse views. However, existing quality assessment methods like PSNR, SSIM, and LPIPS are not tailored for the scenes with dense viewpoints synthesized by NVS and NeRF variants, thus, they often fall short in capturing the perceptual quality, including spatial and angular aspects of NVS-synthesized scenes. Furthermore, the lack of dense ground truth views makes the full reference quality assessment on NVS-synthesized scenes challenging. For instance, datasets such as LLFF provide only sparse images, insufficient for complete full-reference assessments. To address the issues above, we propose NeRF-NQA, the first no-reference quality assessment method for densely-observed scenes synthesized from the NVS and NeRF variants. NeRF-NQA employs a joint quality assessment strategy, integrating both viewwise and pointwise approaches, to evaluate the quality of NVS-generated scenes. The viewwise approach assesses the spatial quality of each individual synthesized view and the overall inter-views consistency, while the pointwise approach focuses on the angular qualities of scene surface points and their compound inter-point quality. Extensive evaluations are conducted to compare NeRF-NQA with 23 mainstream visual quality assessment methods (from fields of image, video, and light-field assessment). The results demonstrate NeRF-NQA outperforms the existing assessment methods significantly and it shows substantial superiority on assessing NVS-synthesized scenes without references. An implementation of this paper are available at https://github.com/VincentQQu/NeRF-NQA. Qiang Qu 0004, Hanxue Liang, Xiaoming Chen 0006, Vera Chung, Yiran Shen 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2023 | LFACon: Introducing Anglewise Attention to No-Reference Quality Assessment in Light Field SpaceabstractLight field imaging can capture both the intensity information and the direction information of light rays. It naturally enables a six-degrees-of-freedom viewing experience and deep user engagement in virtual reality. Compared to 2D image assessment, light field image quality assessment (LFIQA) needs to consider not only the image quality in the spatial domain but also the quality consistency in the angular domain. However, there is a lack of metrics to effectively reflect the angular consistency and thus the angular quality of a light field image (LFI). Furthermore, the existing LFIQA metrics suffer from high computational costs due to the excessive data volume of LFIs. In this paper, we propose a novel concept of "anglewise attention" by introducing a multihead self-attention mechanism to the angular domain of an LFl. This mechanism better reflects the LFI quality. In particular, we propose three new attention kernels, including anglewise self-attention, anglewise grid attention, and anglewise central attention. These attention kernels can realize angular self-attention, extract multiangled features globally or selectively, and reduce the computational cost of feature extraction. By effectively incorporating the proposed kernels, we further propose our light field attentional convolutional neural network (LFACon) as an LFIQA metric. Our experimental results show that the proposed LFACon metric significantly outperforms the state-of-the-art LFIQA metrics. For the majority of distortion types, LFACon attains the best performance with lower complexity and less computational time. Qiang Qu 0004, Xiaoming Chen 0006, Vera Chung, Tom Weidong Cai |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2019 | Improved image classification with 4D light-field and interleaved convolutional neural network
Zhicheng Lu, Henry Wing Fung Yeung, Qiang Qu 0004, Vera Chung, Xiaoming Chen 0006, Zhibo Chen 0001 |
Multim. Tools Appl. | 3 |