Xiaoming Chen 0006

dblp:72/2676-6 · DBLP profile ↗
← Back
39ranked-venue papers
6as first author
17since 2021 · last 2026
0000-0002-7503-3021ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 29 · 6 first-author · 13 since 2021Artificial intelligence and machine learning · 9 · 4 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Towards Event-guided Panoramic HDR Video Reconstruction for Indoor Immersive VR: A Novel Dataset and Approach
abstract
High Dynamic Range (HDR) panoramic video is crucial to enhance immersive experience in Virtual Reality (VR). However, a hurdle is that panoramic cameras often struggle with limited dynamic range and motion blur. Inspired by the event-driven sensing of the human eye, this paper explores the potential of the event cameras to enhance panoramic HDR video reconstruction. As a pioneering research endeavor, we first starts by designing a novel hybrid imaging platform equipped with preprocessing pipelines for event-panorama synchronization, alignment, and HDR ground truth generation. Based on the platform, we then introduce Ev-Pano, the first event-panorama HDR video covering diverse indoor scenes for panoramic HDR video reconstruction. We hope Ev-Pano will establish a foundation to support event-guided panoramic HDR imaging and VR research community. With Ev-Pano, we further propose a novel approach that employs a weighting function-based luminance fusion to enable events to recover missing textures in LDR panoramic videos for panoramic HDR video reconstruction. We conduct extensive experiments to demonstrate the effectiveness of our approach. The results show the best performance of ours than prior arts. Meanwhile, a user study on an HDR-capable head-mounted display (Apple Vision Pro) shows feasible perceptual quality (which is closer to the HDR ground truth) of the reconstructed panoramic HDR videos. The codes and part of the dataset can be accessed via the anonymized link https://anonymous.4open.science/r/Ev-Pano-D2D2/.
Xucheng Guo, Majed Elwardy, Yan Hu 0003, Yuanfeng Zhou, Xiaoming Chen 0006, Yiran Shen 0001
VR8
2026 AMMNet: An asymmetric multi-modal network for earth observation semantic segmentation
abstract
Semantic segmentation of Earth Observation (EO) data has advanced significantly from multi-modal inputs, such as RGB imagery and the Digital Surface Model (DSM), which provides complementary contextual and structural information about ground objects. However, the joint use of RGB and DSM presents two distinct challenges: architectural redundancy , an efficiency concern, and cross-modal representational misalignment , a fusion-quality concern. To overcome these limitations, we propose AMMNet, an Asymmetric Multi-Modal Network designed for robust EO semantic segmentation through asymmetric designs tailored for RGB-DSM. The Asymmetric Dual Encoder (ADE) mitigates architectural redundancy by allocating representational capacity asymmetrically, employing a deeper encoder for RGB imagery to capture rich contextual cues and a lightweight encoder for DSM to extract sparse structural features. Asymmetric Prior Fuser (APF) introduces a modality-aware prior matrix into the fusion process, enabling structure-aware contextual representation and reducing cross-modal representational misalignment. Distribution Alignment (DA) module further enhances compatibility by maintaining task-relevant information. Extensive experiments on the ISPRS Vaihingen and Potsdam benchmarks demonstrate that AMMNet attains superior segmentation performance with reduced computational and memory costs. Results validate the effectiveness of the proposed asymmetric multi-modal design in advancing EO semantic segmentation.
Zexi Hu, Xiaoming Chen 0006, Vera Chung
Neurocomputing4
2026 NVS-SQA: Exploring Self-Supervised Quality Representation Learning for Neurally Synthesized Scenes Without References
abstract
Neural View Synthesis (NVS), such as NeRF and 3D Gaussian Splatting, effectively creates photorealistic scenes from sparse viewpoints, typically evaluated by quality assessment methods like PSNR, SSIM, and LPIPS. However, these full-reference methods, which compare synthesized views to reference views, may not fully capture the perceptual quality of neurally synthesized scenes (NSS), particularly due to the limited availability of dense reference views. Furthermore, the challenges in acquiring human perceptual labels hinder the creation of extensive labeled datasets, risking model overfitting and reduced generalizability. To address these issues, we propose NVS-SQA, a NSS quality assessment method to learn no-reference quality representations through self-supervision without reliance on human labels. Traditional self-supervised learning predominantly relies on the "same instance, similar representation" assumption and extensive datasets. However, given that these conditions do not apply in NSS quality assessment, we employ heuristic cues and quality scores as learning objectives, along with a specialized contrastive pair preparation process to improve the effectiveness and efficiency of learning. The results show that NVS-SQA outperforms 17 no-reference methods by a large margin (i.e., on average 109.5% in SRCC, 98.6% in PLCC, and 91.5% in KRCC over the second best) and even exceeds 16 full-reference methods across all evaluation metrics (i.e., 22.9% in SRCC, 19.1% in PLCC, and 18.6% in KRCC over the second best).
Qiang Qu 0004, Yiran Shen 0001, Xiaoming Chen 0006, Vera Chung, Tom Weidong Cai, Tongliang Liu
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Enhanced VR Learning with Dynamic Somatosensory Feedback to Assist in Understanding the Dynamic Process of Blood Circulation
abstract
With the rapid development of Virtual Reality (VR) and haptic technology, their applications in education have demonstrated tremendous potential. However, traditional VR learning that incorporates haptic feedback usually has a limited body coverage, and there is a weak alignment between the haptic feedback and the educational content. To address the above-mentioned issues, this study established an environment that combines VR with dynamic somatosensory feedback. Through the simulation of the human blood circulation, it explored the impacts of VR and somatosensory feedback on learning effects. Sixty participants were randomly divided into three groups: the computer-based group, the VR-only group, and the VR with somatosensory feedback group. Self-efficacy, perceived enjoyment, and knowledge retention were evaluated through targeted questionnaires. The results showed that VR with somatosensory feedback significantly outperformed the computer-based and VR-only learning modes. This study provides new insights and practical support for the effective integration of VR and somatosensory feedback technology in educational environments.
Fuwei Dong, Anran Meng, Xiaoming Chen 0006, Chen Wang 0043, Vera Chung
ICALT3
2025 Star Operation in Self-Attention for 3D Human Pose Estimation
abstract
Recent transformer-based methods have achieved notable success in 3D human pose estimation. However, the most utilized self-attention mechanisms compute the attention matrix by performing a dot product on inter-vector features, which may overlook finer element-wise interactions. In this paper, we introduce star-attention, which integrates the star operation into the self-attention module. This approach retains the non-linearity and high dimensionality characteristics of matrix multiplication in self-attention while enhancing the feature representation by capturing interactions at a finer granularity. Specifically, we validated the promotion of the proposed star-attention module on MixSTE as an example. Experimental results demonstrate that our approach achieves competitive performance on benchmarks compared to state-of-the-art methods, producing smoother 3D poses across successive frames.
Xiaoming Chen 0006, Hai-Sheng Li 0002
ICASSP4
2025 Can Large Language Models Grasp Event Signals? Exploring Pure Zero-Shot Event-based Recognition
abstract
Recent advancements in event-based zero-shot object recognition have demonstrated promising results. However, these methods heavily depend on extensive training and are inherently constrained by the characteristics of CLIP. To the best of our knowledge, this research is the first study to explore the understanding capabilities of large language models (LLMs) for event-based visual content. We demonstrate that LLMs can achieve event-based object recognition without additional training or fine-tuning in conjunction with CLIP, effectively enabling pure zero-shot event-based recognition. Particularly, we evaluate the ability of GPT-4o / 4turbo and two other open-source LLMs to directly recognize event-based visual content. Extensive experiments are conducted across three benchmark datasets, systematically assessing the recognition accuracy of these models. The results show that LLMs, especially when enhanced with well-designed prompts, significantly improve event-based zero-shot recognition performance. Notably, GPT-4o outperforms the compared models and exceeds the recognition accuracy of state-of-the-art event-based zero-shot methods on N-ImageNet by five orders of magnitude.
Zongyou Yu, Qiang Qu 0004, Xiaoming Chen 0006, Chen Wang 0043
ICASSP3
2025 EEG2Gaussian: Decoding and Visualizing Visual-Evoked EEG for VR Scenes Using 3D Gaussian Splatting
abstract
Decoding and visualizing brain activity evoked by visual stimuli is critical for both understanding neural mechanisms and advancing brain-computer interfaces (BCIs). However, non-invasive signals such as Electroencephalogram (EEG) present significant challenges due to their inherently low signal-to-noise ratios. Although recent deep learning methods have resolved this task, most approaches are confined to 2D visualizations that fail to capture the complexities of real-world 3D perception. In this research, we investigate the relationship between EEG signals and 3D visual stimuli presented in virtual reality (VR) scenes, aiming to extract taskrelevant semantics from the EEG responses elicited by these stimuli. We introduce EEG2Gaussian, a novel framework for decoding and visualizing visual-evoked EEG signals by reconstructing immersive VR scenes using 3D Gaussian Splatting. The framework consists of three stages. The preprocessing stage removes noise and artifacts from raw EEG signals to provide cleaner input for subsequent processing. In the encoding stage, we propose a Neural Temporal-Frequency Encoder (NTF-Encoder) to extract temporal and frequency features using fused channel and band attention mechanisms, and disentangles them into high-level and low-level semantic representations. In the decoding stage, a 3D EEG Decoder takes these multi-level features through separate pathways as conditional inputs to guide the reconstruction of semantically consistent VR scenes. Furthermore, we construct a VR-EEG dataset that pairs real-time EEG recordings with VR scenes, and analyze how different types of scenes affect EEG responses across frequency bands. Our experimental results show that EEG2Gaussian can reconstruct VR scenes that are semantically aligned with the visual stimuli. Ablation studies verify the effectiveness of channel and band attention in EEG feature encoding, and demonstrate that combining high-level and low-level semantic features enhances the consistency and interpretability of the reconstructed scenes.
Qiang Qu 0004, Xiaoming Chen 0006, Longfei Han, Yiran Shen 0001
ISMAR3
2025 VF-Lens: Enhancing Visual Perception of Visually Impaired Users in VR via Adversarial Learning with Visual Field Attention
abstract
This research aims to enhance the image perception of visually impaired users in VR environments. We propose VF-Lens, a model that adaptively compensates for light sensitivity based on the user’s visual field impairment, acting as a virtual lens between the visually impaired users and the VR world. VF-Lens is designed as a tailored generative adversarial learning model with a generator and discriminator, offering applicability to various types of visual impairments while bypassing engineering complexities. The generator creates a "hyperimage" tailored to the user’s visual field impairment, which then undergoes a particular regression process to predict and replicate the real perception of the visually impaired user. The discriminator then evaluates the similarity between the replicated perception and the original image. Through adversarial training, the generator can produce hyperimages that adapt to the user’s visual field parameters, enabling them to perceive the image more similarly to normal-vision users. We further improve VF-Lens by proposing new "visual field attention" mechanisms that prioritize and refine visual information in the user’s visual field. Extensive evaluation, encompassing both visually impaired participants and simulations, has been conducted to demonstrate the effectiveness of VF-Lens in improving visual perception for visually impaired users. Moreover, we establish a standardized evaluation process involving tailored metrics as well as objective and subjective evaluations to promote reusability and comparability for future research in this field.
Xiaoming Chen 0006, Dehao Han, Qiang Qu 0004, Yiran Shen 0001
VR1
2025 Beyond Subspace Isolation: Many-to-Many Transformer for Light Field Image Super-Resolution
abstract
The effective extraction of spatial-angular features plays a crucial role in light field image super-resolution (LFSR) tasks, and the introduction of convolution and Transformers leads to significant improvement in this area. Nevertheless, due to the large 4D data volume of light field images, many existing methods opted to decompose the data into a number of lower-dimensional subspaces and perform Transformers in each sub-space individually. As a side effect, these methods inadvertently restrict the self-attention mechanisms to a One-to-One scheme accessing only a limited subset of LF data, explicitly preventing comprehensive optimization on all spatial and angular cues. In this paper, we identify this limitation as subspace isolation and introduce a novel Many-to-Many Transformer (M2MT) to address it. M2MT aggregates angular information in the spatial subspace before performing the self-attention mechanism. It enables complete access to all information across all sub-aperture images (SAIs) in a light field image. Consequently, M2MT is enabled to comprehensively capture long-range correlation dependencies. With M2MT as the foundational component, we develop a simple yet effective M2MT network for LFSR. Our experimental results demonstrate that M2MT achieves state-of-the-art performance across various public datasets, and it offers a favorable balance between model performance and efficiency, yielding higher-quality LFSR results with substantially lower demand for memory and computation. We further conduct in-depth analysis using local attribution maps (LAM) to obtain visual interpretability, and the results validate that M2MT is empowered with a truly non-local context in both spatial and angular subspaces to mitigate subspace isolation and acquire effective spatial-angular representation.
Zexi Hu, Xiaoming Chen 0006, Vera Chung, Yiran Shen 0001
IEEE Trans. Multim.2
2024 E2HQV: High-Quality Video Generation from Event Camera via Theory-Inspired Model-Aided Deep Learning
abstract
The bio-inspired event cameras or dynamic vision sensors are capable of asynchronously capturing per-pixel brightness changes (called event-streams) in high temporal resolution and high dynamic range. However, the non-structural spatial-temporal event-streams make it challenging for providing intuitive visualization with rich semantic information for human vision. It calls for events-to-video (E2V) solutions which take event-streams as input and generate high quality video frames for intuitive visualization. However, current solutions are predominantly data-driven without considering the prior knowledge of the underlying statistics relating event-streams and video frames. It highly relies on the non-linearity and generalization capability of the deep neural networks, thus, is struggling on reconstructing detailed textures when the scenes are complex. In this work, we propose E2HQV, a novel E2V paradigm designed to produce high-quality video frames from events. This approach leverages a model-aided deep learning framework, underpinned by a theory-inspired E2V model, which is meticulously derived from the fundamental imaging principles of event cameras. To deal with the issue of state-reset in the recurrent components of E2HQV, we also design a temporal shift embedding module to further improve the quality of the video frames. Comprehensive evaluations on the real world event camera datasets validate our approach, with E2HQV, notably outperforming state-of-the-art approaches, e.g., surpassing the second best by over 40% for some evaluation metrics.
Qiang Qu 0004, Yiran Shen 0001, Xiaoming Chen 0006, Vera Chung, Tongliang Liu
AAAI3
2024 Adaptively Augmented Consistency Learning: A Semi-supervised Segmentation Framework for Remote Sensing
Xiaoming Chen 0006, Vera Chung
ICONIP (10)3
2024 EvRepSL: Event-Stream Representation via Self-Supervised Learning for Event-Based Vision
abstract
Event-stream representation is the first step for many computer vision tasks using event cameras. It converts the asynchronous event-streams into a formatted structure so that conventional machine learning models can be applied easily. However, most of the state-of-the-art event-stream representations are manually designed and the quality of these representations cannot be guaranteed due to the noisy nature of event-streams. In this paper, we introduce a data-driven approach aiming at enhancing the quality of event-stream representations. Our approach commences with the introduction of a new event-stream representation based on spatial-temporal statistics, denoted as EvRep. Subsequently, we theoretically derive the intrinsic relationship between asynchronous event-streams and synchronous video frames. Building upon this theoretical relationship, we train a representation generator, RepGen, in a self-supervised learning manner accepting EvRep as input. Finally, the event-streams are converted to high-quality representations, termed as EvRepSL, by going through the learned RepGen (without the need of fine-tuning or retraining). Our methodology is rigorously validated through extensive evaluations on a variety of mainstream event-based classification and optical flow datasets (captured with various types of event cameras). The experimental results highlight not only our approach's superior performance over existing event-stream representations but also its versatility, being agnostic to different event cameras and tasks.
Qiang Qu 0004, Xiaoming Chen 0006, Vera Chung, Yiran Shen 0001
IEEE Trans. Image Process.2
2024 Video2Haptics: Converting Video Motion to Dynamic Haptic Feedback with Bio-Inspired Event Processing
abstract
In cinematic VR applications, haptic feedback can significantly enhance the sense of reality and immersion for users. The increasing availability of emerging haptic devices opens up possibilities for future cinematic VR applications that allow users to receive haptic feedback while they are watching videos. However, automatically rendering haptic cues from real-time video content, particularly from video motion, is a technically challenging task. In this article, we propose a novel framework called "Video2Haptics" that leverages the emerging bio-inspired event camera to capture event signals as a lightweight representation of video motion. We then propose efficient event-based visual processing methods to estimate force or intensity from video motion in the event domain, rather than the pixel domain. To demonstrate the application of Video2Haptics, we convert the estimated force or intensity to dynamic vibrotactile feedback on emerging haptic gloves, synchronized with the corresponding video motion. As a result, Video2Haptics allows users not only to view the video but also to perceive the video motion concurrently. Our experimental results show that the proposed event-based processing methods for force and intensity estimation are one to two orders of magnitude faster than conventional methods. Our user study results confirm that the proposed Video2Haptics framework can considerably enhance the users' video experience.
Xiaoming Chen 0006, Zexi Hu, Guangxin Zhao, Hai-Sheng Li 0002, Vera Chung, Aaron J. Quigley
IEEE Trans. Vis. Comput. Graph.1
2024 NeRF-NQA: No-Reference Quality Assessment for Scenes Generated by NeRF and Neural View Synthesis Methods
abstract
Neural View Synthesis (NVS) has demonstrated efficacy in generating high-fidelity dense viewpoint videos using a image set with sparse views. However, existing quality assessment methods like PSNR, SSIM, and LPIPS are not tailored for the scenes with dense viewpoints synthesized by NVS and NeRF variants, thus, they often fall short in capturing the perceptual quality, including spatial and angular aspects of NVS-synthesized scenes. Furthermore, the lack of dense ground truth views makes the full reference quality assessment on NVS-synthesized scenes challenging. For instance, datasets such as LLFF provide only sparse images, insufficient for complete full-reference assessments. To address the issues above, we propose NeRF-NQA, the first no-reference quality assessment method for densely-observed scenes synthesized from the NVS and NeRF variants. NeRF-NQA employs a joint quality assessment strategy, integrating both viewwise and pointwise approaches, to evaluate the quality of NVS-generated scenes. The viewwise approach assesses the spatial quality of each individual synthesized view and the overall inter-views consistency, while the pointwise approach focuses on the angular qualities of scene surface points and their compound inter-point quality. Extensive evaluations are conducted to compare NeRF-NQA with 23 mainstream visual quality assessment methods (from fields of image, video, and light-field assessment). The results demonstrate NeRF-NQA outperforms the existing assessment methods significantly and it shows substantial superiority on assessing NVS-synthesized scenes without references. An implementation of this paper are available at https://github.com/VincentQQu/NeRF-NQA.
Qiang Qu 0004, Hanxue Liang, Xiaoming Chen 0006, Vera Chung, Yiran Shen 0001
IEEE Trans. Vis. Comput. Graph.3
2024 Swift-Eye: Towards Anti-blink Pupil Tracking for Precise and Robust High-Frequency Near-Eye Movement Analysis with Event Cameras
abstract
Eye tracking has shown great promise in many scientific fields and daily applications, ranging from the early detection of mental health disorders to foveated rendering in virtual reality (VR). These applications all call for a robust system for high-frequency near-eye movement sensing and analysis in high precision, which cannot be guaranteed by the existing eye tracking solutions with CCD/CMOS cameras. To bridge the gap, in this paper, we propose Swift-Eye, an offline precise and robust pupil estimation and tracking framework to support high-frequency near-eye movement analysis, especially when the pupil region is partially occluded. Swift-Eye is built upon the emerging event cameras to capture the high-speed movement of eyes in high temporal resolution. Then, a series of bespoke components are designed to generate high-quality near-eye movement video at a high frame rate over kilohertz and deal with the occlusion over the pupil caused by involuntary eye blinks. According to our extensive evaluations on EV-Eye, a large-scale public dataset for eye tracking using event cameras, Swift-Eye shows high robustness against significant occlusion. It can improve the IoU and F1-score of the pupil estimation by 20% and 12.5% respectively, compared with the second-best competing approach, when over 80% of the pupil region is occluded by the eyelid. Lastly, it provides continuous and smooth traces of pupils in extremely high temporal resolution and can support high-frequency eye movement analysis and a number of potential applications, such as mental health diagnosis, behaviour-brain association, etc. The implementation details and source codes can be found at https://github.com/ztysdu/Swift-Eye.
Tongyu Zhang, Yiran Shen 0001, Guangrong Zhao, Lin Wang 0025, Xiaoming Chen 0006, Lu Bai 0004, Yuanfeng Zhou
IEEE Trans. Vis. Comput. Graph.5
2023 EV-LFV: Synthesizing Light Field Event Streams from an Event Camera and Multiple RGB Cameras
abstract
Light field videos captured in RGB frames (RGB-LFV) can provide users with a 6 degree-of-freedom immersive video experience by capturing dense multi-subview video. Despite its potential benefits, the processing of dense multi-subview video is extremely resource-intensive, which currently limits the frame rate of RGB-LFV (i.e., lower than 30 fps) and results in blurred frames when capturing fast motion. To address this issue, we propose leveraging event cameras, which provide high temporal resolution for capturing fast motion. However, the cost of current event camera models makes it prohibitive to use multiple event cameras for RGB-LFV platforms. Therefore, we propose EV-LFV, an event synthesis framework that generates full multi-subview event-based RGB-LFV with only one event camera and multiple traditional RGB cameras. EV-LFV utilizes spatial-angular convolution, ConvLSTM, and Transformer to model RGB-LFV's angular features, temporal features, and long-range dependency, respectively, to effectively synthesize event streams for RGB-LFV. To train EV-LFV, we construct the first event-to-LFV dataset consisting of 200 RGB-LFV sequences with ground-truth event streams. Experimental results demonstrate that EV-LFV outperforms state-of-the-art event synthesis methods for generating event-based RGB-LFV, effectively alleviating motion blur in the reconstructed RGB-LFV.
Zhicheng Lu, Xiaoming Chen 0006, Vera Chung, Tom Weidong Cai, Yiran Shen 0001
IEEE Trans. Vis. Comput. Graph.2
2023 LFACon: Introducing Anglewise Attention to No-Reference Quality Assessment in Light Field Space
abstract
Light field imaging can capture both the intensity information and the direction information of light rays. It naturally enables a six-degrees-of-freedom viewing experience and deep user engagement in virtual reality. Compared to 2D image assessment, light field image quality assessment (LFIQA) needs to consider not only the image quality in the spatial domain but also the quality consistency in the angular domain. However, there is a lack of metrics to effectively reflect the angular consistency and thus the angular quality of a light field image (LFI). Furthermore, the existing LFIQA metrics suffer from high computational costs due to the excessive data volume of LFIs. In this paper, we propose a novel concept of "anglewise attention" by introducing a multihead self-attention mechanism to the angular domain of an LFl. This mechanism better reflects the LFI quality. In particular, we propose three new attention kernels, including anglewise self-attention, anglewise grid attention, and anglewise central attention. These attention kernels can realize angular self-attention, extract multiangled features globally or selectively, and reduce the computational cost of feature extraction. By effectively incorporating the proposed kernels, we further propose our light field attentional convolutional neural network (LFACon) as an LFIQA metric. Our experimental results show that the proposed LFACon metric significantly outperforms the state-of-the-art LFIQA metrics. For the majority of distortion types, LFACon attains the best performance with lower complexity and less computational time.
Qiang Qu 0004, Xiaoming Chen 0006, Vera Chung, Tom Weidong Cai
IEEE Trans. Vis. Comput. Graph.2
2020 Light field reconstruction using hierarchical features fusion
Zexi Hu, Vera Chung, Wanli Ouyang, Xiaoming Chen 0006, Zhibo Chen 0001
Expert Syst. Appl.4
2019 360SRL: A Sequential Reinforcement Learning Approach for ABR Tile-Based 360 Video Streaming
abstract
Tile-based 360-degree video (360 video) streaming, employed with adaptive bitrate (ABR) algorithms, is a promising approach to offer high video quality of experience (QoE) within limited network bandwidth. Existing ABR algorithms, however, fail to achieve optimal performance in real-world fluctuated network conditions as they heavily rely on unbiased bandwidth predictions. Recently, reinforcement learning (RL) has shown promising potential in generating better ABR algorithms in 2D video streaming. However, unlike existed work in 2D video streaming, directly applying RL in the tile-based 360 video streaming is infeasible due to the resulting exponential decision space. To overcome these limitations, we propose in this paper 360SRL, an improved ABR algorithm employing Sequential RL (360SRL). Firstly, we reduce the decision space of 360SRL from exponential to linear by introducing a sequential ABR decision structure, thus making it feasible to be employed with RL. Secondly, instead of relying on accurate bandwidth predictions, 360SRL learns to make ABR decisions solely through observations of the resulting QoE performance of past decisions. Finally, we compare 360SRL to state-of-the-art ABR algorithms using trace-driven experiments. The experiment results demonstrate that 360SRL outperforms state-of-the-art algorithms with around 12% improvement in average QoE.
Jun Fu 0007, Xiaoming Chen 0006, Zhizheng Zhang 0004, Shilin Wu, Zhibo Chen 0001
ICME2
2019 ImmerTai: Immersive Motion Learning in VR Environments
Xiaoming Chen 0006, Zhibo Chen 0001, Tianyu He, Junhui Hou, Sen Liu 0001, Ying He 0001
J. Vis. Commun. Image Represent.1
2019 Improved image classification with 4D light-field and interleaved convolutional neural network
Zhicheng Lu, Henry Wing Fung Yeung, Qiang Qu 0004, Vera Chung, Xiaoming Chen 0006, Zhibo Chen 0001
Multim. Tools Appl.5
2019 Light Field Spatial Super-Resolution Using Deep Efficient Spatial-Angular Separable Convolution
abstract
Light field (LF) photography is an emerging paradigm for capturing more immersive representations of the real-world. However, arising from the inherent trade-off between the angular and spatial dimensions, the spatial resolution of LF images captured by commercial micro-lens based LF cameras are significantly constrained. In this paper, we propose effective and efficient end-to-end convolutional neural network models for spatially super-resolving LF images. Specifically, the proposed models have an hourglass shape, which allows feature extraction to be performed at the low resolution level to save both computational and memory costs. To fully make use of the four-dimensional (4-D) structure information of LF data in both spatial and angular domains, we propose to use 4-D convolution to characterize the relationship among pixels. Moreover, as an approximation of 4-D convolution, we also propose to use spatialangular separable (SAS) convolutions for more computationallyand memory-efficient extraction of spatial-angular joint features. Extensive experimental results on 57 test LF images with various challenging natural scenes show significant advantages from the proposed models over state-of-the-art methods. That is, an average PSNR gain of more than 3.0 dB and better visual quality are achieved, and our methods preserve the LF structure of the super-resolved LF images better, which is highly desirable for subsequent applications. In addition, the SAS convolutionbased model can achieve 3× speed up with only negligible reconstruction quality decrease when compared with the 4-D convolution-based one. The source code of our method is online available at https://github.com/spatialsr/DeepLightFieldSSR.
Henry Wing Fung Yeung, Junhui Hou, Xiaoming Chen 0006, Jie Chen 0026, Zhibo Chen 0001, Vera Chung
IEEE Trans. Image Process.3
2018 Fast Light Field Reconstruction with Deep Coarse-to-Fine Modeling of Spatial-Angular Clues
Henry Wing Fung Yeung, Junhui Hou, Jie Chen 0026, Vera Chung, Xiaoming Chen 0006
ECCV (6)5
2018 Efficient VR Video Representation and Quality Assessment
Shilin Wu, Xiaoming Chen 0006, Jun Fu 0007, Zhibo Chen 0001
J. Vis. Commun. Image Represent.2
2017 A novel stacked denoising autoencoder with swarm intelligence optimization for stock index prediction
abstract
This paper proposes the use of Stacked Denoising Autoencoder to predict the direction of movement of stock indexes based on the historical and volume data of the underlying stocks. The Stacked Denoising Autoencoder is a deep learning method widely used in the field of computer vision which is capable of learning a compact feature representation of the data for stock index prediction. The Hybrid Gravitational Search Algorithm, a swam intelligence based algorithm, is proposed to optimise the hyper-parameters of the deep network which mitigates the hyper-parameter tuning problem of the network.
Guang Liu 0002, Henry Wing Fung Yeung, Junfu Yin, Vera Chung, Xiaoming Chen 0006
IJCNN6
2017 Immersive and collaborative Taichi motion learning in various VR environments
abstract
Learning “motion” online or from video tutorials is usually inefficient since it is difficult to deliver “motion” information in traditional ways and in the ordinary PC platform. This paper presents ImmerTai, a system that can efficiently teach motion, in particular Chinese Taichi motion, in various immersive environments. ImmerTai captures the Taichi expert's motion and delivers to students the captured motion in multi-modal forms in immersive CAVE, HMD as well as ordinary PC environments. The students' motions are captured too for quality assessment and utilized to form a virtual collaborative learning atmosphere. We built up a Taichi motion dataset with 150 fundamental Taichi motions captured from 30 students, on which we evaluated the learning effectiveness and user experience of ImmerTai. The results show that ImmerTai can enhance the learning efficiency by up to 17.4% and the learning quality by up to 32.3%.
Tianyu He, Xiaoming Chen 0006, Zhibo Chen 0001, Sen Liu 0001, Junhui Hou, Ying He 0001
VR2
2016 Saliency & structure preserving multi-operator image retargeting
abstract
Content-aware image retargeting has attracted substantial research interests in the related research community. However, so far there is still no method can preserve important image contents and structure well without introducing deformation. To address this problem, we propose a Saliency & Structure Preserving Multi-operator (SSPM) method. SSPM classifies images into three categories utilizing SIFT density to improve performance of saliency preservation, helping to mitigate negative influence from center-bias property of most existing saliency detection models. SSPM also employs different principles to improve structure preservation performance, including Earth Mover's Distance (EMD) and Gray-Level Cooccurrence Matrix (GLCM) to get optimal operator sequences for smart content-aware image retargeting. SSPM method not only can well preserve salient contents and structure, but also can greatly improve deformation resilience. Experimental results demonstrated that our method outperforms state-of-art image retargeting methods.
Lingling Zhu, Zhibo Chen 0001, Xiaoming Chen 0006, Ning Liao
ICASSP3
2015 A Novel Optimized Watermark Embedding Scheme for Digital Images
Feng Sha, Felix Lo, Vera Chung, Xiaoming Chen 0006, Wei-Chang Yeh 0001
MMM (2)4
2012 Modeling and Compressing 3-D Facial Expressions Using Geometry Videos
abstract
In this paper, we present a novel geometry video (GV) framework to model and compress 3-D facial expressions. GV bridges the gap of 3-D motion data and 2-D video, and provides a natural way to apply the well-studied video processing techniques to motion data processing. Our framework includes a set of algorithms to construct GVs, such as hole filling, geodesic-based face segmentation, expression-invariant parameterization (EIP), and GV compression. Our EIP algorithm can guarantee the exact correspondence of the salient features (eyes, mouth, and nose) in different frames, which leads to GVs with better spatial and temporal coherence than that of the conventional parameterization methods. By taking advantage of this feature, we also propose a new H.264/AVC-based progressive directional prediction scheme, which can provide further 10%-16% bitrate reductions compared to the original H.264/AVC applied for GV compression while maintaining good video quality. Our experimental results on real-world datasets demonstrate that GV is very effective for modeling the high-resolution 3-D expression data, thus providing an attractive way in expression information processing for gaming and movie industry.
Jiazhi Xia, Dao Thi Phuong Quynh, Ying He 0001, Xiaoming Chen 0006, Steven C. H. Hoi
IEEE Trans. Circuits Syst. Video Technol.4
2011 Modeling 3D articulated motions with conformal geometry videos (CGVs)
abstract
3D articulated motions are widely used in entertainment, sports, military, and medical applications. Among various techniques for modeling 3D motions, geometry videos (GVs) are a compact representation in that each frame is parameterized to a 2D domain, which captures the 3D geometry (x, y, z) to a pixel (r, g, b) in the image domain. As a result, the widely studied image/video processing techniques can be directly borrowed for 3D motion. This paper presents conformal geometry videos (CGVs), a novel extension of the traditional geometry videos by taking into the consideration of the isometric nature of 3D articulated motions. We prove that the 3D articulated motion can be uniquely (up to rigid motion) represented by (»,H), where » is the conformal factor characterizing the intrinsic property of the 3D motion, and H the mean curvature characterizing the extrinsic feature (i.e., embedding or appearance). Furthermore, the conformal factor » is pose-invariant. Thus, in sharp contrast to the GVs which capture 3D motion by three channels, CGVs take only one channel of mean curvature H and the first frame of the conformal factor », i.e., approximately 1/3 the storage of the GVs. In addition, CGVs have strong spatial and temporal coherence, which favors various well studied video compression techniques. Thus, CGVs can be highly compressed by using the state-of the-art video compression techniques, such as H.264/AVC. Our experimental results on real-world 3D motions show that CGVs are a highly compact representation for 3D articulated motions, i.e., given CGVs and GVs of the same file size, CGVs show much better visual quality than GVs.
Dao Thi Phuong Quynh, Ying He 0001, Xiaoming Chen 0006, Jiazhi Xia, Qian Sun 0003, Steven C. H. Hoi
ACM Multimedia3
2011 Sensor-Assisted Video Encoding for Mobile Devices in Real-World Environments
abstract
In this paper, we present a comprehensive study on sensor-assisted video encoding (SaVE) schemes for video capturing on mobile devices in real-world environments. Our purpose is to reduce the computational complexity of video encoding by leveraging sensors that are increasingly available on mobile devices, e.g., accelerometers and digital compasses. Motion estimation is a key component of video encoding. In this paper, SaVE calculates the rotational movement of a camera (on mobile devices) and then infers the global motion in the camera imager. SaVE subsequently employs the estimated global motion as predictors to simplify motion estimation algorithms for state-of-the-art H.264/AVC video coding. We have constructed a prototype of SaVE and evaluated its performance with a pair of accelerometers, a digital compass, and their combination. Our experimental results show that SaVE can significantly reduce the computations of motion estimation while achieving equal or better video quality. Our results also show that SaVE has a strong noise-resistant capability. Therefore, it can be practically employed in real-world environments.
Xiaoming Chen 0006, Zhendong Zhao, Ahmad Rahmati, Ye Wang 0007, Lin Zhong 0001
IEEE Trans. Circuits Syst. Video Technol.1
2010 Modeling 3D facial expressions using geometry videos
abstract
The significant advances in developing high-speed shape acquisition devices make it possible to capture the moving and deforming objects at video speeds. However, due to its complicated nature, it is technically challenging to effectively model and store the captured motion data. In this paper, we present a set of algorithms to construct geometry videos for 3D facial expressions, including hole filling, geodesic-based face segmentation, and expression-invariant parametrization. Our algorithms are efficient and robust, and can guarantee the exact correspondence of the salient features (eyes, mouth and nose). Geometry video naturally bridges the 3D motion data and 2D video, and provides a way to borrow the well-studied video processing techniques to motion data processing. With our proposed intra-frame prediction scheme based on H.264/AVC, we are able to compress the geometry videos into a very compact size while maintaining the video quality. Our experimental results on real-world datasets demonstrate that geometry video is effective for modeling the high-resolution 3D expression data.
Jiazhi Xia, Ying He 0001, Dao Thi Phuong Quynh, Xiaoming Chen 0006, Steven C. H. Hoi
ACM Multimedia4
2009 SaVE: sensor-assisted motion estimation for efficient h.264/AVC video encoding
abstract
Motion estimation is a key component of modern video encoding and is very compute-intensive. We present a novel Sensor-assisted Video Encoding (SaVE) method to reduce the computational complexity of motion estimation in H.264/AVC encoders, leveraging accelerometers and digital compasses that are increasingly available on mobile devices. Using these sensors, SaVE calculates the rotational movement of a camera and then infers the global motion in the camera image sensor; it subsequently employs the estimated global motion to simplify the state-of-the-art motion estimation algorithms, UMHS and EPZS used in H.264/AVC encoders. We have constructed a prototype of SaVE and report extensive evaluation of it. Our experimental results show that SaVE can reduce the computations of UMHS and EPZS algorithms by up to 27% and 18%, respectively, while achieving the same or better video quality.
Xiaoming Chen 0006, Zhendong Zhao, Ahmad Rahmati, Ye Wang 0007, Lin Zhong 0001
ACM Multimedia1
2007 A New Optimized Error-Resilient Coding for Video Applications
abstract
This paper has proposed and implemented a new optimized and adaptive Error-Resilient Coding based on Statistical Parity Pair Embedding (SPPE). The SPPE can locate errors in the compressed I-frame bit stream and this system design can be applied to the MPEG2 wireless video applications. The SPPE is based on the information hiding technique. This SPPE algorithm is proven to have better performance than other traditional algorithms, and is very suitable for use in multipoint communications. In this paper three Error Concealment (EC) algorithms: Previous Coded Block Linear Error Concealment (PCB-EC), Four Neighbors Linear Error Concealment (4N-EC) and Eight Neighbors Linear Error Concealment (8N-EC) were implemented and tested. We found that the proposed SPPE and 4N-EC algorithm together can achieve PSNR increments of 8.2dB after recovery while other algorithms [1-3] can only improve 1-2dB after recovery.
Changseok Bae, Vera Chung, Xiaoming Chen 0006, Mohd Afizi Mohd Shukran
ICME4
2007 A New Multi-Mode Intra-Frame Error Concealment Algorithm for H.264/AVC
abstract
In video communications over error-prone environments, compressed video is fragile to transmission errors. The decoder side error concealment (EC) is an efficient way to recover a damaged video sequence. Typically traditional video EC algorithms for H.264/AVC intra-frames operate in macroblock (MB) level and they offer a single concealment mode for all frames regardless of frame features. This paper proposes a new multi-mode error concealment (MMEC) algorithm for intra-frames of H.264/AVC. The proposed algorithm provides several different EC modes so that the decoder is able to choose the best modes to conceal errors. By applying the proposed method, a MB can be partitioned into smaller 8 times 8 subblocks, each subblock can be recovered using available temporal or spatial information based on the MB characteristics. If there are any subblocks recovered by using temporal information, these subblocks will be further used for directional spatial EC for the rest subblocks. The proposed new MMEC algorithm has been evaluated and compared with the H.264/AVC reference implementation and two other classical EC algorithms. The experimental results show that the proposed MMEC algorithm can achieve up to 5 dB gains in PSNR therefore significantly improves the video quality.
Xiaoming Chen 0006, Vera Chung, Changseok Bae
ICME1
2006 A new error resilient coding schemes for the home entertainment video
abstract
In this paper, a new error resilient coding system based on the Three Layer Error Control Coding (TLECC) techniques for MPEG2[1] coded video is proposed. The results show that the proposed TLECC system can achieve 100% accuracy in the retrieval of a test sequence signal if the transmission noise is below 5%. This system is very useful for the home entertainment video coding system, and the internet or networked video coding system.
Changseok Bae, Vera Chung, Xiaoming Chen 0006, Ahmed Fawzi Otoom
CCNC3
2006 A study of clustering algorithm for wavelet-based image retrieval system
abstract
In this paper we propose a new Two Level Clustering Wavelet-based Image REtrieval System (TLWIRES). The TLWIRES provides a two level clustering and searching. The first level is based on colour variations and the second level is based on colour histograms. The experimental results showed that the proposed TLWIRES can have high accuracy and high speed. This system is very suitable for applications such as home multimedia server to search for images.
Vera Chung, Xiaoming Chen 0006
CCNC2
2005 Implementation of Image Steganographic System Using Wavelet Zerotree
Vera Chung, Penghao Wang 0002, Xiaoming Chen 0006, Changseok Bae
KES (1)4
2005 A Performance Comparison of High Capacity Digital Watermarking Systems
Vera Chung, Penghao Wang 0002, Xiaoming Chen 0006, Changseok Bae, Ahmed Fawzi Otoom, Tich Phuoc Tran
KES (1)3