Mei Guo

dblp:03/1330 · DBLP profile ↗
← Back
26ranked-venue papers
16as first author
20since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 4 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021Systems, architecture and hardware · 7 · 6 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Synchronizing with Attentional Sampling: Modality-Specific Flicker Guidance for Gaze and Hand-Eye Interaction in VR
abstract
Current attentional guidance systems in Virtual Reality (VR) typically treat the user as a static receiver, employing fixed visual cues that fail to account for the fluctuating cognitive state. We address this limitation by positing that effective guidance requires modulating external stimuli to correspond with the brain’s intrinsic attentional sampling mechanism. Using EEG in an immersive environment, we provide direct evidence that this sampling periodicity is not fixed; specifically, the intensified sensorimotor integration load of hand-eye coordination drives the endogenous rhythm to decelerate from the alpha band (∼8 Hz) to the theta band (∼4 Hz). This neural adaptation dictates the temporal requirements for external guidance: identifying optimal parameters via Pareto optimization, we demonstrate that while a 4 Hz cue suffices for gaze, the computationally demanding coordination task requires a higher-frequency (7 Hz) cue to ensure sufficient temporal signal density. Validated in ecological VR scenarios, our adaptive strategy significantly enhanced interaction efficiency without increasing cognitive load. Beyond the specific implementation of flicker, this work establishes a critical design principle for next-generation attention-aware interfaces: maximizing performance by synchronizing information presentation with the user’s task-induced sensorimotor state.
Songyue Yang, Kang Yue, Haolin Gao, Mei Guo, Zhonghao Zhu, Fanlu Zeng, Yu Liu 0081
VR4
2026 A bio-inspired neuromorphic system for fusing visual features and autonomous learning
Mei Guo, Yaoyao Zi, Jikang Liu, Qiye Yang, Jingzhi Xu, Gang Dou, Da Chen 0004
Neural Networks1
2026 A Neuromorphic Circuit With Supramodal Attention Effects Based on Cognitive Resource Limitation
abstract
When organisms face complex environments, the cognitive resource limitation is an important mechanism to ensure the quality of perceived information and prevent information overload. Organisms allocate cognitive resources rationally by regulating attention, thus promoting more important cognitive orientations. However, this phenomenon has been scarcely investigated within the realm of memristive biomimetic circuits. The prefrontal cortex (PFC), as the highest central hub for attention control, achieves goal-oriented attentional selection by assigning the basal ganglion to suppress irrelevant information. Based on this biological mechanism, a neuromorphic circuit has been designed in this paper to implement the attentional regulation function of the PFC between bimodal sensory inputs. When organisms face multi-sensory information input, the enhancement and inhibition effects in the supramodal attention effects are considered. In addition, the circuit realizes biological phenomena such as temporal consistency, semantic consistency, emotional attention, and attention fatigue. The attention regulation mechanism is further extended to more senses with variable sensitivities, providing variable strategies for performing tasks in different scenarios. Performance analysis results demonstrate that the circuit exhibits excellent robustness. This work provides guidance for the further development of information processing in brain-inspired intelligence.
Mei Guo, Chenguang Zheng, Gang Dou, Herbert H. C. Iu
IEEE Trans. Circuits Syst. I Regul. Pap.1
2026 Multisensory Memristive Circuits With Parallel Processing and Dual Adaptive Features
abstract
As the brain-like intelligence develops rapidly, it is urgent to design a more convenient and efficient control framework to cope with the challenge of processing multisensory signals in parallel. Therefore, a multisensory memristive circuit with dual adaptive, parallel processing, and multilevel reinforcement features is proposed. The circuit is mainly composed of modules for receptors, STM and LTM, attention, environmental monitoring and mutual associative memory. Automatic encoding of different sensorial signals is realised by the receptor modules. Dual adaptive regulation of the internal associative memory and external environmental changes on the circuit is implemented by modules of attention and environmental monitoring. Multilevel reinforcement memory is achieved through the interconnection of multiple dimensional features of the same objects. The process of encoding transformation of stimuli, experience memory, and feedback learning is automatically achieved in the brain-inspired neural network structure, which avoids the problems such as encoding difficulties during the conversion of the operating objects, and enables the realization of more brain-like intelligence. The circuit is applied to gripping and recognizing in robotic arms and the scenario memory of different production lines is simulated, which is promising for application in automated factories.
Mei Guo, Xingwei Zhang, Wenhai Guo, Gang Dou, Da Chen 0004, Herbert H. C. Iu
IEEE Trans. Circuits Syst. I Regul. Pap.1
2025 SWAM: Adaptive Sliding Window and Memory-Augmented Attention Model for Rumor Detection
abstract
Detecting rumors on social media has become a critical task in combating misinformation.Existing propagation-based rumor detection methods often focus on the static propagation graph, overlooking that rumor propagation is inherently dynamic and incremental in the real world.Recently propagation-based rumor detection models attempt to use the dynamic graph that is associated with coarse-grained temporal information.However, these methods fail to capture the long-term time dependency and detailed temporal features of propagation.To address these issues, we propose a novel adaptive Sliding Window and memory-augmented Attention Model (SWAM) for rumor detection.The adaptive sliding window divides the sequence of posts into consecutive disjoint windows based on the propagation rate of nodes.We also propose a memory-augmented attention to capture the long-term dependency and the depth of nodes in the propagation graph.Multi-head attention mechanism is applied between nodes in the memorybank and incremental nodes to iteratively update the memorybank, and the depth information of nodes is also considered.Finally, the propagation features of nodes in the memorybank are utilized for rumor detection.Experimental results on two public real-world datasets demonstrate the effectiveness of our model compared with the state-of-the-art baselines.
Mei Guo, Chen Chen 0012, Chunyan Hou, Yike Wu 0002, Xiaojie Yuan
EMNLP1
2025 Improving Pointing Accuracy for 3D Target Selection in Virtual Reality Through Depth Perception Biases Correction
abstract
Accurate 3D target selection in virtual reality (VR) is fundamentally impeded by pointing uncertainty along the depth axis, a challenge that existing 2D pointing models fail to address due to the complexities of depth perception. Near-eye interactions in VR are influenced by binocular depth cues and vergence-accommodation conflicts (VAC), which introduce significant depth perception biases that impair predictive performance. To address this issue, we first investigate these factors and derive a Gaussian distribution to model near-field depth biases within a 2.5m range. Second, to analyze pointing performance across this extended depth range, we classify 3D target motions into three distinct types: motion-indepth, motion-in-plane, and combined motion. Our analysis identifies that interaction depth and motion amplitude are the two most critical factors influencing pointing accuracy. Accordingly, by incorporating these factors alongside our perceptual bias Gaussian into the Ternary-Gaussian framework, we demonstrate significantly improved predictive performance across diverse 3D motion scenarios. These findings enhance the understanding of user perception in virtual environments and support the development of precise, context-aware interaction cues. Future research can extend these models to design real-time adaptive interfaces, thereby elevating user experiences in VR.
Songyue Yang, Kang Yue, Haolin Gao, Yiyi Yang, Mei Guo, Yu Liu 0081, Zhonghao Zhu, Yue Liu 0005
ISMAR5
2025 MvWECM: Multi-view Weighted Evidential C-Means clustering
Kuang Zhou, Mei Guo
Pattern Recognit.3
2025 Design and Application of Brain-Inspired Circuit With Context-Dependent and State-Dependent Memory
abstract
The context and the state of mind are important retrieval cues for long-term memory, which helps information to be retrieved quickly. However, most memristive circuits focus on the process of information memory, few studies consider the process of information retrieval. In this work, a brain-inspired circuit with context-dependent and state-dependent memory is proposed based on the three-level processing model of memory information, which integrates the processes of information memory and information retrieval. The circuit includes sensory memory module, short-term memory module, long-term memory module, information retrieval module, status module, and context module. In the circuit, information, contexts, and states are eventually transferred to long-term memory module for storage and retrieval. Meanwhile, the factors influencing information retrieval are considered, such as the degree of information memory, the time interval between information memory and retrieval, the context, and the state. And the proposed circuit has scalability, which realizes the memory of information in multiple contexts. Finally, based on the characteristics of memristors, the proposed circuit is extended for detecting damage to the machining accuracy of the mobile CNC lathe. Combining brain-inspired circuits with human memory mechanism, this work provides further reference for the research of brain-like intelligence.
Gang Dou, Daoguo Li, Mei Guo, Herbert H. C. Iu
IEEE Trans. Circuits Syst. I Regul. Pap.3
2025 A High-Performance Memristive Circuit Design for DCGAN in Edge Computing
abstract
Edge computing devices based on the von Neumann architecture can’t fulfill the demand for computational resources in Generative Adversarial Networks. This paper proposes a memristive circuit design for a light-weight and efficient Deep Convolutional Generative Adversarial Networks (DCGAN), which can be integrated into edge computing devices for image generation. The DCGAN scheme can perform convolution operations, deconvolution operations, and various activation functions in a fast and low-power way. Moreover, a high-precision segmental approximate linear weight mapping method based on the 2-Memristor crossbar array structure is proposed to improve the precision of memristive neural networks on edge computing devices. Finally, the results show that the DCGAN scheme significantly reduces the power consumption, time consumption, and input ports while keeping the area overhead unchanged. In the Oxford 17 image generation task, the DCGAN scheme achieves faster speed and lower power consumption compared to the traditional structure. The DCGAN scheme based on memristive circuits provides some references for implementing more intelligent applications on edge devices.
Mei Guo, Gang Dou, Herbert H. C. Iu
IEEE Trans. Circuits Syst. I Regul. Pap.1
2025 A Knowledge Distillation Online Training Circuit for Fault Tolerance in Memristor Crossbar Array-Based Neural Networks
abstract
Knowledge distillation is widely used as an effective model compression technique to improve the performance of small models. Most of the current researches on knowledge distillation focus on the algorithmic level and ignore the potential benefits of hardware implementation. In this paper, a multi-loss knowledge distillation online training circuit based on memristor crossbar array is designed, which can improve the inference efficiency and reduce the power consumption of deep learning models on edge devices. The circuit is able to process data in real time, and it can be used to handle stuck-at-faults (SAF) caused by factors such as manufacturing defects in the memristor. Moreover, a fault detection scheme with low time cost is proposed in order to address the low efficiency of stuck-at-fault detection in memristor crossbar arrays. The scheme is combined with a self-compensating pruning method and knowledge distillation online training mechanism, which significantly improves the model training and inference capability of the circuit under fault conditions. Experimental results show that the multi-loss knowledge distillation online training improves the accuracy by 4.15% and 63.48% respectively in two models compared with traditional training schemes. The fault-tolerance scheme reduces the power consumption of the memristor crossbar arrays by 41.2% and 72.6% respectively on the two models, demonstrating its potential and advantages in edge computing.
Mei Guo, Xingwei Zhang, Gang Dou, Herbert H. C. Iu
IEEE Trans. Circuits Syst. I Regul. Pap.1
2025 Adaptive Fuzzy Fixed-Time Control for Uncertain Time-Delay Nonlinear Systems With Output Constraints
abstract
This paper aims to address two complex issues in the control of a class of high-order time-delay nonlinear systems: i) the adaptive fixed-time tracking control and ii) the differential explosion arising from the iterative derivation of the intermediate control law during the design process. The issue is how to construct a tracking controller to ensure that the tracking error converges to an adjustable region around the origin in a fixed time. This article develops an adaptive fixed-time control strategy by employing fuzzy logic system (FLS) to approximate the unknown function terms and the nonlinear growth assumption often used in unknown systems is eliminated. This strategy combines the dynamic surface technology with the first-order filtered signals in recursive design, and effectively addresses the sticky problem of complexity explosion in controller design. Finally, the effectiveness and feasibility of this control scheme are demonstrated through a single-link manipulator system and a numerical example
Gang Dou, Tianliang Zhang 0004, Weihai Zhang, Mei Guo
IEEE Trans. Fuzzy Syst.5
2025 Dynamic Changes of Latency Perception Threshold in Virtual Reality: Behavioral and EEG Evidence
abstract
Virtual Reality (VR) technologies in fields such as telehealth, teleconferencing, and virtual education are significantly affected by end-to-end latency, which notably impacts users' interactive experience and performance. Previous research suggests that a perceptual threshold may exist-once latency is reduced below a certain level, users no longer perceive it, and their interactive performance remains largely unaffected. However, there is no consensus on the exact value of this absolute latency perception threshold. In this study, we employed an experimental design based on Fitts' law to investigate whether interaction strategies and task difficulty can alter the latency perception threshold (LPT), and how variations in this threshold influence users' interactive performance. The results show that the LPT is approximately 130-170 ms, and that when interaction strategies prioritize speed or when tasks become more challenging, users exhibit heightened sensitivity to latency. Due to the presence of the LPT, the effect of latency on interactive performance follows a nonlinear pattern, and building on this finding, we refined a Fitts' law model to incorporate the influence of latency. Notably, electroencephalogram (EEG) signals can still capture users' perception of latency when they are unaware of minor latency, demonstrating a level of sensitivity that exceeds conscious awareness. Our findings provide insights into latency effects on performance and perception, guiding the design of more responsive VR interaction systems.
Songyue Yang, Kang Yue, Haolin Gao, Mei Guo, Yu Liu 0081, Dan Zhang 0014, Yue Liu 0005
IEEE Trans. Vis. Comput. Graph.4
2024 Exploring Depth-based Perception Conflicts in Virtual Reality through Error-Related Potentials
abstract
Virtual Reality (VR) offers a valuable platform for real-life skills training. However, previous research has indicated that human’s perception of depth in VR differs from that of the real world. Such perceptual conflicts can impact immersion and the learning of skills, thus attracting widespread attention. Various methods have been proposed to enhance users’ depth perception, yet the underlying mechanisms of depth perception conflicts still require further research. In this paper, we used Error-Related Potentials (ErrPs) from electroencephalography (EEG) data to investigate the differences in participants’ perceptions at varying depths within the near-field. We designed a within-subjects experiment to successfully introduce depth perception conflicts. From participants exposed to three distinct depths, we collected questionnaire results, performance data, and EEG data. Our findings showed that EEG can effectively detect depth perception conflicts and, following each conflict, participants’ behavioral patterns showed significant changes. In situations with shallower depths, participants exhibited stronger responses to the designed conflicts. This increased sensitivity correlates with their accuracy in depth estimation. This study represents a novel approach to depth perception in VR using ErrPs, setting the stage for further use of physiological signals to measure the granularity of depth perception in VR/AR environments.
Haolin Gao, Kang Yue, Songyue Yang, Yu Liu 0081, Mei Guo, Yue Liu 0005
VR5
2024 Neuromorphic Circuit of Classical and Operant Conditioning Based on Tunable Neural Circuitry Motifs
abstract
Most memristive bionic circuits focus on how to realize bionic functions, few studies consider the biomimetic of the circuit structure and operation rules, so it is difficult to learn, memorize, and make decisions as biological neural networks. In this work, a multifunctional neuromorphic circuit inspired by tunable neural circuitry motifs is proposed. The circuit is more closely with biological characteristics in both structure and functions, which is designed based on neural circuit architectures. By connecting different neural circuitry motifs, the circuit realizes operant conditioning functions such as random exploration, behavioral frequency modulation, and decision-making. Also, the circuit integrated classical conditioning and operant conditioning in order to mimic the decision-making process, which was driven by the association of secondary and primary stimuli. In addition, the factors influencing decision-making are researched, such as the rates of learning and forgetting, and the conversion of short-term to long-term memory. The operational results of the proposed circuits in LTspice show that they can mimic the aforementioned functions, which have advantages in bionicity and scalability. This work can be applied in intelligent robotic platforms to achieve exploration and rescue in complex environments.
Mei Guo, Lingtong Kong, Gang Dou, Herbert H. C. Iu
IEEE Trans. Circuits Syst. I Regul. Pap.1
2023 Segmentation and Grading Method of Potato Late-Blight on field by Improved Mask R-CNN
abstract
Potato late-blight is a severe plant disease caused by the fungus Phytophthora infestans. This disease can quickly spread in a short period, leading to a significant decrease in potato production and posing great harm to the global potato planting industry. Therefore, an effective method for identifying late epidemic diseases has important practical significance. This paper proposed potato late-blight segmentation and grading method for the image with multiple leaves collected on farm. The improved Mask R-CNN network is presented to segment potato leaves at the pixel level and then grade late blight based on the segmentation information. Then, the segmentation effect of the model on the image will determine the accuracy of late blight evaluation. The neural network achieved a Mask recognition rate of 88.99% and 91.35% for leaves and lesions. It is possible to achieve disease grading operations for late epidemic diseases.
Mei Guo, Xiangqian Yin
SECON2
2023 Implementing bionic associate memory based on spiking signal
Mei Guo, Kaixuan Zhao, Junwei Sun 0002, Shiping Wen 0001, Gang Dou
Inf. Sci.1
2022 Chinese Relation Extraction of Apple Diseases and Pests Based on BERT and Entity Information
Mei Guo, Nan Geng, Yaojun Geng
KSEM (3)1
2022 Evidential prototype-based clustering based on transfer learning
Kuang Zhou, Mei Guo, Arnaud Martin 0001
Int. J. Approx. Reason.2
2022 An associative memory circuit based on physical memristors
Mei Guo, Yongliang Zhu, Renyuan Liu, Kaixuan Zhao, Gang Dou
Neurocomputing1
2022 Investigate the Neuro Mechanisms of Stereoscopic Visual Fatigue
abstract
Stereoscopic visual fatigue (SVF) due to prolonged immersion in the virtual environment can lead to negative user experience, thus hindering the development of virtual reality (VR) industry. Previous studies have focused on investigating the evaluation indicators associated with SVF, while few studies have been conducted to reveal the underlying neural mechanism, especially in VR applications. In this paper, a modified Go/NoGo paradigm was adopted to induce SVF in VR environment with Go trials for maintaining participants' attention and NoGo trials for investigating the neural effects under SVF. Random dot stereograms (RDSs) with 11 disparities were presented to evoke the depth-related visual evoked potentials (DVEPs) during 64-channel EEG recordings. EEG datasets collected from 15 participants in NoGo trials were selected to conduct individual processing and group analysis, in which the characteristics of the DVEPs components for various fatigue degrees were compared and independent components were clustered to explore the original cortex areas related to SVF. Point-by-point permutation statistics revealed that DVEPs sample points from 230 ms to 280 ms (component P2) in most brain areas changed significantly when SVF increased. Additionally, independent component analysis (ICA) identified that component P2 which originated from posterior cingulate cortex and precuneus, was associated statistically with SVF. We believe that SVF is rather a conscious status concerning the changes of self-awareness or self-location awareness than the performance reduction of retinal image processing. Moreover, we suggest that indicators representing higher conscious state may be a better indicator for SVF evaluation in VR environments.
Kang Yue, Mei Guo, Yue Liu 0005, Haochen Hu, Danli Wang
IEEE J. Biomed. Health Informatics2
2013 Inter-layer adaptive filtering for scalable extension of HEVC
abstract
This paper introduces a novel inter-layer adaptive filtering method for scalable extension of High Efficiency Video Coding (HEVC) standard, which is being developed by the Joint Collaborative Team on Video Coding (JCT-VC). The scalable extension of HEVC is currently utilizing an interlayer texture prediction method, in which the prediction of texture at enhancement layer is derived from the collocated reconstructions of base layer. The proposed method applies an adaptive filter to the reconstructions of base layer and generates the predictor of enhancement layer. It aims at enhancing the coding efficiency of enhancement layer by further reducing the redundant texture information in interlayer texture prediction. Two techniques are presented in this paper. The first one applies the adaptive filter to the reconstructions of base layer, which is followed by the fixed upsampling in spatial scalability. In the second one, an adaptive upsampling filter is further applied to filtered base layer reconstructions to generate the pixels at interpolated positions. Compared with the Scalable Mode under Consideration (SMuC) version 0.1.1 of HEVC scalable extension, average 0.8% and 3.2% bit-rate reductions are achieved for spatial scalability and SNR scalability respectively.
Mei Guo, Shan Liu 0001, Shawmin Lei
PCS1
2013 Inter-layer intra mode prediction for scalable extension of HEVC
abstract
This paper introduces a novel inter-layer intra mode prediction method for scalable extension of High Efficiency Video Coding (HEVC) standard, which is being developed by the Joint Collaborative Team on Video Coding (JCT-VC). In HEVC and its scalable extension, thirty-five intra prediction modes are adopted to reduce the spatial redundancy of luma texture within one frame, which may introduce noticeable overhead of delivering the intra mode information in the scalable bit-stream. In this paper, the correlation of intra modes between different layers is exploited to improve the efficiency of intra mode coding in enhancement layer. The intra modes in enhancement layer are mainly predicted with the ones of collocated blocks at base layer. Two techniques are presented in this paper with some differences in terms of the derivation of intra mode predictor from base layer and the selection of three most probable modes at enhancement layer. Compared to Test Model version 1.0 of HEVC scalable extension, 0.4% and 0.1% bit-rate reductions can be achieved with technique 2 in All Intra 2x spatial scalability and All Intra 1.5x spatial scalability respectively, while technique 1 can achieve 0.2% and 0.1% reductions. There is no obvious running time increase for both encoder and decoder.
Mei Guo, Shan Liu 0001, Shawmin Lei, Junghye Min, Tammy Lee
PCS1
2011 Witsenhausen-Wyner Video Coding
abstract
Inspired by Witsenhausen and Wyner's 1980 (now expired) patent on “interframe coder for video signals,” this paper presents a Witsenhausen-Wyner video codec, where the motion-compensated previously decoded video frame is used at the decoder as side information for joint decoding. Specifically, we replace predictive Inter coding in H.264/AVC by the syndrome-based coding scheme of Witsenhausen and Wyner, while keeping the Intra and Skip modes of H.264/AVC unchanged. We employ forward motion estimation at the encoder and send the motion vectors to help generate side information at the decoder, since our focus is not on low-complexity encoding. We also examine the tradeoff between the motion vector resolution and coding efficiency. Within the Witsenhausen-Wyner coding mode, we optimize the decision between syndrome coding and entropy coding among different discrete cosine transform (DCT) bands and among different bit-planes within each DCT coefficient. Extensive simulations of video transmission over wireless networks show that Witsenhausen-Wyner video coding is more robust against channel errors than H.264/AVC. The price paid for enhanced error-resilience with Witsenhausen-Wyner coding is a small loss in compression efficiency.
Mei Guo, Zixiang Xiong, Feng Wu 0001, Debin Zhao, Xiangyang Ji, Wen Gao 0001
IEEE Trans. Circuits Syst. Video Technol.1
2008 Wyner-Ziv Switching Scheme for Multiple Bit-Rate Video Streaming
abstract
This paper proposes a Wyner-Ziv (WZ) switching scheme for multiple bit-rate (MBR) video streaming over networks. Identical video content is encoded into a set of normal streams, which are generated by conventional hybrid video coding with multiple bit rates, so that streaming can dynamically switch among these normal streams according to available bandwidth. At encoder side, the WZ codec generates a switching stream by compressing the reconstructed frames of a certain normal stream that will be switched to, no matter which normal stream it switches from. At decoder side, for switching to the same frame, the same WZ switching stream is used to reconstruct the switching-to frame by taking the switching-from frame as the side information. The number of required WZ bits depends on the inherent mutual correlation between two frames switching to and from. Since the WZ switching streams are generated independently of the normal switching-from streams, given normal streams that can switch from any one to another, the proposed scheme reduces the number of switching streams from to . Furthermore, switching streams do not deteriorate the coding efficiency of normal streams when no switching occurs. However a big problem here, similar to requesting bits in distributed video coding, is how many WZ bits should be transmitted when a switching happens because the streaming scenario does not tolerate too much extra delay caused by the requests back and forth. Therefore, a Laplacian model, which is proved in the simplified case, is proposed to characterize the correlation between switching-to and switching-from frames. It can be used to accurately estimate the number of WZ bits at the server side.
Mei Guo, Yan Lu 0001, Feng Wu 0001, Debin Zhao, Wen Gao 0001
IEEE Trans. Circuits Syst. Video Technol.1
2007 Distributed Video Coding with Spatial Correlation Exploited Only at the Decoder
abstract
A new pixel-domain distributed video coding (DVC) scheme is proposed in this paper, in which both the temporal and the spatial correlations are exploited only at the decoder. A video is treated as a collection of data correlated in temporal and spatial directions. Besides splitting a video into frames at different time instants, a frame is further split by spatially sub-sampling. Each yielded part is then encoded individually. At the decoder, the side information signals are from both adjacent frames and the spatially decoupled signals. To utilize these multiple side information signals, a new probability model is proposed, in which the transitional probability is calculated from the conditional probabilities on the multiple side information signals. The coding efficiency is enhanced by further removing the spatial redundancy, while the encoding complexity remains the same as the previous pixel-domain DVC techniques that only consider the temporal correlation
Mei Guo, Yan Lu 0001, Feng Wu 0001, Shipeng Li 0001, Wen Gao 0001
ISCAS1
2006 Practical Wyner-Ziv Switching Scheme for Multiple Bit-Rate Video Streaming
abstract
In this paper, we propose a novel bit-stream switching scheme for the multiple bit-rate (MBR) video streaming, in which a Wyner-Ziv coded frame is used to overcome the mismatch between the MBR streams when the switching occurs. With the proposed technique, the MBR streams can be independently encoded without losing any coding efficiency. Similar to distributed video coding, the proposed Wyner-Ziv switching scheme also faces the challenge of rate allocation at the server side. To solve this problem, we propose a new correlation model based on the analysis on the reconstructed frames from the streams with different bit rates. Accordingly, the number of transmitted bits can be on-line calculated based on the correlation model without any feedback from the decoder. With the proposed technique, the actually transmitted Wyner-Ziv bits are only few more than the truly requested bits. However, the delay due to the bit requesting process can be avoided.
Mei Guo, Yan Lu 0001, Feng Wu 0001, Debin Zhao, Wen Gao 0001
ICIP1