Yin Zhao

dblp:19/8539 · DBLP profile ↗
← Back
35ranked-venue papers
15as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 22 · 10 first-author · 5 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author
YearPublicationVenuePosition
2026 ColorBench: Benchmarking Mobile Agents with Graph-Structured Framework for Complex Long-Horizon Tasks
abstract
The rapid advancement of multimodal large language models has enabled agents to operate mobile devices by directly interacting with graphical user interfaces, opening new possibilities for mobile automation. However, real-world mobile tasks are often complex and allow for multiple valid solutions. This contradicts current mobile agent evaluation standards: offline static benchmarks can only validate a single predefined ''golden path'', while online dynamic testing is constrained by the complexity and non-reproducibility of real devices, making both approaches inadequate for comprehensively assessing agent capabilities. To bridge the gap between offline and online evaluation and enhance testing stability, this paper introduces a novel graph-structured benchmarking framework. By modeling the finite states observed during real-device interactions, it achieves static simulation of dynamic behaviors. Building on this, we develop ColorBench, a benchmark focused on complex long-horizon tasks. It supports evaluation of multiple valid solutions, subtask completion rate statistics, and atomic-level capability analysis. ColorBench contains 175 tasks (74 single-app, 101 cross-app) with an average length of over 13 steps. Each task includes at least two correct paths and several typical error paths, enabling quasi-dynamic interaction.
Yuanyi Song, Heyuan Huang, Qiqiang Lin, Yin Zhao, Xiangmou Qu, Jun Wang 0152, Xingyu Lou, Weiwen Liu, Zhuosheng Zhang 0001, Jun Wang 0020, Zhaoxiang Wang, Yong Yu 0001, Weinan Zhang 0001
WWW4
2026 Constrained economic emission dispatch using gravitational search algorithm integrating multi-trial vectors and hierarchical local optimization
Chaodong Fan, Yin Zhao, Leyi Xiao
Inf. Sci.2
2025 Robust Function-Calling for On-Device Language Model via Function Masking
abstract
Large language models have demonstrated impressive value in performing as autonomous agents when equipped with external tools and API calls. Nonetheless, effectively harnessing their potential for executing complex tasks crucially relies on enhancements in their function-calling capabilities. This paper identifies a critical gap in existing function-calling models, where performance varies significantly across benchmarks, often due to over-fitting to specific naming conventions. To address such an issue, we introduce Hammer, a novel family of foundation models specifically engineered for on-device function calling. Hammer employs an augmented dataset that enhances models’ sensitivity to irrelevant functions and incorporates function masking techniques to minimize over-fitting. Our empirical evaluations reveal that Hammer not only outperforms larger models but also demonstrates robust generalization across diverse benchmarks, achieving state-of-the-art results. Our open-source contributions include a specialized dataset for irrelevance detection, a tuning framework for enhanced generalization, and the Hammer models, establishing a new standard for function-calling performance.
Qiqiang Lin, Muning Wen, Qiuying Peng, Guanyu Nie, Junwei Liao, Xiaoyun Mo, Jiamu Zhou, Yin Zhao, Jun Wang 0012, Weinan Zhang 0001
ICLR10
2025 Learning subjective time-series data via Utopia Label Distribution Approximation
Xuefeng Liang, Hexin Jiang, Yin Zhao
Pattern Recognit.5
2024 Bit Rate Matching Algorithm Optimization in JPEG-AI Verification Model
abstract
The research on neural network (NN) based image compression has shown superior performance compared to classical compression frameworks. Unlike the hand-engineered transforms in the classical frameworks, NN-based models learn the non-linear transforms providing more compact bit represen-tations, and achieve faster coding speed on parallel devices over their classical counterparts. Those properties evoked the attention of both scientific and industrial communities, resulting in the standardization activity JPEG-AI. The verification model for the standardization process of JPEG-AI is already in development and has surpassed the advanced VVC intra codec. To generate reconstructed images with the desired bits per pixel and assess the BD-rate performance of both the JPEG-AI verification model and VVC intra, bit rate matching is employed. However, the current state of the JPEG-AI verification model experiences significant slowdowns during bit rate matching, resulting in suboptimal performance due to an unsuitable model. The proposed methodology offers a gradual algorithmic optimization for matching bit rates, resulting in a fourfold acceleration and over 1% improvement in BD-rate at the base operation point. At the high operation point, the acceleration increases up to sixfold.
Panqi Jia, Ahmet Burakhan Koyuncu, Jue Mao, Ze Cui, Tiansheng Guo, Timofey Solovyev, Alexander Karabutov, Yin Zhao, Jing Wang 0194, Elena Alshina, André Kaup
PCS9
2024 Bit Distribution Study and Implementation of Spatial Quality Map in the JPEG-AI Standardization
abstract
Currently, there is a high demand for neural network-based image compression codecs. These codecs employ non-linear transforms to create compact bit representations and facilitate faster coding speeds on devices compared to the handcrafted transforms used in classical frameworks. The scientific and industrial communities are highly interested in these properties, leading to the standardization effort of JPEG-AI. The JPEG-AI verification model has been released and is currently under development for standardization. Utilizing neural networks, it can outperform the classic codec VVC intra by over 10% BD-rate operating at base operation point. Researchers attribute this success to the flexible bit distribution in the spatial domain, in contrast to VVC intra’s anchor that is generated with a constant quality point. However, our study reveals that VVC intra displays a more adaptable bit distribution structure through the implementation of various block sizes. As a result of our observations, we have proposed a spatial bit allocation method to optimize the JPEG-AI verification model’s bit distribution and enhance the visual quality. Furthermore, by applying the VVC bit distribution strategy, the objective performance of JPEG-AI verification mode can be further improved, resulting in a maximum gain of 0.45 dB in PSNR-Y.
Panqi Jia, Jue Mao, Esin Koyuncu, Ahmet Burakhan Koyuncu, Timofey Solovyev, Alexander Karabutov, Yin Zhao, Elena Alshina, André Kaup
VCIP7
2022 Enlarging the Long-time Dependencies via RL-based Memory Network in Movie Affective Analysis
abstract
Affective analysis of movies heavily depends on the causal understanding of the story with long-time dependencies. Limited by the existing sequence models such as LSTM, Transformer, etc., current works generally split the movies into dependent clips and predict the affective impacts (Valence/Arousal) independently, ignoring the long historical impacts across the clips. In this paper, we introduce a novel Reinforcement learning based Memory Net (RMN) for this task, which facilitates the prediction of the current clip to rely on the possible related historical clips of this movie. Compared with LSTM, the proposed method solves the long-time dependencies from two aspects. First, we introduce a readable and writable memory bank to store useful historical information, which solves the problem of the restricted memory unit for LSTM. However, the traditional parameters' update scheme of the memory network, when applied for long sequence prediction, still needs to store the gradients for long sequences. It suffers from gradient vanishing and exploding, similar to the issues of backpropagation through time (BPTT). For this problem, we introduce a reinforcement learning framework in the memory write operation. The memory updating scheme of the framework is optimized via one-step temporal difference, modeling the long-time dependencies using both the policy and value networks. Experiments on the LIRIS-ACCEDE dataset show that our method achieves significant performance gains over the existing methods. Besides, we also apply our method to other long sequence prediction tasks, such as music emotion recognition and video summarization, and also achieve state-of-the-art on those tasks.
Yin Zhao
ACM Multimedia2
2021 Pairwise Emotional Relationship Recognition in Drama Videos: Dataset and Benchmark
abstract
Recognizing the emotional state of people is a basic but challenging task in video understanding. In this paper, we propose a new task in this field, named Pairwise Emotional Relationship Recognition (PERR). This task aims to recognize the emotional relationship between the two interactive characters in a given video clip. It is different from the traditional emotion and social relation recognition task. Varieties of information, consisting of character appearance, behaviors, facial emotions, dialogues, background music as well as subtitles contribute differently to the final results, which makes the task more challenging but meaningful in developing more advanced multi-modal models. To facilitate the task, we develop a new dataset called Emotional RelAtionship of inTeractiOn (ERATO) based on dramas and movies. ERATO is a large-scale multi-modal dataset for PERR task, which has 31,182 video clips, lasting about 203 video hours. Different from the existing datasets, ERATO contains interaction-centric videos with multi-shots, varied video length, and multiple modalities including visual, audio and text. As a minor contribution, we propose a baseline model composed of Synchronous Modal-Temporal Attention (SMTA) unit to fuse the multi-modal information for the PERR task. In contrast to other prevailing attention mechanisms, our proposed SMTA can steadily improve the performance by about 1%. We expect the ERATO as well as our proposed SMTA to open up a new way for PERR task in video understanding and further improve the research of multi-modal fusion methodology.
Yin Zhao, Longjun Cai
ACM Multimedia2
2021 Reducing the Covariate Shift by Mirror Samples in Cross Domain Alignment
abstract
Eliminating the covariate shift cross domains is one of the common methods to deal with the issue of domain shift in visual unsupervised domain adaptation. However, current alignment methods, especially the prototype based or sample-level based methods neglect the structural properties of the underlying distribution and even break the condition of covariate shift. To relieve the limitations and conflicts, we introduce a novel concept named (virtual) mirror, which represents the equivalent sample in another domain. The equivalent sample pairs, named mirror pairs reflect the natural correspondence of the empirical distributions. Then a mirror loss, which aligns the mirror pairs cross domains, is constructed to enhance the alignment of the domains. The proposed method does not distort the internal structure of the underlying distribution. We also provide theoretical proof that the mirror samples and mirror loss have better asymptotic properties in reducing the domain shift. By applying the virtual mirror and mirror loss to the generic unsupervised domain adaptation model, we achieved consistently superior performance on several mainstream benchmarks.
Yin Zhao, Minquan Wang, Longjun Cai
NeurIPS1
2021 Transform Coding in the VVC Standard
abstract
In the past decade, the development of transform coding techniques has achieved significant progress and several advanced transform tools have been adopted in the new generation Versatile Video Coding (VVC) standard. In this paper, a brief history of transform coding development during VVC standardization is presented, and the transform coding tools in the VVC standard are described in detail together with their initial design, incremental improvements and implementation aspects. To improve coding efficiency, four new transform coding techniques are introduced in VVC, which are namely Multiple Transform Selection (MTS), Low-Frequency Non-separable Secondary Transform (LFNST) and Sub-Block Transform (SBT), as well as a large (64-point) type-2 DCT. The experimental results on VVC reference software (VTM-9.0) show that average 4.5% and 3.6% overall coding gain can be achieved by the VVC transform coding tools for All Intra and Random Access configurations, respectively.
Xin Zhao 0003, Seung-Hwan Kim 0001, Yin Zhao, Hilmi E. Egilmez, Moonmo Koo, Shan Liu 0001, Jani Lainema, Marta Karczewicz
IEEE Trans. Circuits Syst. Video Technol.3
2020 Multi-objective Optimization for Guaranteed Delivery in Video Service Platform
abstract
Guaranteed-Delivery (GD) is one of the important display strategies for the IP videos in video service platform. Different from the traditional recommendation strategy, GD requires the delivery system to guarantee the exposure amount (also called impressions in some works) for the content, where the amount generally comes from the purchase contract or business consideration of the platform. In this paper, we study the problem of how to maximize certain gains, such as video view (VV) or fairness of different contents (CTR variations between contents) under the GD constraints. We formulate such a problem as a constrained nonlinear programming problem, in which the objectives are to maximize the total VVs of contents and the exposure fairness between contents. In order to capture the trends of VV versus the impression number (page views, PV) for each video content, we propose a parameterized ordinary differential equation (ODE) model, and the parameters of the ODE are fitted by the video historical PV and CLICK datas. To solve the constrained nonlinear programming, we use genetic algorithm (GA) with a specific design of coding scheme considering the ODE constraints. The empirical study based on real-world data and online test on Youku.com verifies the effectiveness and superiority of our approach compared with the state of the art in the industry practice.
Yin Zhao, Longjun Cai
KDD2
2020 Video Codec Using Flexible Block Partitioning and Advanced Prediction, Transform and Loop Filtering Technologies
abstract
This paper describes a joint response to the Call for Proposals by Samsung, Huawei, GoPro, and HiSilicon on Video Compression with Capability beyond HEVC/H.265, jointly issued by ITU-T SG16 Q.6 (VCEG) and ISO/IEC JTC1/SC29/WG11 (MPEG). In the proposed codec, the coding framework supports hierarchical splitting with binary and ternary trees and flexible coding order representations. Additionally, novel compression tools on inter/intra prediction, in-loop filtering, and entropy coding have been proposed. The proposed compression scheme provides significantly higher compression capability than the state-of-the-art HEVC/H.265 standard for SDR (Standard Dynamic Range) category while maintaining complexity acceptable for emerging applications. When all the proposed algorithmic tools are used, the proposed video codec achieves approximately 40% bit-saving for the SDR cetegory on average compared to HEVC/H.265 anchor.
Kiho Choi, Jianle Chen, Haitao Yang 0001, Woongil Choi, Sergey Ikonin, Yinji Piao, Semih Esenlik, Minsoo Park, Ye-Kui Wang, Narae Choi, Yin Zhao, Seungsoo Jeong, Anish Tamse, Alexey Filippov, Heechul Yang, Junghye Min, Roman Chernyak, Bora Jin, Anand Meher Kotra, Sunil Lee, Han Gao 0001, Chanyul Kim, Timofey Solovyev, Kwangpyo Choi, Vasily Rufitskiy, Maxim Sychev, Jeonghoon Park
IEEE Trans. Circuits Syst. Video Technol.12
2019 Dense Encoder-Decoder Network based on Two-Level Context Enhanced Residual Attention Mechanism for Segmentation of Breast Tumors in Magnetic Resonance Imaging
abstract
Aiming to effective early detection of breast cancer, automatic tumor segmentation based on breast Magnetic Resonance Imaging (MRI) is concentrated by more and more researchers. This paper proposes a dense encoder-decoder network based on two-level context enhanced residual attention mechanism (TLCRAM-DED). With respect to TLCRAM-DED, we design the encoding structure combining two-level residual attention structure with dense block to extract and refine the features of different layers. Meanwhile, a dense multi-scale atrous convolution is used at the end of the encoder to obtain a larger receptive field and enrich the extracted semantic information. Moreover, residual attention structure (RAS) is also used for the refinement during decoding stage, while a long connection formed with the encoder RAS output is applied to supplement the features and to gradually recover the segmentation details. We validated prosed model in the DCE sequence of challenging breast cancer MRI dataset. The average Dice coefficient is up to 81.04%, which outperforms compared state-of-the-arts.
Ying Gao 0004, Yin Zhao, Xiong-Wen Luo 0001, Xiping Hu, Changhong Liang
BIBM2
2019 PD-GAN: Adversarial Learning for Personalized Diversity-Promoting Recommendation
abstract
This paper proposes Personalized Diversity-promoting GAN (PD-GAN), a novel recommendation model to generate diverse, yet relevant recommendations. Specifically, for each user, a generator recommends a set of diverse and relevant items by sequentially sampling from a personalized Determinantal Point Process (DPP) kernel matrix. This kernel matrix is constructed by two learnable components: the general co-occurrence of diverse items and the user's personal preference to items. To learn the first component, we propose a novel pairwise learning paradigm using training pairs, and each training pair consists of a set of diverse items and a set of similar items randomly sampled from the observed data of all users. The second component is learnt through adversarial training against a discriminator which strives to distinguish between recommended items and the ground-truth sets randomly sampled from the observed data of the target user. Experimental results show that PD-GAN is superior to generate recommendations that are both diverse and relevant.
Qiong Wu 0001, Yong Liu 0020, Chunyan Miao, Binqiang Zhao, Yin Zhao, Lu Guan
IJCAI5
2019 Latency Aware Adaptive Video Streaming using Ensemble Deep Reinforcement Learning
abstract
The development of live broadcasting represents many new technical challenges on adaptive bitrate(ABR) algorithms, which not only requires stable and high-quality transmission but also low end-to-end latency. Reinforcement learning(RL) achieves promising results and can learn ABR algorithms automatically without using any pre-programmed control rules. However, existing methods only consider bitrate control and ignore latency control. Therefore, in order to effectively reduce the end-to-end latency, we propose an independent latency limit model to control the frame skipping. Moreover, a model ensemble algorithm is implemented to reduce performance variance and improve the user quality of experience (QoE). Experimental results show that our model outperforms base- line methods and demonstrate the effectiveness of our model.
Yin Zhao, Qi-Wei Shen, Tong Xu 0013, Wei-Hua Niu, Si-Ran Xu
ACM Multimedia1
2019 RNE: A Scalable Network Embedding for Billion-Scale Recommendation
Jianbin Lin, Daixin Wang, Lu Guan, Yin Zhao, Binqiang Zhao, Jun Zhou 0011, Xiaolong Li 0005, Yuan Qi 0001
PAKDD (2)4
2018 Adaptive Weighted Bi-Prediction based on Template Similarity in Video Coding
abstract
In bi-prediction of merge/skip and advanced motion vector prediction(AMVP) modes, current block is predicted by averaging two uni-reference blocks. In this paper, statistical experiments show that averaged bi-prediction is not a good choice for merge/skip mode, compared with AMVP mode. And we propose an adaptive weighted bi-prediction method for merge/skip mode according to the similarities between current block and its two uni-predictors. The uni-predictor, which is more similar to current block, would be assigned with larger weighting factor. As pixel values of current block are unknown in decoder before conducting motion compensation, the similarity between current block and uni-reference block is estimated by similarity between their corresponding spatial neighboring pixels. The results show that the proposed method achieves 0.55% BD reduction on average over the reference software JEM 6.0.
Jue Mao, Yin Zhao, Lu Yu 0003
VCIP2
2016 Spatial quality index based rate perceptual-distortion optimization for video coding
Xingguo Zhu, Lu Yu 0003, Yin Zhao
J. Vis. Commun. Image Represent.5
2016 An analysis in metal barcode label design for reference
abstract
We employ nondestructive evaluation involving AC field measurement in detecting and identifying metal barcode labels, providing a reference for design. Using the magnetic scalar potential boundary condition at notches in thin-skin field theory and 2D Fourier transform, we introduce an analytical model for the magnetic scalar potential induced by the interaction of a high-frequency inducer with a metal barcode label containing multiple narrow saw-cut notches, and then calculate the magnetic field in the free space above the metal barcode label. With the simulations of the magnetic field, qualitative analysis is given for the effects on detecting and identifying metal barcode labels, which are caused by metal material, notch characteristics, exciting inducer properties, and other factors that can be used in metal barcode label design as reference. Simulation results are in good accordance with experiment results.
Yin Zhao, Hongguang Xu, Qinyu Zhang 0001
Frontiers Inf. Technol. Electron. Eng.1
2016 Control of Large-Scale Boolean Networks via Network Aggregation
abstract
A major challenge to solve problems in control of Boolean networks is that the computational cost increases exponentially when the number of nodes in the network increases. We consider the problem of controllability and stabilizability of Boolean control networks, address the increasing cost problem by partitioning the network graph into several subnetworks, and analyze the subnetworks separately. Easily verifiable necessary conditions for controllability and stabilizability are proposed for a general aggregation structure. For acyclic aggregation, we develop a sufficient condition for stabilizability. It dramatically reduces the computational complexity if the number of nodes in each block of the acyclic aggregation is small enough compared with the number of nodes in the entire Boolean network.
Yin Zhao, Bijoy K. Ghosh, Daizhan Cheng
IEEE Trans. Neural Networks Learn. Syst.1
2014 On controllability and stabilizability of probabilistic Boolean control networks
Yin Zhao, Daizhan Cheng
Sci. China Inf. Sci.1
2014 Block-Based In-Loop View Synthesis for 3-D Video Coding
abstract
View synthesis prediction (VSP) employs a synthesized picture as a reference picture for current-view texture coding, which is an advanced disparity-compensated prediction. However, the picture-based view synthesis demands huge complexity, especially for decoders. Therefore, we propose a block-based in-loop view synthesis scheme which generates VSP samples only for blocks using VSP modes (called target blocks). For a target block, a window in reference view is estimated. Then, pixels within the window are warped to the current view, producing VSP samples for the target block. The proposed method turns the picture-level VSP sample generation into macroblock-level process, and significantly reduces complexity of the VSP module while maintaining coding efficiency.
Yin Zhao, Lu Yu 0003
IEEE Signal Process. Lett.2
2013 Synthesized disparity vectors for 3D video coding
abstract
In 3D video coding, some dependent-view coding tools may utilize derived disparity vectors (DV) to locate inter-view correspondences, such as the Backward block-based View Synthesis Prediction (BVSP) and Depth-based Motion Vector Prediction (DMVP) in ATM. A typical way to derive a DV for a target block, as performed in ATM, is to convert reconstructed depth values associated with the block to a DV. However, this approach only works in depth-first coding order (DFCO) and becomes inapplicable when the dependent-view texture is coded prior to the depth, i.e., in texture-first coding order (TFCO). In this paper, a sparse DV field is synthesized from the depth map of a coded reference view to provide the DVs required by dependent-view coding tools in TFCO. The synthesized sparse DV field is accurate to support disparity-aided coding tools in different coding orders, while only introducing 2% decoding time increase. With the proposed method, BVSP and DMVP can be applied in both TFCO and DFCO, which mitigates the large coding performance gap (around 20% BD rate) between TFCO and DFCO of ATM.
Yin Zhao, Lu Yu 0003
ICIP1
2013 Subjective study of binocular rivalry in stereoscopic images with transmission and compression artifacts
abstract
Binocular rivalry is a visual phenomenon that occurs when two eyes are presented with different patterns. Instead of being fused to form a unitary visual impression, the two patterns may be perceived alternately or evoke a peculiar shimmering effect. In stereoscopic images, corresponding regions in two views may be dissimilar due to artifacts from lossy video processing, e.g., transmission and compression. In this case, binocular rivalry may be introduced, which impairs 3D visual quality. In this paper, we established a database of binocular rivalry artifacts in stereoscopic images with transmission and compression distortions. Ten subjects were engaged to mark binocular rivalry artifacts in ten stereoscopic images. The performance of each subject was analyzed using the hit rate and false alarm rate of the subject compared with the average marking results of the other subjects. The subjective data were then combined into maps which indicate the locations and strength of the rivalry artifacts in the stereo pairs. The database is aimed to facilitate the validation of binocular rivalry artifact detection algorithms, and provides cues for stereoscopic 3D quality assessment and binocular rivalry artifact removal.
Yin Zhao, Lu Yu 0003
ICIP1
2013 System identification for output-dependent bounded noises and its application in learning personalized thermal comfort model
abstract
When the output observation noise is output-dependent, identifying the unknown system parameters becomes challenging. Traditional methods based on Mean Square Error, even the ones with corrections still have biased estimations in this case. Many practical cases such as bounded sensor, uncertainty of expression in human involved system identification, and even in physiological or biological model identification actually have this problem. In this paper, some algorithms were proposed to obtain the unbiased estimation of parameters for input-output-nonlinear but identification-linear system under output-dependent bounded noise. We utilized the truncated probability distribution to model the noise and gave the unbiased estimation algorithms of the system parameters as well as noise parameter if unknown. Asymptotic properties of the algorithms indicate that the algorithms converge to the true parameters. Besides illustrative numerical example, we also utilized the algorithm in a real world application to identify the personalized thermal comfort model using human noisy voting data. Results revealed the effectiveness and applicability of the proposed algorithms.
Yin Zhao, Qianchuan Zhao
ICRA1
2012 Game-based control systems: A semi-tensor product formulation
abstract
A class of control systems, which are emerged from dynamic games, are considered. Using semi-tensor product of matrices, the set of strategies can be described as a set of matrices. Then the dynamics of such systems can be converted from logical type dynamics into standard discrete-time dynamic systems. Hence, the classical techniques for control systems are applicable to such systems. Semi-tensor product formulation of such systems is investigated. Some related optimal control problems are also investigated.
Daizhan Cheng, Yin Zhao
ICARCV2
2011 Cross-view post-filtering for fidelity enhancement on asymmetric coding of 3D video
abstract
3D video employing depth-image-based rendering (DIBR) typically contains a stereo pair plus two associated depth maps. The stereo pair may be compressed asymmetrically (e.g., using mixed resolution coding) to effectively reduce the bit rate while maintaining the stereoscopic visual quality at the same level as that from symmetric coding. With depth information, it is straightforward to locate high-fidelity (HF) inter-view correspondences of pixels in the low-fidelity (LF) view. However, we find that LF-view fidelity cannot be consistently improved by simply substituting LF pixels with their HF counterparts, due to inter-view color/geometric differences and depth errors. In this paper, we propose an effective post-filter which first checks the coherence between local LF and HF waveforms to distinguish unreliable correspondences. Then, the LF view is adaptively rectified by the reliable HF pixels to improve its fidelity. Experimental results show that the proposed cross-view fidelity enhancement scheme can promote Peak Signal-to-Noise Ratio (PSNR) of LF views by up to 1.1 dB, and can effectively suppress ringing and blocking artifacts in flat regions of LF views.
Yin Zhao, Lu Yu 0003
VCIP1
2011 Binocular Just-Noticeable-Difference Model for Stereoscopic Images
abstract
Conventional 2-D Just-Noticeable-Difference (JND) models measure the perceptible distortion of visual signal based on monocular vision properties by presenting a single image for both eyes. However, they are not applicable for stereoscopic displays in which a pair of stereoscopic images is presented to a viewer's left and right eyes, respectively. Some unique binocular vision properties, e.g., binocular combination and rivalry, need to be considered in the development of a JND model for stereoscopic images. In this letter, we propose a binocular JND (BJND) model based on psychophysical experiments which are conducted to model the basic binocular vision properties in response to asymmetric noises in a pair of stereoscopic images. The first experiment exploits the joint visibility thresholds according to the luminance masking effect and the binocular combination of noises. The second experiment examines the reduction of visual sensitivity in binocular vision due to the contrast masking effect. Based on these experiments, the developed BJND model measures the perceptible distortion of binocular vision for stereoscopic images. Subjective evaluations on stereoscopic images validate of the proposed BJND model.
Yin Zhao, Ce Zhu, Yap-Peng Tan, Lu Yu 0003
IEEE Signal Process. Lett.1
2011 Video Quality Assessment Based on Measuring Perceptual Noise From Spatial and Temporal Perspectives
abstract
Video quality assessment (VQA) exploits important properties of the sophisticated human visual system (HVS). In this paper, we study a series of fundamental HVS characteristics for subjective video quality assessment, and incorporate them into a systematic framework to simulate subjective evaluation on impaired videos. Based on this framework, we develop a novel full-reference metric, namely, perceptual quality index (PQI). Specifically, the proposed PQI metric comprises four major modules: 1) visual performance equation for the foveal and extra-foveal vision based on the cortical magnification theory; 2) perceptible noise detection using a spatial-temporal just noticeable difference model, and its quantification in both spatial and temporal channels, considering the varying error sensitivity due to the contrast and motion masking effects; 3) instantaneous error summation with inhibition of weak local distortions, and quality degradation accumulation over time that models the visual persistence and recency effect; and 4) fusion of the spatial and temporal noise intensities into a perceptual quality index. Compared with some state-of-the-art VQA models, the PQI metric, which exploits multiple visual properties, measures video quality more accurately and reliably on two VQA databases.
Yin Zhao, Lu Yu 0003, Ce Zhu
IEEE Trans. Circuits Syst. Video Technol.1
2011 Depth No-Synthesis-Error Model for View Synthesis in 3-D Video
abstract
Currently, 3-D Video targets at the application of disparity-adjustable stereoscopic video, where view synthesis based on depth-image-based rendering (DIBR) is employed to generate virtual views. Distortions in depth information may introduce geometry changes or occlusion variations in the synthesized views. In practice, depth information is stored in 8-bit grayscale format, whereas the disparity range for a visually comfortable stereo pair is usually much less than 256 levels. Thus, several depth levels may correspond to the same integer (or sub-pixel) disparity value in the DIBR-based view synthesis such that some depth distortions may not result in geometry changes in the synthesized view. From this observation, we develop a depth no-synthesis-error (D-NOSE) model to examine the allowable depth distortions in rendering a virtual view without introducing any geometry changes. We further show that the depth distortions prescribed by the proposed D-NOSE profile also do not compromise the occlusion order in view synthesis. Therefore, a virtual view can be synthesized losslessly if depth distortions follow the D-NOSE specified thresholds. Our simulations validate the proposed D-NOSE model in lossless view synthesis and demonstrate the gain with the model in depth coding.
Yin Zhao, Ce Zhu, Lu Yu 0003
IEEE Trans. Image Process.1
2010 Evaluating video quality with temporal noise
abstract
Human can only perceive video distortions beyond certain intensity. Temporal variations of suprathreshold noise in a video induce annoying temporal noise (e.g., flickering artifacts). Prior studies on full-reference video quality assessment (VQA) focused mainly on spatial quality of impaired videos, and ignored or underestimated the impact of temporal noise on visual quality degradation. According to some characteristics of human visual system, we consider temporal noise as the most salient distortion in impaired videos and it can greatly influence the perceived video quality. Thus, we propose a novel and simple full-reference metric that evaluates quality of impaired videos by measuring the energy of temporal noise in them. This metric shows competitive performance with some state-of-the-art objective models on the LIVE VQA database.
Yin Zhao, Lu Yu 0003
ICME1
2010 Temporal consistency enhancement on depth sequences
abstract
Currently, depth sequences generated by automatic depth estimation suffer from the temporal inconsistency problem. Estimated depth values of some objects vary in adjacent frames, whereas the objects actually remain on the same depth planes. These temporal depth errors significantly impair the visual quality of the synthesized virtual view as well as the coding efficiency of the depth sequences. Since depth sequences correspond to texture sequences, some erroneous temporal depth variations can be detected by analyzing temporal variations of the texture sequences. Utilizing this property, we propose a novel solution to enhance the temporal consistency of depth sequences by applying adaptive temporal filtering on them. Experiments demonstrate that the proposed depth filtering algorithm can effectively suppress transient depth errors and generate more stable depth sequences, resulting in notable temporal quality improvement of the synthesized views and higher coding efficiency on the depth sequences.
Deliang Fu, Yin Zhao, Lu Yu 0003
PCS2
2010 Suppressing texture-depth misalignment for boundary noise removal in view synthesis
abstract
During view synthesis based on depth maps, also known as Depth-Image-Based Rendering (DIBR), annoying artifacts are often generated around foreground objects, yielding the visual effects that slim silhouettes of foreground objects are scattered into the background. The artifacts are referred as the boundary noises. We investigate the cause of boundary noises, and find out that they result from the misalignment between texture and depth information along object boundaries. Accordingly, we propose a novel solution to remove such boundary noises by applying restrictions during forward warping on the pixels within the texture-depth misalignment regions. Experiments show this algorithm can effectively eliminate most boundary noises and it is also robust for view synthesis with compressed depth and texture information.
Yin Zhao, Dong Tian, Ce Zhu, Lu Yu 0003
PCS1
2010 A perceptual metric for evaluating quality of synthesized sequences in 3DV system
abstract
3D Video system based on Depth-Image-Based Rendering relies on high quality depth data. Errors distributed randomly in depth map sequences induce annoying temporal noise, such as flickering and object shifting. Prior studies on video quality assessment focused mainly on spatial quality of the tested sequence and often ignored its temporal performance. In synthesized sequences, a large number of tiny geometric distortions and illumination differences are temporally constant and perceptually invisible. The dynamic noise impairs subjective quality of the sequences more greatly than the static spatial noise. Temporal quality plays a dominant role in overall quality assessment on the synthesized sequences with temporal instability problem. We propose a simple full-reference metric, Peak Signal to Perceptible Temporal Noise Ratio, to evaluate quality of synthesized sequences by measuring the perceptible temporal noise in them.
Yin Zhao, Lu Yu 0003
VCIP1
1997 Mortgage data mining
abstract
The paper reports a preliminary investigation of the use of of modern data mining tools for mortgage scoring. Using IBM's Intelligent Miner (a data mining toolbox), the authors built a model of serious delinquency on a sample of data from Mortgage Information Corporation's Loan Performance System, which contains over 20 million loans with a volume of over $1.6 trillion. Currently, two technologies prevail in mortgage scoring: logistic regression, a very old and very simple method, and neural networks, newer and more complex types of models that can be extremely difficult to interpret. The radial basis function (RBF) algorithm in Intelligent Miner combines the mathematical complexity and generality of neural networks with a comprehensible visualization that explains the RBF model. Due to the performance and understandability of the RBF model, as well as other unique technologies not described, the Intelligent Miner should be a useful tool for mortgage bankers, facilitating development of customized systems for mortgage scoring and other mortgage banking applications.
George H. John, Yin Zhao
CIFEr2