EDBT 2026 Demo / reviewers in the wild / expert
Yan Zhao 0012
dblp:88/5320-12
· DBLP profile ↗
29ranked-venue papers
0as first author
22since 2021 · last 2027
0000-0002-5319-0202ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 15 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | MusicIllustrator: A multimodal music-to-image generation approach based on emotion analysis
Wenting Zhao 0003, Qiang Zhao 0005, Yan Zhao 0012 |
Expert Syst. Appl. | 4 |
| 2026 | A Brain-Inspired Saliency Prediction Framework for Human-AI Cognitive Consistency in AIGC Content via Multi-Region Liquid NeuronsabstractIn recent years, human-AI cognitive consistency has emerged as a crucial perspective for evaluating the perceptual quality and interpretability of AIGC (Artificial Intelligence Generated Content). This paper proposes a biologically inspired saliency prediction framework that models six core regions of the human visual system—namely V1, V2, V4, MT, LIP, and FEF—using liquid neurons to capture the dynamic saliency features aligned with human gaze behavior. To enable effective alignment between AIGC models and human cognitive mechanisms, we introduce a cross-domain dual-teacher distillation strategy and construct a large-scale multimodal dataset comprising natural images, eye-tracking data, AIGC-generated images, and their corresponding cross-attention maps. Furthermore, we propose HAMCI (Human-AI Mutual Cognitive Index), a novel metric designed to quantitatively assess the spatial and semantic alignment between predicted saliency maps and model attention distributions. The proposed method demonstrates promising performance across various saliency prediction and cognitive alignment tasks, with results comparable to or surpassing recent state-of-the-art methods in several benchmarks. The code and dataset will be released upon acceptance to facilitate future research on cognitively aligned AIGC evaluation. Yan Zhao 0012, Shigang Wang 0003 |
AAAI | 2 |
| 2026 | A trajectory-based framework for diagnosing and calibrating social order in large language models
Yan Zhao 0012, Peitong Han, Guangtao Zhai |
Expert Syst. Appl. | 2 |
| 2026 | Enhancing skin lesion segmentation via martingale feature fusion and adaptive deep semantic modeling
Yan Zhao 0012, Shigang Wang 0003 |
Multim. Syst. | 2 |
| 2026 | Multimodal behavioral analysis for autism spectrum disorder assessment
Yunxiu Zhao, Shigang Wang 0003, Feiyong Jia, Honghua Li, Yan Zhao 0012 |
Pattern Recognit. | 7 |
| 2025 | SIE: infrared and visible image fusion based on scene information embedding
Yingnan Geng, Weixuan Diao, Yan Zhao 0012 |
Multim. Tools Appl. | 3 |
| 2024 | Fractional Order Spectrum in SAR Image RegistrationabstractSAR image registration is an important processing procedure for change detection and target recognition. However, the registration performance is seriously influenced by Symmetric α Stable (SαS) noises in SAR images. In order to cancel the impact of SαS noises in SAR image applications, a new concept of Fractional Order Spectrum of Cumulant (FOS-C) and SAR image registration based on (FOS-C) are proposed for the first time in this paper. In the proposed method, the images are registered from coarse to precise in three steps. First, the coarse registration based on Fourier Transform is used to estimate the scaling and rotation differences between images. Second, the registration based on FOS-C is designed, and used to provide the rough position of the similar regions. Third, the normalization cross-correlation (NCC) algorithm based on FOS-C is used to achieve the fine registration. Experimental results show that our method outperforms SAR-SIFT and KAZE-SAR. Yan Zhao 0012, Xinbo Li, Shigang Wang 0003 |
ICME | 2 |
| 2024 | TransDiff: medical image segmentation method based on Swin Transformer with diffusion probabilistic model
Yan Zhao 0012, Shigang Wang 0003 |
Appl. Intell. | 2 |
| 2024 | A study on attention-based fine-grained image recognition: Towards musical instrument performing hand shape assessmentabstractAutomatic identification and professional evaluation makes musical instrument learning more intelligent. Since a proper hand shape is the basis of fingerings in playing instruments, this paper explores an integration of intelligent recognition technique into hand shape assessment of instrument players in an attempt of taking Chinese zither (Zheng) as an example. The fine-grained image recognition is novelly applied to automatically assessing basic hand shapes, as a tentative exploration of interdisciplinary research. First, this paper formulates an assessment scales by combining fine-grained image features with hand shape evaluation indicators in musical instrument learning. Then, an image dataset for hand shapes of Chinese zither performance (CZ-Dataset V2) is established based on free multi-view acquisition. Finally, we propose a fine-grained hand shape image recognition method using attention mechanism . Experimental results show that the basic instrumental hand shapes can be effectively recognized and reasonable suggestions for hand shape assessment can be provided. Wenting Zhao 0003, Shigang Wang 0003, Yan Zhao 0012, Yecheng Liang, Jiehua Lin |
Eng. Appl. Artif. Intell. | 3 |
| 2024 | FewarNet: An Efficient Few-Shot View Synthesis Network Based on Trend RegularizationabstractNovel view synthesis from existing inputs remains a research focus in computer vision. Predicting views becomes more challenging when only a limited number of views are available. This challenge is commonly referred to as the few-shot view synthesis problem. Recently, various strategies have emerged for few-shot view synthesis, such as transfer learning, depth supervision, and regularization constraints. However, transfer learning relies on massive scene data, depth supervision is affected by input depth quality, and regularization causes increased computational costs or impaired generalization. To address these issues, we propose a new few-shot view synthesis framework called FewarNet that introduces trend regularization to leverage depth structural features and a warping loss to supervise depth estimation, possessing the advantages of existing few-shot strategies, enabling high-quality novel view prediction with generalization and efficiency. Specifically, FewarNet consists of three stages: fusion, warping, and rectification. In the fusion stage, a fusion network is introduced to estimate depths using scene priors from coarse depths. In the warping stage, the predicted depths are used to guide the warping of the input views, and a distance-weighted warping loss is proposed to correctly guide depth estimation. To further improve prediction accuracy, we propose trend regularization which imposes penalties on depth variation trends to provide depth structural constraints. In the rectification stage, a rectification network is introduced to refine occluded regions in each warped view to generate novel views. Additionally, a rapid view synthesis strategy that leverages depth interpolation is designed to improve efficiency. We validate the method’s effectiveness and generalization on various datasets. Given the same sparse inputs, our method demonstrates superior performance in quality and efficiency over state-of-the-art few-shot view synthesis methods. Chenxi Song, Shigang Wang 0003, Yan Zhao 0012 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Fractional Order Spectrum of Cumulant in SAR Image RegistrationabstractFractional order cumulant (FOC) is a new tool that can suppress symmetric$\alpha $stable (S$\alpha $S) noises in signal processing. Although FOC has obvious advantages for its noise suppression ability, some weaknesses still limit its application in synthetic aperture radar (SAR) image processing. The main challenge is that FOC can only suppress S$\alpha $S noises when all angular frequencies are zeros. In order to provide a statistic with noise robustness for SAR image processing, we propose a new concept of fractional order spectrum of cumulant (FOSC). FOSC can cancel the impact of S$\alpha $S noises without the limitation of angular frequencies in theory, which means FOSC can represent the local features more accurately. Furthermore, a FOSC-based SAR image registration is proposed to verify the advantage of FOSC. First, the 2-D formats of FOSC with different angular frequencies are calculated to provide more image features. Second, a multi-frequency pyramid array is designed to utilize the additional information in FOSC images, which can be used to detect more accurate keypoints. Third, the local descriptors based on FOSC are designed, which are constructed by accumulating orientation histograms of gradients based on FOSC and re-arranging elements in feature vectors regularly based on the orientation assignment. Finally, rotation consistencies are designed and used to eliminate the mismatching points after Euclidean distance-based matching of feature vectors. The proposed method is compared with four state-of-the-art methods. Experimental results show that the proposed method has achieved an impressive SAR image registration with 16.47%–48.37% gains when evaluating using the similarity of stitched areas. Yan Zhao 0012, Xinbo Li, Shigang Wang 0003 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | 5-D Epanechnikov Mixture-of-Experts in Light Field Image CompressionabstractIn this study, we propose a modeling-based compression approach for dense/lenslet light field images captured by Plenoptic 2.0 with square microlenses. This method employs the 5-D Epanechnikov Kernel (5-D EK) and its associated theories. Owing to the limitations of modeling larger image block using the Epanechnikov Mixture Regression (EMR), a 5-D Epanechnikov Mixture-of-Experts using Gaussian Initialization (5-D EMoE-GI) is proposed. This approach outperforms 5-D Gaussian Mixture Regression (5-D GMR). The modeling aspect of our coding framework utilizes the entire EI and the 5D Adaptive Model Selection (5-D AMLS) algorithm. The experimental results demonstrate that the decoded rendered images produced by our method are perceptually superior, outperforming High Efficiency Video Coding (HEVC) and JPEG 2000 at a bit depth below 0.06bpp. Boning Liu 0001, Yan Zhao 0012, Xiaomeng Jiang, Xingguang Ji, Shigang Wang 0003, Yebin Liu |
IEEE Trans. Image Process. | 2 |
| 2023 | A Novel Intelligent Assessment Based on Audio-Visual Data for Chinese Zither Fingerings
Wenting Zhao 0003, Shigang Wang 0003, Yan Zhao 0012, Tianshu Li |
ICIG (4) | 3 |
| 2023 | YOLO-DA: An Efficient YOLO-Based Detector for Remote Sensing Object DetectionabstractIn the past few decades, many efficient object detectors have been proposed for natural scene image object detection. However, due to the complex scenes and high interclass similarity of optical remote sensing (RS) images, applying these detectors to optical RS images directly is not very effective. Most of the recent detectors pursue higher accuracy while ignoring the balance between detection accuracy and speed, which hinders the practical application of these detectors, especially in embedded devices. To meet these challenges, a fast and accurate detector based on YOLO (You Only Look Once) with decoupled attention head (YOLO-DA) is proposed, which effectively improves detection performance while only introducing minimal complexity. Specifically, an attention module at the end of the detector is designed for guiding a neural network to extract more efficient features from the complex background while also minimizing the amount of additional computation. Moreover, a lightweight decoupled detection head with enhanced classification and localization capability is developed to detect objects with high interclass similarity. In the experiments, the proposed method effectively solves the problem of high interclass similarity and improves the mAP by 6.8% on the fine-grained optical RS dataset SIMD, compared with YOLOv5-L. In addition, the proposed method improves the mAP by 1.0%, 1.7% and 0.6% on the other three publicly open optical RS datasets, respectively. Experimental results on detection accuracy and inference time demonstrate that our method achieves the best trade-off between detection performance and speed. Jiehua Lin, Yan Zhao 0012, Shigang Wang 0003 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2023 | 3D Holoscopic Image Compression Based on Gaussian Mixture ModelabstractWe introduce a Gaussian Mixture Model (GMM) framework for 3D holoscopic image compression in this paper. The elemental-images of the 3D holoscopic image are predicted using GMM and the parameters of GMM are estimated using the common Expectation-Maximization (EM) algorithm. GMM Model Optimization (GMO) is used in this framework to select the optimal number of distributions and avoid local optimum of EM at the same time. A three-dimensional distribution-rotation based decomposition is proposed to change covariance parameters to meaningful features and improve the coding efficiency. The features and the remaining parameters of the GMM are encoded using fixed-length bits. A feature-based dictionary is proposed in this framework to match the similar Gaussian distributions utilizing the similar GMM features. And the offsets of the matched distributions are recorded as motion vectors to replace the similar areas in the elemental-images of the 3D holoscopic image. The residual between the original image and the prediction is encoded using Screen Content Coding Extension of High Efficiency Video Coding (HEVC-SCC). Experimental results show that our method performs better than HEVC-SCC, two coding methods based on pseudo-sequences and a state-of-the-art content-based compression method with Gaussian process regression. Yan Zhao 0012, Shigang Wang 0003 |
IEEE Trans. Multim. | 2 |
| 2022 | An improved augmented-reality method of inserting virtual objects into the scene with transparent objectsabstractIn augmented reality, the insertion of virtual objects into the real scene needs to meet the requirements of visual consistency. The virtual objects rendered by the augmented reality system should be consistent with the illumination of the real scene. However, for complex scenes, it is not enough to just complete the illumination estimation. When there are transparent objects in the real scene, the difference in refractive index and roughness of transparent objects will influence the effect of the virtual and real fusion. To tackle this problem, this paper proposes a new approach to jointly estimate the illumination and transparent material for inserting virtual objects into the real scene. We solve for the material parameters of objects and illumination simultaneously by nesting microfacet model and hemispherical area illumination model into inverse path tracing. Although there is no geometry model of light sources in the recovered geometry model, the proposed hemispherical area illumination model can be used to recover scene appearance. Multiple experiments on both virtual and real-world datasets verify that the proposed approach subjectively and objectively performs better than the state-of-the-art method. Yan Zhao 0012, Shigang Wang 0003 |
VR | 2 |
| 2022 | 4D Epanechnikov Mixture Regression in LF Image CompressionabstractWith the emergence of light field imaging in recent years, the compression of its elementary image array (EIA) has become a significant problem. Our coding framework includes modeling and reconstruction. For the modeling, the covariance-matrix form of the 4D Epanechnikov kernel (4D EK) and its correlated statistics were deduced to obtain the 4-D Epanechnikov mixture models (4-D EMMs). A 4D Epanechnikov mixture regression (4D EMR) was proposed based on this 4D EK, and a 4D adaptive model selection (4D AMLS) algorithm was designed to realize the optimal modeling for a pseudo video sequence (PVS) of the extracted key-EIA. A linear function based reconstruction (LFBR) was proposed based on the correlation between adjacent elementary images (EIs). The decoded images realized a clear outline reconstruction and superior coding efficiency compared to high-efficiency video coding (HEVC) and JPEG 2000 below approximately 0.05 bpp. This work realized an unprecedented theoretical application by (1) proposing the 4D Epanechnikov kernel theory, (2) exploiting the 4D Epanechnikov mixture regression and its application in the modeling of the pseudo video sequence of light field images, (3) using 4D adaptive model selection for the optimal number of models, and (4) employing a linear function-based reconstruction according to the content similarity. Boning Liu 0001, Yan Zhao 0012, Xiaomeng Jiang, Shigang Wang 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Aerial Image Object Detection Based on Superpixel-Related Patch
Jiehua Lin, Yan Zhao 0012, Shigang Wang 0003, Meimei Chen, Hongbo Lin, Zhihong Qian |
ICIG (1) | 2 |
| 2021 | 3-D Epanechnikov Mixture Regression in integral imaging compression
Boning Liu 0001, Yan Zhao 0012, Xiaomeng Jiang, Shigang Wang 0003 |
J. Vis. Commun. Image Represent. | 2 |
| 2021 | Three-dimensional Epanechnikov mixture regression in image codingabstractKernel methods have been studied extensively in recent years. We propose a three-dimensional (3-D) Epanechnikov Mixture Regression (EMR) based on our Epanechnikov Kernel (EK) and realize a complete framework for image coding. In our research, we deduce the covariance-matrix form of 3-D Epanechnikov kernels and their correlated statistics to obtain the Epanechnikov mixture models. To apply our theories to image coding, we propose the 3-D EMR which can better model an image in smaller blocks compared with the conventional Gaussian Mixture Regression (GMR). The regressions are all based on our improved Expectation-Maximization (EM) algorithm with mean square error optimization. Finally, we design an Adaptive Mode Selection (AMS) algorithm to realize the best model pattern combination for coding. Our recovered image has clear outlines and superior coding efficiency compared to JPEG below 0.25bpp. Our work realizes an unprecedented theory application by: (1) enriching the theory of Epanechnikov kernel, (2) improving the EM algorithm using MSE optimization, (3) exploiting the EMR and its application in image coding, and (4) AMS optimal modeling combined with Gaussian and Epanechnikov kernel. Boning Liu 0001, Yan Zhao 0012, Xiaomeng Jiang, Shigang Wang 0003 |
Signal Process. | 2 |
| 2021 | Image compression based on Gaussian mixture model constrained using Markov random fieldabstractWe introduce a Gaussian Mixture Model (GMM) constrained by Markov Random Field (MRF) framework for image compression in this paper. The image is predicted using GMM with MRF and the parameters of the GMM are estimated using an adjusted Expectation-Maximization (EM) algorithm. Mixture Model Optimization (MMO) is used in this framework to select the optimal number of distributions and avoid local optimum of EM at the same time. Parameters are encoded using fixed-length bits. A codebook is used to improve the coding efficiency of the covariance parameters. The residual between the original image and the prediction is encoded using High Efficiency Video Coding (HEVC) intra coding. Experimental results show that our method performs better than our previous work, HEVC, JPEG 2000 and Better Portable Graphics (BPG) which is an improved version of HEVC. Yan Zhao 0012, Shigang Wang 0003 |
Signal Process. | 2 |
| 2021 | An Improved Augmented-Reality Framework for Differential Rendering Beyond the Lambertian-World AssumptionabstractIn augmented reality, it is important to achieve visual consistency between inserted virtual objects and the real scene. As specular and transparent objects can produce caustics, which affect the appearance of inserted virtual objects, we herein propose a framework for differential rendering beyond the Lambertian-world assumption. Our key idea is to jointly optimize illumination and parameters of specular and transparent objects. To estimate the parameters of transparent objects efficiently, the psychophysical scaling method is introduced while considering visual characteristics of the human eye to obtain the step size for estimating the refractive index. We verify our technique on multiple real scenes, and the experimental results show that the fusion effects are visually consistent. Yan Zhao 0012, Shigang Wang 0003 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2019 | An Image Coding Approach Based on Mixture-of-experts Regression Using Epanechnikov KernelabstractIn this paper, we propose an optimal modeling framework for image compression using EMM (Epanechnikov Mixture Model). Epanechnikov Kernel and its correlated statistics are basement of our Epanechnikov Mixture Regression (EMR). In our scheme, the stochastic processes of the pixel values are modelled as an EMM with K experts in three-dimensional space and then we use EMR to search for the optimal solution, whose parameters are determined through EM (Expectation-Maximization) algorithm. In the process of regression, the conditional density is the regression kernel function. Experimental results show that the proposed scheme is effective especially for the image with complex texture without consuming extra bits compared to Gaussian Mixture Regression (GMR). Boning Liu 0001, Yan Zhao 0012, Xiaomeng Jiang, Shigang Wang 0003 |
ICASSP | 2 |
| 2019 | Image Compression Using GMM Model OptimizationabstractA Gaussian Mixture Model (GMM)-based framework for image compression is proposed in this paper. The image is predicted using GMM whose parameters are estimated using common Expectation-Maximization (EM) algorithm and encoded with fixed length bits. We introduce a new GMM Model Optimization (GMO) measure to select the optimal number of models and avoid local optimum of EM at the same time. The encoding cost of the residual and parameters are considered in GMO which is demonstrated to be near concave and effective. A parameter dictionary is designed to utilize the correlation of the parameters to improve the coding efficiency. The residual between the original image and the GMM image is encoded using High Efficiency Video Coding (HEVC) intra coding. Experimental results show that our method performs better than HEVC. Yan Zhao 0012, Shigang Wang 0003 |
ICASSP | 2 |
| 2019 | Video-based, Occlusion-robust Multi-view Stereo Using Inner-boundary Depths of Textureless AreasabstractOcclusions and poor textures are two main problems in multi-view stereo reconstruction. This paper presents a video-based solution to address both challenges in depth estimation. We focus on reconstructing accurate inner boundaries of visible textureless areas, particularly for occluded background, by leveraging the reliable depths of object edges. This is done by efficiently respecting two local cues with complementary advantages, i.e. smoothness and density of recovered surfaces. The inner-boundary depths are finally utilized to infer dense geometry without wrong connections between objects. This method only relies on low-level techniques, e.g. intra-view interpolation and inter-view propagation of depths. Experiments indicate its superiority in terms of both depth discontinuities near object silhouettes and surface smoothness in homogeneous regions compared to the state of the art. Shigang Wang 0003, Yan Zhao 0012 |
ICASSP | 3 |
| 2019 | Illumination estimation for augmented reality based on a global illumination model
Yan Zhao 0012, Shigang Wang 0003 |
Multim. Tools Appl. | 2 |
| 2017 | A Feature-Based Coding Algorithm for Face Image
Henan Li, Shigang Wang 0003, Yan Zhao 0012, Chuxi Yang, Aobo Wang |
ICIG (2) | 3 |
| 2017 | Feature-Based Facial Image Coding Method Using Wavelet Transform
Chuxi Yang, Yan Zhao 0012, Shigang Wang 0003 |
ICIG (2) | 2 |
| 2011 | Spatial error concealment for stereoscopic video coding based on pixel matching
Yan Zhao 0012, Shigang Wang 0003, Hexin Chen |
J. Supercomput. | 2 |