VLDB 2026 Research / reviewers in the wild / expert
Li Yang 0014
dblp:09/3925-14
· DBLP profile ↗
20ranked-venue papers
7as first author
19since 2021 · last 2026
0000-0002-1889-3113ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 5 first-author · 13 since 2021Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MSFI: Multi-timescale spatio-temporal features integration in spiking neural networks
Dengfeng Xue, Chunfeng Yuan, Man Yao, Wei Liu 0153, Li Yang 0014, Bing Li 0001, Weiming Hu 0004, Haoliang Sun, Zhetao Li |
Neural Networks | 7 |
| 2025 | Towards More Discriminative Feature Learning in SNNs with Temporal-Self-Erasing SupervisionabstractSpiking Neural Networks (SNNs) are biologically inspired models that process visual inputs over multiple time steps. However, they often struggle with limited feature discrimination along the temporal dimension due to inherent spatiotemporal invariance. This limitation arises from the redundant activation of certain regions and shared supervision for multiple time steps, constraining the network’s ability to adapt and learn diverse features. To address this challenge, we propose a novel Temporal-Self-Erasing (TSE) supervision method that dynamically adapts the learning regions of interest for different time steps. The TSE method operates by identifying highly activated regions from predictions across multiple time steps and adaptively suppressing them during model training, thereby encouraging the network to focus on less activated yet potentially informative regions. This approach not only enhances the feature discrimination capability of SNNs but also facilitates more effective multi-time-step inference by exploiting more semantic information. Experimental results on benchmark datasets demonstrate that our TSE method significantly improves the classification accuracy and robustness of SNNs. Wei Liu 0153, Li Yang 0014, Mingxuan Zhao, Dengfeng Xue, Shuxun Wang, Boyu Cai, Bing Li 0001, Weiming Hu 0004 |
AAAI | 2 |
| 2025 | Deformable Spherical Geometry Transformer For Panoramic Semantic SegmentationabstractThe increasing availability of 360° images has created a demand for effective Panoramic Semantic Segmentation (PASS) to enable comprehensive scene understanding. However, the spherical nature of 360° image introduces significant spatial distortions due to Equirectangular Projection (ERP), making it challenging for traditional 2D methods, which are designed for Euclidean spaces. Existing PASS methods typically mitigate these distortions through developing spherical-to-tangent polyhedron transformations or special-ized convolutional structures. Nevertheless, these approaches still struggle to preserve the spherical geometry and fail to adequately capture the semantic context of 360° images. In this paper, we propose a Deformable Spherical Geometry Transformer (DSGT) network that adapts to spherical distortions through a local-global self-attention mechanism. The local self-attention module captures local semantic information to alleviate distortions, while the global self-attention module integrates spherical geometric priors to enhance predictions. Experimental results on the Stanford2D3D panoramic dataset demonstrate that DSGT outperforms state-of-the-art PASS methods. Boyang Lan, Li Yang 0014, Mai Xu, Lai Jiang 0004, Yufeng Wang 0004 |
ICIP | 2 |
| 2025 | DeepTAGE: Deep Temporal-Aligned Gradient Enhancement for Optimizing Spiking Neural NetworksabstractSpiking Neural Networks (SNNs), with their biologically inspired spatio-temporal dynamics and spike-driven processing, are emerging as a promising low-power alternative to traditional Artificial Neural Networks (ANNs). However, the complex neuronal dynamics and non-differentiable spike communication mechanisms in SNNs present substantial challenges for efficient training. By analyzing the membrane potentials in spiking neurons, we found that their distributions can increasingly deviate from the firing threshold as time progresses, which tends to cause diminished backpropagation gradients and unbalanced optimization. To address these challenges, we propose Deep Temporal-Aligned Gradient Enhancement (DeepTAGE), a novel approach that improves optimization gradients in SNNs from both internal surrogate gradient functions and external supervision methods. Our DeepTAGE dynamically adjusts surrogate gradients in accordance with the membrane potential distribution across different time steps, enhancing their respective gradients in a temporal-aligned manner that promotes balanced training. Moreover, to mitigate issues of gradient vanishing or deviating during backpropagation, DeepTAGE incorporates deep supervision at both spatial (network stages) and temporal (time steps) levels to ensure more effective and robust network optimization. Importantly, our method can be seamlessly integrated into existing SNN architectures without imposing additional inference costs or requiring extra control modules. We validate the efficacy of DeepTAGE through extensive experiments on static benchmarks (CIFAR10, CIFAR100, and ImageNet-1k) and a neuromorphic dataset (DVS-CIFAR10), demonstrating significant performance improvements. Wei Liu 0153, Li Yang 0014, Mingxuan Zhao, Shuxun Wang, Bing Li 0001, Weiming Hu 0004 |
ICLR | 2 |
| 2025 | MI-TRQR: Mutual Information-Based Temporal Redundancy Quantification and Reduction for Energy-Efficient Spiking Neural NetworksabstractBrain-inspired spiking neural networks (SNNs) provide energy-efficient computation through event-driven processing. However, the shared weights across multiple timesteps lead to serious temporal feature redundancy, limiting both efficiency and performance. This issue is further aggravated when processing static images due to the duplicated input. To mitigate this problem, we propose a parameter-free and plug-and-play module named Mutual Information-based Temporal Redundancy Quantification and Reduction (MI-TRQR), constructing energy-efficient SNNs. Specifically, Mutual Information (MI) is properly introduced to quantify redundancy between discrete spike features at different timesteps on two spatial scales: pixel (local) and the entire spatial features (global). Based on the multi-scale redundancy quantification, we apply a probabilistic masking strategy to remove redundant spikes. The final representation is subsequently recalibrated to account for the spike removal. Extensive experimental results demonstrate that our MI-TRQR achieves sparser spiking firing, higher energy efficiency, and better performance concurrently with different SNN architectures in tasks of neuromorphic data classification, static data classification, and time-series forecasting. Notably, MI-TRQR increases accuracy by \textbf{1.7\%} on CIFAR10-DVS with 4 timesteps while reducing energy cost by \textbf{37.5\%}. Our codes are available at https://github.com/dfxue/MI-TRQR. Dengfeng Xue, Yifan Lu 0001, Chunfeng Yuan, Yufan Liu 0001, Wei Liu 0153, Man Yao, Li Yang 0014, Bing Li 0001, Stephen J. Maybank, Weiming Hu 0004, Zhetao Li |
NeurIPS | 8 |
| 2024 | NFT1000: A Cross-Modal Dataset For Non-Fungible Token RetrievalabstractWith the rise of "Metaverse" and "Web 3.0", Non-Fungible Token (NFT) has emerged as a kind of pivotal digital asset, garnering significant attention. By the end of March 2024, more than 1.7 billion NFTs have been minted across various blockchain platforms. To effectively locate a desired NFT, conducting searches within a vast array of NFTs is essential. The challenge in NFT retrieval is heightened due to the high degree of similarity among different NFTs, regarding regional and semantic aspects. In this paper, we will introduce a benchmark dataset named "NFT Top1000 Visual-Text Dataset" (NFT1000), containing 7.56 million image-text pairs, and being collected from 1000 most famous PFP1 NFT collections2 by sales volume on the Ethereum blockchain. Based on this dataset and leveraging the CLIP series of pre-trained models as our foundation, we propose the dynamic masking fine-tuning scheme. This innovative approach results in a 7.4\% improvement in the top1 accuracy rate, while utilizing merely 13\% of the total training data (0.79 million vs. 6.1 million). We also propose a robust metric Comprehensive Variance Index (CVI) to assess the similarity and retrieval difficulty of visual-text pairs data. The dataset will be released as an open-source resource. For more details, please refer to: https://github.com/ShuxunoO/NFT-Net.git. Shuxun Wang, Yunfei Lei, Ziqi Zhang 0010, Wei Liu 0153, Li Yang 0014, Bing Li 0001, Weiming Hu 0004 |
ACM Multimedia | 6 |
| 2024 | Assessing Face Image Quality: A Large-Scale Database and a Transformer MethodabstractThe amount of face images has been witnessing an explosive increase in the last decade, where various distortions inevitably exist on transmitted or stored face images. The distortions lead to visible and undesirable degradation on face images, affecting their quality of experience (QoE). To address this issue, this paper proposes a novel Transformer-based method for quality assessment on face images (named as TransFQA). Specifically, we first establish a large-scale face image quality assessment (FIQA) database, which includes 42,125 face images with diversifying content at different distortion types. Through an extensive crowdsource study, we obtain 712,808 subjective scores, which to the best of our knowledge contribute to the largest database for assessing the quality of face images. Furthermore, by investigating the established database, we comprehensively analyze the impacts of distortion types and facial components (FCs) on the overall image quality. Accordingly, we propose the TransFQA method, in which the FC-guided Transformer network (FT-Net) is developed to integrate the global context, face region and FC detailed features via a new progressive attention mechanism. Then, a distortion-specific prediction network (DP-Net) is designed to weight different distortions and accurately predict final quality scores. Finally, the experiments comprehensively verify that our TransFQA method significantly outperforms other state-of-the-art methods for quality assessment on face images. Shengxi Li, Mai Xu, Li Yang 0014, Xiaofei Wang 0004 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Saliency Prediction on Mobile Videos: A Fixation Mapping-Based Dataset and A Transformer ApproachabstractWith the booming development of smart devices, mobile videos have drawn broad interest when humans surf social media. Different from traditional long-form videos, mobile videos are featured with uncertain human attention behavior so far owing to the specific displaying mode, thus promoting the research on saliency prediction for mobile videos. Unfortunately, the current eye-tracking experiments are not applicable for mobile videos, since the stationary eye-tracker and eye fixation acquisition are dedicated to the videos presented on computers. To tackle this issue, we propose performing the wearable eye-tracker to record viewers’ egocentric fixations and then devising a fixation mapping technique to project the eye fixations from egocentric videos onto mobile videos. Resorting to this technique, the large-scale mobile video saliency (MVS) dataset is established, including 1,007 mobile videos and 5,935,927 fixations. Given this dataset, we exhaustively analyze the characteristics of subjects’ fixations and obtain two findings. Based on the MVS dataset and these findings, we propose a saliency prediction approach on mobile videos upon Video Swin Transformer (MVFormer), wherein long-range spatio-temporal dependency is captured to derive the human attention mechanism on mobile videos. In MVFormer, we develop the selective feature fusion module to balance multi-scale features, and the progressive saliency prediction module to generate saliency maps via progressive aggregation of multi-scale features. Extensive experiments show that our MVFormer approach significantly outperforms other state-of-the-art saliency prediction approaches. Finally, we demonstrate the potential application of our MVFormer approach in the H.265 video coding standard by embedding it into the rate control scheme, such that the perceptual quality of compressed mobile videos can be significantly improved. The dataset and code will be available at https://github.com/wenshijie110/MVFormer. Shijie Wen, Li Yang 0014, Mai Xu, Minglang Qiao, Lin Bai 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Learnt Mutual Feature Compression for Machine VisionabstractRecently, image coding for machines (ICM) has been playing an important role in facilitating intelligent vision tasks. Unfortunately, the existing ICM methods separately compress features at each scale, neglecting the redundancy across multi-scale features. To address this issue, this paper proposes an end-to-end mutual compression framework for the ICM, such that the compression efficiency can be significantly improved by removing the cross-scale redundancy. Specifically, the proposed framework consists of a mutual feature compression network (MFCNet) and a basic feature compression network (BFCNet). The MFCNet predicts large-scale features from basic small-scale features, such that the large amount of bitrates assigned to compress large-scale features can be saved. Moreover, the BFCNet is proposed to compress small-scale features of high quality by removing spatial and channel-wise redundancy. This guarantees superior performances whilst consuming extremely small amount of bit-rates. The experimental results show that our method achieves 90.10% and 74.97% BD-rate saving against the VVC feature anchor and VVC image anchor that have been recently accepted by the moving picture experts group (MPEG). Mai Xu, Shengxi Li, Li Yang 0014, Zhuoyi Lv |
ICASSP | 5 |
| 2023 | Exploiting Contextual Objects and Relations for 3D Visual Groundingabstract3D visual grounding, the task of identifying visual objects in 3D scenes based on natural language inputs, plays a critical role in enabling machines to understand and engage with the real-world environment. However, this task is challenging due to the necessity to capture 3D contextual information to distinguish target objects from complex 3D scenes. The absence of annotations for contextual objects and relations further exacerbates the difficulties. In this paper, we propose a novel model, CORE-3DVG, to address these challenges by explicitly learning about contextual objects and relations. Our method accomplishes 3D visual grounding via three sequential modular networks, including a text-guided object detection network, a relation matching network, and a target identification network. During training, we introduce a pseudo-label self-generation strategy and a weakly-supervised method to facilitate the learning of contextual objects and relations, respectively. The proposed techniques allow the networks to focus more effectively on referred objects within 3D scenes by understanding their context better. We validate our model on the challenging Nr3D, Sr3D, and ScanRefer datasets and demonstrate state-of-the-art performance. Our code will be public at https://github.com/yangli18/CORE-3DVG. Li Yang 0014, Chunfeng Yuan, Ziqi Zhang 0010, Zhongang Qi, Wei Liu 0153, Ying Shan, Bing Li 0001, Weiping Yang, Yan Wang 0153, Weiming Hu 0004 |
NeurIPS | 1 |
| 2023 | Blind VQA on 360° Video via Progressively Learning From Pixels, Frames, and VideoabstractBlind visual quality assessment (BVQA) on 360° video plays a key role in optimizing immersive multimedia systems. When assessing the quality of 360° video, human tends to perceive its quality degradation from the viewport-based spatial distortion of each spherical frame to motion artifact across adjacent frames, ending with the video-level quality score, i.e., a progressive quality assessment paradigm. However, the existing BVQA approaches for 360° video neglect this paradigm. In this paper, we take into account the progressive paradigm of human perception towards spherical video quality, and thus propose a novel BVQA approach (namely ProVQA) for 360° video via progressively learning from pixels, frames and video. Corresponding to the progressive learning of pixels, frames and video, three sub-nets are designed in our ProVQA approach, i.e., the spherical perception aware quality prediction (SPAQ), motion perception aware quality prediction (MPAQ) and multi-frame temporal non-local (MFTN) sub-nets. The SPAQ sub-net first models the spatial quality degradation based on spherical perception mechanism of human. Then, by exploiting motion cues across adjacent frames, the MPAQ sub-net properly incorporates motion contextual information for quality assessment on 360° video. Finally, the MFTN sub-net aggregates multi-frame quality degradation to yield the final quality score, via exploring long-term quality correlation from multiple frames. The experiments validate that our approach significantly advances the state-of-the-art BVQA performance on 360° video over two datasets, the code of which has been public in https://github.com/yanglixiaoshen/ProVQA. Li Yang 0014, Mai Xu, Shengxi Li, Zulin Wang |
IEEE Trans. Image Process. | 1 |
| 2022 | Improving Visual Grounding with Visual-Linguistic Verification and Iterative ReasoningabstractVisual grounding is a task to locate the target indicated by a natural language expression. Existing methods extend the generic object detection framework to this problem. They base the visual grounding on the features from pre-generated proposals or anchors, and fuse these features with the text embeddings to locate the target mentioned by the text. However, modeling the visual features from these predefined locations may fail to fully exploit the visual context and attribute information in the text query, which limits their performance. In this paper, we propose a transformer-based framework for accurate visual grounding by establishing text-conditioned discriminative features and performing multi-stage cross-modal reasoning. Specifically, we develop a visual-linguistic verification module to focus the visual features on regions relevant to the textual descriptions while suppressing the unrelated areas. A language-guided feature encoder is also devised to aggregate the visual contexts of the target object to improve the object's distinctiveness. To retrieve the target from the encoded visual features, we further propose a multi-stage cross-modal decoder to iteratively speculate on the correlations between the image and text for accurate target localization. Extensive experiments on five widely used datasets validate the efficacy of our proposed components and demonstrate state-of-the-art performance. Li Yang 0014, Chunfeng Yuan, Wei Liu 0153, Bing Li 0001, Weiming Hu 0004 |
CVPR | 1 |
| 2022 | TVFormer: Trajectory-guided Visual Quality Assessment on 360° Images with TransformersabstractVisual quality assessment (VQA) on 360° images plays an important role in optimizing immersive multimedia systems. Due to the absence of pristine 360° images in real world, blind VQA (BVQA) on 360° images has drawn much research attention. In subjective VQA on 360^ images, human intuitively make the quality-scoring decisions through the quality degradation of each observed viewport on the head trajectories. Unfortunately, the existing BVQA works for 360° images neglect the dynamic property of head trajectories with viewport interactions, thus failing to obtain human-like quality scores. In this paper, we propose a novel Transformer-based approach for trajectory-guided VQA on 360° images (named TVFormer), in which both the tasks of head trajectory prediction and BVQA can be accomplished for 360° images. In the first task, we develop a trajectory-aware memory updater (TMU) module, for maintaining the coherence and accuracy of predicted head trajectories. To capture the long-range quality dependency across time-ordered viewports, we propose a spatio-temporal factorized self-attention (STF) module in the encoder of TVFormer for the BVQA task. By implanting the predicted head trajectories into the BVQA task, we can obtain the human-like quality scores. Extensive experiments demonstrate the superior BVQA performance of TVFormer over state-of-the-art approaches on three benchmark datasets. Li Yang 0014, Mai Xu, Liangyu Huo, Xinbo Gao 0001 |
ACM Multimedia | 1 |
| 2022 | Hierarchical Bayesian LSTM for Head Trajectory Prediction on Omnidirectional ImagesabstractWhen viewing omnidirectional images (ODIs), viewers can access different viewports via head movement (HM), which sequentially forms head trajectories in spatial-temporal domain. Thus, head trajectories play a key role in modeling human attention on ODIs. In this paper, we establish a large-scale dataset collecting 21,600 head trajectories on 1,080 ODIs. By mining our dataset, we find two important factors influencing head trajectories, i.e., temporal dependency and subject-specific variance. Accordingly, we propose a novel approach integrating hierarchical Bayesian inference into long short-term memory (LSTM) network for head trajectory prediction on ODIs, which is called HiBayes-LSTM. In HiBayes-LSTM, we develop a mechanism of Future Intention Estimation (FIE), which captures the temporal correlations from previous, current and estimated future information, for predicting viewport transition. Additionally, a training scheme called Hierarchical Bayesian inference (HBI) is developed for modeling inter-subject uncertainty in HiBayes-LSTM. For HBI, we introduce a joint Gaussian distribution in a hierarchy, to approximate the posterior distribution over network weights. By sampling subject-specific weights from the approximated posterior distribution, our HiBayes-LSTM approach can yield diverse viewport transition among different subjects and obtain multiple head trajectories. Extensive experiments validate that our HiBayes-LSTM approach significantly outperforms 9 state-of-the-art approaches for trajectory prediction on ODIs, and then it is successfully applied to predict saliency on ODIs. Li Yang 0014, Mai Xu, Xin Deng 0002, Fangyuan Gao, Zhenyu Guan 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | PDNet: Toward Better One-Stage Object Detection With Prediction DecouplingabstractRecent one-stage object detectors follow a per-pixel prediction approach that predicts both the object category scores and boundary positions from every single grid location. However, the most suitable positions for inferring different targets, i.e., the object category and boundaries, are generally different. Predicting all these targets from the same grid location thus may lead to sub-optimal results. In this paper, we analyze the suitable inference positions for object category and boundaries, and propose a prediction-target-decoupled detector named PDNet to establish a more flexible detection paradigm. Our PDNet with the prediction decoupling mechanism encodes different targets separately in different locations. A learnable prediction collection module is devised with two sets of dynamic points, i.e., dynamic boundary points and semantic points, to collect and aggregate the predictions from the favorable regions for localization and classification. We adopt a two-step strategy to learn these dynamic point positions, where the prior positions are estimated for different targets first, and the network further predicts residual offsets to the positions with better perceptions of the object properties. Extensive experiments on the MS COCO benchmark demonstrate the effectiveness and efficiency of our method. With a single ResNeXt-64x4d-101-DCN as the backbone, our detector achieves 50.1 AP with single-scale testing, which outperforms the state-of-the-art methods by an appreciable margin under the same experimental settings. Moreover, our detector is highly efficient as a one-stage framework. Our code will be public. Li Yang 0014, Shaoru Wang, Chunfeng Yuan, Ziqi Zhang 0010, Bing Li 0001, Weiming Hu 0004 |
IEEE Trans. Image Process. | 1 |
| 2021 | LAU-Net: Latitude Adaptive Upscaling Network for Omnidirectional Image Super-ResolutionabstractThe omnidirectional images (ODIs) are usually at low-resolution, due to the constraints of collection, storage and transmission. The traditional two-dimensional (2D) image super-resolution methods are not effective for spherical ODIs, because ODIs tend to have non-uniformly distributed pixel density and varying texture complexity across latitudes. In this work, we propose a novel latitude adaptive upscaling network (LAU-Net) for ODI super-resolution, which allows pixels at different latitudes to adopt distinct upscaling factors. Specifically, we introduce a Laplacian multi-level separation architecture to split an ODI into different latitude bands, and hierarchically upscale them with different factors. In addition, we propose a deep reinforcement learning scheme with a latitude adaptive reward, in order to automatically select optimal upscaling factors for different latitude bands. To the best of our knowledge, LAU-Net is the first attempt to consider the latitude difference for ODI super-resolution. Extensive results demonstrate that our LAU-Net significantly advances the super-resolution performance for ODIs. Codes are available at https://github.com/wangh-allen/LAU-Net. Xin Deng 0002, Hao Wang 0049, Mai Xu, Yuhang Song 0001, Li Yang 0014 |
CVPR | 6 |
| 2021 | A Viewport-Adaptive Rate Control Approach for Omnidirectional Video CodingabstractFor omnidirectional videos (ODVs), the existing off-line coding approaches are designed based on the spatial or perceptual distortion in a whole ODV frame, ignoring the fact that subjects can only access viewports. To improve the subjective quality inside the viewports, this paper proposes an off-line viewport-adaptive rate control (RC) approach for ODVs in high efficiency video coding (HEVC) framework. Specifically, we predict the viewport candidates with importance weights and develop a viewport saliency detection model. Then, the predicted candidates and detected saliency are taken into account in our viewport-adaptive CTU traversal and bit allocation scheme. Finally, the experimental results validate that our approach is effective in saving bit-rates and improving subjective quality for encoding ODVs; meanwhile, our approach is also effective in the auxiliary task of saliency detection in viewports. Mai Xu, Li Yang 0014 |
DCC | 3 |
| 2021 | Spatial Attention-Based Non-Reference Perceptual Quality Prediction Network for Omnidirectional ImagesabstractDue to the strong correlation between visual attention and perceptual quality, many methods attempt to use human saliency information for image quality assessment. Although this mechanism can get good performance, the networks require human saliency labels, which is not easily accessible for omnidirectional images (ODI). To alleviate this issue, we propose a spatial attention-based perceptual quality prediction network for non-reference quality assessment on ODIs (SAP-net). Without any human saliency labels, our network can adaptively estimate human perceptual quality on impaired ODIs through a self-attention manner, which significantly promotes the prediction performance of quality scores. Moreover, our method greatly reduces the computational complexity in quality assessment task on ODIs. Extensive experiments validate that our network outperforms 9 state-of-the-art methods for quality assessment on ODIs. The dataset and code have been available on https://github.com/yanglixiaoshen/SAP-Net. Li Yang 0014, Mai Xu, Xin Deng 0002 |
ICME | 1 |
| 2021 | Saliency Prediction on Omnidirectional Image With Generative Adversarial Imitation LearningabstractWhen watching omnidirectional images (ODIs), subjects can access different viewports by moving their heads. Therefore, it is necessary to predict subjects' head fixations on ODIs. Inspired by generative adversarial imitation learning (GAIL), this paper proposes a novel approach to predict saliency of head fixations on ODIs, named SalGAIL. First, we establish a dataset for attention on ODIs (AOI). In contrast to traditional datasets, our AOI dataset is large-scale, which contains the head fixations of 30 subjects viewing 600 ODIs. Next, we mine our AOI dataset and discover three findings: (1) the consistency of head fixations are consistent among subjects, and it grows alongside the increased subject number; (2) the head fixations exist with a front center bias (FCB); and (3) the magnitude of head movement is similar across the subjects. According to these findings, our SalGAIL approach applies deep reinforcement learning (DRL) to predict the head fixations of one subject, in which GAIL learns the reward of DRL, rather than the traditional human-designed reward. Then, multi-stream DRL is developed to yield the head fixations of different subjects, and the saliency map of an ODI is generated via convoluting predicted head fixations. Finally, experiments validate the effectiveness of our approach in predicting saliency maps of ODIs, significantly better than 11 state-of-the-art approaches. Our AOI dataset and code of SalGAIL are available online at https://github.com/yanglixiaoshen/SalGAIL. Mai Xu, Li Yang 0014, Xiaoming Tao 0001, Yiping Duan, Zulin Wang |
IEEE Trans. Image Process. | 2 |
| 2018 | Rate control schemes for panoramic video coding
Yufan Liu 0001, Li Yang 0014, Mai Xu, Zulin Wang |
J. Vis. Commun. Image Represent. | 2 |