Kao Zhang

dblp:58/1863 · DBLP profile ↗
← Back
14ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0001-9111-0499ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 SphereU-Sal360: Spherical U-Shaped Spatio-Temporal Transformer for $360^\circ $ Video Saliency
Yanzhi Ding, Kao Zhang
KSEM (6)2
2026 CCDM: Continuous Conditional Diffusion Models for Image Generation
abstract
Continuous Conditional Generative Modeling(CCGM) estimates high-dimensional data distributions, such as images, conditioned on scalar continuous variables (aka regression labels). WhileContinuous Conditional Generative Adversarial Networks(CcGANs) were designed for this task, their instability during adversarial learning often leads to suboptimal results.Conditional Diffusion Models(CDMs) offer a promising alternative, generating more realistic images, but their diffusion processes, label conditioning, and model fitting procedures are either not optimized for or incompatible with CCGM, making it difficult to integrate CcGANs' vicinal approach. To address these issues, we introduceContinuous Conditional Diffusion Models(CCDMs), the first CDM specifically tailored for CCGM. CCDMs address existing limitations with specially designed conditional diffusion processes, a novel hard vicinal image denoising loss, a customized label embedding method, and efficient conditional sampling procedures. Through comprehensive experiments on four datasets with resolutions ranging from$64\times 64$to$192\times 192$, we demonstrate that CCDMs outperform state-of-the-art CCGM models, establishing a new benchmark. Ablation studies further validate the model design and implementation, highlighting that some widely used CDM implementations are ineffective for the CCGM task. Our code is publicly available athttps://github.com/UBCDingXin/CCDM.
Xin Ding 0004, Kao Zhang, Z. Jane Wang 0001
IEEE Trans. Multim.3
2025 Prompt-based hybrid supervised contrastive learning for emotion recognition in conversation
Chuangxin Cai, Kao Zhang, Xianxuan Lin
Neurocomputing2
2024 Video saliency prediction for First-Person View UAV videos: Dataset and benchmark
Kao Zhang, Chenxi Jiang
Neurocomputing2
2024 Audio-visual saliency prediction for movie viewing in immersive environments: Dataset and benchmarks
Kao Zhang, Xiaoying Ding, Chenxi Jiang, Zhenzhong Chen 0001
J. Vis. Commun. Image Represent.2
2023 Towards object tracking for quadruped robots
Kao Zhang, Wanping Ouyang, Mingpeng Cui, Chenxi Jiang, Daiqin Yang, Zhenzhong Chen 0001
J. Vis. Commun. Image Represent.2
2022 SalCrop: Spatio-temporal Saliency Based Video Cropping
abstract
Video cropping is a key research task in video processing field. In this paper, a spatio-temporal saliency based video cropping framework (SalCrop) is introduced including four core modules: video scene detection module, video saliency prediction module, adaptive cropping module, and video codec module. It can automatically reframe videos in the desired aspect ratios. In addition, a large-scale video cropping dataset (VCD) is built for training and testing. Experiments on the VCD test dataset show that our SalCrop outperforms the state-of-the-art algorithms with high efficiency. Besides, a FFmpeg video filter is developed based on the framework, which can be widely used in different scenarios. A demo is available at: https://mme.tencent.com/smartcontent/videoCrop (access token: test_token).
Kao Zhang, Yan Shang, Songnan Li, Shan Liu 0001, Zhenzhong Chen 0001
VCIP1
2021 A Spatial-Temporal Recurrent Neural Network for Video Saliency Prediction
abstract
In this paper, a recurrent neural network is designed for video saliency prediction considering spatial-temporal features. In our work, video frames are routed through the static network for spatial features and the dynamic network for temporal features. For the spatial-temporal feature integration, a novel select and re-weight fusion model is proposed which can learn and adjust the fusion weights based on the spatial and temporal features in different scenes automatically. Finally, an attention-aware convolutional long short term memory (ConvLSTM) network is developed to predict salient regions based on the features extracted from consecutive frames and generate the ultimate saliency map for each video frame. The proposed method is compared with state-of-the-art saliency models on five public video saliency benchmark datasets. The experimental results demonstrate that our model can achieve advanced performance on video saliency prediction.
Kao Zhang, Zhenzhong Chen 0001, Shan Liu 0001
IEEE Trans. Image Process.1
2021 Attentive Cross-Modal Fusion Network for RGB-D Saliency Detection
abstract
In this paper, an attentive cross-modal fusion (ACMF) network is proposed for RGB-D salient object detection. The proposed method selectively fuses features in a cross-modal manner and uses a fusion refinement module to fuse output features from different resolutions. Our attentive cross-modal fusion network is built based on residual attention. In each level of ResNet output, both the RGB and depth features are turned into an identity map and a weighted attention map. The identity map is reweighted by the attention map of the paired modality. Moreover, the lower level features with higher resolution are adopted to refine the boundary of detected targets. The entire architecture can be trained end-to-end. The proposed ACMF is compared with state-of-the-art methods on eight recent datasets. The results demonstrate that our model can achieve advanced performance on RGB-D salient object detection.
Di Liu 0007, Kao Zhang, Zhenzhong Chen 0001
IEEE Trans. Multim.2
2019 Two-Stream Refinement Network for RGB-D Saliency Detection
abstract
In this paper, we propose a two-stream refinement network for RGB-D saliency detection. A fusion refinement module is designed to fuse output features from different resolution and modals. The structure information from depth helps distinguish between foreground and background and the lower level features with higher resolution can be adopted to refine the boundary of detected targets. The proposed model predicts high-resolution saliency map and then use a propagation-based module to further refine object boundary. Experimental results demonstrate that the proposed method performs well against to the state of the art methods on the recent RGB-D salient object detection dataset.
Di Liu 0007, Yaosi Hu, Kao Zhang, Zhenzhong Chen 0001
ICIP3
2019 Video Saliency Prediction Based on Spatial-Temporal Two-Stream Network
abstract
In this paper, we propose a novel two-stream neural network for video saliency prediction. Unlike some traditional methods based on hand-crafted feature extraction and integration, our proposed method automatically learns saliency related spatiotemporal features from human fixations without any pre-processing, post-processing, or manual tuning. Video frames are routed through the spatial stream network to compute static or color saliency maps for each of them. And a new two-stage temporal stream network is proposed, which is composed of a pre-trained 2D-CNN model (SF-Net) to extract saliency related features and a shallow 3D-CNN model (Te-Net) to process these features, for temporal or dynamic saliency maps. It can reduce the requirement of video gaze data, improve training efficiency, and achieve high performance. A fusion network is adopted to combine the outputs of both streams and generate the final saliency maps. Besides, a convolutional Gaussian priors (CGP) layer is proposed to learn the bias phenomenon in viewing behavior to improve the performance of the video saliency prediction. The proposed method is compared with state-of-the-art saliency models on two public video saliency benchmark datasets. The results demonstrate that our model can achieve advanced performance on video saliency prediction.
Kao Zhang, Zhenzhong Chen 0001
IEEE Trans. Circuits Syst. Video Technol.1
2018 A saliency prediction model on 360 degree images using color dictionary based sparse representation
Jing Ling, Kao Zhang, Yingxue Zhang 0004, Daiqin Yang, Zhenzhong Chen 0001
Signal Process. Image Commun.2
2008 Performance Analysis of Concurrent Programs Using Ordinary Differential Equations
abstract
Based on Continuous Petri Net, we build differential equation model for concurrent programs. The program behavior can be analyzed from the curves of the solutions of the differential equations. We show that a program state can be measured with a number between 0 and 1, called state measure, indicating how much the state can be reached while the program is in execution. Thus, instead of displaying one state at one time, a program can display all states at one time with state measure attached to each state. This information can help us to estimate where and how much the resources have been used. The advantage of our method is that we can avoid state explosion problem while doing program analysis. Our equations can be solved by Matlab and simulated with a tool: Snoopy.
Zuohua Ding, Kao Zhang
COMPSAC2
2008 A rigorous approach towards test case generation
Zuohua Ding, Kao Zhang, Jueliang Hu
Inf. Sci.2