VLDB 2026 Research / reviewers in the wild / expert
Zhenhao Sun
dblp:251/6703
· DBLP profile ↗
9ranked-venue papers
4as first author
7since 2021 · last 2025
0000-0002-3594-8104ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Compressing Inside Generating: A Latent Domain Codec for AI-Generated ImagesabstractLatent diffusion models (LDMs) have emerged as a prominent framework for image generation, consisting of a diffusion model$\mathcal{M}$and a VAE decoder$\mathcal{D}$. High-quality image generation models are large and computationally intensive. As a result, image generation is typically performed on cloud servers, with the generated images then transmitted to edge devices. Yuxu Chen, Zhenhao Sun, Yuliang Huang, Shiqi Wang 0001 |
DCC | 2 |
| 2024 | Early Determination for Intra Block Copy Prediction with Refined Screen Content DetectionabstractThe intra block copy (IBC) mode introduced can significantly improve the coding efficiency for screen content. However, with the mixed content, the coding performance improvement of IBC is limited and while also introducing the increases of the coding complexity. This paper proposes an early determination scheme for IBC prediction to reduce the encoding complexity over mixed content sequences containing captured content regions. Our method is based on a refined CTU-level screen content detection method, which can better distinguish screen content in mixed scenes. Experimental results show that the proposed method reduces encoding complexity with moderate coding performance loss on mixed content sequences. Zhenhao Sun, Meng Wang 0017, Yingwen Zhang, Shiqi Wang 0001, Sam Kwong |
VCIP | 1 |
| 2024 | LEC-Codec: Learning-Based Genome Data CompressionabstractIn this paper, we propose a Learning-based gEnome Codec (LEC), which is designed for high efficiency and enhanced flexibility. The LEC integrates several advanced technologies, including Group of Bases (GoB) compression, multi-stride coding and bidirectional prediction, all of which are aimed at optimizing the balance between coding complexity and performance in lossless compression. The model applied in our proposed codec is data-driven, based on deep neural networks to infer probabilities for each symbol, enabling fully parallel encoding and decoding with configured complexity for diverse applications. Based upon a set of configurations on compression ratios and inference speed, experimental results show that the proposed method is very efficient in terms of compression performance and provides improved flexibility in real-world applications. Zhenhao Sun, Meng Wang 0017, Shiqi Wang 0001, Sam Kwong |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2024 | Revisiting All-Zero Block Detection for Versatile Video CodingabstractThe Versatile Video Coding (VVC) standard adopts a series of new coding tools in transform and quantization, including multiple transform selection, low-frequency non-separable transform, and trellis quantization. These new technologies, which bring significant coding gain, create daunting challenges to optimizing the VVC codec. In this work, we propose a new all-zero block (AZB) detection scheme tailored for VVC, with the collaboration of genuine all-zero block (GAZB) and pseudo all-zero block (PAZB) detection. First, to accommodate the multiple transform sizes in VVC, we develop a GAZB detection method that is apt for square and non-square residual blocks. Meanwhile, a theoretical upper bound is derived to locate the last significant coefficient and detect the potential frequency domain GAZB. Subsequently, a method tailored for trellis-coded quantization in VVC is devised for detecting PAZB. Finally, the GAZB and PAZB detection methods are collaboratively employed for AZB detection in VVC. The proposed method is implemented on the VVC codec Versatile Video Encoder (VVenC), and extensive experimental results show that the proposed method achieves promising time savings for test sequences of different resolutions with negligible rate-distortion performance loss. Zhenhao Sun, Meng Wang 0017, Peilin Chen 0001, Xu Wang 0006, Shiqi Wang 0001, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Complexity-Configurable Learning-based Genome CompressionabstractIn this paper, we propose the complexity configurable learning-based genome data compression method, in an effort to achieve a good balance between coding complexity and performance in lossless DNA compression. In particular, we first introduce the concept of Group of Bases (GoB), which serves as the foundation and enables the parallel implementation of the learning-based genome data compression. Subsequently, the Markov model is introduced for modeling the initial content, and the learning-based inference is achieved for the remaining base data. The compression is finally achieved with efficient arithmetic coding, and based upon a set of configurations on compression ratios and inference speed, the proposed method is shown to be more efficient and provide more flexibility in real-world applications. Zhenhao Sun, Meng Wang 0017, Shiqi Wang 0001, Sam Kwong |
PCS | 1 |
| 2021 | A Multi-Task Collaborative Network for Light Field Salient Object DetectionabstractBeing able to predict the salient object is of fundamental importance in image processing and computer vision. With numerous approaches proposed for automatic image and video salient object detection, much less work has been dedicated to detecting and segmenting salient objects from light fields. In this article, based on the intrinsic characteristics of light fields, we carefully explore the complementary coherence among multiple cues including spatial, edge and depth information, and elaborately design a multi-task collaborative network for light field salient object detection. More specifically, the correlation mechanisms among edge detection, depth inference and salient object detection are carefully investigated to facilitate the representative saliency features. We first model the coherence among low-level features and heuristic semantic priors, as well as the edge information. Subsequently, the depth-oriented saliency features are derived from the geometry of light fields, in which the 3D convolution operation is leveraged with powerful representation capability to model the disparity correlations among multiple viewpoint images. Finally, a feature-enhanced salient object generator is developed to integrate these complementary saliency features, leading to the final salient object predictions for light fields. Quantitative and qualitative experiments demonstrate the superiority of our proposed model against the state-of-the-art methods over the public light field salient object detection datasets. Qiudan Zhang, Shiqi Wang 0001, Xu Wang 0006, Zhenhao Sun, Sam Kwong, Jianmin Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Geometry Auxiliary Salient Object Detection for Light Fields via Graph Neural NetworksabstractLight field imaging, originated from the availability of light field capture technology, offers a wide range of applications in the field of computational vision. The capability of predicting salient objects of light fields remains technologically challenging due to its complicated geometry structure. In this paper, we propose a light field salient object detection approach that formulates the geometric coherence among multiple views of light fields as graphs, where the angular/central views represent the nodes and their relations compose the edges. The spatial and disparity correlations between multiple views are effectively explored through multi-scale graph neural networks, enabling the more comprehensive understanding of light field content and more representative and discriminative saliency features generation. Moreover, a multi-scale saliency feature consistency learning module is embedded to enhance the saliency features. Finally, an accurate salient object map is produced for the light field based upon the extracted features. In addition, we establish a new light field salient object detection dataset (CITYU-Lytro) that contains 817 light fields with diverse contents and their corresponding annotations, aiming to further promote the research on light field salient object detection. Quantitative and qualitative experiments demonstrate that the proposed method performs favorably compared with the state-of-the-art methods on the benchmark datasets. Qiudan Zhang, Shiqi Wang 0001, Xu Wang 0006, Zhenhao Sun, Sam Kwong, Jianmin Jiang |
IEEE Trans. Image Process. | 4 |
| 2020 | Multi-Exposure Decomposition-Fusion Model for High Dynamic Range Image Saliency DetectionabstractHigh dynamic range (HDR) imaging techniques have witnessed a great improvement in the past few decades. However, saliency detection task on HDR content is still far from well explored. In this paper, we introduce a multi-exposure decomposition-fusion model for HDR image saliency detection inspired by the brightness adaption mechanism. The proposed model is composed of three modules. Firstly, a decomposition module converts the input raw HDR image into a stack of LDR images by uniformly sampling the exposure time range. Secondly, a saliency region proposal network is employed to generate the candidate saliency maps for each LDR image in the exposure stack. Finally, an uncertainty weighting based fusion algorithm is applied to generate the overall saliency map for the input HDR image by merging the obtained LDR saliency maps. Extensive experiments show that our proposed model achieves superior performance compared with the state-of-the-art methods on the existing HDR eye fixation databases. The source code of the proposed model are made publicly available at https://github.com/sunnycia/DFHSal. Xu Wang 0006, Zhenhao Sun, Qiudan Zhang, Yuming Fang 0001, Lin Ma 0002, Shiqi Wang 0001, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Learning to Explore Saliency for Stereoscopic Videos Via Component-Based InteractionabstractIn this paper, we devise a saliency prediction model for stereoscopic videos that learns to explore saliency inspired by the component-based interactions including spatial, temporal, as well as depth cues. The model first takes advantage of specific structure of 3D residual network (3D-ResNet) to model the saliency driven by spatio-temporal coherence from consecutive frames. Subsequently, the saliency inferred by implicit-depth is automatically derived based on the displacement correlation between left and right views by leveraging a deep convolutional network (ConvNet). Finally, a component-wise refinement network is devised to produce final saliency maps over time by aggregating saliency distributions obtained from multiple components. In order to further facilitate research towards stereoscopic video saliency, we create a new dataset including 175 stereoscopic video sequences with diverse content, as well as their dense eye fixation annotations. Extensive experiments support that our proposed model can achieve superior performance compared to the state-of-the-art methods on all publicly available eye fixation datasets. Qiudan Zhang, Xu Wang 0006, Shiqi Wang 0001, Zhenhao Sun, Sam Kwong, Jianmin Jiang |
IEEE Trans. Image Process. | 4 |