VLDB 2026 Research / reviewers in the wild / expert
Xinjue Hu
dblp:213/6065
· DBLP profile ↗
12ranked-venue papers
7as first author
9since 2021 · last 2026
0000-0001-8304-9720ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 6 first-author · 5 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | High-Resolution Image SteganalysisabstractImage steganalysis detects hidden information within images. However, existing methods are primarily designed for low-resolution images and they struggle to address the challenges posed by the widespread use of high-resolution images in real-world scenarios such as communications and social media. Under the secure embedding constrained by the square root law, the steganographic noise of high-resolution images is significantly diluted, thereby making “strong decision regions” critical to detection increasingly scarce, whereas the interference effect of “weak decision regions” is relatively prominent. Existing methods treat all areas equally, making it difficult to fully utilize strong decision regions and suppress the negative impact of weak decision regions. To address these issues, we propose the HRIS, a two-phase collaborative optimization framework for high-resolution steganalysis from “discovery” to “utilization”. In the “discovery” phase, we propose a dynamically contribution-guided decision region recognition mechanism. This mechanism employs a difference amplification module to amplify the steganographic noise differences between regions and then leverages a cooperative game-driven dynamic optimization strategy to compute each sub-region's contribution to the prediction. It accurately identifies and reinforces strong decision regions while suppressing interference from weak decision regions, resulting in significantly improved local detection accuracy. In the “utilization” phase, we propose a decision regions-global steganographic features bidirectional collaborative optimization framework that leverages the identified strong and weak decision regions to direct the extraction of global steganographic noise features. These global features are then fed back to refine the local feature representations, enabling collaborative enhancement between local and global analyses. Extensive experimental results demonstrate that our method achieves state-of-the-art performance on high-resolution images. Xinjue Hu, Zhenshan Tan, Xiang Zhang 0023, Zhangjie Fu 0001 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2026 | MAP-Mamba: Multi-Artifacts Perception Mamba for Generalizable Face Forgery DetectionabstractFace forgery detection suffers from cross-dataset generalization challenges, where performance degradation occurs due to distribution shifts between training and testing data. Recently, pseudo-fake face generation strategy has mitigated models overfitting to specific forgery traces. However, detectors based on this strategy exhibit an overreliance on blending boundary artifacts for their classification decisions. This overreliance significantly limits their ability to generalize to more advanced face manipulation algorithms, such as FaceDancer and InSwap, which are designed to produce smooth and natural transitions in the blending boundary region. To address this, we propose MAP-Mamba, a novel Multi-Artifacts Perception Mamba framework for modeling generalizable artifact representations from “Generation” to “Enrichment” to “Strengthening”. First, we design an attribute-level face blending method that generate pseudo-fake faces containing fine-grained artifacts via three attribute generators. These pseudo-fakes mimic subtle local inconsistencies in advanced forgery algorithms, guiding the MAP-Mamba to learn diverse forgery features beyond the blending boundary artifacts. Second, considering the variability of face artifacts distribution caused by different forgery algorithms, an artifact style mixing strategy is designed to enrich the artifact style distribution in the training phase by mixing and reorganizing the artifact style features, and to enhance the model’s ability to handle unknown forgery methods. Finally, an adaptive artifact guidance mechanism is proposed to dynamically amplify the artifact-related feature to further strengthen the model’s sensitivity to key artifacts. Extensive experiments on several benchmarks show that MAP-Mamba achieves superior robustness and generalization performance. Ziwen He, Xinjue Hu, Weinan Guan, Wei Wang 0025, Zhangjie Fu 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2026 | Toward Functional Testing of Autonomous Ships: A Structured Survey With a Hybrid Virtual-Real Implementation FrameworkabstractThe International Maritime Organization (IMO) proposed the concept of a maritime autonomous surface ship (MASS), which could integrate perception, decision-making, and control functions to implement autonomous navigation in complicated environments. It is recognized as the fundamental component of the future new generation of shipping systems. The development of MASS requires the use of specific test tools and use cases for the evaluation of the safety, reliability, and functional performance of its navigation system. At present, the lack of standardized, modular, and serial testing paradigms and methods for MASS impedes the development and improvement of related products and applications. Therefore, it is essential to develop a practical and applicable testing framework. This paper presents a functional analysis of autonomous navigation systems and provides a comprehensive overview of current research status, potential solutions, and future challenges. Moreover, this paper also takes a step towards facilitating the functional testing of autonomous surface vehicles by proposing a hybrid virtual-real framework consisting of scenario generation, virtual simulation, model-scaled physical experiment, validation and evaluation. The research opportunities and future aspects are also addressed. This work may serve as a reference for academic and industrial researchers to investigate new methods and to develop prototype systems for future autonomous surface ships. Jialun Liu, Zhilin Dong, Shijie Li 0003, Zhouhua Peng, Xinjue Hu |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2025 | Multi-source Domain Adaptation Image Steganalysis for Cover Source Mismatch
Xiang Zhang 0023, Xinjue Hu, Fan Wang 0024, Xu Cheng 0003, Zhangjie Fu 0001 |
PRCV (6) | 3 |
| 2025 | A pre-trained multi-step prediction informer for ship motion prediction with a mechanism-data dual-driven framework
Wenhe Shen, Xinjue Hu, Jialun Liu, Shijie Li 0003, Hongdong Wang |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | Invisible and Steganalysis-Resistant Deep Image Hiding Based on One-Way Adversarial Invertible NetworksabstractDeep image hiding is a challenging image processing task that aims to hide a secret image into a cover image of equal size perfectly. How to improve the imperceptibility of deep image hiding while ensuring high computational efficiency is a primary challenge. Where imperceptibility means not being visually perceived while not being perceived by the steganalysis model. In this paper, we propose a novel deep image hiding framework called DIH-OAIN (Deep Image Hiding based on One-way Adversarial Invertible Networks) to address it. Firstly, an image cascade framework is introduced to extract image semantics and details with dual-resolution branches, and reduces computation complexity by balancing image resolution and model complexity. Secondly, a hidden probability guided module is designed to constrain the secret image to be hidden in the texture region, utilizing the image texture complexity as prior knowledge. The above two points can effectively improve visual imperceptibility. Finally, a one-way adversarial training strategy is proposed to enhance the model imperceptibility. A series of experimental results show that the proposed method is significantly improved in imperceptibility comparing to state-of-the-art deep image hiding algorithms, while maintaining a low computation complexity. Xinjue Hu, Zhangjie Fu 0001, Xiang Zhang 0023 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Multiple Description Coding for Best-Effort Delivery of Light Field Video Using GNN-Based CompressionabstractIn recent years, Light Field (LF) video has grabbed much attention as an emerging form of immersive media. LF collects, through a lens matrix, light information emanating in every direction, and obtains rich information about the scene, providing users with an immersive 6 Degrees of Freedom (DoF) experience. The visual content between different viewpoints is highly homogenized, suggesting the possibility of good compression and encoding. However, most fixed-structure LF coding schemes are difficult to adapt to the real-time requirements of different LF applications and best-effort network conditions causing packet loss. In this paper, we propose a dynamic adaptive LF video transmission scheme that can achieve high compression and yet provide near-distortion-free LF video when the network condition is stable. Additionally, for unstable network conditions a description scheduling algorithm is proposed, which can decode the LF video with the highest possible quality even if partial data cannot be received completely and/or timely. We achieve this by designing a Multiple Description Coding (MDC) based solution to transport the LF video compressed by a Graph Neural Network (GNN) model. Experimental results show that the scheduling algorithm can improve the quality of the decoding results by 3% to 15%. Compared with other similar schemes, our system greatly improves the reliability of the video streaming system against packet loss/error and supports heterogeneous receivers. Xinjue Hu, Yumei Wang, Lin Zhang 0013, Shervin Shirmohammadi |
IEEE Trans. Multim. | 1 |
| 2022 | Angular-spatial analysis of factors affecting the performance of light field reconstructionabstractAbstract As a new VR multimedia format, Light Field (LF) has received more and more attention. LF can support users to move freely within a certain range during the experience. However, the limitation of LF acquisition technology has become the main obstacle to its large‐scale application. Since the current hardware technology cannot produce a lens small enough to meet the needs of LF capture, the intensity of the collected LF image is usually insufficient. Although there are many software LF reconstruction algorithms have been proposed to make captured result denser, they all face the problem of poor generalization in the application process. In this work, it is attempted to locate the factors that affect the performance of LF reconstruction from both the angular and spatial domains. The analysis results show that it is affected not only by the edge and texture of the LF content in the spatial domain, but also by the adjacent disparity and the reflection characteristics in the angular domain. An indicator to quantify the impact of these factors on reconstruction performance is also proposed, which will be helpful to design a new generalized adaptive LF reconstruction algorithm in the future. Xinjue Hu, Lin Zhang 0013 |
IET Image Process. | 1 |
| 2021 | 4DLFVD: A 4D Light Field Video DatasetabstractWe present a 4D Light Field (LF) video dataset, collected by a custom-made camera matrix, to be used for designing and testing algorithms and systems for LF video coding, processing, and streaming. Compared to existing LF datasets, ours provides LF videos, as opposed to only images, and at higher frame resolution, higher number of viewpoints, and/or higher framerate, offering the best visual quality LF video dataset. To achieve this, we built a 10 x 10 LF capture matrix composed of 100 cameras, each with a 1920 x 1056 resolution. We used this matrix to record videos in real and varying illumination and scene dynamics conditions. The dataset contains a total of nine groups of LF videos: eight groups collected with a fixed camera matrix position and orientation recording indoor potted plants, furniture, etc., and the last group collected by rotating around an outdoor environment with roadside vehicles, pedestrians, etc. Each group of LF videos consists of 100 video streams encoded with H.265/HEVC. Scene changes vary from static to slightly dynamic to highly dynamic, providing a good level of diversity. As an example, we present the results of a depth estimation method and show that our dataset can be used for applications such as objection detection, 3D modeling, and others. Xinjue Hu, Yunming Liu, Yumei Wang, Yu Liu 0001, Lin Zhang 0013, Shervin Shirmohammadi |
MMSys | 1 |
| 2020 | Cooperative Tile-Based 360° Panoramic Streaming in Heterogeneous Networks Using Scalable Video CodingabstractThe use of high-quality 360° panoramic video is booming in the video industry. However, existing schemes for smartphones suffer from significant bandwidth consumption as they transmit the entire panoramic views in very high resolutions. This demand for bandwidth becomes even more problematic when multiple adjacent smartphones compete to access the same content, which further challenges a wireless network's capacity, and when the available bandwidth fluctuates much more than wired networks. In this paper, we propose a cooperative streaming scheme for tile-based 360° video using scalable video coding (SVC) to maximize a group of users' quality of experience. We formulate an optimization problem to choose optimal downloading and sharing subsets from a set of all requested SVC layers of tiles to maximize the effective quality of the users' viewport while meeting the feasibility of the bandwidth of heterogeneous networks. We then show that the problem is NP-hard and compose a heuristic approach. In the approach, we rank the SVC layers based on the aggregated group-level preference to guide the devices' downloading and sharing activities. A prototype on the Android platform is developed to test the approach's performance, and the real-world results show that our proposed scheme outperforms baseline alternatives. Xiaoyi Zhang 0001, Xinjue Hu, Shervin Shirmohammadi, Lin Zhang 0013 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | An Adaptive Two-Layer Light Field Compression Scheme Using GNN-Based ReconstructionabstractAs a new form of volumetric media, Light Field (LF) can provide users with a true six degrees of freedom immersive experience because LF captures the scene with photo-realism, including aperture-limited changes in viewpoint. But uncompressed LF data is too large for network transmission, which is the reason why LF compression has become an important research topic. One of the more recent approaches for LF compression is to reduce the angular resolution of the input LF during compression and to use LF reconstruction to recover the discarded viewpoints during decompression. Following this approach, we propose a new LF reconstruction algorithm based on Graph Neural Networks; we show that it can achieve higher compression and better quality compared to existing reconstruction methods, although suffering from the same problem as those methods—the inability to deal effectively with high-frequency image components. To solve this problem, we propose an adaptive two-layer compression architecture that separates high-frequency and low-frequency components and compresses each with a different strategy so that the performance can become robust and controllable. Experiments with multiple datasets 1 show that our proposed scheme is capable of providing a decompression quality of above 40 dB, and can significantly improve compression efficiency compared with similar LF reconstruction schemes. Xinjue Hu, Jingming Shan, Yu Liu 0001, Lin Zhang 0013, Shervin Shirmohammadi |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2019 | Adaptive two-layer light field compression scheme based on sparse reconstructionabstractAs a new form of volumetric media, the technology of light field and its compression has gradually become the research hotspots in academia. The scheme of compressing using the sparsity of the light field is a very promising idea, which has the characteristics of high compression rate and is not affected by the occlusion of scene objects. However, the instability of the reconstruction algorithm's performance on different datasets limits the further application of this solution. Since the quality of the decompression outputs will be limited below the reconstruction result, the poor performance of the reconstruction algorithm on some light field images will result in a very low PSNR upper limit for the compression scheme. This paper finds that the main reason for this performance problem is the poor ability of the algorithm to process the high-frequency components of the light field. And in order to solve it, an adaptive two-layer light field compression scheme is presented. The proposed scheme separates the high-frequency components and the low-frequency components of the light field so that they can be independently compressed. Through the adaptive adjustment, the data of different frequency component can adopt different compression strategies, so that the performance of the proposed scheme can be optimal. Experiments with multiple datasets1 show that the proposed scheme can break the upper limit of PSNR caused by sparse reconstruction and is capable to provide decompression results above 40 dB. It also achieves significant improvement in compression efficiency under diverse requirements. Xinjue Hu, Jingming Shan, Yu Liu 0001, Lin Zhang 0013 |
MMSys | 1 |