VLDB 2026 Research / reviewers in the wild / expert
Qingyang Zhou
dblp:299/9319
· DBLP profile ↗
14ranked-venue papers
8as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 6 first-author · 11 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Non-Local Guided Neural Fields for 4D CT ReconstructionabstractDynamic CT reconstruction plays a crucial role in both medical and industrial applications. However, existing 4D CT reconstruction methods typically rely on complex regularization techniques or external large-scale training datasets, posing challenges for reconstruction quality and generalization when handling complex object motion and varied imaging modes. Neural Radiance Fields (NeRF) offer a promising approach to dynamic CT reconstruction, but existing NeRF-based methods often assume that the scene is low-rank, limiting their representation capabilities. To address these issues, we propose NG-NeRF. First, we combine 3D and 4D hash grids for scene representation, effectively reducing temporal redundancy in static regions of dynamic scenes while improving the model’s representation capabilities and efficiency. Next, we design a non-local hash attention module to establish non-local dependencies between the features of different hash grids. This guides the model to adaptively select features based on hash table load information, significantly alleviating hash collisions and achieving the decoupling of dynamic and static regions. Besides, we introduce global continuity by employing mask positional encoding, which helps reduce the noise often introduced by grid features. Our experimental results on medical and industrial datasets demonstrate that the proposed method outperforms existing state-of-the-art methods by 5.84 dB and 3.4 dB, respectively, and exhibits excellent generalization ability across different 4D CT scenarios. Qingyang Zhou, Yunfan Ye, Zhihuang Liu, Zhiping Cai |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Causally Consistent Normalizing FlowabstractCausal inconsistency arises when the underlying causal graphs captured by generative models like Normalizing Flows are inconsistent with those specified in causal models like Struct Causal Models. This inconsistency can cause unwanted issues including unfairness. Prior works to achieve causal consistency inevitably compromise the expressiveness of their models by disallowing hidden layers. In this work, we introduce a new approach: Causally Consistent Normalizing Flow (CCNF). To the best of our knowledge, CCNF is the first causally consistent generative model that can approximate any distribution with multiple layers. CCNF relies on two novel constructs: a sequential representation of SCMs and partial causal transformations. These constructs allow CCNF to inherently maintain causal consistency without sacrificing expressiveness. CCNF can handle all forms of causal inference tasks, including interventions and counterfactuals. Through experiments, we show that CCNF outperforms current approaches in causal inference. We also empirically validate the practical utility of CCNF by applying it to real-world datasets and show how CCNF addresses challenges like unfairness effectively. Qingyang Zhou, Kangjie Lu, Meng Xu 0001 |
AAAI | 1 |
| 2025 | Spatiotemporal-Aware Neural Fields for Dynamic CT ReconstructionabstractWe propose a dynamic Computed Tomography (CT) reconstruction framework called STNF4D (SpatioTemporal-aware Neural Fields). First, we represent the 4D scene using four orthogonal volumes and compress these volumes into more compact hash grids. Compared to the plane decomposition method, this method enhances the model's capacity while keeping the representation compact and efficient. However, in densely predicted high-resolution dynamic CT scenes, the lack of constraints and hash conflicts in the hash grid features lead to obvious dot-like artifact and blurring in the reconstructed images. To address these issues, we propose the Spatiotemporal Transformer (ST-Former) that guides the model in selecting and optimizing features by sensing the spatiotemporal information in different hash grids, significantly improving the quality of reconstructed images. We conducted experiments on medical and industrial datasets covering various motion types, sampling modes, and reconstruction resolutions. Experimental results show that our method outperforms the second-best by 5.99 dB and 4.11 dB in medical and industrial scenes, respectively. Qingyang Zhou, Yunfan Ye, Zhiping Cai |
AAAI | 1 |
| 2025 | HumanSAM: Classifying Human-Centric Forgery Videos in Human Spatial, Appearance, and Motion Anomaly
Yunfan Ye, Fan Zhang 0144, Qingyang Zhou, Yuchuan Luo, Zhiping Cai |
ICCV | 4 |
| 2025 | Low latency scheme for semi decoupled tree partition in AVMabstractThe latest initiative of Alliance for Open Media Video (AOM), named as AOM Video Model (AVM), is expected to introduce new coding tools to enhance compression benefits. Semi-Decoupled Partitioning (SDP) in AVM decouples the shared tree to support separate block partitioning for the luma and chroma channels from 64×64. Further, Chroma from Luma (CfL) is a chroma-only coding tool in AVM that applies collocated luma reconstructed samples in predicting chroma samples. The dependency of reconstructed luma samples in CfL can result in a delayed decoding process of chroma blocks in separate tree partitioning and introduce a worst-case latency of 4096 luma samples. In response, this study proposes a CfL constrained strategy to reduce the worst-case latency by selectively disallowing the CfL mode in a given chroma partition tree. Detailed latency analysis is also provided to confirm the reduction of worst-case latency to 2048 luma samples. The experiments are implemented on top research-v10.0.0 under Common Test Conditions (CTC) V7. The experimental results show that when the worst-case decoder latency is minimized to 2048 luma samples, the coding loss can be kept to minimal with an average loss of 0.02% for the YUV components in random access configurations with no change in encoder and decoder timings. Jayasingam Adhuran, Madhu Peringassery Krishnan, Qingyang Zhou, Shan Liu 0001 |
VCIP | 5 |
| 2025 | Improved Intra Block Copy Mode Beyond AV1abstractIntra Block Copy (IntraBC) is a key coding tool for screen content video, enabling blocks to reference previously reconstructed regions within the same frame. While IntraBC was introduced in AV1 and further enhanced in VVC, the AVM development framework offers new opportunities to improve its coding efficiency and hardware compatibility. This paper proposes three enhancements adopted into the AVM reference software: (1) a unified and extended local reference buffer design supporting multiple superblock sizes with fixed memory constraints; (2) a decoupled global-local search strategy that improves block vector prediction efficiency; and (3) the integration of intra Block-Adaptive Weighted Prediction (BAWP) into local IntraBC. Experimental results show significant screen content coding gains: -3.53%, -2.25%, -1.43% (YUV-PSNR) for all intra, random access and low delay, respectively, with minimal runtime overhead. These contributions have been adopted into the AVM standard and reference software. Qingyang Zhou, Wei Kuang, Madhu Peringassery Krishnan, Jayasingam Adhuran, Shan Liu 0001 |
VCIP | 1 |
| 2023 | GPCGC: A Green Point Cloud Geometry Coding MethodabstractA low-complexity point cloud compression method called the Green Point Cloud Geometry Codec (GPCGC), is proposed to encode the 3D spatial coordinates of static point clouds efficiently. GPCGC consists of two modules. In the first module, point coordinates of input point clouds are hierarchically organized into an octree structure. Points at each leaf node are projected along one of three axes to yield image maps. In the second module, the occupancy map is clustered into 9 modes while the depth map is coded by a low-complexity high-efficiency image codec, called the green image codec (GIC). GIC is a multi-resolution codec based on vector quantization (VQ). Its complexity is significantly lower than HEVC-Intra. Furthermore, the rate-distortion optimization (RDO) technique is used to select the optimal coding parameters. GPCGC is a progressive codec, and it offers a coding performance competitive with MPEG's V-PCC and G-PCC standards at significantly lower complexity. Qingyang Zhou, Shan Liu 0001, C.-C. Jay Kuo |
ICIP | 1 |
| 2023 | Line-based self-referencing string prediction technique for screen content coding in AVS3
Liping Zhao 0005, Qingyang Zhou, Keli Hu, Sheng Feng, Kailun Zhou, Weixing Wang 0001, Tao Lin 0005 |
Multim. Tools Appl. | 2 |
| 2022 | Non-Distinguishable Inconsistencies as a Deterministic Oracle for Detecting Security BugsabstractSecurity bugs like memory errors are constantly introduced to software programs, and recent years have witnessed an increasing number of reported security bugs. Traditional detection approaches are mainly specification-based---detecting violations against a specified rule as security bugs. This often does not work well in practice because specifications are difficult to specify and generalize, leaving complicated and new types of bugs undetected. Recent research thus leans toward deviation-based detection which finds a substantial number of similar cases and detects deviating cases as potential bugs. This, however, suffers from two other problems. First, it requires enough similar cases to find deviations and thus cannot work for custom code that does not have similar cases. Second, code-similarity analysis is probabilistic and challenging, so the detection can be unreliable. Sometimes, similar cases can normally have deviating behaviors under different contexts. Qingyang Zhou, Qiushi Wu, Dinghao Liu, Shouling Ji, Kangjie Lu |
CCS | 1 |
| 2022 | PCRP: Unsupervised Point Cloud Object Retrieval and Pose EstimationabstractAn unsupervised point cloud object retrieval and pose estimation method, called PCRP, is proposed in this work. It is assumed that there exists a gallery point cloud set that contains point cloud objects with given pose orientation information. PCRP attempts to register the unknown point cloud object with those in the gallery set so as to achieve content-based object retrieval and pose estimation jointly, where the point cloud registration task is built upon an enhanced version of the unsupervised R-PointHop method. Experiments on the ModelNet40 dataset demonstrate the superior performance of PCRP in comparison with traditional and learning based methods. Pranav Kadam, Qingyang Zhou, Shan Liu 0001, C.-C. Jay Kuo |
ICIP | 2 |
| 2022 | Joint-attention feature fusion network and dual-adaptive NMS for object detection
Wentao Ma 0003, Tongqing Zhou, Jiaohua Qin, Qingyang Zhou, Zhiping Cai |
Knowl. Based Syst. | 4 |
| 2022 | MOLS-Net: Multi-organ and lesion segmentation network based on sequence feature pyramid and attention mechanism for aortic dissection diagnosis
Qingyang Zhou, Jiaohua Qin, Xuyu Xiang, Yun Tan |
Knowl. Based Syst. | 1 |
| 2021 | An Intra String Copy Approach for SCC in AVS3abstractAn efficient SCC tool named Intra String Copy (ISC) has been proposed and adopted in AVS3 recently. ISC has two CU-level sub-modes: FPSP (fully-matching-string and partially-matching-string based string prediction) sub-mode and EUSP (equal-value-string, unit-basis-vector-string and unmatched-pixel-string based string prediction) sub-mode. Compared with the latest AVS3 reference software HPM with SCC tools disabled, using AVS3 SCC Common Test Condition and YUV test sequences in text and graphics with motion (TGM) and mixed content (MC) categories, the proposed tool achieves an average Y BD-rate reduction of 57.7%/39.5% and 77.2%/57.9% for TGM and MC in All Intra (AI)/Low Delay B(LDB) configurations, respectively, with low additional encoding complexity and almost the same decoding complexity. Liping Zhao 0005, Kailun Zhou, Qingyang Zhou, Tao Lin 0005 |
VCIP | 3 |
| 2021 | String Prediction for 4: 2: 0 Format Screen Content Coding and Its Implementation in AVS3abstractIn the past, string prediction (also known as string matching) was applied only to RGB and YUV 4:4:4 format screen content coding. This paper proposes a string prediction approach to 4:2:0 format screen content coding implemented in the third generation of Audio Video Standard (AVS3) in China. String prediction is applied to both YUV CU and Y CU. To further improve the coding performance, several improved technicals of string prediction are presented, including a mixed string searching strategy for finding the optimal reference string, a joint picture-level, CU-level, and pixel-level early termination strategy to reduce coding complexity, and two effective coding methods for string prediction parameters. For low-complexity hardware implementation of string prediction decoder, the memory access bandwidth is reduced by introducing string constraints. Meanwhile, string prediction reuses the reference pixel buffer of intra block copy (IBC). Compared with the newest AVS3 reference software HPM7.0 with string prediction disabled, the proposed string prediction approach achieves up to 18.48% Y BD-rate reduction. Using AVS3 Screen Content Coding (SCC) Common Test Condition and YUV test sequences in Text and Graphics with Motion category, the proposed technique achieves an average Y BD-rate reduction of 10.33%, 8.47%, 6.91% for All Intra (AI), Random Access (RA) and Low Delay (LD) configurations, respectively, with low additional encoding and decoding complexity. The proposed string prediction approach has been adopted in the newest AVS3 reference software HPM7.0. Qingyang Zhou, Liping Zhao 0005, Kailun Zhou, Tao Lin 0005, Shuhui Wang, Mengcao Jiao |
IEEE Trans. Multim. | 1 |