VLDB 2026 Research / reviewers in the wild / expert
Haocheng Tang
dblp:219/2232
· DBLP profile ↗
7ranked-venue papers
2as first author
6since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient Feedforward Human-Centric Video Compression via 3D Gaussian GenerationabstractIn this paper, we propose a feed-forward framework (Fig. 1) for human video compression based on 3D generative reconstruction. Recent surveys [1], [2] highlight the need for more efficient and semantically aligned solutions. Our approach disentangles video content into complementary structural and motion layers: the structural layer encodes regularized texture from a single frame, while the motion layer leverages the SMPL-X prior to represent complex dynamics with a compact set of pose and shape parameters. A hierarchical coding scheme exploits the heterogeneity of these representations for improved efficiency. After decoding, a feed-forward 3D reconstruction pipeline with facial feature extraction is employed, in which a multimodal transformer and a Gaussian head synthesize parametric cues that are fused with motion signals for accurate animation and high-fidelity rendering. Experiments show over$1000 \times$compression while preserving structural and semantic fidelity. The method consistently outperforms strong baselines (especially at$0.04-0.1 \text{kbpp})$, with significant gains in rate-distortion, FVD, and perceptual quality, as well as robust generalization across identities and scenes. Haocheng Tang, Ruoke Yan, Jiaqi Zhang 0007, Siwei Ma 0001 |
DCC | 1 |
| 2026 | Lightweight CNN-Based In-Loop Filtering for Video Coding with Hardware-Aware OptimizationsabstractNeural network-based in-loop filtering significantly enhances video compression efficiency. However, high computational complexity hinders their deployment in real-time and ultra-high-definition scenarios. To address this, we propose a lightweight CNN-based in-loop filter for the luma component. In terms of model design, we utilize a U-Net-like architecture to learn the residual signal, incorporating depthwise separable 3 × 3 convolutions and 1 × 1 convolutions to reduce computational complexity, which results in a low complexity of only$37.707 \text{kMACs} /$pixel. For deployment optimization, we implement memory linearization to improve cache efficiency and combine blocked matrix multiplication with SIMD to maximize parallelism, ensuring cross-platform compatibility without third-party libraries. Experimental results on AVS4 EVM-0.9 (All-Intra) on a CPU platform show that the proposed method achieves BD-rate reductions of$1.36 \%, 0.35 \%$, and 0.34% for$\mathrm{Y}, \mathrm{U}$, and V components, respectively. Furthermore, the optimizations lead to a 91.5% reduction in decoding time, resulting in a decoding complexity of 7757% compared to the anchor. Yanchen Zhao, Xuewei Meng, Jiaqi Zhang 0007, Haocheng Tang, Lin Li 0062, Siwei Ma 0001 |
DCC | 5 |
| 2026 | Fusion-regularized alignment modality-adaptive audio-visual network for audio-visual zero-shot learning
Siteng Ma, Haocheng Tang, Jisheng Chu, Wenrui Li 0001 |
Neurocomputing | 3 |
| 2025 | HGC-Avatar: Hierarchical Gaussian Compression for Streamable Dynamic 3D AvatarsabstractRecent advances in 3D Gaussian Splatting (3DGS) have enabled fast, photorealistic rendering of dynamic 3D scenes, showing strong potential in immersive communication. However, in digital human encoding and transmission, the compression methods based on general 3DGS representations are limited by the lack of human priors, resulting in suboptimal bitrate efficiency and reconstruction quality at the decoder side, which hinders their application in streamable 3D avatar systems. We propose HGC-Avatar, a novel Hierarchical Gaussian Compression framework designed for efficient transmission and high-quality rendering of dynamic avatars. Our method disentangles the Gaussian representation into a structural layer, which maps poses to Gaussians via a StyleUNet-based generator, and a motion layer, which leverages the SMPL-X model to represent temporal pose variations compactly and semantically. This hierarchical design supports layer-wise compression, progressive decoding, and controllable rendering from diverse pose inputs such as video sequences or text. Since people are most concerned with facial realism, we incorporate a facial attention mechanism during StyleUNet training to preserve identity and expression details under low-bitrate constraints. Experimental results demonstrate that HGC-Avatar provides a streamable solution for rapid 3D avatar rendering, while significantly outperforming prior methods in both visual quality and compression efficiency. Haocheng Tang, Ruoke Yan, Xinhui Yin, Qi Zhang 0042, Xinfeng Zhang 0001, Siwei Ma 0001, Wen Gao 0001, Chuanmin Jia |
ACM Multimedia | 1 |
| 2025 | NNKcat: deep neural network to predict catalytic constants (Kcat) by integrating protein sequence and substrate structure with enhanced data imbalance handlingabstractCatalytic constant (Kcat) is to describe the efficiency of catalyzing reactions. The Kcat value of an enzyme-substrate pair indicates the rate an enzyme converts saturated substrates into product during the catalytic process. However, it is challenging to construct robust prediction models for this important property. Most of the existing models, including the one recently published by Nature Catalysis (Li et al.), are suffering from the overfitting issue. In this study, we proposed a novel protocol to construct Kcat prediction models, introducing an intermedia step to separately develop substrate and protein processors. The substrate processor leverages analyzing Simplified Molecular Input Line Entry System (SMILES) strings using a graph neural network model, attentive FP, while the protein processor abstracts protein sequence information utilizing long short-term memory architecture. This protocol not only mitigates the impact of data imbalance in the original dataset but also provides greater flexibility in customizing the general-purpose Kcat prediction model to enhance the prediction accuracy for specific enzyme classes. Our general-purpose Kcat prediction model demonstrates significantly enhanced stability and slightly better accuracy (R2 value of 0.54 versus 0.50) in comparison with Li et al.'s model using the same dataset. Additionally, our modeling protocol enables personalization of fine-tuning the general-purpose Kcat model for specific enzyme categories through focused learning. Using Cytochrome P450 (CYP450) enzymes as a case study, we achieved the best R2 value of 0.64 for the focused model. The high-quality performance and expandability of the model guarantee its broad applications in enzyme engineering and drug research & development. Jingchen Zhai, Xiguang Qi, Lianjin Cai, Haocheng Tang, Junmei Wang |
Briefings Bioinform. | 5 |
| 2023 | A Global View-Guided Autoregressive Residual Network for Irregular Time Series Classification
Jianping Zhu 0002, Haocheng Tang, Liang Zhang 0031, Bo Jin 0001, Xiaopeng Wei |
PAKDD (4) | 2 |
| 2018 | Deep Sparse Informative Transfer SoftMax for Cross-Domain Image Classification
Hanfang Yang, Bo Yao 0007, Zijing Tan, Haocheng Tang, Yingjie Tian 0002 |
DASFAA (2) | 6 |