VLDB 2026 Research / reviewers in the wild / expert
Chao Song 0001
dblp:59/1815-1
· DBLP profile ↗
16ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0002-9415-4929ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 4 since 2021Theory of computation · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OctMamba: Mamba-based octree context entropy model for point cloud geometry compressionabstractExisting learned point cloud compression frameworks face two major limitations: (1) they focus almost exclusively on spatial redundancy and (2) rely on architectures built around local-global transformers or global Mamba blocks. Transformers incur quadratic complexity, while global Mamba lacks the granularity to capture structured correlations across multiple dimensions. We propose OctMamba, the first unified framework to jointly exploit spatial, channel, and topological redundancies, dimensions previously overlooked in point cloud geometry compression. Our approach introduces a new architectural principle: embedding Mamba modules within specialized subcomponents rather than applying them globally, challenging existing design paradigms. OctMamba combines two modules: Spatial-Channel Coupled Grouping Mamba (SCCGM) for spatial-channel fusion and Local Graph CNN-Mamba (LGCM) for topological encoding. This design enables efficient long-range modeling with linear complexity, delivering a smaller model and faster decoding while outperforming transformer-based and global Mamba baselines. On SemanticKITTI, OctMamba reduces bitrate by 60.2% over GPCC (D1 PSNR) and achieves state-of-the-art performance across LiDAR and dynamic human point cloud benchmarks with practical speed and scalability. By introducing multi-dimensional redundancy modeling, OctMamba has the potential to influence future research on efficient point cloud compression. The code is available at https://github.com/ZjgsVMC/OctMamba . Zhaoyi Jiang, Frederick W. B. Li, Gary K. L. Tam, Chao Song 0001, Bailin Yang |
Pattern Recognit. | 5 |
| 2025 | 3D data augmentation and dual-branch model for robust face forgery detectionabstractWe propose Dual-Branch Network (DBNet), a novel deepfake detection framework that addresses key limitations of existing works by jointly modeling 3D-temporal and fine-grained texture representations. Specifically, we aim to investigate how to (1) capture dynamic properties and spatial details in a unified model and (2) identify subtle inconsistencies beyond localized artifacts through temporally consistent modeling. To this end, DBNet extracts 3D landmarks from videos to construct temporal sequences for an RNN branch, while a Vision Transformer analyzes local patches. A Temporal Consistency-aware Loss is introduced to explicitly supervise the RNN. Additionally, a 3D generative model augments training data. Extensive experiments demonstrate our method achieves state-of-the-art performance on benchmarks, and ablation studies validate its effectiveness in generalizing to unseen data under various manipulations and compression. Changshuang Zhou, Frederick W. B. Li, Chao Song 0001, Bailin Yang |
Graph. Model. | 3 |
| 2025 | WDFSR: Normalizing Flow Based on the Wavelet-Domain for Super-ResolutionabstractWe propose a normalizing flow based on the wavelet framework for super-resolution (SR) called WDFSR. It learns the conditional distribution mapping between low-resolution images in the RGB domain and high-resolution images in the wavelet domain to simultaneously generate high-resolution images of different styles. To address the issue of some flow-based models being sensitive to datasets, which results in training fluctuations that reduce the mapping ability of the model and weaken generalization, we designed a method that combines a T-distribution and QR decomposition layer. Our method alleviates this problem while maintaining the ability of the model to map different distributions and produce higher-quality images. Good contextual conditional features can promote model training and enhance the distribution mapping capabilities for conditional distribution mapping. Therefore, we propose a Refinement layer combined with an attention mechanism to refine and fuse the extracted condition features to improve image quality. Extensive experiments on several SR datasets demonstrate that WDFSR outperforms most general CNN- and flow-based models in terms of PSNR value and perception quality. We also demonstrated that our framework works well for other low-level vision tasks, such as low-light enhancement. The pretrained models and source code with guidance for reference are available at https://github.com/Lisbegin/WDFSR. Chao Song 0001, Shaobang Li, Frederick W. B. Li, Bailin Yang |
Comput. Vis. Media | 1 |
| 2024 | Non-autoregressive transformer with fine-grained optimization for user-specified indoor layout
Chao Song 0001, Shujie Chen 0001, Zhaoyi Jiang, Bailin Yang |
Eng. Appl. Artif. Intell. | 1 |
| 2024 | Multi-feature fusion enhanced monocular depth estimation with boundary awareness
Chao Song 0001, Qingjie Chen, Frederick W. B. Li, Zhaoyi Jiang, Yuliang Shen, Bailin Yang |
Vis. Comput. | 1 |
| 2024 | CVAE-LAYOUT: automatic furniture layout with constraints
Yixin Xuan, Chao Song 0001, Jianqiu Jin, Bailin Yang |
Vis. Comput. | 2 |
| 2023 | HSE: Hybrid Species Embedding for Deep Metric LearningabstractDeep metric learning is crucial for finding an embedding function that can generalize to training and testing data, including unknown test classes. However, limited training samples restrict the model’s generalization to downstream tasks. While adding new training samples is a promising solution, determining their labels remains a significant challenge. Here, we introduce Hybrid Species Embedding (HSE), which employs mixed sample data augmentations to generate hybrid species and provide additional training signals. We demonstrate that HSE outperforms multiple state-of-the-art methods in improving the metric Recall@K on the CUB-200 , CAR-196 and SOP datasets, thus offering a novel solution to deep metric learning’s limitations. Bailin Yang, Haoqiang Sun, Frederick W. B. Li, Jianlu Cai, Chao Song 0001 |
ICCV | 6 |
| 2023 | Transformer-Based Video Deinterlacing Method
Chao Song 0001, Zhaoyi Jiang, Bailin Yang |
ICONIP (5) | 1 |
| 2023 | C2SPoint: A classification-to-saliency network for point cloud saliency detectionabstractPoint cloud saliency detection is an important technique that support downstream tasks in 3D graphics and vision, like 3D model simplification, compression, reconstruction and viewpoint selection. Existing approaches often rely on hand-crafted features and are only applicable to specific datasets. In this paper, we propose a novel weakly supervised classification network, called C2SPoint, which directly performs saliency detection on the point clouds. Unlike previous methods that require per-point saliency annotations, C2SPoint only requires category labels of the point clouds during training. The network consists of two branches: a Classification branch and a Saliency branch. The former branch is composed of two Adaptive Set Abstraction layers for feature extraction and a Saliency Transform layer for learning saliency knowledge from the classification network. The latter branch introduces a multi-scale point-cluster similarity matrix for propagating the cluster saliency to each point within it, resulting in the prediction of point-level saliency. Experimental results demonstrate the effectiveness of our method in point cloud saliency detection, with improvements of 2% in both AUC and NSS compared to state-of-the-art methods. Zhaoyi Jiang, Luyun Ding, Gary K. L. Tam, Chao Song 0001, Frederick W. B. Li, Bailin Yang |
Comput. Graph. | 4 |
| 2021 | Automatic interior layout with user-specified furniture
Bailin Yang, Liuliu Li, Chao Song 0001, Zhaoyi Jiang |
Comput. Graph. | 3 |
| 2020 | Joint temporal context exploitation and active learning for video segmentation
Guohua Cheng, Judith Gelernter, Shihao Yu, Chao Song 0001, Bailin Yang |
Pattern Recognit. | 5 |
| 2019 | Automatic Furniture Layout Based on Functional Area DivisionabstractWe propose an automatic indoor furniture layout scheme based on functional area division and furniture filling. According to the function, we suppose each kind of furniture may be laid out in one or several functional areas, for example, a sofa may be located in the meeting area of a living room and a bed may be located in the sleeping area of a bedroom, etc.. Our automatic layout method divides an empty room region into several functional areas by using conditional generative adversarial networks (CGAN). We expound the learning process of the algorithm in the process of functional areas division, including the objective function construct and training process. Moreover, in order to fill furniture into a specific functional area, a learning-based furniture filling algorithm is proposed by training a fully connected network model for different kinds of functional area. Experiments show our automatic furniture layout method has its advantages in performance and effect compared with the existing methods. Bailin Yang, Liuliu Li, Chao Song 0001, Zhaoyi Jiang |
CW | 3 |
| 2019 | Compressed dynamic mesh sequence for progressive streamingabstractAbstract Dynamic mesh sequence (DMS) is a simple and accurate representation for precisely recording a 3D animation sequence. Despite its simplicity, this representation is typically large in data size, making storage and transmission expensive. This paper presents a novel framework that allows effective DMS compression and progressive streaming by eliminating spatial and temporal redundancy. To explore temporal redundancy, we propose a temporal frame‐clustering algorithm to organize DMS frames by their motion trajectory changes, eliminating intracluster redundancy by principal component analysis dimensionality reduction. To eliminate spatial redundancy, we propose an algorithm to transform the coordinates of mesh vertex trajectory into a decorrelated trajectory space, generating a new spatially nonredundant trajectory representation. We finally apply a spectral graph wavelet transform with color set partitioning embedded block encoding to turn the resultant DMS into a multiresolution representation to support progressive streaming. Experiment results show that our method outperforms several existing methods in terms of storage requirement and reconstruction quality. Bailin Yang, Zhaoyi Jiang, Jiantao Shangguan, Frederick W. B. Li, Chao Song 0001, Yibo Guo, Mingliang Xu 0001 |
Comput. Animat. Virtual Worlds | 5 |
| 2014 | Fast corotational simulation for example-driven deformation
Chao Song 0001, Hongxin Zhang 0001, Xun Wang 0007, Jianwei Han, Huiyan Wang 0002 |
Comput. Graph. | 1 |
| 2008 | Cutting and Fracturing Models without Remeshing
Chao Song 0001, Hongxin Zhang 0001, Hujun Bao |
GMP | 1 |
| 2008 | Space-Time Curve Analogies for Motion Editing
Hongxin Zhang 0001, Chao Song 0001, Hujun Bao |
GMP | 3 |