Jin Zeng 0004

dblp:52/331-4 · DBLP profile ↗
← Back
20ranked-venue papers
5as first author
8since 2021 · last 2026
0000-0002-0180-7733ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Adaptive Spatial-Channel Masked Reconstruction Knowledge Distillation for Dense Prediction
Ziniu Liu, Shuheng Zhou 0001, Jin Zeng 0004, Mingqing Liu 0002, Hao Deng 0002
ICMR3
2025 Consistent Time-of-Flight Depth Denoising via Graph-Informed Geometric Attention
Weida Wang, Changyong He, Jin Zeng 0004, Di Qiu
ICCV3
2025 SAVE-GSL: Scalable and Expressive Graph Structure Learning for Large Graphs
abstract
Graph structure learning (GSL) has emerged as a promising approach for optimizing graph structures to enhance downstream task performance. However, the quadratic complexity of GSL renders it impractical for large graphs, such as social networks. While several attempts have been made to mitigate the scalability issue of GSL, they struggle to ensure efficiency and expressiveness simultaneously. In this paper, we propose a novel scalable and expressive graph structure learning framework (SAVE-GSL), that models all-pair interactions with linear complexity. Specifically, we design cluster and expander sparse patterns for efficient expressiveness enhancement, and adaptively fuse them to generate the optimized graph. Moreover, leveraging these sparse patterns, we theoretically prove that SAVE-GSL is an efficient universal approximator of permutation-equivariant functions, providing a formal justification for its superior expressiveness. Extensive experiments demonstrate that SAVE-GSL outperforms state-of-the-art schemes in both efficiency and accuracy on social network datasets and other graph datasets of various scales.
Manxin Xu, Shengjie Zhao 0001, Jin Zeng 0004, Weichao Chen 0001, Shilong Dong
ICME3
2025 SpatialGeo: Boosting Spatial Reasoning in Multimodal LLMs via Geometry-Semantics Fusion
abstract
Multimodal large language models (MLLMs) have achieved significant progress in image and language tasks due to the strong reasoning capability of large language models (LLMs). Nevertheless, most MLLMs suffer from limited spatial reasoning ability to interpret and infer spatial arrangements in three-dimensional space. In this work, we propose a novel vision encoder based on hierarchical fusion of geometry and semantics features, generating spatial-aware visual embedding and boosting the spatial grounding capability of MLLMs. Specifically, we first unveil that the spatial ambiguity shortcoming stems from the lossy embedding of the vision encoder utilized in most existing MLLMs (e.g., CLIP), restricted to instance-level semantic features. This motivates us to complement CLIP with the geometry features from vision-only self-supervised learning via a hierarchical adapter, enhancing the spatial awareness in the proposed SpatialGeo. The network is efficiently trained using pretrained LLaVA model and optimized with random feature dropping to avoid trivial solutions relying solely on the CLIP encoder. Experimental results show that SpatialGeo improves the accuracy in spatial reasoning tasks, enhancing state-of-the-art models by at least 8.0% in SpatialRGPT-Bench with ∼50% less memory cost during inference. The source code is available via https://ricky-plus.github.io/SpatialGeoPages/.
Jiajie Guo, Qingpeng Zhu, Jin Zeng 0004, Changyong He, Weida Wang
MMSP3
2025 Deep Unrolled Weighted Graph Laplacian Regularization for Depth Completion
Jin Zeng 0004, Qingpeng Zhu, Tongxuan Tian, Wenxiu Sun, Lin Zhang 0014, Shengjie Zhao 0001
Int. J. Comput. Vis.1
2025 Deep Unrolled Graph Laplacian Regularization for Robust Time-of-Flight Depth Denoising
abstract
Depth images captured by Time-of-Flight (ToF) sensors are subject to severe noise. Recent approaches based on deep neural networks achieve good depth denoising performance in synthetic data, but the application to real-world data is limited, due to the complexity of actual depth noise characteristics and the difficulty in acquiring ground truth. In this paper, we propose a novel ToF depth denoising network based on unrolled graph Laplacian regularization to “robustify” the network against both noise complexity and dataset deficiency. Unlike previous schemes that are ignorant of underlying ToF imaging mechanism, we formulate a fidelity term in the optimization problem to adapt to the depth probabilistic distribution with spatially-varying noise variance. Then, we add quadratic graph Laplacian regularization as the smoothness prior, leading to a maximum a posteriori problem that is optimized efficiently by solving a linear system of equations. We unroll the solution into iterative filters so that parameters used in the optimization and graph construction are amendable to data-driven tuning. Because the resulting network is built using domain knowledge of ToF imaging principle and graph prior, it is robust against overfitting to synthetic training data. Experimental results demonstrate that the proposal outperforms existing schemes in ToF depth denoising on synthetic FLAT dataset and generalizes well to real Kinectv2 dataset.
Jingwei Jia, Changyong He, Gene Cheung, Jin Zeng 0004
IEEE Signal Process. Lett.5
2024 CASTNet: Convolution Augmented Graph Sampling Transformer Network for Traffic Flow Forecasting
abstract
Accurate and efficient traffic flow forecasting plays an essential role in urban management which is conductive to traffic safety and travel experience. Despite the tremendous progress of attention-based traffic flow forecasting methods, complicated spatial and temporal correlations of traffic data pose difficulty for the network in simultaneously capturing the long-range contexts and the local feature details, leading to limited performance. To this end, we propose Convolution Augmented Graph Sampling Transformer Network (CASTNet) for accurate traffic flow prediction. Specifically, we design a hybrid graph transformer module to utilize the graph convolutions to perceive local structural details, which are incorporated into the transformer to augment the attention computation for more accurate global dependency modeling. Moreover, we develop a sampling-based attention layer based on the approximated sampling strategy to further optimize the graph structure, which is implemented with linear complexity. Finally, the spatial features from the graph transformer module are fused with the temporal features from the sample convolution and interaction network. Experimental results on three real-world traffic flow datasets demonstrate the superiority of the proposed CASTNet over existing schemes.
Shengjie Zhao 0001, Jin Zeng 0004, Shilong Dong, Geyunqian Zu
CSCWD3
2024 SEA-GNN: Sequence Extension Augmented Graph Neural Network for Sequential Recommendation
abstract
Sequential recommendation aims to anticipate the next preference of users by examining their recent interactions. Recently, graph neural networks (GNNs) have been widely utilized in sequential recommendation, but existing schemes focus on interactions within individual sequences and tend to connect irrelevant items in case of insufficient historical data. In this work, we propose Sequence Extension Augmented GNN (SEA-GNN) which augments the node representation learning with inter-sequence global context aggregation while maintaining intra-sequence local preference. Specifically, we augment the graph construction with sequence extension that diversifies the item connections to exploit global context and robustify the node representation against data insufficiency. Meanwhile, we extract local preference based on the intra-sequence user-item graph to enhance the node representation with user-specific interest. Experimental results demonstrate the superiority of the proposed algorithm compared with existing schemes in recommendation performance.
Geyunqian Zu, Shengjie Zhao 0001, Jin Zeng 0004, Shilong Dong
ICASSP3
2020 Towards Geometry Guided Neural Relighting with Flash Photography
abstract
Previous image based relighting methods require capturing multiple images to acquire high frequency lighting effect under different lighting conditions, which needs nontrivial effort and may be unrealistic in certain practical use scenarios. While such approaches rely entirely on cleverly sampling the color images under different lighting conditions, little has been done to utilize geometric information that crucially influences the high-frequency features in the images, such as glossy highlight and cast shadow. We therefore propose a framework for image relighting from a single flash photograph with its corresponding depth map using deep learning. By incorporating the depth map, our approach is able to extrapolate realistic high-frequency effects under novel lighting via geometry guided image decomposition from the flashlight image, and predict the cast shadow map from the shadow-encoding transformed depth map. Moreover, the single-image based setup greatly simplifies the data capture process. We experimentally validate the advantage of our geometry guided approach over state-of-the-art image-based approaches in intrinsic image decomposition and image relighting, and also demonstrate our performance on real mobile phone photo examples.
Di Qiu, Jin Zeng 0004, Zhanghan Ke, Wenxiu Sun, Chengxi Yang
3DV2
2020 Deep Surface Normal Estimation on the 2-Sphere with Confidence Guided Semantic Attention
Quewei Li, Jie Guo 0001, Qinyu Tang, Wenxiu Sun, Jin Zeng 0004, Yanwen Guo 0001
ECCV (24)6
2020 3D Point Cloud Denoising Using Graph Laplacian Regularization of a Low Dimensional Manifold Model
abstract
3D point cloud-a new signal representation of volumetric objects-is a discrete collection of triples marking exterior object surface locations in 3D space. Conventional imperfect acquisition processes of 3D point cloud-e.g., stereo-matching from multiple viewpoint images or depth data acquired directly from active light sensors-imply non-negligible noise in the data. In this paper, we extend a previously proposed low-dimensional manifold model for the image patches to surface patches in the point cloud, and seek self-similar patches to denoise them simultaneously using the patch manifold prior. Due to discrete observations of the patches on the manifold, we approximate the manifold dimension computation defined in the continuous domain with a patch-based graph Laplacian regularizer, and propose a new discrete patch distance measure to quantify the similarity between two same-sized surface patches for graph construction that is robust to noise. We show that our graph Laplacian regularizer leads to speedy implementation and has desirable numerical stability properties given its natural graph spectral interpretation. Extensive simulation results show that our proposed denoising scheme outperforms state-of-the-art methods in objective metrics and better preserves visually salient structural features like edges.
Jin Zeng 0004, Gene Cheung, Michael Kwok-Po Ng, Jiahao Pang, Cheng Yang 0003
IEEE Trans. Image Process.1
2019 Deep Surface Normal Estimation With Hierarchical RGB-D Fusion
abstract
The growing availability of commodity RGB-D cameras has boosted the applications in the field of scene understanding. However, as a fundamental scene understanding task, surface normal estimation from RGB-D data lacks thorough investigation. In this paper, a hierarchical fusion network with adaptive feature re-weighting is proposed for surface normal estimation from a single RGB-D image. Specifically, the features from color image and depth are successively integrated at multiple scales to ensure global surface smoothness while preserving visually salient details. Meanwhile, the depth features are re-weighted with a confidence map estimated from depth before merging into the color branch to avoid artifacts caused by input depth corruption. Additionally, a hybrid multi-scale loss function is designed to learn accurate normal estimation given noisy ground-truth dataset. Extensive experimental results validate the effectiveness of the fusion strategy and the loss design, outperforming state-of-the-art normal estimation schemes.
Jin Zeng 0004, Yanfeng Tong, Yunmu Huang, Qiong Yan, Wenxiu Sun, Jing Chen 0018, Yongtian Wang
CVPR1
2018 Zoom and Learn: Generalizing Deep Stereo Matching to Novel Domains
abstract
Despite the recent success of stereo matching with convolutional neural networks (CNNs), it remains arduous to generalize a pre-trained deep stereo model to a novel domain. A major difficulty is to collect accurate ground-truth disparities for stereo pairs in the target domain. In this work, we propose a self-adaptation approach for CNN training, utilizing both synthetic training data (with ground-truth disparities) and stereo pairs in the new domain (without ground-truths). Our method is driven by two empirical observations. By feeding real stereo pairs of different domains to stereo models pre-trained with synthetic data, we see that: i) a pre-trained model does not generalize well to the new domain, producing artifacts at boundaries and ill-posed regions; however, ii) feeding an up-sampled stereo pair leads to a disparity map with extra details. To avoid i) while exploiting ii), we formulate an iterative optimization problem with graph Laplacian regularization. At each iteration, the CNN adapts itself better to the new domain: we let the CNN learn its own higher-resolution output; at the meanwhile, a graph Laplacian regularization is imposed to discriminatively keep the desired edges while smoothing out the artifacts. We demonstrate the effectiveness of our method in two domains: daily scenes collected by smart-phone cameras, and street views captured in a driving car.
Jiahao Pang, Wenxiu Sun, Chengxi Yang, Jimmy S. J. Ren, Ruichao Xiao, Jin Zeng 0004, Liang Lin 0004
CVPR6
2016 Subpixel-Based Image Scaling for Grid-like Subpixel Arrangements: A Generalized Continuous-Domain Analysis Model
abstract
Subpixel-based image scaling can improve the apparent resolution of displayed images by controlling individual subpixels rather than whole pixels. However, improved luminance resolution brings chrominance distortion, making it crucial to suppress color error while maintaining sharpness. Moreover, it is challenging to develop a scheme that is applicable for various subpixel arrangements and for arbitrary scaling factors. In this paper, we address the aforementioned issues by proposing a generalized continuous-domain analysis model, which considers the low-pass nature of the human visual system (HVS). Specifically, given a discrete image and a grid-like subpixel arrangement, the signal perceived by the HVS is modeled as a 2D continuous image. Minimizing the difference between the perceived image and the continuous target image leads to the proposed scheme, which we call continuous-domain analysis for subpixel-based scaling (CASS). To eliminate the ringing artifacts caused by the ideal low-pass filtering in CASS, we propose an improved scheme, which we call CASS with Laplacian-of-Gaussian filtering. Experiments show that the proposed methods provide sharp images with negligible color fringing artifacts. Our methods are comparable with the state-of-the-art methods when applied on the RGB stripe arrangement, and outperform existing methods when applied on other subpixel arrangements.
Jiahao Pang, Lu Fang 0001, Jin Zeng 0004, Yuanfang Guo, Ketan Tang
IEEE Trans. Image Process.3
2016 Subpixel Image Quality Assessment Syncretizing Local Subpixel and Global Pixel Features
abstract
The subpixel rendering technology increases the apparent resolution of an LCD/OLED screen by exploiting the physical property that a pixel is composed of RGB individually addressable subpixels. Due to the intrinsic intercoordination between apparent luminance resolution and color fringing artifact, a common method of subpixel image assessment is subjective evaluation. In this paper, we propose a unified subpixel image quality assessment metric called subpixel image assessment (SPA), which syncretizes local subpixel and global pixel features. Specifically, comprehensive subjective studies are conducted to acquire data of user preferences. Accordingly, a collection of low-level features is designed under extensive perceptual validation, capturing subpixel and pixel features, which reflect local details and global distance from the original image. With the features and their measurements as the basis, the SPA is obtained, which leads to a good representation of the subpixel image characteristics. The experimental results justify the effectiveness and the superiority of the SPA. The SPA is also successfully adopted in a variety of applications, including content adaptive sampling and metric-guided image compression.
Jin Zeng 0004, Lu Fang 0001, Jiahao Pang, Houqiang Li, Feng Wu 0001
IEEE Trans. Image Process.1
2015 Image colorization via color propagation and rank minimization
abstract
Image colorization aims to add colors to grayscale images, which used to be a time-consuming and tedious task that requires lots of human efforts. In this paper, we present a novel colorization method based on color propagation and rank minimization. Given a small portion of chrominance values and a grayscale image, we firstly propagate the known color values to other pixels to be colorized. As the colorized image after color propagation is not accurate, we then define a confidence matrix to measure the propagation fidelity. Finally, pixels that have propagated chrominance values with confidence are colorized by rank minimization, which exploits the redundancy of natural images. Experimental results on real data set show that our proposed method achieves state-of-the-art colorization quality.
Yonggen Ling, Oscar C. Au, Jiahao Pang, Jin Zeng 0004, Yuan Yuan 0002, Amin Zheng
ICIP4
2014 Analysis of sampling pattern and Luma-Chroma filter design for subpixel-based image downsampling
abstract
Subpixel-based image downsampling is attractive in that it produces higher apparent resolution of down-sampled images on LCD displays. However increased luminance resolution is achieved at the price of color fringing artifacts. In this paper, we propose an algorithm to find a pleasing balance between increased resolution and color fidelity. We separate the subpixel-based downsampling into two stages, shifting followed by downsampling with anti-aliasing filtering. In stage one, we find special characteristics of the luminance and chrominance spectra of the shifted image, based on which the optimal sampling pattern is found. In stage two, anti-aliasing filters for luminance and chrominance are designed respectively. Experimental results verify that the proposed method manages to suppress color artifacts while maintaining high luminance sharpness.
Jin Zeng 0004, Oscar C. Au, Yuanfang Guo, Jiahao Pang, Ketan Tang, Yonggen Ling
ICASSP1
2014 Palette-based compound image compression in HEVC by exploiting non-local spatial correlation
abstract
Non-camera captured images (also known as compound image) contain a mixture of camera-captured natural images and computer-generated graphics and texts. Nowadays, there are more and more applications calling for non-camera captured image/video compression scheme. However, current video coding standards, which are designed for natural video, treat non-camera captured video less carefully. For example, the state-of-the-art video coding standard High Efficiency Video Coding (HEVC) may blur or even remove edges in text/graphic region. A lot of schemes are proposed to preserve direction property of texts and graphics, such as palette-based intra coding. In this paper, a novel palette coding scheme is proposed for palette-based intra coding in HEVC. The palette in a block is predicted from an adaptive palette template, which records the statistical non-local spatial correlation of an image. Every block chooses its own palette using the palette template as the prediction in a rate-distortion optimized manner. Experimental results show that the proposed scheme can achieve up to 5.2% bit-rate saving compared to the state-of-the-art palette-based coding scheme in HEVC.
Oscar C. Au, Wei Dai 0002, Haitao Yang 0001, Luheng Jia, Jin Zeng 0004, Pengfei Wan 0001
ICASSP7
2014 Self-similarity-based image colorization
abstract
In this work, we tackle the problem of coloring black-and-white images, which is image colorization. Existing image colorization algorithms can be categorized into two types: scribble-based colorization algorithms and example-based colorization algorithms. Differently, we propose a hybrid scheme that combines the advantages of both categories. Given the grayscale image to be colorized and a few color scribbles (or scattered color labels) as input, the proposed method manages to colorize the grayscale image with high quality. Similar to the mechanisms in example-based colorization methods, our algorithm firstly propagates chrominance information based on the assumption that similar image patches should have similar colors. Therefore colors of some pixels can be transferred from similar patches with known colors. After that, we apply scribble-based colorization algorithm to fully colorize the grayscale image, with different confidences assigned onto the transferred color labels. Experimental results show that, the proposed method effectively utilizes the known chrominance, and provides pleasant colorizations with very few user interventions.
Jiahao Pang, Oscar C. Au, Yukihiko Yamashita, Yonggen Ling, Yuanfang Guo, Jin Zeng 0004
ICIP6
2013 An analytical study of subpixel-based image down-sampling patterns in frequency domain
abstract
Subpixel-based image down-sampling is a class of methods that can provide improved apparent resolution of the down-scaled image compared to the pixel-based methods. The frequency characteristics of all possible subpixel-based down-sampling patterns for RGB vertical stripes are analytically studied in this paper. Our proposed algorithm reveals that there are merely seven equivalent energy distributions in the luminance frequency spectrum. To achieve higher luminance resolution, we then calculate and choose the optimal down-sampling pattern with anti-aliasing low-pass filter designed for it so as to maximize the energy of the luminance component within the cut-off shape. Experimental results show that the proposed method provides sharper images compared to the state-of-art subpixel-based methods, with little color distortion.
Yonggen Ling, Oscar C. Au, Ketan Tang, Jiahao Pang, Jin Zeng 0004, Lu Fang 0001
VCIP5