VLDB 2026 Research / reviewers in the wild / expert
Tianyu Geng
dblp:181/0425
· DBLP profile ↗
15ranked-venue papers
5as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 6 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hierarchical Information Embeddings With Neural ODEs for Personalized Federated LearningabstractPersonalized federated learning (PFL) plays a pivotal role in ensuring efficient privacy preservation and secure collaborative learning. However, PFL faces significant challenges due to data heterogeneity and device diversity. To enhance personalization and robustness in PFL, we propose a novel model called FedNODE, which leverages hierarchical embeddings. FedNODE incorporates personalized, pseudo-generic, and fusion embeddings to facilitate hierarchical information representation. We utilize a hypernetwork based on neural ordinary differential equations (ODEs) within the server to generate backbone parameters for different clients, enabling the creation of personalized embeddings. Additionally, we introduce a pseudo-generic embedding based on a learnable vector to balance personalized and generic information. A neural ODE-based network follows the backbone module for each client, integrating personalized and pseudo-generic embeddings. To validate the efficacy of FedNODE, we conduct extensive evaluations across various classification datasets, encompassing diverse statistically heterogeneous settings and noisy scenarios. The results demonstrate that FedNODE achieves state-of-the-art performance. Rui She 0001, Qiyu Kang, Kai Zhao 0010, Tianyu Geng, Yanan Zhao 0003, Wenfei Liang 0001, Wee-Peng Tay |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | PBECount: Prompt-Before-Extract Paradigm for Class-Agnostic CountingabstractIn the field of class-agnostic counting (CAC), counting only objects of interest that are similar to exemplars in multi-class scenarios has been a challenging task. To address this challenge, recent research has proposed the extract-and-match paradigm based on the vision transformer (ViT) architecture. However, although this paradigm can improve the accuracy of exemplar-similar object identification, it overly emphasizes the role of the ViT structure. To address this shortcoming, this work introduces a more generalized prompt-before-extract paradigm on top of the extract-and-match paradigm and designs a pure convolutional neural network (CNN) model named PBECount. In addition, an innovative loss function, a post-processing strategy, and a dynamic threshold method are proposed to enhance the detection performance of the proposed model when the probability maps are used as ground truth during model training. The experimental results on the FSC-147 and CARPK datasets demonstrate that the proposed PBECount can identify whether unknown class objects are similar to exemplars and outperform the state-of-the-art CAC methods in terms of accuracy and generalization. Canchen Yang, Tianyu Geng, Jian Peng 0002 |
AAAI | 2 |
| 2025 | Multi-Modal Aerial-Ground Cross-View Place Recognition with Neural ODEsabstractPlace recognition (PR) aims at retrieving the query place from a database and plays a crucial role in various applications, including navigation, autonomous driving, and augmented reality. While previous multi-modal PR works have mainly focused on the same-view scenario in which ground-view descriptors are matched with a database of ground-view descriptors during inference, the multi-modal cross-view scenario, in which ground-view descriptors are matched with aerial-view descriptors in a database, remains underexplored. We propose AGPlace, a model that effectively integrates information from multi-modal ground sensors (cameras and LiDARs) to achieve accurate aerial-ground PR. AGPlace achieves effective aerial-ground cross-view PR by leveraging a manifold-based neural ordinary differential equation (ODE) framework with a multi-domain alignment loss. It outperforms existing state-of-the-art cross-view PR models on large-scale datasets. As most existing PR models are designed for ground-ground PR, we adapt these baselines into our cross-view pipeline. Experiments demonstrate that this direct adaptation performs worse than our overall model architecture AGPlace. AGPlace represents a significant advancement in multi-modal aerial-ground PR, with promising implications for real-world applications. Rui She 0001, Qiyu Kang, Disheng Li, Tianyu Geng, Shangshu Yu, Wee-Peng Tay |
CVPR | 6 |
| 2025 | Modulo Video Recovery via Selective Spatiotemporal Vision TransformerabstractConventional image sensors have limited dynamic range, causing saturation in high-dynamic-range (HDR) scenes. Modulo cameras address this by folding incident irradiance into a bounded range, yet require specialized unwrapping algorithms to reconstruct the underlying signal. Unlike HDR recovery, which extends dynamic range from conventional sampling, modulo recovery restores actual values from folded samples. Despite being introduced over a decade ago, progress in modulo image recovery has been slow, especially in the use of modern deep learning techniques. In this work, we demonstrate that standard HDR methods are unsuitable for modulo recovery. Transformers, however, can capture global dependencies and spatial-temporal relationships crucial for resolving folded video frames. Still, adapting existing Transformer architectures for modulo recovery demands novel techniques. To this end, we present Selective Spatiotemporal Vision Transformer (SSViT), the first deep learning framework for modulo video reconstruction. SSViT employs a token selection strategy to improve efficiency and concentrate on the most critical regions. Experiments confirm that SSViT produces high-quality reconstructions from 8-bit folded videos and achieves state-of-the-art performance in modulo video recovery. Tianyu Geng, Wee-Peng Tay |
IJCNN | 1 |
| 2024 | PosDiffNet: Positional Neural Diffusion for Point Cloud Registration in a Large Field of View with PerturbationsabstractPoint cloud registration is a crucial technique in 3D computer vision with a wide range of applications. However, this task can be challenging, particularly in large fields of view with dynamic objects, environmental noise, or other perturbations. To address this challenge, we propose a model called PosDiffNet. Our approach performs hierarchical registration based on window-level, patch-level, and point-level correspondence. We leverage a graph neural partial differential equation (PDE) based on Beltrami flow to obtain high-dimensional features and position embeddings for point clouds. We incorporate position embeddings into a Transformer module based on a neural ordinary differential equation (ODE) to efficiently represent patches within points. We employ the multi-level correspondence derived from the high feature similarity scores to facilitate alignment between point clouds. Subsequently, we use registration methods such as SVD-based algorithms to predict the transformation using corresponding point pairs. We evaluate PosDiffNet on several 3D point cloud datasets, verifying that it achieves state-of-the-art (SOTA) performance for point cloud registration in large fields of view with perturbations. The implementation code of experiments is available at https://github.com/AI-IT-AVs/PosDiffNet. Rui She 0001, Qiyu Kang, Kai Zhao 0010, Yang Song 0012, Wee-Peng Tay, Tianyu Geng, Xingchao Jian |
AAAI | 7 |
| 2024 | PointDifformer: Robust Point Cloud Registration With Neural Diffusion and TransformerabstractPoint cloud registration is a fundamental technique in 3-D computer vision with applications in graphics, autonomous driving, and robotics. However, registration tasks under challenging conditions, under which noise or perturbations are prevalent, can be difficult. We propose a robust point cloud registration approach that leverages graph neural partial differential equations (PDEs) and heat kernel signatures. Our method first uses graph neural PDE modules to extract high-dimensional features from point clouds by aggregating information from the 3-D point neighborhood, thereby enhancing the robustness of the feature representations. Then, we incorporate heat kernel signatures into an attention mechanism to efficiently obtain corresponding keypoints. Finally, a singular value decomposition (SVD) module with learnable weights is used to predict the transformation between two point clouds. Empirical experiments on a 3-D point cloud dataset demonstrate that our approach not only achieves state-of-the-art performance for point cloud registration but also exhibits better robustness to additive noise or 3-D shape perturbations. Rui She 0001, Qiyu Kang, Wee-Peng Tay, Kai Zhao 0010, Yang Song 0012, Tianyu Geng, Yi Xu 0014, Diego Navarro Navarro, Andreas Hartmannsgruber |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2023 | Modulo EEG Signal Recovery Using TransformerabstractTime series signals such as EEG signals may have large variability across different individuals, making it difficult to sample without distortion or clipping using the same sensor for different individuals. Modulo sampling allows one to overcome the problem of signal clipping in the case where the signal has a very high dynamic range. This paper studies the problem of recovering a time series signal from under-determined modulo observations. We propose a deep learning method for modulo signal recovery, which can be applied to recover folded EEG signals. We make the first attempt to introduce the Transformer framework to modulo signal recovery. In addition, for efficiency and robustness, we introduce a modification of the Transformer module by inserting a learnable pre-estimation. The experiment on the real data demonstrates the superior performance of the proposed algorithm. Tianyu Geng, Pratibha, Wee-Peng Tay |
ICASSP | 1 |
| 2023 | Robust Graph Neural Diffusion for Image MatchingabstractImage matching identifies matching street landmark patches between the images captured by a vehicular camera and those stored in a database. Applications include autonomous driving perception and localization. However, in practical scenarios, challenging conditions such as changing weather, illumination, and dynamic objects result in perturbations of the captured images, leading to inaccurate matching. To achieve robust landmark patch matching, we present a method, named GRAND-Mat, which leverages a neural diffusion over graph embeddings to counteract perturbations. We first extract high-dimensional features of landmark patches using a ResNet. Then, we utilize graph neural diffusion models to aggregate the self and cross-graph information from these features. Furthermore, we apply feature similarity learning to acquire the final matching score. We evaluate the performance of our model on a street scene dataset, which demonstrates state-of-the-art matching performance under additive perturbations. Rui She 0001, Qiyu Kang, Kai Zhao 0010, Yang Song 0012, Yi Xu 0014, Tianyu Geng, Wee-Peng Tay, Diego Navarro Navarro, Andreas Hartmannsgruber |
ICIP | 7 |
| 2023 | SHNN: A single-channel EEG sleep staging model based on semi-supervised learning
Yongqing Zhang 0001, Wenpeng Cao, Lixiao Feng, Manqing Wang, Tianyu Geng, Jiliu Zhou, Dongrui Gao |
Expert Syst. Appl. | 5 |
| 2021 | Deep Shearlet Residual Learning Network for Single Image Super-ResolutionabstractRecently, the residual learning strategy has been integrated into the convolutional neural network (CNN) for single image super-resolution (SISR), where the CNN is trained to estimate the residual images. Recognizing that a residual image usually consists of high-frequency details and exhibits cartoon-like characteristics, in this paper, we propose a deep shearlet residual learning network (DSRLN) to estimate the residual images based on the shearlet transform. The proposed network is trained in the shearlet transform-domain which provides an optimal sparse approximation of the cartoon-like image. Specifically, to address the large statistical variation among the shearlet coefficients, a dual-path training strategy and a data weighting technique are proposed. Extensive evaluations on general natural image datasets as well as remote sensing image datasets show that the proposed DSRLN scheme achieves close results in PSNR to the state-of-the-art deep learning methods, using much less network parameters. Tianyu Geng, Xiao-Yang Liu, Xiaodong Wang 0001, Guiling Sun |
IEEE Trans. Image Process. | 1 |
| 2020 | TSGYE: Two-Stage Grape Yield Estimation
Geng Deng, Tianyu Geng, Chengxin He, Xinao Wang, Bangjun He, Lei Duan |
ICONIP (4) | 2 |
| 2020 | A Supervised Learning Algorithm for Learning Precise Timing of Multispike in Multilayer Spiking Neural Networks
Rong Xiao 0001, Tianyu Geng |
ICONIP (5) | 2 |
| 2018 | Multiscale overlapping blocks binarized statistical image features descriptor with flip-free distance for face verification in the wild
Tianyu Geng, Menglong Yang, Zhisheng You, Ying Cai 0002, Feihu Huang 0002 |
Neural Comput. Appl. | 1 |
| 2018 | Truncated Nuclear Norm Minimization Based Group Sparse Representation for Image RestorationabstractGroup sparse representation has shown great potential in image restoration, which can be considered as a low-rank matrix approximation problem. The nuclear norm minimization method, as a convex relaxation of the rank minimization, shrinks all the singular values simultaneously. Recent advances have suggested the truncated nuclear norm minimization method to better approximate the rank of the matrix. In this paper, we connect group sparse representation with truncated nuclear norm minimization with the application to image restoration. Then, an implementation of fast convergence via the alternating direction method of multipliers is developed to solve the proposed problem. Moreover, an effective dictionary for each group is learned from the recovery image itself rather than a dataset with a large number of natural images. Experimental results demonstrate that the proposed GSR-TNNM method achieves a good convergence performance and is able to improve image quality significantly compared with the state-of-the-art methods. Tianyu Geng, Guiling Sun, Yi Xu 0014, Jingfei He |
SIAM J. Imaging Sci. | 1 |
| 2016 | Data recovery in heterogeneous wireless sensor networks based on low-rank tensorsabstractAn effective way to reduce the energy consumption of energy constrained wireless sensor networks is reducing the number of collected data, which causes the recovery problem. In this paper, we propose a novel data recovery method based on low-rank tensors for the heterogeneous wireless sensor networks with various sensor types. The proposed method represents the collected high-dimensional data as low-rank tensors to effectively exploit the spatiotemporal correlation that exists in the various data. Furthermore, an algorithm based on the alternating direction method of multipliers is developed to solve the resultant optimization problem efficiently. Experimental results demonstrate that the proposed method significantly outperforms the sparsity constraint method and matrix completion method for each type of signals. Jingfei He, Guiling Sun, Tianyu Geng |
ISCC | 4 |