EDBT 2026 Demo / reviewers in the wild / expert
Tingyu Wang 0002
dblp:08/9892-2
· DBLP profile ↗
21ranked-venue papers
5as first author
21since 2021 · last 2026
0000-0002-4169-1595ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 12 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Feature boosting and scale-aware network with multi-modal information for underwater salient object detection
Tingyu Wang 0002, Junzhe Lu 0002, Bin Wan, Rongfeng Lu, Yaoqi Sun, Duanpo Wu, Chenggang Yan 0001 |
Eng. Appl. Artif. Intell. | 1 |
| 2026 | AND-GS: Adaptive supervision of normal and depth in Gaussian splatting for accurate and efficient surface reconstruction
Xiang Le, Qiang Zhao 0005, Haofan Ren, Zhongtian Zheng, Tingyu Wang 0002, Jiyong Zhang 0001, Chenggang Yan 0001 |
Neurocomputing | 5 |
| 2026 | ThermalGaussian++: Improving Alignment and Resolution for ThermalGaussianabstractThermography is especially valuable for the military and other users of surveillance cameras. Some recent methods based on Neural Radiance Fields (NeRF) have been proposed to reconstruct thermal scenes in 3D from a set of thermal and RGB images. However, unlike NeRF, 3D Gaussian splatting (3DGS) prevails due to its rapid training and real-time rendering. In this work, we propose ThermalGaussian, the first thermal 3DGS approach capable of rendering high-quality images in RGB and thermal modalities. We first calibrate the RGB camera and the thermal camera to ensure that both modalities are accurately aligned. Subsequently, we use the registered images to learn the multimodal 3D Gaussians. To prevent the overfitting of any single modality, we introduce several multimodal regularization constraints. We also develop smoothing constraints tailored to the physical characteristics of the thermal modality. Besides, we contribute a real-world dataset named RGBT-Scenes, captured by a handheld thermal-infrared camera, facilitating future research on thermal scene reconstruction. Based on ThermalGaussian, we further introduce ThermalGaussian++ to improve the alignment and resolution of ThermalGaussian. To improve multimodal alignment, we design a multimodal pose optimization module. This module enables direct processing of non-aligned multimodal image pairs, reducing the need for professional calibration before each use. To improve thermal resolution, we also propose a multimodal joint super-resolution reconstruction module, which enhances the quality of low-resolution thermal fields. Additionally, we contribute a new dataset: RGBT-Scenes++, which offers higher-resolution thermal images. We conduct comprehensive experiments demonstrating that ThermalGaussian++ achieves photorealistic thermal rendering and improves RGB rendering quality. It significantly enhances both alignment and resolution, enabling better practical deployment. In addition, our multimodal regularization constraints reduce the model's storage requirements. The code and datasets will be released. Rongfeng Lu, Ming Lu 0002, Tingyu Wang 0002, Haofan Ren, Yitian Xue, Chenggang Yan 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2026 | LLFeat: Noise-Aware Feature Matching Under Various Low-Light Conditions
Longjian Zeng, Zunjie Zhu, Ming Lu 0002, Bolun Zheng, Rongfeng Lu, Tingyu Wang 0002, Zhongtian Zheng, Yaoqi Sun, Chenggang Yan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | Enhance Panoramic Object Detection Using Planar Image DatasetsabstractPanoramic images have been used in various applications because of their ability to provide comprehensive spatial information. However, the high cost of obtaining panoramic images and the complexity of annotation pose serious obstacles to enhancing the performance of panoramic tasks by restricting the size and quality of datasets. The use of annotated planar images to synthesize panoramic images proves to be an effective approach to narrow this gap. Prior synthesis methods introduced new distortions that lead to inconsistent object shapes before and after synthesis. Concurrently, since the angle of view of the planar image is much smaller than that of the panoramic image, there is an issue of missing spatial information in the synthesized panoramic image. To address these challenges, we introduce a novel approach for converting planar images into panoramic ones with reduced deformation. For annotating targets in synthetic images, we develop a new algorithm based on determining the minimum spherical area to calculate spherical bounding boxes that closely adhere to object boundaries, rather than relying on an estimated target center point which results in inevitable calculation errors as in previous studies. Subsequently, we propose a new method for effectively filling any blank areas in the synthetic panoramic images to compensate for the loss of precision caused by absence of spatial information. Additionally, with a focus on the distortion characteristics of panoramic images, an innovative data augmentation strategy is devised to further enhance the model's ability to recognize objects in different positions. These methods are conducive to a more effective utilization of the rich planar image datasets for panoramic object detection tasks. In the experiments, we generated two synthetic panoramic datasets based on COCO for training models. Experimental results demonstrate that training with these synthetic datasets significantly improves prediction accuracy, far surpasses the state-of-the-art methods, and helps to unleash the full potential of panoramic object detection models. Source code is available athttps://github.com/longlong-yu/official-panorama-coco.git. Longlong Yu 0001, Qiang Zhao 0005, Xinyuan Liu 0003, Tingyu Wang 0002, Chenggang Yan 0001 |
IEEE Trans. Multim. | 6 |
| 2025 | Lightweight three-stream encoder-decoder network for multi-modal salient object detection
Junzhe Lu 0002, Tingyu Wang 0002, Bin Wan, Qiang Zhao 0005, Shuai Wang 0003, Yaoqi Sun, Yang Zhou 0052, Chenggang Yan 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2025 | P2FCN: Environment-Independent UAV-View Geo-Localization via Pixel-to-Feature Co-EnhancementabstractThis paper investigates the challenges of UAV-view geo-localization under extreme environmental changes, where significant cross-domain style differences can lead to degradation of model performance. Existing methods primarily focus on mitigating domain shift issues caused by environmental factors but generally overlook the direct interference of environmental noise. We argue that mitigating environmental noise is equally critical for extracting discriminative cross-view features and introduce a pixel-to-feature co-enhancement network (P2FCN). P2FCN comprises a style-noise dual suppression module (SNDS) and a part-based multi-dimensional feature learning strategy (PMDFL). Specifically, the SNDS module mitigates the stylistic discrepancies between cross-environment images through pixel-level dynamic adjustment, while reducing the introduction of environmental noise to a certain degree. PMDFL improves the representation generalization by purifying discriminative information within partitioned channel-and spatial-wise feature subspaces. Extensive experiments on two widely used benchmarks,i.e., University-1652 and SUES-200, demonstrate that the proposed method achieves state-of-the-art performance in multiple environments compared to existing methods. Qiang Zhao 0005, Tingyu Wang 0002, Rongfeng Lu, Chenggang Yan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Quad Bayer Joint Demosaicing and Denoising Based on Dual Encoder Network with Joint Residual LearningabstractThe recent imaging technology Quad Bayer CFA brings better imaging PSNR and higher visual quality compared to traditional Bayer CFA, but also serious challenges for demosaicing and denoising during the ISP pipeline. In this paper, we propose a novel dual encoder network, namely DRNet, to achieve joint demosaicing and denoising for Quad Bayer CFA. The dual encoders are carefully designed in that one is mainly constructed by a joint residual block to jointly estimate the residuals for demosaicing and denoising separately. In contrast, the other one is started with a pixel modulation block which is specially designed to match the characteristics of Quad Bayer pattern for better feature extraction. We demonstrate the effectiveness of each proposed component through detailed ablation investigations. The comparison results on public benchmarks illustrate that our DRNet achieves an apparent performance gain~(0.38dB to the 2nd best) from the state-of-the-art method and balances performance and efficiency well. The experiments on real-world images show that the proposed method could enhance the reconstruction quality from the native ISP algorithm. Bolun Zheng, Haoran Li 0025, Tingyu Wang 0002, Xiaofei Zhou 0003, Chenggang Yan 0001 |
AAAI | 4 |
| 2024 | Towards Natural Language-Guided Drones: GeoText-1652 Benchmark with Spatial Relation Matching
Meng Chu, Zhedong Zheng, Wei Ji 0008, Tingyu Wang 0002, Tat-Seng Chua |
ECCV (11) | 4 |
| 2024 | Domain Shared and Specific Prompt Learning for Incremental Monocular Depth Estimation
Zhiwen Yang 0003, Liang Li 0003, Tingyu Wang 0002, Yaoqi Sun, Chenggang Yan 0001 |
ACM Multimedia | 4 |
| 2024 | Aerial-view geo-localization based on multi-layer local pattern cross-attention network
Haoran Li 0025, Tingyu Wang 0002, Qiang Zhao 0005, Shaowei Jiang, Chenggang Yan 0001, Bolun Zheng |
Appl. Intell. | 2 |
| 2024 | ADNet: Anti-noise dual-branch network for road defect detection
Bin Wan, Xiaofei Zhou 0003, Yaoqi Sun, Tingyu Wang 0002, Chengtao Lv, Shuai Wang 0003, Haibing Yin, Chenggang Yan 0001 |
Eng. Appl. Artif. Intell. | 4 |
| 2024 | Coplane-constrained sparse depth sampling and local depth propagation for depth estimation
Zhiwen Yang 0003, Chuqiao Chen, Hongkui Wang, Tingyu Wang 0002, Chenggang Yan 0001, Yihong Gong |
Image Vis. Comput. | 5 |
| 2024 | GoLDFormer: A global-local deformable window transformer for efficient image restoration
Bolun Zheng, Chenggang Yan 0001, Zunjie Zhu, Tingyu Wang 0002, Gregory Slabaugh, Shanxin Yuan |
J. Vis. Commun. Image Represent. | 5 |
| 2024 | Multiple-environment Self-adaptive Network for aerial-view geo-localization
Tingyu Wang 0002, Zhedong Zheng, Yaoqi Sun, Chenggang Yan 0001, Yi Yang 0001, Tat-Seng Chua |
Pattern Recognit. | 1 |
| 2024 | Rethinking Pooling for Multi-Granularity Features in Aerial-View Geo-LocalizationabstractVision-based aerial-view geo-localization aims to match drone- and satellite-views of the same geographical location. Several feature partition strategies divide spatial features to mine contextual information. However, the compression from fine-grained features to visual descriptors is ill-considered, that is, classical pooling destroys discriminative features while increasing the sensitivity of networks to contextual information. In order to clarify this, we first review existing pooling layer and analyze their pros and cons when applied in feature compression. Inspired by the appearance of aerial views, we then summarize an ideal feature compression operation, i.e., precisely highlighting the central target while maximizing the use of environmental information in a feature-smoothing manner. To achieve the above process, we propose a distance-dependent parameter initialization strategy and form a novel pooling called$D^{2}$-GeM pooling, which can explicitly guide the network to compress fine-grained features in multiple patterns. Extensive experiments on public benchmark University-1652 substantiate that our strategy attains more appealing results without additional costs. Tingyu Wang 0002, Yaoqi Sun, Chenggang Yan 0001 |
IEEE Signal Process. Lett. | 1 |
| 2024 | SDPL: Shifting-Dense Partition Learning for UAV-View Geo-LocalizationabstractCross-view geo-localization aims to match images of the same target from different platforms, e.g., drone and satellite. It is a challenging task due to the changing appearance of targets and environmental content from different views. Most methods focus on obtaining more comprehensive information through feature map segmentation, while inevitably destroying the image structure, and are sensitive to the shifting and scale of the target in the query. To address the above issues, we introduce simple yet effective part-based representation learning, shifting-dense partition learning (SDPL). We propose a dense partition strategy (DPS), dividing the image into multiple parts to explore contextual information while explicitly maintaining the global structure. To handle scenarios with non-centered targets, we further propose the shifting-fusion strategy, which generates multiple sets of parts in parallel based on various segmentation centers, and then adaptively fuses all features to integrate their anti-offset ability. Extensive experiments show that SDPL is robust to position shifting, and performs competitively on two prevailing benchmarks, University-1652 and SUES-200. In addition, SDPL shows satisfactory compatibility with a variety of backbone networks (e.g., ResNet and Swin).https://github.com/C-water/SDPL_release. Tingyu Wang 0002, Haoran Li 0025, Rongfeng Lu, Yaoqi Sun, Bolun Zheng, Chenggang Yan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Learning Cross-View Geo-Localization Embeddings via Dynamic Weighted Decorrelation RegularizationabstractIn the domain of cross-view geo-localization, the challenge lies in accurately matching images captured from distinct perspectives, such as aerial drone imagery and satellite imagery of the same geographical location. Existing methods predominantly concentrate on minimizing distances between feature embeddings in the representational space, inadvertently overlooking the significance of reducing embedding redundancy. This oversight potentially hampers the extraction of diverse and distinctive visual patterns critical for precise localization. This work argues that minimizing embedding redundancy is a pivotal factor in enhancing a model’s ability to discriminate diverse scene characteristics. To support this claim, we introduce a straightforward yet effective regularization technique, termed dynamic weighted decorrelation regularization (DWDR). DWDR serves to actively promote the learning of orthogonal feature channels within neural networks. By dynamically adjusting weights, DWDR targets the minimization of interchannel correlations, guiding the correlation matrix toward diagonality, indicative of independence among channels. The dynamic weighting mechanism adaptively prioritizes the decorrelation of channels that remain highly correlated throughout training. Additionally, we devise a symmetrical sampling strategy for cross-view scenarios to ensure that the training examples are balanced across different imaging platforms in a batch. Despite its simplicity, the integration of DWDR and the proposed sampling scheme yields remarkable performance across four extensive benchmark datasets: University-1652, CVUSA, CVACT, and VIGOR. Notably, in stringent conditions, such as when constrained to exceedingly compact feature dimensions of 64, our methodology significantly outperforms conventional baselines, thereby affirming its efficacy and robustness under challenging constraints. Tingyu Wang 0002, Zhedong Zheng, Zunjie Zhu, Yaoqi Sun, Chenggang Yan 0001, Yi Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | MFFNet: Multi-Modal Feature Fusion Network for V-D-T Salient Object DetectionabstractThis article discusses the limitations of single- and two-modal salient object detection (SOD) methods and the emergence of multi-modal SOD techniques that integrate Visible, Depth, or Thermal information. However, current multi-modal methods often rely on simple fusion techniques such as addition, multiplication and concatenation, to combine the different modalities, which is ineffective for challenging scenes, such as low illumination and background messy. To address this issue, we propose a novel multi-modal feature fusion network (MFFNet) for V-D-T salient object detection, where the two key points are the triple-modal deep fusion encoder and the progressive feature enhancement decoder. The MFFNet's triple-modal deep fusion (TDF) module is designed to integrate the features of the three modalities and explore their complementarity by utilizing mutual optimization during the encoding phase. In addition, the progressive feature enhancement decoder consists of the weighted context-enhanced feature (WCF) module, region optimization (RO) module and boundary perception (BP) module to produce region-aware and contour-aware features. After that, a multi-scale fusion (MF) module is proposed to integrate these features and generate high-quality saliency maps. We conduct extensive experiments on the VDT-2048 dataset, and our results show that the proposed MFFNet outperforms 12 state-of-the-art multi-modal methods. Bin Wan, Xiaofei Zhou 0003, Yaoqi Sun, Tingyu Wang 0002, Chengtao Lv, Shuai Wang 0003, Haibing Yin, Chenggang Yan 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | UAVM '23: 2023 Workshop on UAVs in Multimedia: Capturing the World from a New PerspectiveabstractUnmanned Aerial Vehicles (UAVs), also known as drones, have become increasingly popular in recent years due to their ability to capture high-quality multimedia data from the sky. With the rise of multimedia applications, such as aerial photography, cinematography, and mapping, UAVs have emerged as a powerful tool for gathering rich and diverse multimedia content. This workshop aims to bring together researchers, practitioners, and enthusiasts interested in UAV multimedia to explore the latest advancements, challenges, and opportunities in this exciting field. The workshop covers various topics related to UAV multimedia, including aerial image and video processing, machine learning for UAV data analysis, UAV swarm technology, and UAV-based multimedia applications. In the context of the ACM Multimedia conference, this workshop is highly relevant as multimedia data from UAVs is becoming an increasingly important source of content for many multimedia applications. The workshop provides a platform for researchers to share their work and discuss potential collaborations, as well as an opportunity for practitioners to learn about the latest developments in UAV multimedia technology. Overall, this workshop provides a unique opportunity to explore the exciting and rapidly evolving field of UAV multimedia and its potential impact on the wider multimedia community. Zhedong Zheng, Yujiao Shi 0002, Tingyu Wang 0002, Jun Liu 0036, Jianwu Fang, Yunchao Wei, Tat-Seng Chua |
ACM Multimedia | 3 |
| 2022 | Each Part Matters: Local Patterns Facilitate Cross-View Geo-LocalizationabstractCross-view geo-localization is to spot images of the same geographic target from different platforms,e.g., drone-view cameras and satellites. It is challenging in the large visual appearance changes caused by extreme viewpoint variations. Existing methods usually concentrate on mining the fine-grained feature of the geographic target in the image center, but underestimate the contextual information in neighbor areas. In this work, we argue that neighbor areas can be leveraged as auxiliary information, enriching discriminative clues for geo-localization. Specifically, we introduce a simple and effective deep neural network, called Local Pattern Network (LPN), to take advantage of contextual information in an end-to-end manner. Without using extra part estimators, LPN adopts a square-ring feature partition strategy, which provides the attention according to the distance to the image center. It eases the part matching and enables the part-wise representation learning. Owing to the square-ring partition design, the proposed LPN has good scalability to rotation variations and achieves competitive results on three prevailing benchmarks,i.e., University-1652, CVUSA and CVACT. Besides, we also show the proposed LPN can be easily embedded into other frameworks to further boost performance. Tingyu Wang 0002, Zhedong Zheng, Chenggang Yan 0001, Jiyong Zhang 0001, Yaoqi Sun, Bolun Zheng, Yi Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |