VLDB 2026 Research / reviewers in the wild / expert
Yitong Jiang
dblp:131/1954
· DBLP profile ↗
10ranked-venue papers
4as first author
7since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Token-Efficient VLM: High-Resolution Image Understanding Via Dynamic Region Proposal
Yitong Jiang, Jinwei Gu, Tianfan Xue, Ka Chun Cheung, Pavlo Molchanov 0001, Hongxu Yin, Sifei Liu |
ICCV | 1 |
| 2025 | GSPN-2: Efficient Parallel Sequence ModelingabstractEfficient vision transformer remains a bottleneck for high-resolution images and long-video related real-world applications. Generalized Spatial Propagation Network (GSPN) \cite{wang2025parallel} addresses this by replacing quadratic self-attention with a line-scan propagation scheme, bringing the cost close to linear in the number of rows or columns, while retaining accuracy. Despite this advancement, the existing GSPN implementation still suffers from (i) heavy overhead due to repeatedly launching GPU kernels, (ii) excessive data transfers from global GPU memory, and (iii) redundant computations caused by maintaining separate propagation weights for each channel. We introduce GSPN-2, a joint algorithm–system redesign. In particular, we eliminate thousands of micro-launches from the previous implementation into one single 2D kernel, explicitly pin one warp to each channel slice, and stage the previous column's activations in shared memory. On the model side, we introduce a set of channel-shared propagation weights that replace per-channel matrices, trimming parameters, and align naturally with the affinity map used in transformer attention. Experiments demonstrate GSPN-2's effectiveness across image classification and text-to-image synthesis tasks, matching transformer-level accuracy with significantly lower computational cost. GSPN-2 establishes a new efficiency frontier for modeling global spatial context in vision applications through its unique combination of structured matrix transformations and GPU-optimized implementation. Yitong Jiang, Collin McCarthy, David Wehr, Hanrong Ye, Ka Chun Cheung, Wonmin Byeon, Jinwei Gu, Kai Han 0001, Hongxu Yin, Pavlo Molchanov 0001, Jan Kautz, Sifei Liu |
NeurIPS | 2 |
| 2024 | AutoDIR: Automatic All-in-One Image Restoration with Latent Diffusion
Yitong Jiang, Tianfan Xue, Jinwei Gu |
ECCV (40) | 1 |
| 2023 | Real-Time Controllable Denoising for Image and VideoabstractControllable image denoising aims to generate clean samples with human perceptual priors and balance sharpness and smoothness. In traditional filter-based denoising methods, this can be easily achieved by adjusting the filtering strength. However, for NN (Neural Network)-based models, adjusting the final denoising strength requires performing network inference each time, making it almost impossible for real-time user interaction. In this paper, we introduce Real-time Controllable Denoising (RCD), the first deep image and video denoising pipeline that provides a fully controllable user interface to edit arbitrary denoising levels in real-time with only one-time network inference. Unlike existing controllable denoising methods that require multiple denoisers and training stages, RCD replaces the last output layer (which usually outputs a single noise map) of an existing CNN-based model with a lightweight module that outputs multiple noise maps. We propose a novel Noise Decorrelation process to enforce the orthogonality of the noise feature maps, allowing arbitrary noise level control through noise map interpolation. This process is network-free and does not require network inference. Our experiments show that RCD can enable real-time editable image and video denoising for various existing heavy-weight models without sacrificing their original performance. Zhaoyang Zhang 0004, Yitong Jiang, Wenqi Shao, Xiaogang Wang 0001, Ping Luo 0002, Kaimo Lin, Jinwei Gu |
CVPR | 2 |
| 2023 | Learning Image-Adaptive Codebooks for Class-Agnostic Image RestorationabstractRecent work on discrete generative priors, in the form of codebooks, has shown exciting performance for image reconstruction and restoration, as the discrete prior space spanned by the codebooks increases the robustness against diverse image degradations. Nevertheless, these methods require separate training of codebooks for different image categories, which limits their use to specific image categories only (e.g. face, architecture, etc.), and fail to handle arbitrary natural images. In this paper, we propose AdaCode for learning image-adaptive codebooks for class-agnostic image restoration. Instead of learning a single codebook for each image category, we learn a set of basis codebooks. Given an input image, AdaCode learns a weight map with and computes a weighted combination of these basis codebooks for adaptive image restoration. Intuitively, AdaCode is a more flexible and expressive discrete generative prior than previous work. Experimental results demonstrate that AdaCode achieves state-of-the-art performance on image reconstruction and restoration tasks, including image super-resolution and inpainting. Codes are released at https://github.com/kechunl/AdaCode. Kechun Liu, Yitong Jiang, Inchang Choi, Jinwei Gu |
ICCV | 2 |
| 2021 | STAR: A Structure-aware Lightweight Transformer for Real-time Image EnhancementabstractImage and video enhancement such as color constancy, low light enhancement, and tone mapping on smartphones is challenging, because high-quality images should be achieved efficiently with a limited resource budget. Unlike prior works that either used very deep CNNs or large Trans-former models, we propose a structure-aware lightweight Transformer, termed STAR, for real-time image enhancement. STAR is formulated to capture long-range dependencies between image patches, which naturally and implicitly captures the structural relationships of different regions in an image. STAR is a general architecture that can be easily adapted to different image enhancement tasks. Extensive experiments show that STAR can effectively boost the quality and efficiency of many tasks such as illumination enhancement, auto white balance, and photo retouching, which are indispensable components for image processing on smartphones. For example, STAR reduces model complexity and improves image quality compared to the recent state-of-the-art [19] on the MIT-Adobe FiveK dataset [7] (i.e., 1.8dB PSNR improvements with 25% parameters and 13% float operations.) Zhaoyang Zhang 0004, Yitong Jiang, Xiaogang Wang 0001, Ping Luo 0002, Jinwei Gu |
ICCV | 2 |
| 2021 | Revisiting Shadow Detection: A New Benchmark Dataset for Complex WorldabstractShadow detection in general photos is a nontrivial problem, due to the complexity of the real world. Though recent shadow detectors have already achieved remarkable performance on various benchmark data, their performance is still limited for general real-world situations. In this work, we collected shadow images for multiple scenarios and compiled a new dataset of 10,500 shadow images, each with labeled ground-truth mask, for supporting shadow detection in the complex world. Our dataset covers a rich variety of scene categories, with diverse shadow sizes, locations, contrasts, and types. Further, we comprehensively analyze the complexity of the dataset, present a fast shadow detection network with a detail enhancement module to harvest shadow details, and demonstrate the effectiveness of our method to detect shadows in general situations. Xiaowei Hu 0001, Tianyu Wang 0003, Chi-Wing Fu, Yitong Jiang, Qiong Wang 0001, Pheng-Ann Heng |
IEEE Trans. Image Process. | 4 |
| 2019 | Mask-ShadowGAN: Learning to Remove Shadows From Unpaired DataabstractThis paper presents a new method for shadow removal using unpaired data, enabling us to avoid tedious annotations and obtain more diverse training samples. However, directly employing adversarial learning and cycle-consistency constraints is insufficient to learn the underlying relationship between the shadow and shadow-free domains, since the mapping between shadow and shadow-free images is not simply one-to-one. To address the problem, we formulate Mask-ShadowGAN, a new deep framework that automatically learns to produce a shadow mask from the input shadow image and then takes the mask to guide the shadow generation via re-formulated cycle-consistency constraints. Particularly, the framework simultaneously learns to produce shadow masks and learns to remove shadows, to maximize the overall performance. Also, we prepared an unpaired dataset for shadow removal and demonstrated the effectiveness of Mask-ShadowGAN on various experiments, even it was trained on unpaired data. Xiaowei Hu 0001, Yitong Jiang, Chi-Wing Fu, Pheng-Ann Heng |
ICCV | 2 |
| 2015 | Downscaling GOES Land Surface Temperature for Assessing Heat Wave Health RisksabstractRecent years have witnessed an emerging concern of the health impact of heat waves. A common approach to investigate heat waves is to resort to the geostationary thermal infrared imagery, such as those from the Geostationary Operational Environmental Satellite (GOES) and Meteosat Second Generation. However, coarse spatial resolutions of geostationary images cannot meet the need of assessing and monitoring heat waves in complex urban settings. To address the spatial and temporal variability of heat waves in urban areas, this letter presented a study of analyzing heat wave risk in Los Angeles, USA, by the synergistic use of GOES land surface temperature (LST), auxiliary geospatial, and census data within the framework of Crichton's Risk Triangle (i.e., hazard, exposure, and vulnerability). Principal component analysis and regression analysis were employed to downscale the original GOES LST imagery from 4 to 1 km. The resultant subhourly 1-km LST data was used to characterize and quantify heat hazard. The census population represented the exposure, while existing health, socioeconomic, and physical environmental conditions were used to describe the vulnerabilities. The risk map of heat wave was computed using the weighted indices of hazard, exposure, and vulnerability. The map was further overlaid with a zip-code data layer to generate statistics. The derived risk map showed that areas with high risk were identified in the central city, part of western LA County, and the desert area, based on a 10-point scale rank. Yitong Jiang, Peng Fu 0004, Qihao Weng |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2013 | Estimating LST Using a Vegetation-Cover-Based Thermal Sharpening TechniqueabstractVegetation-cover-based thermal sharpening techniques have mostly been developed and tested in agricultural areas. Overlooking the impact of soil moisture on surface temperatures is a common problem in these algorithms. This letter developed a vegetation thermal sharpening method for the City of Indianapolis, Indiana, USA, and estimated land surface temperature by disaggregated Landsat Thematic Mapper thermal infrared data from 120 to 30 m. The root-mean-square error was yielded at 1.90°C and 1.91°C using NDVI and fractional vegetation cover as predictors, respectively. The error of the estimation was overlaid with a soil moisture map, which was derived based on the surface energy balance modeling. The pixels with large errors were largely distributed in the areas with low soil moisture. These areas were covered by impervious surfaces such as major roads, commercial land, and the airport. This result suggested that in the urban areas, besides vegetation cover and soil moisture, impervious surfaces must be incorporated in developing any future thermal sharpening techniques. The incorporation of population density and per capita consumption of energy may provide further improvements in the estimation. Yitong Jiang, Qihao Weng |
IEEE Geosci. Remote. Sens. Lett. | 1 |