Yi Tang 0008

dblp:88/3775-8 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
7since 2021 · last 2026
0000-0002-4882-5234ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 6 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 4 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 Learning Underwater Image Enhancement Iteratively Without Reference Images
abstract
Since high-fidelity reference images are difficult to obtain in real underwater scenes, most deep models trained by synthetic paired data cannot match real-world data exactly. In this paper, we propose an unsupervised training framework for underwater image enhancement (UIE) by leveraging an iterative training strategy and quantification of specific neural units. Specifically, to eliminate the heavy color cast and distortion in the underwater images, we decompose the unsupervised image enhancement as two targeted sub-tasks, namely colorization and color compensation. First, a diffusion model is introduced for colorization to correct the green and blue color casts. Then, to intensify the learning ability of balanced color information, we introduce an extra network branch and propose a quantification mechanism for color compensation. The extra branch encodes style information from normal images into the generative model, while the quantification mechanism identifies and adjusts neural units relevant to warm colors, improving the model’s ability to learn balanced color feature representations for robust generation. In the end, through iterative training, color cast and distortion are progressively reduced, leading to a gradual improvement in the quality of the generated images. Experimental results on various widely used underwater datasets demonstrate that our approach achieves excellent performance, even when compared to recent supervised methods.
Yi Tang 0008, Hiroshi Kawasaki, Takafumi Iwaguchi, Hiroshi Masui
AAAI1
2026 Control-Lit: Illumination Controllable Backlit Image Enhancement
abstract
Backlit image enhancement (BIE) aims to address image degradation caused by challenging lighting conditions. By enhancing the illumination of underexposed areas and restoring image details while avoiding overexposure, it achieves an overall harmonious luminance. In contrast to traditional BIE methods that apply global enhancements with limited effectiveness, we propose a controllable Mamba-based enhancement method, termed Control-Lit. Our method not only delivers effective backlit image enhancement but also allows users to interactively adjust the illumination of specific regions. Control-Lit achieves superior global enhancement performance by employing the Dark Channel Prior (DCP)-based Finite Scalar Quantization (DFSQ) module that provides a high-quality image prior. Additionally, the Dark Channel Prior Enhancement (DCPE) module is designed to guide the network in pixel-wise illumination adjustment at the feature level, thereby achieving enhancement of backlit regions. Furthermore, to address the loss of content and details in backlit regions, the Global-Local Vision State Space (GLVSS) module is incorporated to extract both global and local features for BIE. To enable customizable, controllable enhancement, we introduce an illumination control vector. By adjusting the coefficient map elements that are multiplied by this vector, users can achieve precise regional illumination adjustments. To further validate the generalization of different light enhancement methods, we contribute a synthetic backlit image dataset using a relighting generative model. Along with current widely used datasets, the experimental results demonstrate that our method achieves state-of-the-art performance on all datasets while enabling high-quality and controllable image illumination adjustment. The code is available at https://github.com/wuhj43/Control-Lit.
Hongjun Wu 0003, Yi Tang 0008, Chongyi Li, Zhi Jin 0002
IEEE Trans. Circuits Syst. Video Technol.2
2023 Underwater Image Enhancement by Transformer-based Diffusion Model with Non-uniform Sampling for Skip Strategy
abstract
In this paper, we present an approach to image enhancement with diffusion model in underwater scenes. Our method adapts conditional denoising diffusion probabilistic models to generate the corresponding enhanced images by using the underwater images and the Gaussian noise as the inputs. Additionally, in order to improve the efficiency of the reverse process in the diffusion model, we adopt two different ways. We firstly propose a lightweight transformer-based denoising network, which can effectively promote the time of network forward per iteration. On the other hand, we introduce a skip sampling strategy to reduce the number of iterations. Besides, based on the skip sampling strategy, we propose two different non-uniform sampling methods for the sequence of the time step, namely piecewise sampling and searching with the evolutionary algorithm. Both of them are effective and can further improve performance by using the same steps against the previous uniform sampling. In the end, we conduct a relative evaluation of the widely used underwater enhancement datasets between the recent state-of-the-art methods and the proposed approach. The experimental results prove that our approach can achieve both competitive performance and high efficiency. Our code is available at https://github.com/piggy2009/DM_underwater.
Yi Tang 0008, Hiroshi Kawasaki, Takafumi Iwaguchi
ACM Multimedia1
2022 AutoEnhancer: Transformer on U-Net Architecture Search for Underwater Image Enhancement
Yi Tang 0008, Takafumi Iwaguchi, Hiroshi Kawasaki, Ryusuke Sagawa, Ryo Furukawa 0001
ACCV (3)1
2021 Temporal Pyramid Network for Pedestrian Trajectory Prediction with Multi-Supervision
abstract
Predicting human motion behavior in a crowd is important for many applications, ranging from the natural navigation of autonomous vehicles to intelligent security systems of video surveillance. All the previous works model and predict the trajectory with a single resolution, which is relatively ineffective and difficult to simultaneously exploit the long-range information (e.g., the destination of the trajectory), and the short-range information (e.g., the walking direction and speed at a certain time) of the motion behavior. In this paper, we propose a temporal pyramid network for pedestrian trajectory prediction through a squeeze modulation and a dilation modulation. Our hierarchical framework builds a feature pyramid with increasingly richer temporal information from top to bottom, which can better capture the motion behavior at various tempos. Furthermore, we propose a coarse-to-fine fusion strategy with multi-supervision. By progressively merging the top coarse features of global context to the bottom fine features of rich local context, our method can fully exploit both the long-range and short-range information of the trajectory. Experimental results on two benchmarks demonstrate the superiority of our method. Our code and models will be available upon acceptance.
Rongqin Liang, Yuanman Li, Xia Li 0006, Yi Tang 0008, Jiantao Zhou 0001, Wenbin Zou
AAAI4
2021 VI-eye: semantic-based 3D point cloud registration for infrastructure-assisted autonomous driving
abstract
Infrastructure-assisted autonomous driving is an emerging paradigm that aims to make affordable autonomous vehicles a reality. A key technology for realizing this vision is real-time point cloud registration which allows a vehicle to fuse the 3D point clouds generated by its own LiDAR and those on roadside infrastructures such as smart lampposts, which can deliver increased sensing range, more robust object detection, and centimeter-level navigation. Unfortunately, the existing methods for point cloud registration assume two clouds to share a similar perspective and large overlap, which result in significant delay and inaccuracy in real-world infrastructure-assisted driving settings. This paper proposes VI-Eye - the first system that can align vehicle-infrastructure point clouds at centimeter accuracy in real-time. Our key idea is to exploit traffic domain knowledge by detecting a set of key semantic objects including road, lane lines, curbs, and traffic signs. Based on the inherent regular geometries of such semantic objects, VI-Eye extracts a small number of saliency points and leverage them to achieve real-time registration of two point clouds. By allowing vehicles and infrastructures to extract the semantic information in parallel, VI-Eye leads to a highly scalable architecture for infrastructure-assisted autonomous driving. To evaluate the performance of VI-Eye, we collect two new multiview LiDAR point cloud datasets on an indoor autonomous driving testbed and a campus smart lamppost testbed, respectively. They contain total 915 point cloud pairs and cover three roads of 1.12km. Experiment results show that VI-Eye achieves centimeter-level accuracy within around 0.2s, and delivers a 5X improvement in accuracy and 2X speedup over state-of-the-art baselines.
Zhehao Jiang, Yi Tang 0008, Guoliang Xing
MobiCom4
2021 STA3D: Spatiotemporally attentive 3D network for video saliency prediction
Wenbin Zou, Shengkai Zhuo, Yi Tang 0008, Shishun Tian, Xia Li 0006, Chen Xu 0004
Pattern Recognit. Lett.3
2020 Video salient object detection via spatiotemporal attention neural networks
Yi Tang 0008, Wenbin Zou, Yang Hua 0001, Zhi Jin 0002, Xia Li 0006
Neurocomputing1
2019 An Efficient Quality Enhancement Solution for Stereo Images
Yingqing Peng, Zhi Jin 0002, Wenbin Zou, Yi Tang 0008, Xia Li 0006
ICIG (3)4
2019 Weakly Supervised Salient Object Detection With Spatiotemporal Cascade Neural Networks
abstract
Recently, deep learning techniques have substantially boosted the performance of salient object detection in still images. However, the salient object detection in videos by using traditional handcrafted features or deep learning features is not fully investigated, probably due to the lack of sufficient manually labeled video data for saliency modeling, especially for the data-driven deep learning. This paper proposes a novel weakly supervised approach to the salient object detection in a video, which can learn a robust saliency prediction model by using very limited manually labeled data and a large amount of weakly labeled data that could be easily generated in a supervised approach. Furthermore, we propose a spatiotemporal cascade neural network architecture for saliency modeling, in which two fully convolutional networks are cascaded to evaluate the visual saliency from both spatial and temporal cues to lead the optimal video saliency prediction. The proposed approach is extensively evaluated on the widely used challenging data sets, and the experiments demonstrate that our proposed approach substantially outperforms the state-of-the-art salient object detection models.
Yi Tang 0008, Wenbin Zou, Zhi Jin 0002, Yuhuan Chen, Yang Hua 0001, Xia Li 0006
IEEE Trans. Circuits Syst. Video Technol.1
2018 Multi-Scale Spatiotemporal Conv-LSTM Network for Video Saliency Detection
abstract
Recently, deep neural networks have been crucial techniques for image salient detection. However, two difficulties prevent the development of deep learning in video saliency detection. The first one is that the traditional static network cannot conduct a robust motion estimation in videos. The other is that the data-driven deep learning is in lack of sufficient manually annotated pixel-wise ground truths for video saliency network training. In this paper, we propose a multi-scale spatiotemporal convolutional LSTM network (MSST-ConvLSTM) to incorporate spatial and temporal cues for video salient objects detection. Furthermore, as manually pixel-wised labeling is very time-consuming, we sign lots of coarse labels, which are mixed with fine labels to train a robust saliency prediction model. Experiments on the widely used challenging benchmark datasets (e.g., FBMS and DAVIS) demonstrate that the proposed approach has competitive performance of video saliency detection compared with the state-of-the-art saliency models.
Yi Tang 0008, Wenbin Zou, Zhi Jin 0002, Xia Li 0006
ICMR1
2018 MS-CapsNet: A Novel Multi-Scale Capsule Network
abstract
Capsule network is a novel architecture to encode the properties and spatial relationships of the feature in an image, which shows encouraging results on image classification. However, the original capsule network is not suitable for some classification tasks, where the target objects are complex internal representations. Hence, we propose a multi-scale capsule network that is more robust and efficient for feature representation in image classification. The proposed multi-scale capsule network consists of two stages. In the first stage, structural and semantic information are obtained by multi-scale feature extraction. In the second stage, the hierarchy of features is encoded to multi-dimensional primary capsules. Moreover, we propose an improved dropout to enhance the robustness of the capsule network. Experimental results show that our method has a competitive performance on FashionMNIST and CIFAR10 datasets.
Canqun Xiang, Lu Zhang 0037, Yi Tang 0008, Wenbin Zou, Chen Xu 0004
IEEE Signal Process. Lett.3
2018 SCOM: Spatiotemporal Constrained Optimization for Salient Object Detection
abstract
This paper presents a novel model for video salient object detection called spatiotemporal constrained optimization model (SCOM), which exploits spatial and temporal cues, as well as a local constraint, to achieve a global saliency optimization. For a robust motion estimation of salient objects, we propose a novel approach to modeling the motion cues from optical flow field, the saliency map of the prior video frame and the motion history of change detection, which is able to distinguish the moving salient objects from diverse changing background regions. Furthermore, an effective objectness measure is proposed with intuitive geometrical interpretation to extract some reliable object and background regions, which provided as the basis to define the foreground potential, background potential, and the constraint to support saliency propagation. These potentials and the constraint are formulated into the proposed SCOM framework to generate an optimal saliency map for each frame in a video. The proposed model is extensively evaluated on the widely used challenging benchmark data sets. Experiments demonstrate that our proposed SCOM substantially outperforms the state-of-the-art saliency models.
Yuhuan Chen, Wenbin Zou, Yi Tang 0008, Xia Li 0006, Chen Xu 0004, Nikos Komodakis
IEEE Trans. Image Process.3
2017 Multi-modal metric learning for vehicle re-identification in traffic surveillance environment
abstract
Vehicle re-identification (Re-Id) aims to retrieve the same vehicle captured by disjoint cameras at different time instants from different locations, and is a challenging task mainly due to the high similarity among the captured vehicle images in surveillance environment. With the rapid development of Convolutional Neural Network (CNN), learning-based deep features have been adopted to combine with hand-crafted features to re-identify vehicles in traffic surveillance environment. However, the two kinds of features are in different feature space, and if they are fused directly together, their complementary correlation is not able to be fully explored. To address such an issue, this paper proposes a multi-modal metric learning architecture to fuse deep features and hand-crafted ones in an end-to-end optimization network, which achieves a more robust and discriminative feature representation for vehicle re-identification. The extensive experiments on a large-scale traffic surveillance vehicle dataset demonstrate that our proposed approach substantially outperforms the state-of-the-art methods on vehicle Re-Id.
Yi Tang 0008, Di Wu 0009, Zhi Jin 0002, Wenbin Zou, Xia Li 0006
ICIP1
2017 A CNN cascade for quality enhancement of compressed depth images
abstract
Transmitting depth images along with the corresponding textures enables a wide range of receiver-side 3D applications. Since each pixel on the depth images represents a corresponding 3D scene geometric information, when compressed during transmission the compression artifacts will lead to severe geometry distortions and visual perceptual degradation. To solve this problem, in this paper we proposed a convolutional neural network (CNN) cascade for suppressing the compression artifacts on depth images. According to the feature of depth images, we furthermore, adopt a weighted loss function for network training which can adaptively improve the learning efficiency and accuracy. Meanwhile, in order to overcome the limited training data problem, we audaciously trained our network on textures first and then finetune on the target depth images. To our best knowledge, few works have applied CNN on depth images targeting for compression artifacts reduction (CAR). Through extensive experiments, our proposed solution achieves higher quality for both reconstructed depth images and synthesized virtual views than the state-of-the-art methods.
Zhi Jin 0002, Lei Luo 0003, Yi Tang 0008, Wenbin Zou, Xia Li 0006
VCIP3