Senmao Tian

dblp:312/8130 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
11since 2021 · last 2026
0009-0001-2459-0290ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Sampling Control for Imbalanced Calibration in Semi-Supervised Learning
abstract
Class imbalance remains a critical challenge in semi-supervised learning (SSL), especially when distributional mismatches between labeled and unlabeled data lead to biased classification. Although existing methods address this issue by adjusting logits based on the estimated class distribution of unlabeled data, they often handle model imbalance in a coarse-grained manner, conflating data imbalance with bias arising from varying class-specific learning difficulties. To address this issue, we propose a unified framework, SC-SSL, which suppresses model bias through decoupled sampling control. During training, we identify the key variables for sampling control under ideal conditions. By introducing a classifier with explicit expansion capability and adaptively adjusting sampling probabilities across different data distributions, SC-SSL mitigates feature-level imbalance for minority classes. In the inference phase, we further analyze the weight imbalance of the linear classifier and apply post-hoc sampling control with an optimization bias vector to directly calibrate the logits. Extensive experiments across various benchmark datasets and distribution settings validate the consistency and state-of-the-art performance of SC-SSL.
Senmao Tian
AAAI1
2026 GaitAdapt: Continual learning for evolving gait recognition
Shunli Zhang 0005, Senmao Tian
Pattern Recognit.4
2026 Compression-Oriented Video Super-Resolution
abstract
Current compressed video super-resolution methods have achieved promising performance, but they often assume that an input video is compressed under low-delay configurations. However, under random access configurations, those methods might struggle to leverage the metadata effectively due to the large variations of metadata in different compression configurations. In this work, we propose a Compression-Oriented Video Super-Resolution (COVSR) method that can address video super-resolution for both low-delay and random-access configurations. Specifically, we first introduce an efficient compression-aware propagation (ECAP) module that dynamically adjusts propagation routes in accordance with the compression configurations. Since existing methods require reconstructing frames in a frame-by-frame manner, it is difficult to achieve efficient parallelization. However, we find that by slightly relaxing sequential dependencies, our ECAP can significantly improve inference speed. Furthermore, existing methods typically perform alignment between adjacent frames or adjacent features. However, since ECAP may propagate features along non-adjacent reference routes, it introduces new challenges for accurate cross-frame feature alignment. In response, we propose a metadata-driven alignment (MDA) module that refines cross-frame motion vectors into dense, feature-level flow offsets, enabling precise alignment across temporally distant features. Extensive experimental results demonstrate that our COVSR not only achieves efficient and superior super-resolution performance but also is generalizable to various compression configurations. Our code will be available at https://covsr.github.io.
Yanbin Liu 0003, Ming Lu 0002, Zhuojie Wu, Senmao Tian, Yandong Guo, Xin Yu 0002
IEEE Trans. Image Process.5
2025 QGait: Toward Accurate Quantization for Gait Recognition
abstract
Existing deep learning methods have made significant progress in gait recognition. Quantization can facilitate the application of gait models as a model-agnostic general compression technique. Typically, appearance-based models binarize inputs into silhouette sequences. However, mainstream quantization methods prioritize minimizing task loss over quantization error, which is detrimental to gait recognition with binarized inputs. To address this, we propose a differentiable soft quantizer, which better simulates the gradient of the round function during backpropagation. This enables the network to learn from subtle input perturbations. However, our theoretical analysis and empirical studies reveal that directly applying the soft quantizer can hinder network convergence. We addressed this issue by adopting a two-stage training strategy, introducing a soft quantizer during the fine-tuning phase. However, in the first stage of training, we observed a significant change in the output distribution of different samples in the feature space compared to the full-precision network. It is this change that led to a loss in performance. Based on this, we propose an Inter-class Distance-guided Calibration (IDC) strategy to preserve the relative distance between the embeddings of samples with different labels. Extensive experiments validate the effectiveness of our approach, demonstrating state-of-the-art accuracy across various settings and datasets. The code will be made publicly available.
Senmao Tian, Gangyi Hong, JingJie Wang, Xin Yu 0002, Shunli Zhang 0005
IJCB1
2025 FAST: Facial Avatar Animation via Spatial-Temporal Aggregation
abstract
Facial avatar animation methods animate virtual characters based on natural human performance and have been widely used in film and game production. Despite its practical relevance, academic research in this area has been scarce, particularly in the deep learning era. To fill this gap, we propose Facial Avatar Animation via Spatial-Temporal Aggregation (FAST), which leverages Memory-based Spatial-Temporal Aggregation (MSTA) to capture both spatial and temporal dependencies in facial animation. However, directly regressing blendshape and pose coefficients introduces uncertainty and reduces interpretability. Thus, we introduce Implicit Latent Representations (ILRs), which learn the semantic correspondence between the predicted results and blendshapes/poses, enhancing both model interpretability and the vivid tracking of facial expressions. Additionally, monocular RGB-based pose estimation suffers from depth ambiguity that destabilizes animation. To address this, we incorporate a Semantic-Aware Rigid Prior (SRP) to enhance the rigid stability of the animation. To tackle the lack of blendshape coefficient annotations in existing datasets and support the advancement of avatar animation methods, we present the BS500 dataset, which includes 500 individuals and over 4.5 million frames with diverse demographic features such as gender, age, expression, and head pose. Extensive experiments show the superiority of the FAST method over existing approaches. The code and dataset will be released.
Gangyi Hong, Senmao Tian, Xiangyi Chen, Hui Zhang 0013
ICME3
2024 Adaptive Quantization with Mixed-Precision Based on Low-Cost Proxy
abstract
It is critical to deploy complicated neural network models on hardware with limited resources. This paper proposes a novel model quantization method, named the Low-Cost Proxy-Based Adaptive Mixed-Precision Model Quantization (LCPAQ), which contains three key modules. The hardware-aware module is designed by considering the hardware limitations, while an adaptive mixed-precision quantization module is developed to evaluate the quantization sensitivity by using the Hessian matrix and Pareto frontier techniques. Integer linear programming is used to fine-tune the quantization across different layers. Then the low-cost proxy neural architecture search module efficiently explores the ideal quantization hyperparameters. Experiments on the ImageNet demonstrate that the proposed LCPAQ achieves comparable or superior quantization accuracy to existing mixed-precision models. Notably, LCPAQ achieves 1/200 of the search time compared with existing methods, which provides a shortcut in practical quantization use for resource-limited devices.
Senmao Tian
ICASSP3
2024 Dual Attention Enhanced Transformer for Image Defocus Deblurring
abstract
Image Defocus Deblurring remains a challenging problem due to the uncertainty of the blurred region and the varying depth of field. Although the convolutional neural network (CNN) has achieved promising results on this task, its limited receptive field and static weights hinder the restoration performance. In contrast, Transformer models are able to mitigate the weaknesses of CNN. However, recent Transformer-based models that deal with Image Defocus Deblurring only utilize self-attention from either spatial or channel dimension, which neglects cross-dimensional information essential for restoration. In this paper, we propose a novel Transformer model, Dual Attention Enhanced Transformer (DAEformer), for Image Defocus Deblurring. DAEformer combines self-attention from both spatial and channel dimensions, meanwhile applying auxiliary enhanced attention modules. We present Spatial Attention Enhanced Block (SAEB) and Channel Attention Enhanced Block (CAEB), which not only fuse the spatial and channel information within blocks but also enhance details. Furthermore, we design a progressive hierarchical architecture that applies SAEB/CAEB at different levels to model distinct information and facilitate fusion across blocks. Experimental results demonstrate that DAEformer can achieve state-of-the-art results on the dual-pixel dataset.
Yuhang He 0002, Senmao Tian
ICIP2
2023 CABM: Content-Aware Bit Mapping for Single Image Super-Resolution Network with Large Input
abstract
With the development of high-definition display devices, the practical scenario of Super-Resolution (SR) usually needs to super-resolve large input like 2K to higher resolution (4K/8K). To reduce the computational and memory cost, current methods first split the large input into local patches and then merge the SR patches into the output. These methods adaptively allocate a subnet for each patch. Quantization is a very important technique for network acceleration and has been used to design the subnets. Current methods train an MLP bit selector to determine the propoer bit for each layer. However, they uniformly sample subnets for training, making simple subnets overfitted and complicated subnets underfitted. Therefore, the trained bit selector fails to determine the optimal bit. Apart from this, the introduced bit selector brings additional cost to each layer of the$SR$network. In this paper, we propose a novel method named Content-Aware Bit Mapping (CABM), which can remove the bit selector without any performance loss. CABM also learns a bit selector for each layer during training. After training, we analyze the relation between the edge information of an input patch and the bit of each layer. We observe that the edge information can be an effective metric for the selected bit. Therefore, we design a strategy to build an Edge-to-Bit lookup table that maps the edge score of a patch to the bit of each layer during inference. The bit configuration of SR network can be determined by the lookup tables of all layers. Our strategy can find better bit configuration, resulting in more efficient mixed precision networks. We conduct detailed experiments to demonstrate the generalization ability of our method. The code will be released.
Senmao Tian, Ming Lu 0002, Jiaming Liu 0003, Yandong Guo, Yurong Chen 0001, Shunli Zhang 0005
CVPR1
2023 A Comprehensive Comparison of Projections in Omnidirectional Super-Resolution
abstract
Super-Resolution (SR) has gained increasing research attention over the past few years. With the development of Deep Neural Networks (DNNs), many super-resolution methods based on DNNs have been proposed. Although most of these methods are aimed at ordinary frames, there are few works on super-resolution of omnidirectional frames. In these works, omnidirectional frames are projected from the 3D sphere to a 2D plane by Equi-Rectangular Projection (ERP). Although ERP has been widely used for projection, it has severe projection distortion near poles. Current DNN-based SR methods use 2D convolution modules, which is more suitable for the regular grid. In this paper, we find that different projection methods have great impact on the performance of DNNs. To study this problem, a comprehensive comparison of projections in omnidirectional super-resolution is conducted. We compare the SR results of different projection methods. Experimental results show that Equi-Angular cube map projection (EAC), which has minimal distortion, achieves the best result in terms of WS-PSNR compared with other projections. Code and data will be released.
Huicheng Pi, Senmao Tian, Ming Lu 0002, Jiaming Liu 0003, Yandong Guo, Shunli Zhang 0005
ICASSP2
2023 HQRetouch: Learning Professional Face Retouching Via Masked Feature Fusion and Semantic-Aware Modulation
abstract
Face retouching is a crucial technique for many consumer-level products. The goal of face retouching is to remove skin imperfections and preserve facial details simultaneously. However, it usually requires tedious manual work to achieve professional retouching effect. With the advent of Deep Neural Networks (DNNs), some methods were recently proposed to complete the task of face retouching automatically by using DNNs. They divide a portrait photo into local patches and train a DNN for face retouching. Although they can produce professional results automatically, there are still some limitations. Firstly, the network architecture fails to preserve sufficient facial details. Secondly, the facial semantic information is ignored when dividing a photo into some local patches. In this paper, we propose a novel method to solve these limitations. We first introduce the Masked Feature Fusion (MFF) module to a UNet, enabling the network to better preserve details in facial regions. Then, we exploit the semantic information by the Semantic-Aware Modulation (SAM) module, further boosting the retouching performance. Experiments on the recent public dataset Flickr-Faces-HQ-Retouched (FFHQR) demonstrate the effectiveness of our method. The code will be released.
Gangyi Hong, Fangshi Wang, Senmao Tian, Ming Lu 0002, Jiaming Liu 0003, Shunli Zhang 0005
ICIP3
2021 Blind Image Deblurring Based on Dual Attention Network and 2D Blur Kernel Estimation
abstract
In the problem of image deblurring, the restoration of details in severely blurred images has always been difficult. In this paper, we focus on effectively eliminating the ringing artifact and wrinkles that appear after deburring, and propose a novel blind debluring method based on dual attention deep image prior (DADIP) network and 2-dimensional (2D) blur kernel estimation with convolutional neural network (CNN). In the DADIP network, the dual attention mechanism is firstly combined with squeeze and excitation network (SENet), which greatly improves the restoration effect of image details. More importantly, the 2D blur kernel estimation approach via CNN is developed to suppress the ringing artifact of the image, which significantly outperforms previous fully connected network based methods. Experiments show that our deblurring approach achieves superior performance compared with most existing methods.
Senmao Tian, Shunli Zhang 0005, Beibei Lin
ICIP1