Tiantian Yan

dblp:268/8178 · DBLP profile ↗
← Back
17ranked-venue papers
6as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 11 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Cross-modal local fine-grained feature localization and alignment for text-to-image person re-identification
Tiantian Yan, Jiexiang Fang
Multim. Syst.1
2026 Part-aware attention calibration and dynamic enhancement for low-resolution fine-grained image recognition
Tiantian Yan, Jiexiang Fang
Pattern Recognit.1
2025 Enhanced Feature Representations for Low-Resolution Fine-Grained Image Recognition via Categorical Knowledge Guidance
Tiantian Yan, Bao-Li Sun, Jie Mu, Wei Wang 0077
J. Comput. Sci. Technol.1
2025 Structure-preserving dental plaque segmentation via dynamically complementary information interaction
Rui Xu 0002, Baoli Sun, Tiantian Yan, Zhihui Wang 0001
Multim. Syst.4
2025 An unsupervised medical image registration network for intelligent medical education
Jie Mu, Jing Zhang 0037, Tiantian Yan, Wei Wang 0335, Hua Zhang 0008, Wenqi Ren
Neural Comput. Appl.5
2025 LarTap: A Luminance-Aware Framework With Text-Correlation Priors for Multi-Exposure Image Fusion
abstract
Conventional imaging devices often struggle to produce high-dynamic-range (HDR) images that accurately represent natural scenes. To overcome this limitation, multi-exposure image fusion (MEF) techniques have been introduced as a viable solution. Existing MEF approaches aim to enhance performance by optimizing or searching architectures. However, they face challenges in precise feature extraction and scene reconstruction, leading to distortion in the fused images. Additionally, most methods do not adequately address luminance variations across different image regions, which may result in the loss of essential details. To address these challenges, we present a novel luminance-aware MEF framework that integrates text-correlation priors (LarTap). By embedding textual information into fusion process, the proposed framework enhances content extraction and comprehension. Specifically, it consist of two key components: the text-image correlation network (N1) and the multi-exposure fusion network (N2). First, N1 performs correlation training to achieve a holistic alignment between text and image pairs. Its iterative vision encoders (VEs) generate text-correlated prior knowledge to facilitate the fusion process in N2. Second, N2 leverages these priors for scene reconstruction and dynamically adjusts luminance based on comparative perception. Extensive experiments on multiple datasets demonstrate that LarTap outperforms state-of-the-art methods. The source code is available at https://github.com/EnLong-wang/LarTap.
Enlong Wang, Jiawei Li 0016, Tiantian Yan, Jia Lei 0001, Shihua Zhou, Bin Wang 0005, Jinyuan Liu 0001, Nikola K. Kasabov
IEEE Trans. Circuits Syst. Video Technol.3
2025 Multimodal Large Language Model with LoRA Fine-Tuning for Multimodal Sentiment Analysis
abstract
Multimodal sentiment analysis has become a popular research topic in recent years. However, existing methods have two unaddressed limitations: (1) they use limited supervised labels to train models, which makes it impossible for model to fully learn sentiments in different modal data; (2) they employ text and image pre-trained models trained in different unimodal tasks to extract different modal features, so that the extracted features cannot take into account the interactive information between image and text. To solve these problems, in this paper we propose a Vision-Language Contrastive Learning network (VLCLNet). First, we introduce a pre-trained Large Language Model (LLM), which is trained from vast quantities of multimodal data, has better understanding ability for image and text contents, thus being effectively applied to different tasks while requiring few amount of labelled training data. Second, we adapt a Multimodal Large Language Model (MLLM), BLIP-2 (Bootstrapping Language-Image Pre-training) network, to extract multimodal fusion feature. Such MLLM can fully consider the correlation between images and texts when extracting features. In addition, due to the discrepancy between the pre-training task and the sentiment analysis task, the pre-trained model will output the suboptimal prediction results. We use Low-Rank Adaptation (LoRA) fine-tuning strategy to update the model parameters on sentiment analysis task, which avoids the issue of inconsistent task between pre-training task and downstream task. Experiments verify that the proposed VLCLNet is superior to other strong baselines.
Jie Mu, Wei Wang 0335, Tiantian Yan, Guanglu Wang
ACM Trans. Intell. Syst. Technol.4
2024 Segmentation-Driven Infrared and Visible Image Fusion Via Transformer-Enhanced Architecture Searching
abstract
A series of infrared and visible image fusion (IVIF) methods have emerged to improve the performance of segmentation task. However, existing perception-focused IVIF methods take visual effects and semantic information as a unified goal for training, ignoring the task conflicts. Moreover, these methods often involve manually designed modules, which are laborious and suboptimal. To solve the problems, we propose a collaborative feature learning framework based on neural architecture search (NAS). Specifically, we extract shared features of fusion and segmentation tasks into a unified space and separately process task objectives through a dual decoder. In light of the essential role that semantic information plays in the segmentation task, we construct a hybrid search space with transformers incorporated to enhance context dependence handling. Our method undergoes extensive experiments, showcasing exceptional visual effects and significant enhancements in segmentation tasks compared to other state-of-the-art methods.
Hongming Fu, Guanyao Wu, Zhu Liu 0004, Tiantian Yan, Jinyuan Liu 0001
ICASSP4
2024 Adaptive Multi-Exposure Fusion for Enhanced Neural Radiance Fields
abstract
Neural Radiance Fields (NeRF) have revolutionized 3D scene modeling and rendering. However, their performance dips when handling images with diverse exposure levels, mainly due to the intricate luminance dynamics. Addressing this, we present an innovative method that proficiently models and renders images across a spectrum of exposure conditions. Our approach utilizes an unsupervised classifier-generator structure for HDR fusion, significantly enhancing NeRF’s ability to comprehend and adjust to light variations, leading to the generation of images with appropriate brightness. Extensive evaluations on the LOM[1] and LOL[2] datasets underscore our method’s edge. Our approach significantly improves the task of novel view synthesis for multi-exposure images, attaining state-of-the-art results.
Yang Zou 0004, Xingyuan Li 0005, Zhiying Jiang, Tiantian Yan, Jinyuan Liu 0001
ICASSP4
2024 Learning a Holistic-Specific color transformer with Couple Contrastive constraints for underwater image enhancement and beyond
Debin Wei, Hongji Xie, Zengxi Zhang, Tiantian Yan
J. Vis. Commun. Image Represent.4
2024 Semantic-Aware Detail Search and Feature Constraint for Cross-Resolution Person Re-Identification
abstract
Cross-resolution person re-identification (CRReID) task devotes to identifying the same person from cross-resolution and cross-camera images. Existing CRReID methods learn the identity features of persons by jointly training the super-resolution (SR) and recognition models. These methods achieve sub-optimal results because the design of SR techniques is mostly oriented towards the visual quality of images rather than recognition tasks. To address this deficiency, we propose a Semantic-Aware detail Search and feature Constraint Network (SASC-Net). Specifically, we propose the semantic-aware detail search (SDS) module that is used to customize an SR module by perceiving identity-related semantic information. Then, we devise an Intra-Scale and Inter-Scale Feature Constraint loss function. It ensures that the affinity relationships of the semantic features of repaired images are close to that of high-resolution (HR) images at the scale level, reducing the solution space of the SDS module and promoting the identification module to focus on more discriminative pedestrian features. The effectiveness of our proposed method is validated by the experimental results on five cross-resolution person datasets.
Tiantian Yan, Xin Yang 0011, Qiang Zhang 0008
IEEE Signal Process. Lett.2
2024 Digging into Depth and Color Spaces: A Mapping Constraint Network for Depth Super-Resolution
abstract
Scene depth super-resolution (DSR) poses an inherently ill-posed problem due to the extremely large space of one-to-many mapping functions from a given low-resolution (LR) depth map, which possesses limited depth information, to multiple plausible high-resolution (HR) depth maps. This characteristic renders the task highly challenging, as identifying an optimal solution becomes significantly intricate amidst this multitude of potential mappings. While simplistic constraints have been proposed to address the DSR task, the relationship between LR and HR depth maps and the color image has not been thoroughly investigated. In this paper, we introduce a novel mapping constraint network (MCNet) that incorporates additional constraints derived from both LR depth maps and color images. This integration aims to optimize the space of mapping functions and enhance the performance of DSR. Specifically, alongside the primary DSR network (DSRNet) dedicated to learning LR-to-HR mapping, we have developed an auxiliary degradation network (ADNet) that operates in reverse, generating the LR depth map from the reconstructed HR depth map to obtain depth features in LR space. To enhance the learning process of DSRNet in LR-to-HR mapping, we introduce two mapping constraints in LR space: (1) the cycle-consistent constraint, which offers additional supervision by establishing a closed loop between LR-to-HR and HR-to-LR mappings, and (2) the region-level contrastive constraint, aimed at reinforcing region-specific HR representations by explicitly modeling the consistency between LR and HR spaces. To leverage the color image effectively, we introduce a feature screening module to adaptively fuse color features at different layers, which can simultaneously maintain strong structural context and suppress texture distraction through subspace generation and image projection. Comprehensive experimental results across synthetic and real-world benchmark datasets unequivocally demonstrate the superiority of our proposed method over state-of-the-art DSR methods. Our MCNet achieves an average MAD reduction of 3.7% and 7.5% over state-of-the-art DSR method for ×8 and ×16 cases on Milddleburry dataset, respectively, without incurring additional costs during inference.
Baoli Sun, Tiantian Yan, Xinchen Ye, Zhihui Wang 0001, Zhiyong Wang 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2024 Discriminative Segment Focus Network for Fine-grained Video Action Recognition
abstract
Fine-grained video action recognition aims at identifying minor and discriminative variations among fine categories of actions. While many recent action recognition methods have been proposed to better model spatio-temporal representations, how to model the interactions among discriminative atomic actions to effectively characterize inter-class and intra-class variations has been neglected, which is vital for understanding fine-grained actions. In this work, we devise a Discriminative Segment Focus Network (DSFNet) to mine the discriminability of segment correlations and localize discriminative action-relevant segments for fine-grained video action recognition. Firstly, we propose a hierarchic correlation reasoning (HCR) module which explicitly establishes correlations between different segments at multiple temporal scales and enhances each segment by exploiting the correlations with other segments. Secondly, a discriminative segment focus (DSF) module is devised to localize the most action-relevant segments from the enhanced representations of HCR by enforcing the consistency between the discriminability and the classification confidence of a given segment with a consistency constraint. Finally, these localized segment representations are combined with the global action representation of the whole video for boosting final recognition. Extensive experimental results on two fine-grained action recognition datasets, i.e., FineGym and Diving48, and two action recognition datasets, i.e., Kinetics400 and Something-Something, demonstrate the effectiveness of our approach compared with the state-of-the-art methods.
Baoli Sun, Xinchen Ye, Tiantian Yan, Zhihui Wang 0001, Zhiyong Wang 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2022 Fine-grained Action Recognition with Robust Motion Representation Decoupling and Concentration
abstract
Fine-grained action recognition is a challenging task that requires identifying discriminative and subtle motion variations among fine-grained action classes. Existing methods typically focus on spatio-temporal feature extraction and long-temporal modeling to characterize complex spatio-temporal patterns of fine-grained actions. However, the learned spatio-temporal features without explicit motion modeling may emphasize more on visual appearance than on motion, which could compromise the learning of effective motion features required for fine-grained temporal reasoning. Therefore, how to decouple robust motion representations from the spatio-temporal features and further effectively leverage them to enhance the learning of discriminative features still remains less explored, which is crucial for fine-grained action recognition. In this paper, we propose a motion representation decoupling and concentration network (MDCNet) to address these two key issues. First, we devise a motion representation decoupling (MRD) module to disentangle the spatio-temporal representation into appearance and motion features through contrastive learning from video and segment views. Next, in the proposed motion representation concentration (MRC) module, the decoupled motion representations are further leveraged to learn a universal motion prototype shared across all the instances of each action class. Finally, we project the decoupled motion features onto all the motion prototypes through semantic relations to obtain the concentrated action-relevant features for each action class, which can effectively characterize the temporal distinctions of fine-grained actions for improved recognition performance. Comprehensive experimental results on four widely used action recognition benchmarks, i.e., FineGym, Diving48, Kinetics400 and Something-Something, clearly demonstrate the superiority of our proposed method in comparison with other state-of-the-art ones.
Baoli Sun, Xinchen Ye, Tiantian Yan, Zhihui Wang 0001, Zhiyong Wang 0001
ACM Multimedia3
2022 Discriminative information restoration and extraction for weakly supervised low-resolution fine-grained image recognition
Tiantian Yan, Zhongxuan Luo, Zhihui Wang 0001
Pattern Recognit.1
2022 Discriminative Feature Mining and Enhancement Network for Low-Resolution Fine-Grained Image Recognition
abstract
Existing fine-grained image recognition methods are difficult to learn complete discriminative features from low-resolution (LR) data, because the original subtle inter-class distinctions become slimmer with the reduction of the image resolution. Besides, existing methods of LR fine-grained image recognition and general LR image recognition only consider the restoration and extraction of global discriminative features, ignoring unreliable local fine-grained details can be detrimental to final recognition. To address the above problems, we propose a multi-tasking framework, discriminative feature mining and enhancement network (DME-Net), for the LR fine-grained image recognition task, which aims to capture the reliable object descriptions from macro and micro perspectives, respectively. Macroscopically, we train the framework’s ability to recover and extract global discriminative features based on the whole images. Microscopically, we purposefully reinforce the framework’s ability to repair and capture the local discriminative details on the mined informative parts. To precisely excavate the most potential parts, we design an informative part mining (IPM) module, in which we firstly employ a part generation layer to predict several part masks that focus on different discriminative parts under the guidance of discrepancy loss and discriminant loss. Then we introduce a part selection (PS) submodule to further screen out a group of most informative parts from the predicted part masks according to their corresponding scores, which measure the semantic correlation degree of each part to the others. Experimental results on three benchmark datasets and one retail product dataset consistently show that our proposed framework can significantly boost the performance of the baseline model. Besides, extensive ablation studies are conducted, which further prove the effectiveness of each component of our designs.
Tiantian Yan, Baoli Sun, Zhihui Wang 0001, Zhongxuan Luo
IEEE Trans. Circuits Syst. Video Technol.1
2020 Progressive learning for weakly supervised fine-grained classification
Tiantian Yan, Shijie Wang 0003, Zhihui Wang 0001, Zhongxuan Luo
Signal Process.1