Lichao Zhang 0001

dblp:126/8027-1 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0003-4580-2317ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2026 FracMix: A Fractional Fourier Based Augmentation for Generalizable Person Re-Identification
abstract
Domain generalization in person re-identification (DG re-ID) aims to build models that can accurately retrieve a targeted person across cameras in arbitrary unseen domains, without access to those domains during training. Data augmentation (DA) has become a de facto solution to DG re-ID, and one line of approaches realizes this goal from a frequency-centric perspective, where the Fourier Transform (FT) is employed to mix amplitudes for novel sample generation. However, FT-based strategies are fundamentally limited by the assumption of signal stationarity, which may be inadequate to explore a broader perturbation space. To overcome this limitation, this work investigates the potential of the Fractional Fourier Transform (FrFT) for DA and proposes FracMix (Fractional Fourier based MixUp), which yields an intermediate representation that bridges spatial and frequency domains with a fractional order α, enabling the generation of richer samples. Furthermore, an Adaptive Token Pruning strategy (ATP) is designed to dynamically select the top-K salient tokens and perform perturbations exclusively to corresponding patches. This simple yet effective mechanism alleviates over-reliance on a small set of dominant patches and encourages broader contextual utilization, thereby improving robustness and generalization. Experiments on multiple DG re-ID benchmarks demonstrate the effectiveness of FracMix.
Jieru Jia, Huidi Xie, Yantao Song, Lichao Zhang 0001, Chao Li 0070
IEEE Signal Process. Lett.6
2024 Tailored Visions: Enhancing Text-to-Image Generation with Personalized Prompt Rewriting
abstract
Despite significant progress in the field, it is still challenging to create personalized visual representations that align closely with the desires and preferences of individ-ual users. This process requires users to articulate their ideas in words that are both comprehensible to the models and accurately capture their vision, posing difficul-ties for many users. In this paper, we tackle this challenge by leveraging historical user interactions with the system to enhance user prompts. We propose a novel approach that involves rewriting user prompts based on a newly collected large-scale text-to-image dataset with over 300k prompts from 3115 users. Our rewriting model enhances the expressiveness and alignment of user prompts with their intended visual outputs. Experimental results demonstrate the superiority of our methods over baseline approaches, as evidenced in our new offline evaluation method and online tests. Our code and dataset are available at https://github.com/zzjchen/Tailored-Visions
Lichao Zhang 0001, Fangsheng Weng, Lili Pan 0001, Zhen-Zhong Lan
CVPR2
2024 Benchmark for Detecting Child Pornography in Open Domain Dialogues Using Large Language Models
abstract
As large language models become increasingly prevalent, their safe and secure application, particularly in preventing the generation of child pornographic content in real-world open-domain dialogues, has become a crucial concern. Despite the urgency of this issue, research efforts are hindered by the lack of dedicated datasets for this area. Addressing this gap, we introduce a pioneering benchmark dataset specifically designed for the detection of child pornography in open-domain dialogues. Recognizing the intrinsic complexities involved in labling such data, we developed a novel Distillation-Based Recurrent Extraction method. This approach enables us to efficiently gather, annotate, and refine the data collection process. Our dataset categorizes dialogues into three distinct sections: non-pornographic, child pornographic, and adult pornographic, ensuring clear differentiation between child and adult content. Through extensive experiments, we demonstrate that LLMs including BERT, RoBERTa, LLaMA, among others, significantly enhance their detection capabilities when fine-tuned with our dataset. This improvement not only attests to the dataset’s immediate utility but also highlights its importance for guiding future research. Furthermore, our findings indicate substantial potential for further advancements in detection performance, emphasizing the critical role of our benchmark dataset in ongoing research efforts.
Zhiwei Huang 0006, Lichao Zhang 0001, Yuming Yan, Zhenyang Xiao, Zhen-Zhong Lan
IJCNN3
2023 Learning Robust Self-Attention Features for Speech Emotion Recognition with Label-Adaptive Mixup
abstract
Speech Emotion Recognition (SER) is to recognize human emotions in a natural verbal interaction scenario with machines, which is considered as a challenging problem due to the ambiguous human emotions. Despite the recent progress in SER, state-of-the-art models struggle to achieve a satisfactory performance. We propose a self-attention based method with combined use of label-adaptive mixup and center loss. By adapting label probabilities in mixup and fitting center loss to the mixup training scheme, our proposed method achieves a superior performance to the state-of-the-art methods.
Lei Kang 0002, Lichao Zhang 0001, Dazhi Jiang
ICASSP2
2022 Information Lossless Multi-modal Image Generation for RGB-T Tracking
Yufei Zha, Lichao Zhang 0001, Peng Zhang 0005, Lang Chen
PRCV (4)3
2022 Semantic-aware spatial regularization correlation filter for visual tracking
abstract
Abstract Correlation filters with convolutional neural network (CNN) features have been successfully applied to visual tracking owing to their impressive combined capability for object representation. Unfortunately, further performance improvement is limited due to unwanted boundary effects of the circular structure. In this work, through an in‐depth study of the features’ characteristics, the authors propose a novel tracking strategy to achieve simultaneous filter matching and regularization with CNN features when tracking is on the fly. With a feature decomposed transform matrix, a spatial semantic regularization is generated to reduce the boundary effect effectively during filter optimization. Before each output, the regularized filter is then back performed to match with the extracted features of a search region to find the optimum candidate. Specifically, the most important advantage of the proposed spatial semantic map is to initialize only in the first frame as all the other tracking strategies. Besides, the authors design a novel updating strategy to tackle the cases where the object is occluded or disappeared in the scene. At this time, the maximum of the map is small, even negative. A substantial experiment has been carried out on the popular benchmark tracking datasets; the reliable results have demonstrated that the authors’ method is able to outperform most of the state‐of‐the‐art tracking works in both accuracy and robustness.
Yufei Zha, Peng Zhang 0005, Lei Pu, Lichao Zhang 0001
IET Comput. Vis.4
2021 Unsupervised Cross-Modal Distillation for Thermal Infrared Tracking
abstract
The target representation learned by convolutional neural networks plays an important role in Thermal Infrared (TIR) tracking. Currently, most of the top-performing TIR trackers are still employing representations learned by the model trained on the RGB data. However, this representation does not take into account the information in the TIR modality itself, limiting the performance of TIR tracking.
Jingxian Sun 0003, Lichao Zhang 0001, Yufei Zha, Abel Gonzalez-Garcia, Peng Zhang 0005, Wei Huang 0013, Yanning Zhang 0001
ACM Multimedia2
2019 Learning the Model Update for Siamese Trackers
abstract
Siamese approaches address the visual tracking problem by extracting an appearance template from the current frame, which is used to localize the target in the next frame. In general, this template is linearly combined with the accumulated template from the previous frame, resulting in an exponential decay of information over time. While such an approach to updating has led to improved results, its simplicity limits the potential gain likely to be obtained by learning to update. Therefore, we propose to replace the handcrafted update function with a method which learns to update. We use a convolutional neural network, called UpdateNet, which given the initial template, the accumulated template and the template of the current frame aims to estimate the optimal template for the next frame. The UpdateNet is compact and can easily be integrated into existing Siamese trackers. We demonstrate the generality of the proposed approach by applying it to two Siamese trackers, SiamFC and DaSiamRPN. Extensive experiments on VOT2016, VOT2018, LaSOT, and TrackingNet datasets demonstrate that our UpdateNet effectively predicts the new target template, outperforming the standard linear update. On the large-scale TrackingNet dataset, our UpdateNet improves the results of DaSiamRPN with an absolute gain of 3.9% in terms of success score.
Lichao Zhang 0001, Abel Gonzalez-Garcia, Joost van de Weijer 0001, Martin Danelljan, Fahad Shahbaz Khan
ICCV1
2019 Synthetic Data Generation for End-to-End Thermal Infrared Tracking
abstract
The usage of both off-the-shelf and end-to-end trained deep networks have significantly improved the performance of visual tracking on RGB videos. However, the lack of large labeled datasets hampers the usage of convolutional neural networks for tracking in thermal infrared (TIR) images. Therefore, most state-of-the-art methods on tracking for TIR data are still based on handcrafted features. To address this problem, we propose to use image-to-image translation models. These models allow us to translate the abundantly available labeled RGB data to synthetic TIR data. We explore both the usage of paired and unpaired image translation models for this purpose. These methods provide us with a large labeled dataset of synthetic TIR sequences, on which we can train end-to-end optimal features for tracking. To the best of our knowledge, we are the first to train end-to-end features for TIR tracking. We perform extensive experiments on the VOT-TIR2017 dataset. We show that a network trained on a large dataset of synthetic TIR data obtains better performance than one trained on the available real TIR data. Combining both data sources leads to further improvement. In addition, when we combine the network with motion features, we outperform the state of the art with a relative gain of over 10%, clearly showing the efficiency of using synthetic data to train end-to-end TIR trackers.
Lichao Zhang 0001, Abel Gonzalez-Garcia, Joost van de Weijer 0001, Martin Danelljan, Fahad Shahbaz Khan
IEEE Trans. Image Process.1
2018 Joint Identification-Verification Model for Visual Tracking
abstract
Similarity algorithms determine the location of the target by the similarity between the template and the candidate, the most similar candidate to the template is considered as the target in visual tracking. Similarity algorithms search the most similar candidate to the template as the current estimation for visual object. In practice, most trackers only take usage of the intra-class similarity, yet the inter-class semantic separability is ignored. In this paper, a joint identification-verification model is proposed to learn the similarity with the category attribute for visual tracking. This approach constructs the cost function both on the inter-class semantic separability and intra-class similarity, firstly. Then, the training dataset is fed into the network. To the end, the discriminative features are learned in the embedding space. During tracking phase, the template and candidates are fed into the network simultaneously. Thereforce, the target will be located correctly by the similarity metric between the template and candidates in the learned embedding space. We evaluate the proposed approach on the open benchmark: OTB50 and UAV123 dataset. A large number of experimental results show that the inter-class semantic separability can increase the discrimination for the similar distractors effectively, and bootstrap the tracking performances of the trackers based on the similarity learning.
Yufei Zha, Yuanqiang Zhang, Tao Ku, Lichao Zhang 0001
ICPR5
2018 Beyond Eleven Color Names for Image Understanding
Lu Yu 0004, Lichao Zhang 0001, Joost van de Weijer 0001, Fahad Shahbaz Khan, Yongmei Cheng, C. Alejandro Párraga
Mach. Vis. Appl.2
2016 Robust and fast visual tracking via spatial kernel phase correlation filter
Lichao Zhang 0001, Duyan Bi, Yufei Zha, Hongxun Wang, Tao Ku
Neurocomputing1