VLDB 2026 Research / reviewers in the wild / expert
Lihua Tian
dblp:27/7827
· DBLP profile ↗
38ranked-venue papers
5as first author
19since 2021 · last 2026
0000-0002-7206-4332ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 13 · 12 since 2021Systems, architecture and hardware · 4 · 1 first-author · 3 since 2021Computer networks · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A HEVC video steganalysis algorithm for transform unit partition modes
Mingyuan Cao, Lihua Tian, Chen Li 0033 |
Expert Syst. Appl. | 2 |
| 2026 | No-reference image quality assessment via bidirectional feature fusion and regional distortion extraction
Xiong Nie, Lihua Tian, Chen Li 0033 |
Neurocomputing | 2 |
| 2026 | Force everything to one: Targeted output redirection against diffusion-based customization
Xingzhi Xu, Lihua Tian, Shuhui Wang |
Neurocomputing | 2 |
| 2026 | Robust image watermarking algorithm based on generative adversarial networks and dynamic attention
Jiahao Yan, Lihua Tian, Xingzhi Xu, Chen Li 0033 |
Neurocomputing | 2 |
| 2025 | Research on Authenticity Identification of Cigarettes Based on YOLOv5-ResNet34abstractChina is the largest producer and consumer of cigarettes in the world. However, counterfeit and substandard cigarettes frequently appear in the market, which not only causes financial losses to the country but also disrupts normal market order and poses serious threats to public health. Traditional methods for authenticity identification of cigarettes are plagued by low efficiency, a high reliance on manual labor, and time-consuming procedures. These limitations render them increasingly ineffective against complex and diverse counterfeiting techniques. In this paper, we propose a method for authenticity identification of cigarettes by integrating YOLOv5 and ResNet34, aiming to improve automation and accuracy. The method operates in two stages. In the first stage, a pre-trained YOLOv5 network rapidly detects and crops cigarette pack images, removing distracting backgrounds to produce cleaner data. In the second stage, a ResNet34 network extracts finegrained features from these processed images to classify their authenticity. We construct a dataset comprising images of genuine and counterfeit cigarettes from 12 mainstream brands to train a ResNet34 model. Experiments demonstrate that our method achieves a classification accuracy of 94.464 % on the test set. This result validates the effectiveness and adaptability of our method in different environments. Furthermore, the method significantly reduces the need for human intervention, providing reliable technical support for both regulatory agencies and consumers. Baohong Gao, Shuai Wei, Shesheng Zhang, Jiang Jia, Wenlong Hu, Lihua Tian, Libo Xie, Xingzhi Xu |
ICPADS | 8 |
| 2025 | Robust watermarking algorithm based on dual-branch and adaptive noise weighting
Tilin Gu, Jiahao Yan, Lihua Tian |
Multim. Tools Appl. | 4 |
| 2024 | LaWNet: Audio-Visual Emotion Recognition by Listening and WatchingabstractAudio-Visual Emotion Recognition (AVER) represents an emerging and highly innovative research domain that exploits the combined utilization of audio and visual cues to enhance the comprehension and interpretation of human emotions. In this study, we introduce a novel model for Audio-Visual Emotion Recognition, referred to as LaWNet, which effectively encodes and fuses the audio signal and the Mel spectrogram derived from the audio signal as unified audio features, circumventing the loss of latent features that typically occurs during audio feature extraction in previous methodologies. Furthermore, we propose a Random Sampling method that addresses the issues of inadequate robustness and suboptimal utilization of video data observed in prior techniques. Additionally, we devise a cross-attention based fusion module to facilitate the fusion of audio-visual features. Through comprehensive experiments, we substantiate the efficacy of our proposed methodology, as evidenced by the outstanding performance achieved by the LaWNet model on two widely employed datasets, namely RAVDESS and CREMA-D, attaining accuracy rates of 99.31% and 91.53% respectively. These remarkable results significantly surpass the performance of previous methods and establish a new state-of-the-art benchmark within the realm of Audio-Visual Emotion Recognition. Kailei Cheng, Lihua Tian |
IJCNN | 2 |
| 2024 | A scheduling algorithm based on critical factors for heterogeneous multicore processorsabstractSummary As the development of chip manufacturing technology slows down, high‐performance processors often have high energy consumption and high heat generation. Therefore, heterogeneous multi‐core processors become more and more popular, and the heterogeneous multi‐core processors is adopted to execute programs. At present, the general program consists of multiple threads. To reach goals of accelerating program execution and reducing energy consumption and heat generation of system, a suitable thread scheduling algorithm for heterogeneous multi‐core processors is needed. In this article, a thread scheduling algorithm based on multiple critical scheduling factors is proposed. First, a prediction model of thread performance and energy consumption is used to predict the core sensitivity of threads. Then, critical threads are judged and accelerated by collecting the synchronization information between threads. Finally, the load balancing method based on the computing power of cores and the core sensitivity of threads is employed to perform system load balancing, which ensures the fairness of the scheduling. Several experiments are provided, and the results show that the proposed algorithm can obtain better performance of thread schedule. Chen Li 0033, Ziniu Lin, Lihua Tian, Bin Zhang 0022 |
Concurr. Comput. Pract. Exp. | 3 |
| 2024 | Generative adversarial network for semi-supervised image captioning
Lihua Tian |
Comput. Vis. Image Underst. | 3 |
| 2024 | Clustering-based mask recovery for image captioning
Chen Li 0033, Lihua Tian |
Neurocomputing | 3 |
| 2024 | A multi-embedding domain video steganography algorithm based on TU partitioning and intra prediction modeabstractWith High Efficiency Video Coding (HEVC) becoming a popular video coding standard in the world, combing the HEVC with data hiding methods is a complex task. In this paper, a multi-embedded video steganography algorithm based on Intra Prediction Mode (IPM) and Transform Unit (TU) partition is proposed. By analyzing the coding characteristics of the TU partition and the IPM, we select embedded regions that are relatively independent as much as possible. Then, suitable distortion functions are designed for different embedding domains. Using the minimum distortion framework of the STC model to embed messages. The experimental results indicate that the proposed method achieves a maximum embedding capacity 1.5 times larger than that of traditional single domain embedding methods, while maintaining good visual quality. Heyu Xing, Lihua Tian, Mingyuan Cao, Chen Li 0033 |
Neurocomputing | 2 |
| 2024 | A high capacity video steganography based on intra luma and chroma modes
Heyu Xing, Lihua Tian, Chen Li 0033 |
Multim. Tools Appl. | 2 |
| 2024 | A Dialogues Summarization Algorithm Based on Multi-task LearningabstractAbstract With the continuous advancement of social information, the number of texts in the form of dialogue between individuals has exponentially increased. However, it is very challenging to review the previous dialogue content before initiating a new conversation. In view of the above background, a new dialogue summarization algorithm based on multi-task learning is first proposed in the paper. Specifically, Minimum Risk Training is used as the loss function to alleviate the problem of inconsistent goals between the training phase and the testing phase. Then, in order to deal with the problem that the model cannot effectively distinguish gender pronouns, a gender pronoun discrimination auxiliary task based on contrast learning is designed to help the model learn to distinguish different gender pronouns. Finally, an auxiliary task of reducing exposure bias is introduced, which involves incorporating the summary generated during inference into another round of training to reduce the difference between the decoder inputs during the training and testing stages. Experimental results show that our model outperforms strong baselines on three public dialogue summarization datasets: SAMSUM, DialogSum, and CSDS. Chen Li 0033, Jiajing Liang, Lihua Tian |
Neural Process. Lett. | 4 |
| 2024 | A Domain Adaptive Semantic Segmentation Method Using Contrastive Learning and Data AugmentationabstractAbstract For semantic segmentation tasks, it is expensive to get pixel-level annotations on real images. Domain adaptation eliminates this process by transferring networks trained on synthetic images to real-world images. As one of the mainstream approaches to domain adaptation, most of the self-training based domain adaptive methods focus on how to select high confidence pseudo-labels, i.e., to obtain domain invariant knowledge indirectly. A more direct means to explicitly align the data of the source and target domains globally and locally is lacking. Meanwhile, the target features obtained by traditional self-training methods are relatively scattered and cannot be aggregated in a relatively compact space. We offer an approach that utilizes data augmentation and contrastive learning in this paper to perform more effective knowledge migration with the basis of self-training. Specifically, the style migration and image mixing modules are first introduced for data augmentation to cope with the problem of large domain gaps in the source and target domains. To assure the aggregation of features from the same class and the discriminability of features from other classes during the training process, we propose a multi-scale pixel-level contrastive learning module. What’s more, a cross-scale contrastive learning module is proposed to help each level of the model gain the capability to obtain more information on the basis of its own original task. Experiments show that our final trained model can effectively classify the images from target domain. Yixiao Xiang, Lihua Tian, Chen Li 0033 |
Neural Process. Lett. | 2 |
| 2024 | Variable length deep cross-modal hashing based on Cauchy probability function
Chen Li 0033, Zhuotong Liu, Ziniu Lin, Lihua Tian |
Wirel. Networks | 5 |
| 2022 | Singing Melody Extraction Based on Combined Frequency-Temporal Attention and Attentional Feature Fusion with Self-AttentionabstractThe main melody extraction of polyphonic music is a challenging task for music information retrieval. Traditional convolutional neural networks, recurrent neural networks have effectively improved this task. In recent years, with the development of attention mechanism in neural networks, the frequency and time attention information of audio has been fully exploited, and the amplitude properties of audio can also be better integrated with a good fusion module. This paper improves the frequency-temporal attention based on others’ prior work. By extracting the attention information with the frequency-temporal attention and performing additive fusion of features, the combined frequency-temporal attention is obtained. Then we apply attentional feature fusion based on multi-scale channel attention, and finally the temporal dependencies are learned through the self-attention module. Our experimental results on four datasets demonstrate that our model outperforms existing models. Xi Qi, Lihua Tian, Chen Li 0033, Jiahui Yan |
ISM | 2 |
| 2022 | A Stagewise Deep Learning Framework for Tooth Instance Segmentation in CBCT Images
Lihua Tian, Chen Li 0033, Jianwei Ye, Weimin Yu |
PRICAI (1) | 2 |
| 2021 | Deep hashing based on triplet labels and quantitative regularization term with exponential convergenceabstractSummary Due to the outstanding performance of the deep network architecture–based hash for data storage and retrieval in recent years, it has been widely applied in massive image retrieval. Most previous approaches have not paid attention to the significant effect on the hash model of quantization error during the learning process. Furthermore, the saturated loss function may result in similar hash codes being generated by large‐difference images. The underuse of classification information in the training process also brings about poor performance by hash codes in retrieval assignments. In this paper, we propose a novel quantitative regularization term with an exponential convergence rate to minimize the impact of quantization error on the model and accelerate the convergence speed of the network. In the training process, to resolve the dilemma caused by saturation loss functions, a new sigmoid function with a slope parameter that can be changed automatically according to the number of iterations is proposed. For the sufficient application of classification information, triplet labels and image labels are used in parallel under the same framework by integrating image labels into the output layer. The experiment results indicate that our algorithm is superior to several advanced hash methods on two standard datasets. Zhuotong Liu, Chen Li 0033, Lihua Tian |
Concurr. Comput. Pract. Exp. | 3 |
| 2021 | DEAttack: A differential evolution based attack method for the robustness evaluation of medical image segmentation
Xiangxiang Cui, Shi Chang, Chen Li 0033, Bin Kong 0001, Lihua Tian, Meng Yang 0026, Yenan Wu, Zhongyu Li 0002 |
Neurocomputing | 5 |
| 2020 | Audio Steganography Algorithm Based on Genetic Algorithm for MDCT Coefficient Adjustment for AACabstractAn AAC steganography algorithm based on genetic algorithm and MDCT coefficient adjustment is proposed. Our algorithm selects the small value region of MDCT coefficient as the embedding bit and the coefficients in codebook 1/2 are designed to change. In order to be against steganalysis better, genetic algorithm is used to optimize the change of the coefficient. The experiment results show that the algorithm has good embedding capacity, high steganography and good imperceptibility. Chen Li 0033, Lihua Tian |
ISM | 4 |
| 2020 | Action temporal detection method based on confidence curve analysis
Hanjian Song, Lihua Tian, Chen Li 0033 |
Multim. Tools Appl. | 2 |
| 2020 | A Semi-fragile Video Watermarking Algorithm Based On Chromatic Residual DCT
Lihua Tian, Hangtao Dai, Chen Li 0033 |
Multim. Tools Appl. | 1 |
| 2020 | A Semi-Fragile Video Watermarking Algorithm Based on H.264/AVCabstractWith the increasing application of advanced video coding (H.264/AVC) in the multimedia field, a great significance to research in video watermarking based on this video compression standard has been established. We propose a semifragile video watermarking algorithm, which can simultaneously implement frame attack and video tamper detection, herein. In this paper, the frame number is selected as the watermark information, and the relationship of the discrete cosine transform (DCT) nonzero coefficients is used as the authentication code. The 4×4 subblocks, whose DCT nonzero coefficients are sufficiently complex, are selected to embed the watermark. The parities of these nonzero coefficients in the medium frequency are modulated to embed watermarks. The experimental results show that the visual quality of the embedded watermarked video is virtually unaffected, and the algorithm exhibits good robustness. Furthermore, the algorithm can correctly implement frame attack and video tamper detection. Chen Li 0033, Lihua Tian |
Wirel. Commun. Mob. Comput. | 4 |
| 2018 | Pixel-wise binary classification network for salient object detectionabstractDeep convolutional neural networks have led to significant improvement over the previous salient object detection systems. The existing deep models are trained end-to-end and predicts salient objects by calculating pixel values, which results saliency maps are typically blurry. Our Pixel-wise Binary Classification Network (PBCN) focuses on binary classification in pixel level for salient object detection: saliency and background. In order to increase the resolution of output feature maps and get denser feature maps, Hybrid dilation convolution (HDC) is employed into PBCN. Then, Hybrid Dilation Spatial Pyramid Pooling (HDSPP) is proposed to extract denser multi-scale image representations. In HDSPP, it contains one 1×1 convolution and several dilated convolutions, with different rates and the output feature maps of the convolutions will be fused. Finally, softmax is introduced to implement the binary classification instead of sigmoid. Experiment, on four datasets, show that PBCN significantly improves the state-of-the-art. Lihua Tian, Chen Li 0033 |
ICMV | 2 |
| 2018 | 3D Convolutional Network Based Foreground Feature FusionabstractWith explosion of videos, action recognition has become an important research subject. This paper makes a special effort to investigate and study 3D Convolutional Network. Focused on the problem of ConvNet dependence on multiple large scale dataset, we propose a 3D ConvNet structure which incorporate the original 3D-ConvNet features and foreground 3D-ConvNet features fused by static object and motion detection. Our architecture is trained and evaluated on the standard video actions benchmarks of UCF-101 and HMDB-51, experimental results demonstrate that with merely 50% pixels utilization, foreground ConvNet achieves satisfying performance as same as origin. With feature fusion, we achieve 83.7% accuracy on UCF-101 exceeding original ConvNet. Hanjian Song, Lihua Tian, Chen Li 0033 |
ISM | 2 |
| 2017 | FPFH-based graph matching for 3D point cloud registrationabstractCorrespondence detection is a vital step in point cloud registration and it can help getting a reliable initial alignment. In this paper, we put forward an advanced point feature-based graph matching algorithm to solve the initial alignment problem of rigid 3D point cloud registration with partial overlap. Specifically, Fast Point Feature Histograms are used to determine the initial possible correspondences firstly. Next, a new objective function is provided to make the graph matching more suitable for partially overlapping point cloud. The objective function is optimized by the simulated annealing algorithm for final group of correct correspondences. Finally, we present a novel set partitioning method which can transform the NP-hard optimization problem into a O(n3)-solvable one. Experiments on the Stanford and UWA public data sets indicates that our method can obtain better result in terms of both accuracy and time cost compared with other point cloud registration methods. Jiapeng Zhao, Chen Li 0033, Lihua Tian, Jihua Zhu |
ICMV | 3 |
| 2017 | Hand Gesture Recognition Based on Wavelet Invariant MomentsabstractIn this paper, a new method of hand gesture recognition is proposed. First, the hand region is separated based on the depth information. Then the wavelet feature is calculated by enforcing the wavelet invariant moments of the hand region, and the distance feature is extracted by calculating the distance from fingers to hand centroid. Next, a feature vector which is composed of wavelet invariant moments and distance feature is generated. Finally, a support vector machine classifier based on the feature vectors is used to identify these hand gestures. Experimental results show that our method can achieve high accuracy, and can distinguish similar gestures well. Chen Li 0033, Lihua Tian |
ISM | 3 |
| 2017 | LBP-SVD Based Copy Move Forgery Detection AlgorithmabstractWith the extensive use of sophisticated image editing software, it has become easy to manipulate digital images without any visually visible clue. Copy-move is a special type of image forgery performed by copying a part of the image and pasting anywhere else in the same image. We proposed a passive image authentication technique to determine the copy-move forgery. First, the method divides the image into overlapping blocks. It use LBP (Local Binary Pattern) to label each block. Then, the biggest N of SVD values are extracted on the labeled blocks. N SVD values plus average Y, Cb, Cr values constitutes the feature vector for the block. Finally, the feature vectors are lexicographically sorted and element-by-element similarity measurement is used to determine the forged blocks. Experiment results demonstrate commendable performance in image copy-move forgery detection. Lihua Tian, Chen Li 0033 |
ISM | 2 |
| 2017 | A Fast Video Shot Boundary Detection Employing OTSU's Method and Dual Pauta CriterionabstractVideo shot boundary detection is a fundamental step towards video information processing in e-learning scenarios. In the field of shot boundary detection, there still exists difficulty in choosing suitable thresholds for different videos, and empirical thresholds usually lead to low precision. Thus, we propose an original method to generate video-based threshold which is calculated by video itself. In the process of generating video-based threshold, we employ the OTSU's method [11] to eliminate non-boundary segments, and dual Pauta criterion to identify the cut shots and gradual shots. The experiments show that video-based threshold has better performance than empirical threshold. Lihua Tian, Chen Li 0033 |
ISM | 2 |
| 2017 | Key Frame Extraction Based on Entropy Difference and Perceptual HashabstractKey frame extraction is a crucial step in content-based video retrieval. To accurately describe character of frames, various features like color, texture, shape can be integrated and used for key frame extraction. In this paper, we proposed an improved two-phase approach of key frame extraction based on entropy and perceptual hash. It weakens the threshold's direct influence on final results, and solves the problem of fading, sunlight and other information easily resulting in redundant key frames. Firstly, candidate key frames are selected with the use of golden-section partition and weighted histogram. Next, key frames are determined by the entropy values of candidate frames. Finally, a new method of perceptual hash is applied to remove redundant key frames. Experimental data set is created with videos from different domains like movie, cartoon, news etc. Results show that the proposed method is accurate and effective for key frame extraction. The selected key frames can be a good representative of main content. Lihua Tian, Chen Li 0033 |
ISM | 2 |
| 2016 | Multiple Cartesian K-Medoids for a Fine QuantizationabstractK-means is a widely used method for the process of vector quantization in image retrieval, and its results will directly affect the subsequent retrieval quality. Although k-means is popular in image retrieval, it has some obvious disadvantages, such as randomness and sensitivity to outliers. This paper presents a new model, namely Multiple Cartesian K-medoids, to replace k-means for quantization and retrieval. The proposed model proceeds in two steps. The first step is to establish multiple K-medoids model to finely quantize feature vectors to codewords. Then, the second step establishes local linear search: adopt an inverted file for efficiently searching candidate nearest neighbors of a given query, and finally obtains accurate neighbors of the query by re-ranking these candidate neighbors with Euclidean distances of the original feature vectors. Experimental results show that the proposed method is effective, and substantially improves the search accuracy of the returned nearest neighbors. Lihua Tian, Shanmin Pang, Chen Li 0033 |
ICPADS | 1 |
| 2015 | Authentication and copyright protection watermarking scheme for H.264 based on visual saliency and secret sharing
Lihua Tian, Nanning Zheng 0001, Jianru Xue, Ce Li 0001 |
Multim. Tools Appl. | 1 |
| 2015 | A robust approach to detect digital forgeries by exploring correlation patterns
Lu Li 0009, Jianru Xue, Lihua Tian |
Pattern Anal. Appl. | 4 |
| 2011 | A Robust Approach to Detect Tampering by Exploring Correlation Patterns
Lu Li 0009, Jianru Xue, Lihua Tian |
CAIP (2) | 4 |
| 2011 | An integrated visual saliency-based watermarking approach for synchronous image authentication and copyright protection
Lihua Tian, Nanning Zheng 0001, Jianru Xue, Ce Li 0001 |
Signal Process. Image Commun. | 1 |
| 2010 | Hash key-based video encryption scheme for H.264/AVC
Nanning Zheng 0001, Lihua Tian |
Signal Process. Image Commun. | 3 |
| 2009 | A Peer-to-Peer Architecture for Live Streaming with DRMabstractDRM is becoming more and more important for P2P live streaming. In this paper, a manageable overlay network architecture with DRM, is proposed for live streaming. The system consists of register server, index servers, supernodes and peers. The register server authorizes the peers, assign the key and index server list; The index server acts as the centralized index server to store the peer list, program list, and buffer information of peers. Supernodes are special peers which store the bigger buffer of live streaming. Each peer periodically exchanges data availability information with the assigned index server. The peer can retrieve correspondingly unavailable data from partners given by index server using the proposed scheduling algorithm. It also supports DRM in register server. The proposed system has been demonstrated based on CERNET of China. Good streaming quality can be achieved due to its global optimization and the digital right of video contents can be protected to some extent. Xuguang Lan, Jianru Xue, Lihua Tian, Nanning Zheng 0001 |
CCNC | 3 |
| 2008 | A CAVLC-Based Blind Watermarking Method for H.264/AVC Compressed VideoabstractStreaming media services have been applied in many applications. At the same time, the security of it should be considered. To protect the video content, watermarking data (image) is usually embedded into the quantized DCT coefficients of video. At the same time, the watermarked video should keep the fidelity and the bit-rate. While for high efficient compression video such as H.264/AVC it is very difficult, because just one bit alteration may widely affect the video content and the bit-rate. A CAVLC-based blind watermarking method for H.264/AVC compressed video is proposed. The watermarking data is only embedded into the last non-zero and non-trailing AC coefficient in context adaptive variable length coding (CAVLC) of H.264/AVC. With this kind of embedding, the artifact due to the embedding could be reduced efficiently by CAVLC. Experimental results show that on average, the introduced distortion by watermark embedding is less than 0.5 dB, and the increased stream bit rate is only 0.1%. Lihua Tian, Nanning Zheng 0001, Jianru Xue |
APSCC | 1 |