Litao Yu

dblp:62/10939 · DBLP profile ↗
← Back
23ranked-venue papers
11as first author
11since 2021 · last 2025
0000-0001-5260-885XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 7 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Hierarchical Multi-Prototype Discrimination: Boosting Support-Query Matching for Few-Shot Segmentation
abstract
Few-shot segmentation (FSS) aims at training a model on base classes with sufficient annotations and then tasking the model with predicting a binary mask to identify novel class pixels with limited labeled images. Mainstream FSS methods adopt a support-query matching paradigm that activates target regions of the query image according to their similarity with a single support class prototype. However, this prototype vector is inclined to overfit the support images, leading to potential under-matching in latent query object regions and incorrect mismatches with base class features in the query image. To address these issues, this study reformulates conventional single foreground prototype matching to a multi-prototype matching paradigm. In this paradigm, query features exhibiting high confidence with non-target prototypes will be categorized as background. Specifically, the target query features are drawn closer to the novel class prototype through a Masked Cross-Image Encoding (MCE) module and a Semantic Multi-prototype Matching (SMM) module is incorporated to collaboratively filter unexpected base class regions on multi-scale features. Furthermore, we devise an adaptive class activation map, termed target-aware class activation map (TCAM) to preserve semantically coherent regions that might be inadvertently suppressed under pixel-wise matching guidance. Experimental results on PASCAL-5$^{i}$and COCO-20$^{i}$datasets demonstrate the advantage of the proposed novel modules, with the holistic approach outperforming compared state-of-the-art methods.
Wenbo Xu 0004, Huaxi Huang, Yongshun Gong, Litao Yu, Qiang Wu 0001, Jian Zhang 0002
IEEE Trans. Multim.4
2023 Masked Cross-image Encoding for Few-shot Segmentation
abstract
Few-shot segmentation (FSS) is a dense prediction task that aims to infer the pixel-wise labels of unseen classes using only a limited number of annotated images. The key challenge in FSS is to classify the labels of query pixels using class prototypes learned from the few labeled support exemplars. Prior approaches to FSS have typically focused on learning class-wise descriptors independently from support images, thereby ignoring the rich contextual information and mutual dependencies among support-query features. To address this limitation, we propose a joint learning method termed Masked Cross-Image Encoding (MCE), which is designed to capture common visual properties that describe object details and to learn bidirectional inter-image dependencies that enhance feature interaction. MCE is more than a visual representation enrichment module; it also considers cross-image mutual dependencies and implicit guidance. Experiments on FSS benchmarks PASCAL-5iand COCO-20idemonstrate the advanced meta-learning ability of the proposed method.
Wenbo Xu 0004, Huaxi Huang, Litao Yu, Qiang Wu 0001, Jian Zhang 0002
ICME4
2023 Automated Flock Density and Activity Recognition for Welfare Monitoring on Commercial Egg Farms
abstract
Monitoring poultry behaviour provides the opportunity to aid egg production and animal welfare. With the current development in machine learning and computer vision, automated content analysis has become a practical way for low-cost and continuous monitoring of animal behaviours. In this demo, we will show a simple yet effective flock monitoring system based on computer vision and machine learning techniques for egg farmers that allows them to reduce labour yet improve performance. This demo shows that it is possible to auto-analyse flock activities thereby providing early warning of welfare issues, by applying object detection, tracking and crowd-counting techniques. Summaries of individual bird activity and their distribution are closely related to the flock behaviour, which in turn reflects the welfare status. Specifically, the density and movement patterns of birds provide reliable information on the welfare status of the flock. For example, the real-time monitoring of density and movement can give early warnings of pile-ups. To observe these and other important flock activities, we developed a low-cost and easy-use system based on recent computer vision techniques to auto-estimate the density and movement of birds on commercial egg farms.
Litao Yu, Wenbo Xu 0004, Qiang Wu 0001, Jian Zhang 0002
MMSP1
2022 Distribution-Aware Margin Calibration for Semantic Segmentation in Images
Litao Yu, Zhibin Li 0002, Min Xu 0001, Yongsheng Gao 0001, Jiebo Luo 0001, Jian Zhang 0002
Int. J. Comput. Vis.1
2022 Blockchain-Enabled Fish Provenance and Quality Tracking System
abstract
Accurate assessment of fish quality is difficult in practice due to the lack of trusted fish provenance and quality tracking information. Working with Sydney Fish Market (SFM), we develop a Blockchain-enabled fish provenance and quality tracking (BeFAQT) system. A multilayer Blockchain architecture based on attribute-based encryption (ABE) is proposed to tackle the privacy issue caused by applying Blockchain to secure supply chain data and achieve trusted and confidential data sharing among parties in fish supply chains. An Internet-of-Things (IoT) chain saves encrypted fish provenance and quality tracking data, and an ABE chain is specifically designed for the access control to the data in the IoT chain. Latest IoT and artificial intelligence (AI) technologies, including NarrowBand-IoT, image processing, and biosensing, are developed for fish origin proof, supply chain tracking, and objective fish quality assessment. As proven by field trials with SFM and a local fish supply chain, the BeFAQT is able to provide trusted and comprehensive fish provenance and quality tracking information in real time.
Xu Wang 0004, Guangsheng Yu, Ren Ping Liu 0001, Jian Zhang 0002, Qiang Wu 0001, Steven W. Su, Ying He 0011, Zongjian Zhang, Litao Yu, Taoping Liu, Wentian Zhang, Peter Loneragan, Eryk Dutkiewicz, Erik Poole, Nick Paton
IEEE Internet Things J.9
2022 TOAN: Target-Oriented Alignment Network for Fine-Grained Image Categorization With Few Labeled Samples
abstract
In this paper, we study the fine-grained categorization problem under the few-shot setting, i.e., each fine-grained class only contains a few labeled examples, termed Fine-Grained Few-Shot classification (FGFS). The core predicament in FGFS is the high intra-class variance yet low inter-class fluctuations in the dataset. In traditional fine-grained classification, the high intra-class variance can be somewhat relieved by conducting the supervised training on the abundant labeled samples. However, with few labeled examples, it is hard for the FGFS model to learn a robust class representation with the significantly higher intra-class variance. Moreover, the inter- and intra-class variance are closely related. The significant intra-class variance in FGFS often aggravates the low inter-class variance issue. To address the above challenges, we propose a Target-Oriented Alignment Network (TOAN) to tackle the FGFS problem from both intra- and inter-class perspective. To reduce the intra-class variance, we propose a target-oriented matching mechanism to reformulate the spatial features of each support image to match the query ones in the embedding space. To enhance the inter-class discrimination, we devise discriminative fine-grained features by integrating local compositional concept representations with the global second-order pooling. We conducted extensive experiments on four public datasets for fine-grained categorization, and the results show the proposed TOAN obtains the state-of-the-art.
Huaxi Huang, Junjie Zhang 0002, Litao Yu, Jian Zhang 0002, Qiang Wu 0001, Chang Xu 0002
IEEE Trans. Circuits Syst. Video Technol.3
2022 Dual Attention on Pyramid Feature Maps for Image Captioning
abstract
Generating natural sentences from images is a fundamental learning task for visual-semantic understanding in multimedia. In this paper, we propose to apply dual attention on pyramid image feature maps to fully explore the visual-semantic correlations and improve the quality of generated sentences. Specifically, with the full consideration of the contextual information provided by the hidden state of the RNN controller, the pyramid attention can better localize the visually indicative and semantically consistent regions in images. On the other hand, the contextual information can help re-calibrate the importance of feature components by learning the channel-wise dependencies, to improve the discriminative power of visual features for better content description. We conducted comprehensive experiments on three well-known datasets: Flickr8K, Flickr30 K and MS COCO, which achieved impressive results in generating descriptive and smooth natural sentences from images. Using either convolution visual features or more informative bottom-up attention features, the composite model can boost the performance of image-to-sentence translation, with a limited computational resource overhead. The proposed pyramid attention and dual attention methods are highly modular, which can be inserted into various image captioning modules to further improve the performance.
Litao Yu, Jian Zhang 0002, Qiang Wu 0001
IEEE Trans. Multim.1
2022 Multimodal Marketing Intent Analysis for Effective Targeted Advertising
abstract
People’s daily information sharing and acquisition through the Internet has become more and more popular. The comprehensive multimodal marketing advertorial generated by ‘We Media’ accounts besides the normal social news is gaining its importance on social media platforms. In order to achieve effective advertising, the marketing intent understanding is a key step towards generating targeted advertising strategies (push advertorials to specific people at a specific time). However, advertorials in real are usually designed to pretend as normal social news with a wide range of contents. This poses big challenges to the platforms on accurately recognizing and analyzing the marketing intents behind the advertorials. As a pioneering study, we address this new problem of multimodal-based marketing intent analysis and answer three core questions: (1) does a piece of social news contain marketing intent? (2) what is the topic of marketing intent? (3) what is the extent of marketing intent? Towards this end, we propose a novel Multimodal-based Marketing Intent Analysis scheme (MMIA) to estimate the marketing intent embedded in the multimodal contents. Specifically, a novel supervised neural autoregressive model (SmiDocNADE) is proposed to enhance the discriminative capacity of the learned hidden features so that a single system is capable of solving the three questions. In order to effectively model inter-correlations between images and text in advertorials, we fuse multimodal data and extract features by Graph Convolution Networks as an enhancement to SmiDocNADE. The extensive evaluations demonstrate the advantages of our proposed system in multimodal-based marketing intent analysis from multiple aspects.
Lu Zhang 0062, Jialie Shen 0001, Jian Zhang 0002, Jingsong Xu, Zhibin Li 0002, Yazhou Yao, Litao Yu
IEEE Trans. Multim.7
2022 Unsupervised Image and Text Fusion for Travel Information Enhancement
abstract
With the explosive growth of the shared information on social media platforms, people are increasingly interested in sharing and making their travel plans by referring to others’ travel experiences. However, different social media sources render the heterogeneity of these valuable data, bringing difficulties for data collection and fusion. Thus, facing massive information online, one of the biggest challenges to enhance travel information is how to integrate and match these multi-source data without clear labels. In this paper, we propose an unsupervised method to fuse and match images and travelogues. We first use the three textual components (title, tag, and description) of the descriptive texts of images as three criteria to embed travelogues and the descriptive texts of images, and further introduce images into our method by joint embedding texts and images. Finally, a multiple kernel clustering approach is adopted for matching travelogues and images. Extensive experiments conducted on the real dataset crawled from two websites (Flickr and TripAdvisor) demonstrate the effectiveness and robustness of our proposed method.
Lu Zhang 0062, Jingsong Xu, Yongshun Gong, Litao Yu, Jian Zhang 0002, Jialie Shen 0001
IEEE Trans. Multim.4
2021 Incorporating Multimodal Cues for Advertorial Discovery
abstract
Commercial advertorials shared on websites are usually designed to pretend as normal social news for commercial benefits. The analysis of the commercial intents embedded in advertorials can greatly help media platforms personalize content. However, commercial intents are not only concealed in news texts but also conveyed by news images explicitly or implicitly. Consequently, how to effectively extract and incorporate the crucial cues of multiple modalities has been emerging as an important but challenging problem. Motivated by this observation, we propose a framework Multimodal Advertorial Discovery Model (MADM) to estimate the commercial intents embedded in the multimodal social news. Specifically, a novel Cross-graph Fusion (CGF) strategy is developed to achieve a soft assignment to incorporate images and text and generate comprehensive multimodal representations. The extensive evaluations demonstrate the superiority of our proposed system in multimodal-based advertorial detection and analysis.
Lu Zhang 0062, Jian Zhang 0002, Jialie Shen 0001, Jingsong Xu, Zhibin Li 0002, Litao Yu
ICME6
2021 Parameter-Efficient Deep Neural Networks With Bilinear Projections
abstract
Recent research on deep neural networks (DNNs) has primarily focused on improving the model accuracy. Given a proper deep learning framework, it is generally possible to increase the depth or layer width to achieve a higher level of accuracy. However, the huge number of model parameters imposes more computational and memory usage overhead and leads to the parameter redundancy. In this article, we address the parameter redundancy problem in DNNs by replacing conventional full projections with bilinear projections (BPs). For a fully connected layer with D input nodes and D output nodes, applying BP can reduce the model space complexity fromO(D2) toO(2D), achieving a deep model with a sublinear layer size. However, the structured projection has a lower freedom of degree compared with the full projection, causing the underfitting problem. Therefore, we simply scale up the mapping size by increasing the number of output channels, which can keep and even boosts the model accuracy. This makes it very parameter-efficient and handy to deploy such deep models on mobile systems with memory limitations. Experiments on four benchmark data sets show that applying the proposed BP to DNNs can achieve even higher accuracies than conventional full DNNs while significantly reducing the model size.
Litao Yu, Yongsheng Gao 0001, Jun Zhou 0001, Jian Zhang 0002
IEEE Trans. Neural Networks Learn. Syst.1
2020 Automatic Sheep Counting by Multi-object Tracking
abstract
Animal counting is a highly skilled yet tedious task in livestock transportation and trading. To effectively free up the human labour and provide accurate counts for sheep loading/unloading, we develop an auto sheep counting system based on multi-object detection, tracking and extrapolation techniques. Our system has demonstrated more than 99.9% accuracy with sheep moving freely in a race under optimal visual conditions.
Jingsong Xu, Litao Yu, Jian Zhang 0002, Qiang Wu 0001
VCIP2
2020 A Vision Based Fish Processing System
abstract
The digital fish provenance and quality tracking system is essential for the seafood supply chain. As a part of this system, we develop a vision-based fish processing system to automatically perform fish freshness estimation, size measurement and species classification. Under the constrained illumination environment, our system is able to auto-process the fish selection, thus greatly reduce the human labour and bring trust and efficiency to the seafood supply chain from catch to market.
Zongjian Zhang, Litao Yu, Jian Zhang 0002, Qiang Wu 0001
VCIP2
2018 Generative Adversarial Product Quantisation
abstract
Product Quantisation (PQ) has been recognised as an effective encoding technique for scalable multimedia content analysis. In this paper, we propose a novel learning framework that enables an end-to-end encoding strategy from raw images to compact PQ codes. The system aims to learn both PQ encoding functions and codewords for content-based image retrieval. In detail, we first design a trainable encoding layer that is pluggable into neural networks, so the codewords can be trained in back-forward propagation. Then we integrate it into a Deep Convolutional Generative Adversarial Network (DC-GAN). In our proposed encoding framework, the raw images are directly encoded by passing through the convolutional and encoding layers, and the generator aims to use the codewords as constrained inputs to generate full image representations that are visually similar to the original images. By taking the advantages of the generative adversarial model, our proposed system can produce high-quality PQ codewords and encoding functions for scalable multimedia retrieval tasks. Experiments show that the proposed architecture GA-PQ outperforms the state-of-the-art encoding techniques on three public image datasets.
Litao Yu, Yongsheng Gao 0001, Jun Zhou 0001
ACM Multimedia1
2017 Bilinear Optimized Product Quantization for Scalable Visual Content Analysis
abstract
Product quantization (PQ) has been recognized as a useful technique to encode visual feature vectors into compact codes to reduce both the storage and computation cost. Recent advances in retrieval and vision tasks indicate that high-dimensional descriptors are critical to ensuring high accuracy on large-scale data sets. However, optimizing PQ codes with high-dimensional data is extremely time-consuming and memory-consuming. To solve this problem, in this paper, we present a novel PQ method based on bilinear projection, which can well exploit the natural data structure and reduce the computational complexity. Specifically, we learn a global bilinear projection for PQ, where we provide both non-parametric and parametric solutions. The non-parametric solution does not need any data distribution assumption. The parametric solution can avoid the problem of local optima caused by random initialization, and enjoys a theoretical error bound. Besides, we further extend this approach by learning locally bilinear projections to fit underlying data distributions. We show by extensive experiments that our proposed method, dubbed bilinear optimization product quantization, achieves competitive retrieval and classification accuracies while having significant lower time and space complexities.
Litao Yu, Zi Huang, Fumin Shen, Jingkuan Song, Heng Tao Shen, Xiaofang Zhou 0001
IEEE Trans. Image Process.1
2017 Graph PCA Hashing for Similarity Search
abstract
This paper proposes a new hashing framework to conduct similarity search via the following steps: first, employing linear clustering methods to obtain a set of representative data points and a set of landmarks of the big dataset; second, using the landmarks to generate a probability representation for each data point. The proposed probability representation method is further proved to preserve the neighborhood of each data point. Third, PCA is integrated with manifold learning to lean the hash functions using the probability representations of all representative data points. As a consequence, the proposed hashing method achieves efficient similarity search (with linear time complexity) and effective hashing performance and high generalization ability (simultaneously preserving two kinds of complementary similarity structures, i.e., local structures via manifold learning and global structures via PCA). Experimental results on four public datasets clearly demonstrate the advantages of our proposed method in terms of similarity search, compared to the state-of-the-art hashing methods.
Xiaofeng Zhu 0001, Xuelong Li 0001, Shichao Zhang 0001, Zongben Xu, Litao Yu, Can Wang 0004
IEEE Trans. Multim.5
2016 Robust spatial-temporal deep model for multimedia event detection
Litao Yu, Xiaoshuai Sun, Zi Huang
Neurocomputing1
2016 Web Video Event Recognition by Semantic Analysis From Ubiquitous Documents
abstract
In recent years, the task of event recognition from videos has attracted increasing interest in multimedia area. While most of the existing research was mainly focused on exploring visual cues to handle relatively small-granular events, it is difficult to directly analyze video content without any prior knowledge. Therefore, synthesizing both the visual and semantic analysis is a natural way for video event understanding. In this paper, we study the problem of Web video event recognition, where Web videos often describe large-granular events and carry limited textual information. Key challenges include how to accurately represent event semantics from incomplete textual information and how to effectively explore the correlation between visual and textual cues for video event understanding. We propose a novel framework to perform complex event recognition from Web videos. In order to compensate the insufficient expressive power of visual cues, we construct an event knowledge base by deeply mining semantic information from ubiquitous Web documents. This event knowledge base is capable of describing each event with comprehensive semantics. By utilizing this base, the textual cues for a video can be significantly enriched. Furthermore, we introduce a two-view adaptive regression model, which explores the intrinsic correlation between the visual and textual cues of the videos to learn reliable classifiers. Extensive experiments on two real-world video data sets show the effectiveness of our proposed framework and prove that the event knowledge base indeed helps improve the performance of Web video event recognition.
Litao Yu, Yang Yang 0002, Zi Huang, Peng Wang 0023, Jingkuan Song, Heng Tao Shen
IEEE Trans. Image Process.1
2016 Scalable Video Event Retrieval by Visual State Binary Embedding
abstract
With the exponential increase of media data on the web, fast media retrieval is becoming a significant research topic in multimedia content analysis. Among the variety of techniques, learning binary embedding (hashing) functions is one of the most popular approaches that can achieve scalable information retrieval in large databases, and it is mainly used in the near-duplicate multimedia search. However, till now most hashing methods are specifically designed for near-duplicate retrieval at the visual level rather than the semantic level. In this paper, we propose a visual state binary embedding (VSBE) model to encode the video frames, which can preserve the essential semantic information in binary matrices, to facilitate fast video event retrieval in unconstrained cases. Compared with other video binary embedding models, one advantage of our proposed VSBE model is that it only needs a limited number of key frames from the training videos for hash function training, so the computational complexity is much lower in the training phase. At the same time, we apply the pairwise constraints generated from the visual states to sketch the local properties of the events at the semantic level, so accuracy is also ensured. We conducted extensive experiments on the challenging TRECVID MED dataset, and have proved the superiority of our proposed VSBE model.
Litao Yu, Zi Huang, Jiewei Cao, Heng Tao Shen
IEEE Trans. Multim.1
2015 Max-margin adaptive model for complex video pattern recognition
Litao Yu, Jie Shao 0001, Xin-Shun Xu, Heng Tao Shen
Multim. Tools Appl.1
2013 Content integrity and non-repudiation preserving audio-hiding scheme based on robust digital signature
abstract
ABSTRACT Current secure communication schemes do not take together traffic security and data security (content integrity and non‐repudiation) of the secret message into consideration, making the content prone to blind tampering and compromised party cheating attacks. In this paper, we present a scheme that hides secret audio in cover audio on the basis of robust digital signature to preserve not only hidden communication but also content integrity and non‐repudiation of the secret audio. Furthermore, instead of traditional binary authentication that only outputs yes or no, the authentication of our scheme is flexibly measurable, and the measurement value is in correspondence with the sense of human hearing precisely. Experimental results show that the proposed scheme provides highly robust authentication against content‐preserving degradations with 99.03% of test audios having the strongest authenticity (1.00) and high level of distinct authentication between content‐destructive degradations with 95.01% of test audios having relatively weak authenticity (less than 0.15). As the authentication is flexibly measureable, there is no false alarm in the semantic aspect. Copyright © 2013 John Wiley & Sons, Ltd.
Liehuang Zhu, Dan Liu 0002, Litao Yu, Yuzhou Xie, Mingzhong Wang
Secur. Commun. Networks3
2012 Transfer clustering via constraints generated from topics
abstract
Clustering technique is widely used in data mining like gene-microarray analysis and natural language processing. When there are sufficient data samples and good representations, traditional clustering algorithms such as K-means can work well. But when the number of samples is small and the data representation is bad, direct use of clustering may yield bad results. In this paper we propose a new algorithm TCTC(Topic-Constraint Transfer Clustering), which is an instance of unsupervised transfer learning, to cluster a small number of unlabeled data with the help of sufficient and better represented auxiliary data. First several latent topics are extracted from the clusters of the auxiliary data. Then the affinities between target data samples and topics are discovered to “guide” the disseminated data clustering. Finally semi-supervised clustering algorithm is applied on target data. The experiments demonstrate our method is quite effective to solve the problem of disseminated and ill-presented data clustering.
Litao Yu, Yanzhong Dang, Guangfei Yang
SMC1
2011 A New Domain Adaptation Method Based on Rules Discovered from Cross-Domain Features
Yanzhong Dang, Litao Yu, Guangfei Yang, Ming-Zheng Wang
KSEM2