Shuili Zhang

dblp:181/5091 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Fashion Microscope: Pixel-Level Attribute Perception via Optimal Transport and Neural Semantic Aggregation
abstract
Attribute-specific fashion retrieval aims to enhance fine-grained image retrieval by emphasizing the similarity of specific attributes. Current methods primarily rely on attention mechanisms to extract attribute-related visual features but face two key challenges: the limitations of coarse-grained localization in achieving fine-grained accuracy, and an imbalance between global and local perception, where excessive focus on local features can undermine overall performance. To address these issues, we propose the fashion microscope ProFashion, which achieves pixel-level attribute awareness through optimal transport and neural semantic aggregation. The framework begins by employing optimal transport to align semantic attributes with visual patterns from a global perspective, generating an attribute-visual value map that highlights distinctive regions while reducing interference. This is followed by simulating the human brain's perception of attribute feature patterns through superpixel generation and aggregation, capturing attribute-related features at the pixel semantic level and forming key semantic clusters that preserve microstructures. Building on this, an attribute graph is constructed to facilitate feature clustering, significantly enhancing the framework's capability to handle overlapping features and cross-scale relationships. Comprehensive experiments on the FashionAI, DeepFashion, and DARN datasets demonstrate the framework's effectiveness, achieving overall MAP improvements of 3.11%, 3.70%, and 3.49%, respectively. Additionally, the framework delivers relative average throughput gains of 26.94%, 22.22%, and 24.78% on the FashionAI, DeepFashion, and DARN datasets, respectively.
Shuili Zhang, Hongzhang Mu, Jiawei Sheng, Qianqian Tong 0001, Wenyuan Zhang 0002, Quangang Li, Tingwen Liu
AAAI1
2026 Beyond Patches: Superpixel Token-based Transformers for Attribute-Specific Fashion Retrieval
abstract
Attribute-Specific Fashion Retrieval (ASFR) aims to improve fine-grained image retrieval by focusing on specific attributes. However, existing patch-based attention and Transformer methods often misalign with irregular attribute regions and are prone to background noise, limiting their ability to capture subtle, pixel-level microstructures. To tackle these challenges, we propose Super Fashion., the first ASFR framework that adopts superpixel tokens within a Transformer architecture. Super Fashion initially employs an attribute-guided attention mechanism to extract attribute-related features, which in turn guide the cropping of semantically meaningful image regions. Superpixel segmentation is then leveraged on these regions to generate compact, semantically coherent superpixel tokens. By incorporating modality-specific embeddings for both attribute and superpixel tokens, the superpixel token-based Transformer facilitates adaptive interaction and fusion, thereby enhancing attribute localization and discrimination. Extensive experiments on FashionAI, DARN, and DeepFashion demonstrate relative overall MAP improvements of 1.84%, 9.27%, and 9.35% over prior SOTA. Super Fashion offers a new solution for web-based image retrieval.
Shuili Zhang, Hongzhang Mu, Wenyuan Zhang 0002, Duohe Ma, Tingwen Liu
WWW1
2024 MSKR: Advancing Multi-modal Structured Knowledge Representation with Synergistic Hard Negative Samples
abstract
Despite the notable progress achieved by large-scale vision-language pre-training models in a wide range of multi-modal tasks, their performance often falls short in image-text matching challenges that require an in-depth understanding of structured representations. For instance, when distinguishing between texts or images that are generally similar but have distinct structured knowledge (such as entities and relationships in text, or objects and object attributes in images), the model's capabilities are limited. In this paper, we propose a advancing Multi-modal Structured Knowledge Representation with synergistic hard negative samples (MSKR), thereby significantly improving the model's matching capability for such data. Specifically, our model comprises a structured knowledge-enhanced encoder designed to bolster the structured knowledge inherent in textual data, such as entities, their attributes, and the relationships among these entities as well as structured knowledge within images, focusing on elements like objects and their attributes. To further refine the model's learning process, we produce both image and text challenging negative samples. Extensive experimental evaluations on the Winoground, InpaintCOCO, and MSCOCO benchmark reveal that MSKR significantly outperforms the baseline model, showcasing marked improvements 2.66% on average in structured representation learning compared to the baseline. Moreover, general representation results illustrate that our model not only excels in structured representation learning but also maintains its proficiency in general representation learning.
Shuili Zhang, Hongzhang Mu, Tingwen Liu, Qianqian Tong 0001, Jiawei Sheng
CIKM1
2024 Fine-Grained Features Alignment and Fusion for Text-Video Cross-Modal Retrieval
abstract
Text-video cross-modal retrieval is an increasingly prominent and challenging task that has garnered significant attention. Traditional models typically embed videos and texts into global vectors, aiming to capture the global features of these modalities. While the models often fall short in capturing fine-grained semantic details. Relying solely on global features proves insufficient to address this challenge. Hence, there is a pressing need to bridge the gap between different modalities by incorporating fine-grained features. In light of this, we propose a highly efficient model designed to capture the fine-grained features of videos and texts including question answer semantic alignment, object alignment and text-video feature fusion. For texts, our model includes the incorporation of entity information and part-of-speech information including adjectives, nouns and verbs information, while for videos, the identification of objects plays a crucial role in facilitating text-video retrieval. Our model undergoes extensive training on the WebVid and CC3M datasets, yielding unequivocal evidence of its superior performance over baseline models. It excels particularly in zero-shot text-video cross-modal retrieval tasks, offering substantial reductions in required computational resources.
Shuili Zhang, Hongzhang Mu, Quangang Li, Chenglong Xiao, Tingwen Liu
ICASSP1
2024 Dynamic Multi-Modal Representation Learning For Topic Modeling
abstract
Topic modeling aims to identify and group related topics within collections of text documents. As multi-modal data becomes increasingly effective and prevalent, researchers are delving into the integration of diverse data types (e.g., images) into topic modeling. However, existing methods for topic modeling struggle with unrelated modal features (e.g., visual features) and efficiency. In this paper, we propose a dynamic multi-modal representation learning method (DMMR) that adaptively integrates multi-modal features to enhance the effectiveness and efficiency of modeling multi-modal data. Concretely, a gating network controls the modality-level decision to choose text, image, or both. Based on the sample-wise choice predicted by the gating network, DMMR performs the fusion of multi-modal features encoded by modality-related expert network. With the dynamic modality selection and fusion, the representative modal features can be chosen to exclude irrelevant modality and speed the inference. Extensive experiments on public datasets demonstrate that the proposed method significantly improves the topic quality (e.g., coherence and diversity) and increases the efficiency by 20.37%.
Hongzhang Mu, Shuili Zhang, Quangang Li, Tingwen Liu
ICME2
2024 TRGNN: Text-Rich Graph Neural Network for Few-Shot Document Filtering
abstract
The internet contains a vast amount of textual information, and a crucial step is to filter out irrelevant information and extract relevant topics of interest. However, supervised methods for filtering irrelevant information require a large amount of annotated data, which is time-consuming and labor-intensive. In the real world, it is also impractical to have annotated data covering all the text. Therefore, we propose a joint training framework for few-shot document filtering, which is based on a text-rich graph. Specifically, to filter out irrelevant information, we construct a text-rich graph by treating unlabeled documents, seed words, and seed document pairs as nodes, and their associations as edges. On this basis, we build a network learning module that combines neighborhood sampling based on PageRank and attention mechanisms to get graph embedding. Additionally, we construct a text embedding module by representing the original data using BERT for vector representation. Next, we leverage the advantages of the text representation from the text embedding module to enrich the network learning module through feature sharing, enhancing deep semantic features, and through joint training and label merging, we achieve the filtering of unlabeled documents. Extensive experiments conducted on two real-world datasets consistently demonstrate that our model outperforms existing technical alternatives. These alternatives include traditional classification and retrieval baselines for filtering documents with fewer frames. Furthermore, we conducted ablation studies to demonstrate the effectiveness of each component in enhancing filtering performance.
Hongzhang Mu, Shuili Zhang
IJCNN2
2024 Improving Accuracy and Generalizability via Multi-Modal Large Language Models Collaboration
abstract
With the growing interest in Large Language Models (LLMs), integrating visual tasks has led to the development of Multi-Layer Language Models (MLLMs). Despite their advancements, MLLMs face challenges in accuracy and generalization, often due to resource and time constraints. Addressing these issues, our paper introduces a novel Multi-Agent Collaborative Network for MLLMs (MLLM network). This framework harnesses collective intelligence and cooperation among multiple agents to enhance the accuracy and generalizability of MLLMs. The collaborative nature of our MLLMs—featuring inter-layer neuron interaction and information exchange—facilitates superior processing and integration of multi-modal data. This leads to marked improvements in performance. The findings underscore the efficacy and potential of our proposed framework, presenting a robust solution for complex multi-modal challenges in machine learning and artificial intelligence. Our experimental evaluations demonstrate that this approach significantly surpasses traditional single MLLM architectures in task accuracy and generalization.
Shuili Zhang, Hongzhang Mu, Tingwen Liu
IJCNN1
2024 A Knowledge-Driven Approach to Enhance Topic Modeling with Multi-Modal Representation Learning
abstract
multi-modal topic models strive to integrate semantic information from multi-modal data to generate more precise topics. Topic modeling methods encounter challenges in terms of topic diversity and effectiveness. To address this issue, the majority of current approaches focus on modeling the correlation among numerous multi-modal sources. Nevertheless, little emphasis has been placed on fine-grained feature representation and structured knowledge. In this regard, we propose a fine-grained Prompt representation method. Specifically, we adopt a dual-stream structure where a pre-trained language model and an image model are parallelly combined to construct a multi-modal model. We then enhance the structured representation by integrating fine-grained scene graph knowledge through a Knowledge-Enhanced Encoder, which is constructed based on the scene graph. To validate the effectiveness of the proposed framework, we significantly improve topic quality (such as coherence and diversity) using the aforementioned approach. On publicly available datasets, our approach outperforms state-of-the-art multi-modal topic models respectively.
Hongzhang Mu, Shuili Zhang
ICMR2
2023 Adapt-to-Learn Policy Network for Abstractive Multi-document Summarization
abstract
Abstractive multi-document summarization (MDS) aims to generate a summary for a set of topic-related documents, in which different documents may contain trivial and redundant information and present complementary or contradictory content. It is essential to extract salient information and detect redundancy across documents for abstractive MDS compared with single-document summarization. However, it is challenging for models to exploit salient information and generate concise summaries. In this paper, we propose an Adapt-to-Learn Policy (ALP) network to seamlessly adapt the key sentence selection and word generation over retrieval and generative agents, guided by the multi-document summarization reward. Moreover, our model learns latent dependencies among textual units and explicitly takes advantage of critical information by focusing on semantic similarity or discourse relations. Extensive experiments on WikiSum and Multi-News datasets confirm that our method is superior to prior competitive baselines, and experimental analyses show that higher-quality summaries and more fluent word order can be generated in our ALP network.
Hongzhang Mu, Shuili Zhang, Quangang Li, Tingwen Liu
IJCNN2