Jiazhen Hu

dblp:262/6128 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
4since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 DyFuLM: A Dynamic Collaborative Dual-Encoder Network for Fine-Grained Sentiment Analysis
Jiachen Yuan, Ruohan Zhou, Wenzheng Huang, Churui Yang, Guoyan Zhang, Shiyao Wei, Jiazhen Hu, Ning Xin, Md Maruf Hasan
ICIC (24)7
2025 VideoAVE: A Multi-Attribute Video-to-Text Attribute Value Extraction Dataset and Benchmark Models
abstract
Attribute Value Extraction (AVE) is important for structuring product information in e-commerce. However, existing AVE datasets are primarily limited to text-to-text or image-to-text settings, lacking support for product videos, diverse attribute coverage, and public availability. To address these gaps, we introduce VideoAVE, the first publicly available video-to-text e-commerce AVE dataset across 14 different domains and covering 172 unique attributes. To ensure data quality, we propose a post-hoc CLIP-based Mixture of Experts filtering system (CLIP-MoE) to remove the mismatched video-product pairs, resulting in a refined dataset of 224k training data and 25k evaluation data. In order to evaluate the usability of the dataset, we further establish a comprehensive benchmark by evaluating several state-of-the-art video vision language models (VLMs) under both attribute-conditioned value prediction and open attribute-value pair extraction tasks. Our results analysis reveals that video-to-text AVE remains a challenging problem, particularly in open settings, and there is still room for developing more advanced VLMs capable of leveraging effective temporal information. The dataset and benchmark code for VideoAVE are available at: https://github.com/gjiaying/VideoAVE.
Jiazhen Hu, Jiaying Gong, Hoda Eldardiry
CIKM3
2025 Hypergraph-based Zero-shot Multi-modal Product Attribute Value Extraction
abstract
It is essential for e-commerce platforms to provide accurate, complete, and timely product attribute values, in order to improve the search and recommendation experience for both customers and sellers. In the real-world scenario, it is difficult for these platforms to identify attribute values for the newly introduced products given no similar product history records for training or retrieval. Besides, how to jointly learn the product representation given various product information in multiple modalities, such as textual modality (e.g., product titles and descriptions) and visual modality (e.g., product images), is also a challenging task. To address these limitations, we propose a novel method for extracting multi-label product attribute-value pairs from multiple modalities in the zero-shot scenario, where labeled data is absent during training. Specifically, our method constructs heterogeneous hypergraphs, where product information from different modalities is represented by different types of nodes, and the text and image nodes are embedded and learned through CLIP encoders to effectively capture and integrate multi-modal product information. Then, the complex interrelations among these nodes are modeled through the hyperedges. By learning informative node representations, our method can accurately predict links between unseen product nodes and attribute-value nodes, enabling zero-shot attribute value extraction. We conduct extensive experiments and ablation studies on several categories of the public MAVE dataset and the results demonstrate that our proposed method significantly outperforms several state-of-the-art generative model baselines in multi-label, multi-modal product attribute value extraction in the zero-shot setting.
Jiazhen Hu, Jiaying Gong, Hongda Shen, Hoda Eldardiry
WWW1
2024 CoT-Tree: An Efficient Prompting Strategy for Large Language Model Classification Reasoning
abstract
Prompt engineering in Large Language Models (LLMs) has become pivotal for task adaptation, especially in complex reasoning tasks where traditional methods fall short. This paper introduces an innovative method to optimize prompt efficiency by constructing a decision tree of the Chain of Thought (CoT) for classification tasks, namely CoT-Tree. Initially, we demonstrate two advantages of using CoT as features for decision-tree prompts: task adaptability and semantic dissimilarity between features. These advantages theoretically ensure that the CoT-Tree prompting strategy achieves higher performance in classification tasks. Subsequently, we propose the CoT-Tree algorithm and apply it on three classic classification scenarios. By leveraging the CoT-Tree prompting strategy, we select distinctive and efficient prompts that maximize task relevance while minimizing redundancy and token consumption, allowing for a dynamic pruning mechanism, effectively controlling the depth and complexity of the prompts. The results indicate that our method achieves higher accuracy with fewer LLM interaction rounds. Additionally, a case study is conducted to further demonstrate the practicality of our method.
Jiazhen Hu, Zongwei Luo
IEEE Big Data2
2020 Improving the Learning Efficiency of DQN by an Unsupervised Auxiliary Task
Jiazhen Hu, Chenglu Wu, Miaoyu Yang
ICAART (2)1