Huijing Zhan

dblp:168/3015 · DBLP profile ↗
← Back
18ranked-venue papers
12as first author
10since 2021 · last 2025
0000-0003-3171-2214ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 10 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 1 since 2021Systems, architecture and hardware · 3 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 MDE: Modality Discrimination Enhancement for Multi-modal Recommendation
abstract
Multi-modal recommendation systems aim to enhance performance by integrating an item’s content features across various modalities with user behavior data. Effective utilization of features from different modalities requires addressing two challenges: preserving semantic commonality across modalities (modality-shared) and capturing unique characteristics for each modality (modality-specific). Most existing approaches focus on aligning feature spaces across modalities, which helps represent modality-shared features. However, modality-specific distinctions are often neglected, especially when there are significant semantic variations between modalities. To address this, we propose a Modality Distinctiveness Enhancement (MDE) framework that prioritizes extracting modality-specific information to improve recommendation accuracy while maintaining shared features. MDE enhances differences across modalities through a novel multi-modal fusion module and introduces a node-level trade-off mechanism to balance cross-modal alignment and differentiation. Extensive experiments on three public datasets show that our approach significantly outperforms other state-of-the-art methods, demonstrating the effectiveness of jointly considering modality-shared and modality-specific features.
Huijing Zhan
ISCAS3
2025 DiSCo: Disentangled Attribute Manipulation Retrieval via Semantic Reconstruction and Consistency Regularization
abstract
The rapid evolution of the online fashion industry has intensified the demand for interactive fashion retrieval systems capable of precise and flexible searches based on user-specified attribute modifications. However, prevailing fashion retrieval methods often overlook the distinctive distributional properties of fashion images and struggle to preserve semantic consistency during attribute manipulation. To address these limitations, we propose DiSCo, a novel disentangled attribute manipulation retrieval framework via semantic reconstruction and consistency regularization. Our approach comprises three key components: (1) An attribute-aware manipulation network that constructs target fashion embeddings through cross-modal attribute modification deltas, leveraging dedicated fashion attribute encoders; (2) A cross-modal semantic reconstruction network that synthesizes target images directly from modified attribute descriptions, supervised by adversarial and attribute classification losses to ensure interpretable edits; (3) An adaptive fusion mechanism that dynamically integrates attribute-modified embeddings with reconstructed image features. Extensive evaluations on two benchmark datasets (DeepFashion and Shopping100K) demonstrate that DiSCo achieves superior retrieval accuracy over state-of-the-arts while maintaining high-fidelity editing. Quantitative and qualitative analyses further confirm that DiSCo generates more realistic fashion representations, underscoring its effectiveness in attribute-aware retrieval tasks.
Min Tan 0005, Guanhao Liu, Huijing Zhan, Yuyu Yin, Zhou Yu 0001, Jiajun Ding, Yinfu Feng
ACM Multimedia3
2025 Diffusion-EXR: Controllable Review Generation for Explainable Recommendation via Diffusion Models
abstract
Denoising Diffusion Probabilistic Model (DDPM) has shown great competence in image and audio generation tasks. However, there exist few attempts to employ DDPM in the text generation, especially review generation under recommendation systems. Fueled by the predicted reviews' explainability that justifies recommendations could assist users better understand the recommended items and increase the transparency of recommendation system, we propose a Diffusion Model-based Review Generation towards EXplainable Recommendation named Diffusion-EXR. Diffusion-EXR corrupts the sequence of review embeddings by incrementally introducing varied levels of Gaussian noise to the sequence of word embeddings and learns to reconstruct the original word representations in the reverse process. The nature of DDPM enables our lightweight Transformer backbone to perform excellently in recommendation review generation. Instead of relying solely on texture descriptions, we propose integrating visual tokens into the framework to better capture semantics and promote greater diversity in the results. Extensive experimental results have demonstrated that Diffusion-EXR can achieve state-of-the-art review generation for recommendation on two publicly available benchmark datasets.
June Tay, Huijing Zhan
VCIP4
2025 Improving multi-modal brain tumor segmentation via pre-training and knowledge distillation based post-training
Weide Liu, Jingwen Hou, Xiaoyang Zhong, Huijing Zhan, Jun Cheng 0003, Yuming Fang 0001, Guanghui Yue 0001
Neurocomputing4
2024 Contrastive Learning with Bidirectional Transformers for Knowledge Tracing
abstract
Knowledge tracing aims to predict students’ probability of correctly answering the next question based on their interaction history. Previous methods employ left-to-right unidirectional transformers to encode the historical behaviors into hidden representations, especially with contrastive learning methods. Using uni-directional models to model student behaviors can only learn the hidden representation from its previous items, which restricts the power of the representation capability. Inspired by the success of BERT in text understanding, we propose a novel Bidirectional Transformer encoder guided Contrastive Learning framework for deep Knowledge Tracing system, named as Bi-CL4KT to generate the accurate response predictions for the next question. We incorporate the Cloze task and carefully design data augmentation methods to generate high-quality positive and negative instances for contrastive learning. Extensive experiments conducted on three real-world education datasets show that the proposed method significantly outperforms the state-of-the-art methods.
Huijing Zhan, Jung-Jae Kim 0001, Guimei Liu
ICASSP1
2024 Question Difficulty Consistent Knowledge Tracing
abstract
Knowledge tracing aims to estimate knowledge states of students over a set of skills based on students' past learning activities. Deep learning based knowledge tracing models show superior performance to traditional knowledge tracing approaches. Early works like DKT use skill IDs and student responses only. Recent works also incorporate questions IDs into their models and achieve much improved performance in the next question correctness prediction task. However, predictions made by these models are thus on specific questions, and it is not straightforward to translate them to estimation of students' knowledge states over skills. In this paper, we propose to replace question IDs with question difficulty levels in deep knowledge tracing models. The predictions made by our model can be more readily translated to students' knowledge states over skills. Furthermore, by using question difficulty levels to replace question IDs, we can also alleviate the cold-start problem in knowledge tracing as online learning platforms are updated frequently with new questions. We further use two techniques to smooth the predicted scores. One is to combine embeddings of nearby difficulty levels using the Hann function. The other is to constrain the predicted probabilities to be consistent with question difficulties by imposing a penalty if they are not consistent. We conduct extensive experiments to study the performance of the proposed model. Our experimental results show that our model outperforms the state-of-the-art knowledge tracing models in terms of both accuracy and consistency with question difficulty levels.
Guimei Liu, Huijing Zhan, Jung-Jae Kim 0001
WWW2
2023 Towards Explainable Recommendation Via Bert-Guided Explanation Generator
abstract
Explainable recommender system has recently drawn increasing attention due to its capability of providing justification to recommendation. Rather than focusing on certain topics or specific item features, the explanation generated by existing works are too general without the guidance of aspects. However, such information is not given in the practical scenario. To address this issue, we propose a novel Explainable recommender system with BERT-guided explanation generator, named ExBERT to generate reliable explanation with finer granularity. More specifically, a multi-head self-attention based encoder is employed to incorporate pseudo user and item profiles into semantic representation. Moreover, we propose a novel matched explanation prediction task with discriminative ability to enable personalization of the generated sentence. Extensive experiments conducted on two real-world explainable recommendation datasets significantly outperform the state-of-the-art in generation.
Huijing Zhan, Ling Li 0012, Weide Liu, Manas Gupta, Alex Chichung Kot
ICASSP1
2022 $A^3$-FKG: Attentive Attribute-Aware Fashion Knowledge Graph for Outfit Preference Prediction
abstract
With the booming development of the online fashion industry, effective personalized recommender systems have become indispensable for the convenience they brought to the customers and the profits to the e-commercial platforms. Estimating the user’s preference towards the outfit is at the core of a personalized recommendation system. Existing works on fashion recommendation are largely centering on modelling the clothing compatibility without considering the user factor or characterizing the user’s preference over the single item. However, how to effectively model the outfits with either few or even none interactions, is yet under-explored. In this paper, we address the task of personalized outfit preference prediction via a novelAttentiveAttribute-AwareFashionKnowledgeGraph ($A^3$-FKG), which is incorporated to build the association between different outfits with both outfit- and item- level attributes. Additionally, a two-level attention mechanism is developed to capture the user’s preference: 1) User-specific relation-aware attention layer, which captures the user’s fine-grained preferences with different focus on relations for learning outfit representation; 2) Target-aware attention layer, which characterizes the user’s latent diverse interests from his/her behavior sequences for learning user representation. Extensive experiments conducted on a large-scale fashion outfit dataset demonstrate significant improvements over other methods, which verify the excellence of our proposed framework.
Huijing Zhan, Jie Lin 0001, Kenan E. Ak, Boxin Shi, Ling-Yu Duan, Alex Chichung Kot
IEEE Trans. Multim.1
2021 PAN: Personalized Attention Network For Outfit Recommendation
abstract
Recent years have witnessed the dramatic development of e-fashion industry, it becomes essential to build an intelligent fashion recommender system. Most of existing works on fashion recommendation focus on modeling the general compatibility while ignoring the user preferences. In this paper, we present a Personalized Attention Network (PAN) for fashion recommendation. The key component of PAN includes a user encoder, an item encoder and a preference predictor. To modeling users’ diverse interests, we develop an attention network to incorporate the learnt user representation into the item encoder component. More specifically, the attention module consists of a sequential user-aware channel-level and a spatial-level sub-module. Moreover, a novel ranking an user-specific loss, is proposed to capture the interest of different users on the same outfit. To make the training more effective and efficient, a novel user-aware online hard negative mining strategy is proposed. Extensive experiments on Polyvore-U dataset demonstrate the excellence of the proposed system and the effectiveness of different modules.
Huijing Zhan, Jie Lin 0001
ICIP1
2021 Pose-Normalized and Appearance-Preserved Street-to-Shop Clothing Image Generation and Feature Learning
abstract
We tackle the task of street-to-shop clothing image synthesis. Given a daily person image with a particular clothing item captured in the street scenario, we aim to synthesize the frontal facing view of that item in the shop scenario. This problem has the following challenges: 1) the distinct visual discrepancy between the street and shop scenario; 2) the severe shape deformation of clothing in the presence of an arbitrary human pose; 3) the preservation of fine-grained details during the process of clothing image generation. In this paper, we jointly solve these difficulties by proposing a Pose-Normalized and Appearance-Preserved Generative Adversarial Network (PNAP-GAN). More specifically, conditioned on the clothing-agnostic representation (i.e., clothing landmarks and semantic parsing map), we disentangle the shape and appearance synthesis in a coarse-to-fine framework. Moreover, a semantic embedding loss is introduced to guide the domain transfer in the semantic level (i.e., keeping the clothing attributes). With the synthesized frontal shop image, a pose-normalized representation in complementary to the domain-invariant feature learnt from the original street image are integrated to facilitate the problem of street-to-shop clothing retrieval. Extensive experiments conducted demonstrate the effectiveness of the proposed PNAP-GAN on generating high quality frontal-view images and the excellence of the learnt pose-normalized features on the retrieval task than existing methods. In addition, we demonstrate that the pose-normalized retrieval feature benefits the cross-scenario (i.e., street-to-shop) clothing image generation in a semantic-preserved manner.
Huijing Zhan, Chenyu Yi, Boxin Shi, Jie Lin 0001, Ling-Yu Duan, Alex Chichung Kot
IEEE Trans. Multim.1
2020 A*3D Dataset: Towards Autonomous Driving in Challenging Environments
abstract
With the increasing global popularity of self-driving cars, there is an immediate need for challenging real-world datasets for benchmarking and training various computer vision tasks such as 3D object detection. Existing datasets either represent simple scenarios or provide only day-time data. In this paper, we introduce a new challenging A*3D dataset which consists of RGB images and LiDAR data with a significant diversity of scene, time, and weather. The dataset consists of high-density images (≈ 10 times more than the pioneering KITTI dataset), heavy occlusions, a large number of nighttime frames (≈ 3 times the nuScenes dataset), addressing the gaps in the existing datasets to push the boundaries of tasks in autonomous driving research to more challenging highly diverse environments. The dataset contains 39K frames, 7 classes, and 230K 3D object annotations. An extensive 3D object detection benchmark evaluation on the A*3D dataset for various attributes such as high density, day-time/night-time, gives interesting insights into the advantages and limitations of training and testing 3D object detection in real-world setting.
Quang-Hieu Pham, Pierre Sevestre, Ramanpreet Singh Pahwa, Huijing Zhan, Chun Ho Pang, Yuda Chen, Armin Mustafa, Vijay Chandrasekhar 0001, Jie Lin 0001
ICRA4
2019 Fashion Recommendation on Street Images
abstract
Learning the compatibility relationship is of vital importance to a fashion recommendation system, while existing works achieve this merely on product images but not on street images in the complex daily life scenario. In this paper, we propose a novel fashion recommendation system: Given a query item of interest in the street scenario, the system can return the compatible items. More specifically, a two-stage curriculum learning scheme is developed to transfer the semantics from the product to street outfit images. We also propose a domain-specific missing item imputation method based on style and color similarity to handle the incomplete outfits. To support the training of deep recommendation model, we collect a large dataset with street outfit images. The experiments on the dataset demonstrate the advantages of the proposed method over the state-of-the-art approaches on both the street images and the product images.
Huijing Zhan, Boxin Shi, Ling-Yu Duan, Alex Chichung Kot
ICIP1
2019 DeepShoe: An improved Multi-Task View-invariant CNN for street-to-shop shoe retrieval
Huijing Zhan, Boxin Shi, Ling-Yu Duan, Alex Chichung Kot
Comput. Vis. Image Underst.1
2017 DeepShoe: A Multi-Task View-Invariant CNN for Street-to-Shop Shoe Retrieval
Huijing Zhan, Boxin Shi, Alex Chichung Kot
BMVC1
2017 Street-to-shop shoe retrieval with multi-scale viewpoint invariant triplet network
abstract
In this paper we aim to find exactly the same shoes given a daily shoe photo (street scenario) that matches the online shop shoe photo (shop scenario). There are large visual differences between the street and shop scenario shoe images. To handle the discrepancy of different scenarios, we learn a feature embedding for shoes via a viewpoint-invariant triplet network, the feature activations of which reflect the inherent similarity between any two shoe images. Specifically, we propose a new loss function that minimizes the distances between images of the same shoes captured from different viewpoints. Moreover, we train the proposed triplet network at two different scales so that the representation of shoes incorporates different levels of invariance at different scales. To support training the multi-scale triplet networks, we collect a large dataset with shoe images from the daily life and online shopping websites. Experiments on the dataset show excellence over state-of-the-art approaches, which demonstrate the effectiveness of our proposed method.
Huijing Zhan, Boxin Shi, Alex Chichung Kot
ICIP1
2017 Fashion analysis with a subordinate attribute classification network
abstract
In this paper we deal with two image-based object search tasks in the fashion domain, clothing attribute prediction and cross-domain shoe retrieval. Clothing attribute prediction is about describing the appearances of clothes via semantic attributes and cross-domain shoe retrieval aims at retrieving the same shoe items from online stores given a daily life shoe photo. We jointly solve these two problems by a novel Subordinate Attribute Convolutional Neural Network (SA-CNN), with the newly designed loss function that systematically merges semantic attributes of closer visual appearance to prevent images with obvious visual differences being confused with each other. A three-level feature representation is further developed based on SA-CNN for shoes from different domains. The experimental results demonstrate that the clothing attribute prediction using the proposed SA-CNN achieves better performance than that using traditional features and fine-tuned conventional CNN. Moreover, for the task of cross-domain shoe retrieval, the top-20 retrieval accuracy with deep features extracted from SA-CNN has a significant improvement of 43% compared to that with the pretrained CNN features.
Huijing Zhan, Boxin Shi, Alex Chichung Kot
ICME1
2017 Cross-domain shoe retrieval using a three-level deep feature representation
abstract
In this paper, we address the problem of matching the shoes from the daily life photos to exactly the same shoes from online shops. The problem is extremely challenging because of the significant visual differences between street domain images (shoe images captured in the daily life scenario) and online domain photos (images from online shops taken in the controlled environment). This paper presents a semantic Shoe Attribute-Guided Convolutional Neural Network (SAG-CNN) to extract the deep features. Moreover, we develop a three-level feature representation based on SAG-CNN. The deep features extracted from the image, region and part levels effectively match the images across different domains. We collect a novel shoe dataset, which consists of 8021 street domain and 5821 online domain images. The experimental results on our dataset show that the top-20 retrieval accuracy of our approach improves over that using the pre-trained CNN features by about 40%.
Huijing Zhan, Boxin Shi, Alex Chichung Kot
ISCAS1
2017 Cross-Domain Shoe Retrieval With a Semantic Hierarchy of Attribute Classification Network
abstract
Cross-domain shoe image retrieval is a challenging problem, because the query photo from the street domain (daily life scenario) and the reference photo in the online domain (online shop images) have significant visual differences due to the viewpoint and scale variation, self-occlusion, and cluttered background. This paper proposes the semantic hierarchy of attribute convolutional neural network (SHOE-CNN) with a three-level feature representation for discriminative shoe feature expression and efficient retrieval. The SHOE-CNN with its newly designed loss function systematically merges semantic attributes of closer visual appearances to prevent shoe images with the obvious visual differences being confused with each other; the features extracted from image, region, and part levels effectively match the shoe images across different domains. We collect a large-scale shoe data set composed of 14341 street domain and 12652 corresponding online domain images with fine-grained attributes to train our network and evaluate our system. The top-20 retrieval accuracy improves significantly over the solution with the pre-trained CNN features.
Huijing Zhan, Boxin Shi, Alex Chichung Kot
IEEE Trans. Image Process.1