VLDB 2026 Research / reviewers in the wild / expert
Huasong Zhong
dblp:227/3501
· DBLP profile ↗
11ranked-venue papers
2as first author
10since 2021 · last 2026
0000-0001-7172-0556ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UniAPO: Unified Multimodal Automated Prompt OptimizationabstractPrompting is fundamental to unlocking the full potential of large language models. To automate and enhance this process, automatic prompt optimization (APO) has been developed, demonstrating effectiveness primarily in text-only input scenarios. However, extending existing APO methods to multimodal tasks—such as video-language generation—introduces two core challenges: (i) visual token inflation, where long visual-token sequences restrict context capacity and result in insufficient feedback signals; (ii) a lack of process-level supervision, as existing methods focus on outcome-level supervision and overlook intermediate supervision, limiting prompt optimization. We present UniAPO: Unified Multimodal Automated Prompt Optimization, the first framework tailored for multimodal APO. UniAPO adopts an EM-inspired optimization process that decouples feedback modeling and prompt refinement, making the optimization more stable and goal-driven. To further address the aforementioned challenges, we introduce a short-long term memory mechanism: historical feedback mitigates context limitations, while historical prompts provide directional guidance for effective prompt optimization. UniAPO achieves consistent gains across text, image, and video benchmarks, establishing a unified framework for efficient and transferable prompt optimization. Qipeng Zhu, Yanzhe Chen, Huasong Zhong, Junping Zhang, Zhenheng Yang |
AAAI | 3 |
| 2024 | FashionERN: Enhance-and-Refine Network for Composed Fashion Image RetrievalabstractThe goal of composed fashion image retrieval is to locate a target image based on a reference image and modified text. Recent methods utilize symmetric encoders (e.g., CLIP) pre-trained on large-scale non-fashion datasets. However, the input for this task exhibits an asymmetric nature, where the reference image contains rich content while the modified text is often brief. Therefore, methods employing symmetric encoders encounter a severe phenomenon: retrieval results dominated by reference images, leading to the oversight of modified text. We propose a Fashion Enhance-and-Refine Network (FashionERN) centered around two aspects: enhancing the text encoder and refining visual semantics. We introduce a Triple-branch Modifier Enhancement model, which injects relevant information from the reference image and aligns the modified text modality with the target image modality. Furthermore, we propose a Dual-guided Vision Refinement model that retains critical visual information through text-guided refinement and self-guided refinement processes. The combination of these two models significantly mitigates the reference dominance phenomenon, ensuring accurate fulfillment of modifier requirements. Comprehensive experiments demonstrate our approach's state-of-the-art performance on four commonly used datasets. Yanzhe Chen, Huasong Zhong, Xiangteng He, Yuxin Peng 0001, Jiahuan Zhou, Lele Cheng |
AAAI | 2 |
| 2023 | Real20M: A Large-scale E-commerce Dataset for Cross-domain RetrievalabstractIn e-commerce, products and micro-videos serve as two primary carriers. Introducing cross-domain retrieval between these carriers can establish associations, thereby leading to the advancement of specific scenarios, such as retrieving products based on micro-videos or recommending relevant videos based on products. However, existing datasets only focus on retrieval within the product domain while neglecting the micro-video domain and often ignore the multi-modal characteristics of the product domain. Additionally, these datasets strictly limit their data scale through content alignment and use a content-based data organization format that hinders the inclusion of user retrieval intentions. To address these limitations, we propose the PKU Real20M dataset, a large-scale e-commerce dataset designed for cross-domain retrieval. We adopt a query-driven approach to efficiently gather over 20 million e-commerce products and micro-videos, including multimodal information. Additionally, we design a three-level entity prompt learning framework to align inter-modality information from coarse to fine. Moreover, we introduce the Query-driven Cross-Domain retrieval framework (QCD), which leverages user queries to facilitate efficient alignment between the product and micro-video domains. Extensive experiments on two downstream tasks validate the effectiveness of our proposed approaches. The dataset and source code are available at https://github.com/PKU-ICST-MIPL/Real20M_ACMMM2023. Yanzhe Chen, Huasong Zhong, Xiangteng He, Yuxin Peng 0001, Lele Cheng |
ACM Multimedia | 2 |
| 2023 | Unsupervised graph-level representation learning with hierarchical contrasts
Wei Ju 0001, Yiyang Gu, Xiao Luo 0001, Yifan Wang 0014, Huasong Zhong, Ming Zhang 0004 |
Neural Networks | 6 |
| 2022 | On Mitigating Hard Clusters for Face Clustering
Yingjie Chen 0002, Huasong Zhong, Chong Chen 0002, Chen Shen 0003, Jianqiang Huang 0001, Tao Wang 0004, Yun Liang 0001, Qianru Sun |
ECCV (12) | 2 |
| 2021 | Graph Contrastive ClusteringabstractRecently, some contrastive learning methods have been proposed to simultaneously learn representations and clustering assignments, achieving significant improvements. However, these methods do not take the category information and clustering objective into consideration, thus the learned representations are not optimal for clustering and the performance might be limited. Towards this issue, we first propose a novel graph contrastive learning framework, and then apply it to the clustering task, resulting in the Graph Constrastive Clustering (GCC) method. Different from basic contrastive clustering that only assumes an image and its augmentation should share similar representation and clustering assignments, we lift the instance-level consistency to the cluster-level consistency with the assumption that samples in one cluster and their augmentations should all be similar. Specifically, on the one hand, we propose the graph Laplacian based contrastive loss to learn more discriminative and clustering-friendly features. On the other hand, we propose a novel graph-based contrastive learning strategy to learn more compact clustering assignments. Both of them incorporate the latent category information to reduce the intra-cluster variance as well as increase the inter-cluster variance. Experiments on six commonly used datasets demonstrate the superiority of our proposed approach over the state-of-the-art methods.1 Huasong Zhong, Jianlong Wu, Chong Chen 0002, Jianqiang Huang 0001, Minghua Deng, Liqiang Nie, Zhouchen Lin, Xian-Sheng Hua 0001 |
ICCV | 1 |
| 2021 | Deep Unsupervised Hashing by Distilled Smooth GuidanceabstractHashing has been widely used in approximate nearest neighbor search recently. Deep supervised hashing methods are not widely-used because of the lack of labeled data, especially when the domain is transferred. Meanwhile, unsupervised deep hashing models can hardly achieve satisfactory performance due to the lack of reliable similarity signals. Here, we propose a novel deep unsupervised hashing method, namely Distilled Smooth Guidance (DSG), which can learn a distilled dataset consisting of similarity signals as well as smooth confidence signals. Specifically, we obtain the similarity confidence weights based on the initial noisy similarity signals learned from local structures and construct a priority loss function for smooth similarity-preserving learning. Besides, global information based on clustering is utilized to distill the image pairs by removing contradictory similarity signals. Extensive experiments on three widely used bench-mark datasets show that the proposed DSG consistently out-performs the state-of-the-art search methods. Xiao Luo 0001, Zeyu Ma 0001, Daqing Wu, Huasong Zhong, Chong Chen 0002, Jinwen Ma, Minghua Deng |
ICME | 4 |
| 2021 | Deep Supervised Hashing by Classification for Image Retrieval
Xiao Luo 0001, Yuhang Guo 0002, Zeyu Ma 0001, Huasong Zhong, Tao Li 0040, Wei Ju 0001, Chong Chen 0002, Minghua Deng |
ICONIP (4) | 4 |
| 2021 | Simulated annealing for optimization of graphs and sequences
Xianggen Liu, Pengyong Li, Fandong Meng, Hao Zhou 0012, Huasong Zhong, Jie Zhou 0016, Lili Mou, Sen Song |
Neurocomputing | 5 |
| 2021 | Self-Adaptive Neural Module Transformer for Visual Question AnsweringabstractVision and language understanding is one of the most fundamental and difficult tasks in Multimedia Intelligence. Simultaneously Visual Question Answering (VQA) is even more challenging since it requires complex reasoning steps to the correct answer. To achieve this, Neural Module Network (NMN) and its variants rely on parsing the natural language question into a module layout (i.e., a problem-solving program). In particular, this process follows a feedforward encoder-decoder pipeline: the encoder embeds the question into a static vector and the decoder generates the layout. However, we argue that such conventional encoder-decoder neglects the dynamic nature of question comprehension (i.e., we should attend to different words from step to step) and per-module intermediate results (i.e., we should discard module performing badly) in the reasoning steps. In this paper, we present a novel NMN, called Self-Adaptive Neural Module Transformer (SANMT), which adaptively adjusts both of the question feature encoding and the layout decoding by considering intermediate Q&A results. Specifically, we encode the intermediate results with the given question features by a novel transformer module to generate dynamic question feature embedding which evolves over reasoning steps. Besides, the transformer utilizes the intermediate results from each reasoning step to guide subsequent layout arrangement. Extensive experimental evaluations demonstrate the superiority of the proposed SANMT over NMN and its variants on four challenging benchmarks, including CLEVR, CLEVR-CoGenT, VQAv1.0, and VQAv2.0 (on average the relative improvement over NMN are 1.5, 2.3, 0.7 and 0.5 points with respect to accuracy). Huasong Zhong, Jingyuan Chen 0003, Chen Shen 0003, Hanwang Zhang, Jianqiang Huang 0001, Xian-Sheng Hua 0001 |
IEEE Trans. Multim. | 1 |
| 2019 | AddressNet: Shift-Based Primitives for Efficient Convolutional Neural NetworksabstractWe propose a collection of three shift-based primitives for building efficient compact CNN-based networks. These three primitives (channel shift, address shift, shortcut shift) can reduce the inference time on GPU while maintains the prediction accuracy. These shift-based primitives only moves the pointer but avoids memory copy, thus very fast. For example, the channel shift operation is 12.7× faster compared to channel shuffle in ShuffleNet but achieves the same accuracy. The address shift and channel shift can be merged into the point-wise group convolution and invokes only a single kernel call, taking little time to perform spatial convolution and channel shift. Shortcut shift requires no time to realize residual connection through allocating space in advance. We blend these shift-based primitives with point-wise group convolution and built two inference-efficient CNN architectures named AddressNet and Enhanced AddressNet. Experiments on CIFAR100 and ImageNet datasets show that our models are faster and achieve comparable or better accuracy. Yihui He, Xianggen Liu, Huasong Zhong, Yuchun Ma |
WACV | 3 |