Haofeng Zhang 0001

dblp:17/4297 · also Hao-Feng Zhang 0001, Hao-feng Zhang 0001 · DBLP profile ↗
← Back
9ranked-venue papers in the field
1as first author
7since 2021 · last 2025
0000-0002-4039-7618ORCID · conflict

Domains — venue-derived; a paper can count in several

Knowledge Engineering, Semantic Web & Information Systems · 4 (1 first)Other / Interdisciplinary · 4Information Retrieval & Web Search · 1
YearPublicationVenuePosition
2025 Test-time Filtering Boosts Training-free Zero-shot Composed Image Retrieval
abstract
Composed Image Retrieval (CIR) is an image retrieval task where users provide a reference image along with modification text to retrieve a target image. Zero-shot CIR (ZS-CIR) attracts significant research interest owing to its strong generalization capability and independence from labeled training data. Most ZS-CIR methods employ late fusion and textual inversion, but these approaches fail to precisely convert the reference image and modification text into a target-aligned query. Consequently, the final query may retain redundant information—such as elements present in the reference image but absent in the target image. To address these limitations, a training-free ZS-CIR method called Test-time Filtering (TTF) is proposed in this paper, leveraging Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs). Specifically, an MLLM generates captions for reference images, which, along with the modification text, are fed into an LLM to produce target image descriptions. This converts the final query into plain text. Using these results, an MLLM filters the selected candidate images to select positive and negative samples. Finally, the similarity scores of candidate images matching positive samples are amplified, while those of negative samples are suppressed. The proposed method is evaluated on three ZS-CIR datasets—CIRR, FashionIQ, and CIRCO—with experimental results demonstrating superior performance over prior ZS-CIR approaches. The source code is available at https://github.com/After-lifes/TTF.
Haoyue Chong, Lunbo Li, Haofeng Zhang 0001
MMAsia3
2025 Leveraging Pseudo-triplet and Flexible Prompt for Zero-shot Composed Image Retrieval
abstract
Compared to supervised composed image retrieval (CIR), which requires a large number of manually labeled triplets (reference image, modification text, target image) for training, Zero-shot Composed Image Retrieval (ZS-CIR) only needs easily available image-caption pairs for training. Prior works for ZS-CIR have primarily used textual inversion to convert images from image-caption pairs into pseudo-words using a pre-trained Vision-language Model (VLM). These pseudo-words are then concatenated to a template, such as “a photo of $”. We find that this fixed template representation hinders the generalization ability of ZS-CIR, and that converting an image into a single pseudo-token fails to accurately represent the information contained in the image. Additionally, given the high performance demonstrated by traditional supervised CIR after triplet training, we aim to achieve a similar structure in ZS-CIR. In this paper, we propose a novel ZS-CIR paradigm based on pseudo-triplets and flexible prompts, rather than fixed templates. Specifically, we fine-tune the vision encoder and split the caption to generate pseudo-triplets using random patch masking. The obtained masked image tokens and corresponding split captions are fed into a Q-former, which is guided by the learned queries and generates queries containing both image and caption information. Finally, the learned queries are concatenated with the split caption and fed into the text encoder to obtain a composed query. Extensive experiments on commonly used ZS-CIR benchmarks such as CIRR, FashionIQ, and CIRCO demonstrate that our method outperforms previous ZS-CIR approaches. The source code is available at https://github.com/After-lifes/PFP.
Haoyue Chong, Lunbo Li, Haofeng Zhang 0001
MMAsia3
2024 Extrinsic Calibration of Camera and LiDAR Systems With Three-Dimensional Towered Checkerboards
abstract
With the increasing utilization of cameras and three‐dimensional Light Detection and Ranging (LiDAR) systems in perception tasks, the fusion of these two sensor modalities has emerged as a prominent research focus in the fields of robotics and unmanned systems. While various extrinsic calibration methods have been developed, they often suffer from limited accuracy when using low‐resolution LiDAR sensors and require the placement of calibration targets at multiple locations. This paper introduces a novel calibration target known as the Three‐Dimensional Towered Checkerboard (3TC), along with a precise and straightforward extrinsic calibration approach for camera‐LiDAR systems. The 3TC consists of stacked cubes adorned with planar or 2D checkerboards, which provide the known positions of checkerboard corner points in three‐dimensional space. Leveraging the Iterative Closest Point (ICP) algorithm, the proposed method calculates the spatial relationship between LiDAR point cloud data and the 3TC model to infer the positions of checkerboard corner points in the LiDAR coordinate system. Subsequently, the Perspective‐n‐Point (PnP) algorithm is employed to establish the correlation between corner positions in the LiDAR coordinate system and the camera image, given the intrinsic parameters of the camera. By ensuring an adequate number of cubes and 2D checkerboards on a specific 3TC, along with accurately estimated corner point positions in LiDAR, a single frame of data from both the camera and LiDAR facilitates their extrinsic calibration. Experimental validations conducted across diverse camera and LiDAR systems, achieving minimal error close to the theoretical limit of the devices, attest to the robustness and precision of the 3TC and the proposed calibration methodology.
Dexin Ren, Mingwu Ren, Haofeng Zhang 0001
Int. J. Intell. Syst.3
2023 Multi-Scale Superpoint Network for 3D Point Cloud Semantic Segmentation
abstract
3D point cloud semantic segmentation is a fundamental task for 3D scene understanding. However, most existing pipelines usually use k-NN or ball query operation to form hard neighborhoods, which may cross different semantic objects, resulting low-quality local features. To address this issue, we propose a multi-scale superpoint network that gradually generates multi-scale soft neighborhoods to extract geometric local features, thereby boosting the 3D semantic segmentation performance. Specifically, we present a simple yet efficient superpoint merging module that merge small-scale superpoints to obtain large-scale superpoint by considering the feature similarity of superpoints, so that we can obtain multi-scale geometric features of point clouds. We also develop a superpoint upsampling module that adopt inverse mapping function to propagate multi-scale features from low-resolution point cloud to high-resolution point cloud. By integrating our multi-scale superpoint network into a simple point based semantic segmentation network, our method can obtain SOTA results on S3DIS Area 5 and 6-fold, and competitive results on ScanNet v2.
Ft Zheng, Le Hui, Jin Xie 0001, Haofeng Zhang 0001
MMAsia4
2023 Data driven recurrent generative adversarial network for generalized zero shot image classification
Jie Zhang 0005, Shengbin Liao, Haofeng Zhang 0001, Yang Long 0001, Zheng Zhang 0006, Li Liu 0004
Inf. Sci.3
2022 Semi-supervised cross-modal hashing with multi-view graph representation
Haofeng Zhang 0001, Lunbo Li, Wankou Yang, Li Liu 0004
Inf. Sci.2
2021 Sparse graph based self-supervised hashing for scalable image retrieval
Haofeng Zhang 0001, Zheng Zhang 0006, Li Liu 0004, Ling Shao 0001
Inf. Sci.2
2020 Semantic-rebased cross-modal hashing for scalable unsupervised text-visual retrieval
Yuming Shen, Haofeng Zhang 0001, Li Liu 0004
Inf. Process. Manag.3
2019 Dual-verification network for zero-shot learning
Haofeng Zhang 0001, Yang Long 0001, Wankou Yang, Ling Shao 0001
Inf. Sci.1