Haoruo Zhang

dblp:189/7592 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
2since 2021 · last 2026
0000-0002-4508-096XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
1 paper
Multimedia analysis and retrieval · 70% Audio and music processing · 30%
Artificial intelligence
1 paper
3D vision · 67% Robot manipulation · 33%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Multimedia analysis and retrieval
cross-modal retrieval
0.512021
Deep Music Retrieval for Fine-Grained Videos by Exploiting Cross-Modal-Encoded Voice-Overs · SIGIR 2021
Audio and music processing
music information retrieval
0.512021
Deep Music Retrieval for Fine-Grained Videos by Exploiting Cross-Modal-Encoded Voice-Overs · SIGIR 2021
Multimedia analysis and retrieval › cross-modal retrieval
video-music retrieval
0.512021
Deep Music Retrieval for Fine-Grained Videos by Exploiting Cross-Modal-Encoded Voice-Overs · SIGIR 2021
Robotics › Robot manipulation › grasping › grasping in clutter
bin picking
0.412019
Detect in RGB, Optimize in Edge: Accurate 6D Pose Estimation for Texture-less Industrial Parts · ICRA 2019
Computer vision › 3D vision
object pose estimation
0.412019
Detect in RGB, Optimize in Edge: Accurate 6D Pose Estimation for Texture-less Industrial Parts · ICRA 2019
Computer vision › 3D vision › object pose estimation
texture-less object pose estimation
0.412019
Detect in RGB, Optimize in Edge: Accurate 6D Pose Estimation for Texture-less Industrial Parts · ICRA 2019
Multimedia analysis and retrieval
video understanding
0.112021
Deep Music Retrieval for Fine-Grained Videos by Exploiting Cross-Modal-Encoded Voice-Overs · SIGIR 2021

Methods — techniques the papers use, named apart from their topics

triplet loss · 0.5self-attention · 0.5cross-modal attention · 0.5edge-based optimization · 0.4convolutional neural network · 0.4
YearPublicationVenuePosition
2026 DMV - CLIP : Disentangled Multimodal Visual Adaptation for Text-Driven Face Editing
abstract
ABSTRACT Text‐driven face editing has attracted widespread interest due to its intuitive control and user‐friendly interaction. However, current state‐of‐the‐art (SOTA) methods face two main challenges: (1) they utilize unfinetuned general image‐text encoders for modality fusion, making it difficult to comprehend domain‐specific knowledge in facial attribute editing (dozens of fine‐grained facial attributes such as moustache and lipsticks); (2) they roughly optimize all attributes simultaneously using a cross‐entropy loss, leading to severe mutual interference among attributes. To this end, we propose Disentangled Multimodal Visual Adaptation for CLIP (DMV‐CLIP). First, DMV‐CLIP incorporates learnable context tokens to inject facial domain knowledge into the CLIP model via multimodal prompt learning (MPL). Second, it employs directional contrastive learning (DCL) to disentangle facial attributes and enable precise editing. Finally, DMV‐CLIP utilizes a vision‐language consistency model (VLCM) to maintain identity consistency while ensuring that the generated images strictly adhere to the semantic instructions.
Xin Wei 0002, Huan Wan, Haoruo Zhang, Xuhui Huang
Expert Syst. J. Knowl. Eng.5
2021 Deep Music Retrieval for Fine-Grained Videos by Exploiting Cross-Modal-Encoded Voice-Overs
abstract
Recently, the witness of the rapidly growing popularity of short videos on different Internet platforms has intensified the need for a background music (BGM) retrieval system. However, existing video-music retrieval methods only based on the visual modality cannot show promising performance regarding videos with fine-grained virtual contents. In this paper, we also investigate the widely added voice-overs in short videos and propose a novel framework to retrieve BGM for fine-grained short videos. In our framework, we use the self-attention (SA) and the cross-modal attention (CMA) modules to explore the intra- and the inter-relationships of different modalities respectively. For balancing the modalities, we dynamically assign different weights to the modal features via a fusion gate. For paring the query and the BGM embeddings, we introduce a triplet pseudo-label loss to constrain the semantics of the modal embeddings. As there are no existing virtual-content video-BGM retrieval datasets, we build and release two virtual-content video datasets HoK400 and CFM400. Experimental results show that our method achieves superior performance and outperforms other state-of-the-art methods with large margins.
Tingtian Li, Zixun Sun, Haoruo Zhang, Ziming Wu, Hui Zhan, Yipeng Yu, Hengcan Shi
SIGIR3
2019 Detect in RGB, Optimize in Edge: Accurate 6D Pose Estimation for Texture-less Industrial Parts
abstract
In order to solve robotic bin-picking problem in many industrial applications, accurate 6D object pose estimation is one of fundamental technologies. This paper presents a method for accurate 6D pose estimation from a single RGB image for texture-less industrial parts. These objects are common but still challenging to deal with, due to the fact that poor surface texture and brightness makes difficult to compute discriminative local appearance descriptors. The proposed method mainly consists of two stages, which ranges from the detection stage to the optimization stage. Firstly, all known objects in the RGB image are detected with 2D bounding box via a tiny convolutional neural network. Then, the second stage will optimize the 6D pose in the Edge image given several coarse initializations. These coarse initializations are generated from the Edge image via a hypothesis-evaluation scheme. Furthermore, the proposed method is validated by achieving state-of-the-art results of texture-less industrial parts for RGB input. According to practical experiments, the proposed method is accurate and robust enough to be applied on the robotic manipulation platform to complete a simple assembly task.
Haoruo Zhang, Qixin Cao
ICRA1
2019 Fast 6D object pose refinement in depth images
Haoruo Zhang, Qixin Cao
Appl. Intell.1
2019 Holistic and local patch framework for 6D object pose estimation in RGB-D images
Haoruo Zhang, Qixin Cao
Comput. Vis. Image Underst.1