VLDB 2026 Research / reviewers in the wild / expert
Qi He 0007
dblp:51/6972-7
· DBLP profile ↗
9ranked-venue papers
4as first author
9since 2021 · last 2025
0000-0002-6109-9417ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Dual-Rate Dynamic Teacher for Source-Free Domain Adaptive Object Detection
Qi He 0007, Xiao Wu 0001, Jun-Yan He, Shuai Li 0014 |
ICCV | 1 |
| 2025 | MetaDesigner: Advancing Artistic Typography through AI-Driven, User-Centric, and Multilingual WordArt SynthesisabstractMetaDesigner introduces a transformative framework for artistic typography synthesis, powered by Large Language Models (LLMs) and grounded in a user-centric design paradigm. Its foundation is a multi-agent system comprising the Pipeline, Glyph, and Texture agents, which collectively orchestrate the creation of customizable WordArt, ranging from semantic enhancements to intricate textural elements. A central feedback mechanism leverages insights from both multimodal models and user evaluations, enabling iterative refinement of design parameters. Through this iterative process, MetaDesigner dynamically adjusts hyperparameters to align with user-defined stylistic and thematic preferences, consistently delivering WordArt that excels in visual quality and contextual resonance. Empirical evaluations underscore the system's versatility and effectiveness across diverse WordArt applications, yielding outputs that are both aesthetically compelling and context-sensitive. Jun-Yan He, Zhi-Qi Cheng, Chenyang Li 0007, Jingdong Sun, Qi He 0007, Wangmeng Xiang, Jin-Peng Lan, Xianhui Lin, Kang Zhu, Bin Luo 0008, Yifeng Geng, Xuansong Xie, Alex Hauptmann 0001 |
ICLR | 5 |
| 2025 | HOOI Detection: Cascade-Clue Integrated Modeling over Multiple Temporal SegmentsabstractTo fully comprehend a visual scene, recognizing and localizing interaction actions are essential components. Recently, significant advances have been made in detecting human-object interaction actions, which aim to capture pairwise relations between entities in the scene. Although these methods have made significant progress, they ignore the human-object-object interaction (HOOI) actions that frequently occur between a human and two objects in the real world. To advance related research, a new task named HOOI detection is introduced. It aims to accurately localize the humans in each video frame and identify the HOOI actions they perform. For this purpose, two novel HOOI datasets oriented to industrial production and daily life are constructed. These new datasets provide essential data support for in-depth research of HOOI detection. Furthermore, a cutting-edge method named Cascade-Clue Integrated Modeling over Multiple Temporal Segments (C2TS) is proposed for effectively detecting HOOI actions. Specifically, considering the phased characteristics of the action, C2TS comprehensively considers the HOOI information in the preceding, neighborhood, and subsequent temporal segments. For each temporal segment, the Cascaded Modeling and Clue Augmentation methods are applied to extract the corresponding HOOI features. The final detection result is obtained by classifying the aggregated HOOI features from the three temporal segments. Experiments conducted on the two proposed HOOI-related datasets vividly demonstrate that our method outperforms state-of-the-art approaches by achieving remarkable improvements of approximately 3% and 5%, which powerfully validates its efficacy in tackling the given challenge. Mingxuan Zhang 0001, Qi He 0007, Zhaoquan Yuan, Tingquan He |
ICMR | 2 |
| 2025 | DualEnhance: External Multimodal Foundation Models Guidance and Internal Fast-Slow Teacher RegulationabstractSource-Free Domain Adaptive Object Detection addresses cross-domain detection on an unlabeled target domain without accessing source data. Existing methods implement self-training with Mean Teacher but are bottlenecked by error accumulation from noisy pseudo-labels generated via recursive teacher-student updates. This issue is handled through the proposed dual enhancements: (1) External Guidance via Multimodal Foundation Models (FMs); (2) Internal Regulation through Fast-Slow Teacher. First, despite FMs' multimodal comprehension, their semantic misalignment with a specific task introduces noise during adaptation. Bidirectional Distillation mitigates this by calibrating the FM using task-specific knowledge transferred from the source detector. The aligned cross-modal knowledge then propagates through high-quality pseudo-label generation. Second, the conventional Mean Teacher suffers from plasticity-stability dilemma, where rapid adaptation corrupts historical knowledge. Fast-Slow Teacher introduces dual-velocity knowledge consolidation: The Fast Teacher dynamically captures emerging domain features, while the Slow Teacher preserves stable historical knowledge and periodically resets the Fast Teacher, establishing an error-correcting dynamic equilibrium. Experiments show our method achieves significant improvements over SOTA. Qi He 0007, Xiao Wu 0001, Jun-Yan He, Wei Li 0110, Zhaoquan Yuan |
ACM Multimedia | 1 |
| 2024 | MagicCartoon: 3D Pose and Shape Estimation for Bipedal Cartoon CharactersabstractThe 3D model can be estimated by regressing the pose and shape parameters from the image data of the digital model. The reconstruction of 3D cartoon characters poses a challenging task due to diverse visual representations and postural variations. This paper proposes a dual-branch structure named MagicCartoon for 3D bipedal cartoon character estimation, which models pose and shape independently through feature decoupling. Considering the correlation between category difference and shape parameters, a hybrid feature fusion technique is introduced, which integrates the global features of the original image with the corresponding local features expressed by the puzzle image, reducing the abstractness of understanding shape parameter differences. To semantically align image and geometric between feature space, a geometric-guided feedback loop is proposed in an iterative way, so that the pose of modeling results can be expressed consistently with the image. Moreover, a feature consistency loss is designed to augment the training data by incorporating the same character with different postures and the same posture of different characters. It enhances the correlation between the features extracted by the backbone network and the specific task. Experiments conducted on the 3DBiCar dataset demonstrate that MagicCartoon outperforms the state-of-the-art methods. Yu-Pei Song, Yuantong Liu, Xiao Wu 0001, Qi He 0007, Zhaoquan Yuan, Ao Luo |
ACM Multimedia | 4 |
| 2023 | Human-Object-Object Interaction: Towards Human-Centric Complex Interaction DetectionabstractLocalizing and recognizing interactive actions in videos is a pivotal yet intricate task that paves the way towards profound video comprehension. Recent advancements in Human-Object Interaction (HOI) detection, which involve detecting and localizing the interactions between human and object pairs, have undeniably marked significant progress. However, the realm of human-object-object interaction, an essential aspect of real-world industrial applications, remains largely uncharted. In this paper, we introduce a novel task referred to as Human-Object-Object Interaction (HOOI) detection and present a cutting-edge method named the Human-Object-Object Interaction Network (H2O-Net). The proposed H2O-Net is comprised of two principal modules: sequential motion feature extraction and HOOI modeling. The former module delves into the gradually evolving visual characteristics of entities throughout the HOOI process, harnessing spatial-temporal features across multiple fine-grained partitions. Conversely, the latter module aspires to encapsulate HOOI actions through intricate interactions between entities. It commences by capturing and amalgamating two sub-interaction features to extract comprehensive HOOI features, subsequently refining them using the interaction cues embedded within the long-term global context. Furthermore, we contribute to the research community by constructing a new video dataset, dubbed the HOOI dataset. The actions encompassed within this dataset pertain to pivotal operational behaviors in industrial manufacturing, imbuing it with substantial application potential and serving as a valuable addition to the existing repertoire of interaction action detection datasets. Experimental evaluations conducted on the proposed HOOI and widely-used AVA datasets demonstrate that our method outperforms existing state-of-the-art techniques by margins of 6.16 mAP and 1.9 mAP, respectively, thus substantiating its effectiveness. Mingxuan Zhang 0001, Xiao Wu 0001, Zhaoquan Yuan, Qi He 0007, Xiang Huang 0004 |
ACM Multimedia | 4 |
| 2023 | Improving Anomaly Segmentation with Multi-Granularity Cross-Domain AlignmentabstractAnomaly segmentation plays a crucial role in identifying anomalous objects within images, which facilitates the detection of road anomalies for autonomous driving. Although existing methods have shown impressive results in anomaly segmentation using synthetic training data, the domain discrepancies between synthetic training data and real test data are often neglected. To address this issue, Multi-Granularity Cross-Domain Alignment (MGCDA) framework is proposed for anomaly segmentation in complex driving environments. It uniquely combines a new Multi-source Domain Adversarial Training (MDAT) module and a novel Cross-domain Anomaly-aware Contrastive Learning (CACL) method to boost the generality of the model, seamlessly integrating multi-domain data at both scene and sample levels. Multi-source domain adversarial loss and a dynamic label smoothing strategy are integrated into MDAT module to facilitate the acquisition of domain-invariant features at the scene level, through adversarial training across multiple stages. CACL aligns sample-level representations with contrastive loss on cross-domain data, which utilizes an anomaly-aware sampling strategy to efficiently sample hard samples and anchors. The proposed framework has decent properties of parameter-free during the inference stage and is compatible with other anomaly segmentation networks. Experimental conducted on Fishyscapes and RoadAnomaly datasets demonstrate that the proposed framework achieves the state-of-the-art performance. Ji Zhang 0027, Xiao Wu 0001, Zhi-Qi Cheng, Qi He 0007, Wei Li 0110 |
ACM Multimedia | 4 |
| 2022 | Domain-Specific Conditional Jigsaw Adaptation for Enhancing transferability and DiscriminabilityabstractUnsupervised Domain Adaptation (UDA) aims to transfer knowledge from a label-rich source domain to a target domain where the label is unavailable. Existing approaches tend to reduce the distribution discrepancy between the source and target domains or assign the pseudo target labels to implement a self-training strategy. However, the transferability or discriminability lackage of the traditional methods results in the limited ability to generalize on the target domain. To remedy this issue, a novel unsupervised domain adaptation framework called Domain-specific Conditional Jigsaw Adaptation Network (DCJAN) is proposed for UDA, which simultaneously encourages the network to extract transferable and discriminative features. To improve the discriminability, a conditional jigsaw module is presented to reconstruct class-aware features of the original images by reconstructing that of corresponding shuffled images. Moreover, in order to enhance the transferability, a domain-specific jigsaw adaptation is proposed to deal with the domain gaps, which utilizes the prior knowledge of jigsaw puzzles to reduce mismatching. It trains conditional jigsaw modules for each domain and updates the shared feature extractor to make the domain-specific conditional jigsaw modules could perform well not only on the corresponding domain but also on the other domain. A consistent conditioning strategy is proposed to ensure the safe training of conditional jigsaw. Experiments conducted on the widely-used Office-31, Office-Home, VisDA-2017, and DomainNet datasets demonstrate the effectiveness of the proposed approach, which outperforms the state-of-the-art methods. Qi He 0007, Zhaoquan Yuan, Xiao Wu 0001, Jun-Yan He |
ACM Multimedia | 1 |
| 2021 | A novel class restriction loss for unsupervised domain adaptation
Qi He 0007, Qi Dai 0001, Xiao Wu 0001, Jun-Yan He |
Neurocomputing | 1 |