Qichen He

dblp:212/6853 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2025
0000-0003-1973-9529ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Transfer learning and domain adaptation · 30% Generative modeling · 24% Segmentation and scene understanding · 19%
Databases, data mining, and information retrieval
1 paper
Recommender systems · 100%
Computer graphics and multimedia
1 paper
Image and video processing · 100%

Topics — the 14 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation
1.422024
Adversarial Experts Model for Black-box Domain Adaptation · ACM Multimedia 2024
Independent Feature Decomposition and Instance Alignment for Unsupervised Domain Adaptation · IJCAI 2023
Computer vision › Vision and language
vision-language model
1.122025
CLIP-MT: Multi-Modal Knowledge-Driven Adaptive Scale Feature Allocation for Multi-Task Dense Prediction · ACM Multimedia 2025
Adversarial Experts Model for Black-box Domain Adaptation · ACM Multimedia 2024
Computer vision › Segmentation and scene understanding
dense prediction
0.912025
CLIP-MT: Multi-Modal Knowledge-Driven Adaptive Scale Feature Allocation for Multi-Task Dense Prediction · ACM Multimedia 2025
Computer vision › Segmentation and scene understanding › dense prediction
multi-task dense prediction
0.912025
CLIP-MT: Multi-Modal Knowledge-Driven Adaptive Scale Feature Allocation for Multi-Task Dense Prediction · ACM Multimedia 2025
Recommender systems
multimodal recommendation
0.912025
Generating Negative Samples for Multi-Modal Recommendation · ACM Multimedia 2025
Image and video processing › image reconstruction
tomographic reconstruction
0.912025
Discretized Gaussian Representation for Tomographic Reconstruction · ICCV 2025
Machine learning › Transfer learning and domain adaptation › domain adaptation › source-free domain adaptation
black-box domain adaptation
0.812024
Adversarial Experts Model for Black-box Domain Adaptation · ACM Multimedia 2024
Machine learning › Trustworthy machine learning › robustness › learning with noisy labels
pseudo-label denoising
0.812024
Adversarial Experts Model for Black-box Domain Adaptation · ACM Multimedia 2024
Machine learning › Transfer learning and domain adaptation
domain-invariant representation learning
0.712023
Independent Feature Decomposition and Instance Alignment for Unsupervised Domain Adaptation · IJCAI 2023
Machine learning › Representation and self-supervised learning › feature transformation
feature decomposition
0.712023
Independent Feature Decomposition and Instance Alignment for Unsupervised Domain Adaptation · IJCAI 2023
Machine learning › Generative modeling › normalizing flow
injective flow
0.712023
Independent Feature Decomposition and Instance Alignment for Unsupervised Domain Adaptation · IJCAI 2023
Machine learning › Generative modeling
normalizing flow
0.712023
Independent Feature Decomposition and Instance Alignment for Unsupervised Domain Adaptation · IJCAI 2023
Recommender systems
collaborative filtering
0.312025
Generating Negative Samples for Multi-Modal Recommendation · ACM Multimedia 2025
Recommender systems › collaborative filtering
implicit feedback
0.312025
Generating Negative Samples for Multi-Modal Recommendation · ACM Multimedia 2025

Methods — techniques the papers use, named apart from their topics

discretized gaussian representation · 1.7prompt engineering · 0.9multimodal learning · 0.9multimodal large language model · 0.9gating mechanism · 0.9feature fusion · 0.9causal learning · 0.9pseudo-label denoising · 0.8knowledge distillation · 0.8consistency regularization · 0.8adversarial learning · 0.8instance alignment · 0.7feature swapping · 0.7
YearPublicationVenuePosition
2025 Discretized Gaussian Representation for Tomographic Reconstruction
Shaokai Wu, Yapan Guo, Suizhi Huang, Shalayiding Sirejiding, Qichen He, Jing Tong, Yanbiao Ji, Yue Ding 0001, Hongtao Lu 0001
ICCV8
2025 Generating Negative Samples for Multi-Modal Recommendation
abstract
Multi-modal recommender systems (MMRS) have gained significant attention due to their ability to leverage information from various modalities to enhance recommendation quality. However, existing negative sampling techniques often struggle to effectively utilize the multi-modal data, leading to suboptimal performance. In this paper, we identify two key challenges in negative sampling for MMRS: (1) producing cohesive negative samples contrasting with positive samples and (2) maintaining a balanced influence across different modalities. To address these challenges, we propose NegGen, a novel framework that utilizes multi-modal large language models (MLLMs) to generate balanced and contrastive negative samples. We design three different prompt templates to enable NegGen to analyze and manipulate item attributes across multiple modalities, and then generate negative samples that introduce better supervision signals and ensure modality balance. Furthermore, NegGen employs a causal learning module to disentangle the effect of intervened key features and irrelevant item attributes, enabling fine-grained learning of user preferences. Extensive experiments on real-world datasets demonstrate the superior performance of NegGen compared to state-of-the-art methods in both negative sampling and multi-modal recommendation.
Yanbiao Ji, Dan Luo 0004, Chang Liu 0078, Shaokai Wu, Jing Tong, Qichen He, Deyi Ji, Hongtao Lu 0001, Yue Ding 0001
ACM Multimedia6
2025 CLIP-MT: Multi-Modal Knowledge-Driven Adaptive Scale Feature Allocation for Multi-Task Dense Prediction
abstract
Recent advancements in visual multi-task learning (MTL) have sparked significant interest. However, existing dense prediction MTL methods predominantly rely on single-modality image data, limiting their performance due to the absence of complementary knowledge from other modalities. Additionally, different dense tasks exhibit heterogeneous preferences during information decoding, posing a critical challenge in effectively allocating multi-scale encoded features. To address these limitations, we propose CLIP-MT, a Multi-Modal Knowledge-Driven Adaptive Scale Feature Allocation for Multi-Task Dense Prediction. Specifically, to enrich task-shared image features with multi-modal knowledge, we introduce a novel CLIP-Guided Global Feature Enhancer (CGGF), which leverages aligned text-image information to augment object-level representations through a dual-path feature fusion architecture. Furthermore, to tackle the task-specific scale preference problem, we design an Adaptive Scale Selection Gate (ASSG), a learnable gating mechanism that dynamically selects high- or low-scale features based on task-specific demands. Finally, we integrate multi-modal and multi-scale information through a Task-Aware Feature Fusion Module (TAFF). Extensive experiments on the NYUDv2 and PASCAL-Context datasets demonstrate that CLIP-MT achieves state-of-the-art performance, outperforming existing methods across multiple dense prediction tasks.
Shalayiding Sirejiding, Yue Ding 0001, Xinyi Hou, Shaokai Wu, Qichen He, Hongtao Lu 0001
ACM Multimedia6
2024 Adversarial Experts Model for Black-box Domain Adaptation
abstract
Black-box domain adaptation treats the source domain model as a black box. During the transfer process, the only available information about the target domain is the noisy labels output by the black-box model. This poses significant challenges for domain adaptation. Conventional approaches typically tackle the black-box noisy label problem from two aspects: self-knowledge distillation and pseudo-label denoising, both achieving limited performance due to limited knowledge information. To mitigate this issue, we explore the potential of off-the-shelf vision-language (ViL) multimodal models with rich semantic information for black-box domain adaptation by introducing an Adversarial Experts Model (AEM). Specifically, our target domain model is designed as one feature extractor and two classifiers, trained over two stages: In the knowledge transferring stage, with a shared feature extractor, the black-box source model and the ViL model act as two distinct experts for joint knowledge contribution, guiding the learning of one classifier each. While contributing their respective knowledge, the experts are also updated due to their own limitation and bias. In the adversarial alignment stage, to further distill expert knowledge to the target domain model, adversarial learning is conducted between the feature extractor and the two classifiers. A new consistency-max loss function is proposed to measure two classifier consistency and further improve classifier prediction certainty. Extensive experiments on multiple datasets demonstrate the effectiveness of our approach. Code is available at https://github.com/singinger/AEM.
Siying Xiao, Mao Ye 0001, Qichen He, Shuaifeng Li, Song Tang 0001, Xiatian Zhu
ACM Multimedia3
2023 Independent Feature Decomposition and Instance Alignment for Unsupervised Domain Adaptation
abstract
Existing Unsupervised Domain Adaptation (UDA) methods typically attempt to perform knowledge transfer in a domain-invariant space explicitly or implicitly. In practice, however, the obtained features is often mixed with domain-specific information which causes performance degradation. To overcome this fundamental limitation, this article presents a novel independent feature decomposition and instance alignment method (IndUDA in short). Specifically, based on an invertible flow, we project the base features into a decomposed latent space with domain-invariant and domain-specific dimensions. To drive semantic decomposition independently, we then swap the domain-invariant part across source and target domain samples with the same category and require their inverted features are consistent in class-level with the original features. By treating domain-specific information as noise, we replace it by Gaussian noise and further regularize source model training by instance alignment, i.e., requiring the base features close to the corresponding reconstructed features, respectively. Extensive experiment results demonstrate that our method achieves state-of-the-art performance on popular UDA benchmarks. The appendix and code are available at https://github.com/ayombeach/IndUDA.
Qichen He, Siying Xiao, Mao Ye 0001, Xiatian Zhu, Ferrante Neri, Dongde Hou
IJCAI1