Yuhao Tang

dblp:242/0013 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
12since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Round-Efficient Composable Two-Party Quantum Computation
Vipul Goyal, Xiao Liang 0014, Omkant Pandey, Yuhao Tang, Takashi Yamakawa
ASIACRYPT (8)4
2025 Almost-Total Puzzles and Their Applications
Xiao Liang 0014, Omkant Pandey, Yuhao Tang, Takashi Yamakawa
ASIACRYPT (8)3
2025 Sparse Causal Discovery with Generative Intervention for Unsupervised Graph Domain Adaptation
abstract
Unsupervised Graph Domain Adaptation (UGDA) leverages labeled source domain graphs to achieve effective performance in unlabeled target domains despite distribution shifts. However, existing methods often yield suboptimal results due to the entanglement of causal-spurious features and the failure of global alignment strategies. We propose SLOGAN (Sparse Causal Discovery with Generative Intervention), a novel approach that achieves stable graph representation transfer through sparse causal modeling and dynamic intervention mechanisms. Specifically, SLOGAN first constructs a sparse causal graph structure, leveraging mutual information bottleneck constraints to disentangle sparse, stable causal features while compressing domain-dependent spurious correlations through variational inference. To address residual spurious correlations, we innovatively design a generative intervention mechanism that breaks local spurious couplings through cross-domain feature recombination while maintaining causal feature semantic consistency via covariance constraints. Furthermore, to mitigate error accumulation in target domain pseudo-labels, we introduce a category-adaptive dynamic calibration strategy, ensuring stable discriminative learning. Extensive experiments on multiple real-world datasets demonstrate that SLOGAN significantly outperforms existing baselines.
Junyu Luo 0002, Yuhao Tang, Yiwei Fu, Xiao Luo 0001, Zhizhuo Kou, Zhiping Xiao 0001, Wei Ju 0001, Wentao Zhang 0001, Ming Zhang 0004
ICML2
2025 Two-Stage Feature Generation with Transformer and Reinforcement Learning
abstract
Feature generation is a critical step in machine learning, aiming to enhance model performance by capturing complex relationships within the data and generating meaningful new features. Traditional feature generation methods heavily rely on domain expertise and manual intervention, making the process labor-intensive and challenging to adapt to different scenarios. Although automated feature generation techniques address these issues to some extent, they often face challenges such as feature redundancy, inefficiency in feature space exploration, and limited adaptability to diverse datasets and tasks. To address these problems, we propose a Two-Stage Feature Generation (TSFG) framework, which integrates a Transformer-based encoder-decoder architecture with Proximal Policy Optimization (PPO). The encoder-decoder model in TSFG leverages the Transformer’s self-attention mechanism to efficiently represent and transform features, capturing complex dependencies within the data. PPO further enhances TSFG by dynamically adjusting the feature generation strategy based on task-specific feedback, optimizing the process for improved performance and adaptability. TSFG dynamically generates high-quality feature sets, significantly improving the predictive performance of machine learning models. Experimental results demonstrate that TSFG outperforms existing state-of-the-art methods in terms of feature quality and adaptability.
Wanfu Gao, Zengyao Man, Zebin He, Yuhao Tang, Kunpeng Liu 0001
IJCAI4
2024 Multi- View Teacher with Curriculum Data Fusion for Robust Unsupervised Domain Adaptation
abstract
Graph Neural Networks (GNNs) have emerged as an effective tool for graph classification, yet their reliance on extensive labeled data poses a significant challenge, especially when such labels are scarce. To address this challenge, this paper presents a novel framework, denoted as Multi-View Teacher with Curriculum Data Fusion (MTDF). MTDF achieves robust unsupervised domain adaptation in both the model and data perspectives. On the one hand, MTDF utilizes a multi-teacher framework with diverse update strategies for robust adaptation. Moreover, it employs a complementary perspective consistency model from local implicit representation and global explicit graph structure. On the other hand, MTDF generates source-mimicry data at the target domain to serve as a bridge to overcome the challenge of domain shift. MTDF achieves stable unsupervised domain adaptation through bi-directional processes from the perspective of both the model and the data. We have conducted comprehensive experimental evaluations across multiple real-world datasets with a range of baseline methods to demonstrate the superior performance of our proposed method.
Yuhao Tang, Junyu Luo 0002, Ling Yang 0006, Xiao Luo 0001, Wentao Zhang 0001, Bin Cui 0001
ICDE1
2024 Work like a doctor: Unifying scan localizer and dynamic generator for automated computed tomography report generation
Yuhao Tang, Haichen Yang, Liyan Zhang 0001
Expert Syst. Appl.1
2024 NDAM-YOLOseg: a real-time instance segmentation model based on multi-head attention mechanism
Chengang Dong, Yuhao Tang, Liyan Zhang 0001
Multim. Syst.2
2024 Higher efficient YOLOv7: a one-stage method for non-salient object detection
Chengang Dong, Yuhao Tang, Liyan Zhang 0001
Multim. Tools Appl.2
2024 Fashion item captioning via grid-relation self-attention and gated-enhanced decoder
Yuhao Tang, Liyan Zhang 0001
Multim. Tools Appl.1
2023 Improving fashion captioning via attribute-based alignment and multi-level language model
Yuhao Tang, Liyan Zhang 0001
Appl. Intell.1
2023 Polyp-Mixer: An Efficient Context-Aware MLP-Based Paradigm for Polyp Segmentation
abstract
Precise and efficient polyp segmentation plays a crucial role in colonoscopy, which is important for the prevention of colorectal cancer. Despite CNN-based methods have achieved great progress in the polyp segmentation task, they are incapable of modeling long-range dependencies. Transformer-based models utilize self-attention mechanism to overcome this problem while suffering from heavy computing cost. Benefiting from simple structures, MLP-based models seem to be an alternative. However, they struggle with dealing with flexible input scales and modeling long-term dependencies. Both of these two factors are important for image segmentation, which could explain why the MLP architecture performs poorly compared to the Transformer. To remedy this issue, we propose a novel Polyp-Mixer, which utilizes MLP-based structures in both encoder and decoder. In particular, we use CycleMLP as the encoder to overcome the fixed input scale issue. Besides, we propose a Multi-head Mixer by converting the current CycleMLP into a Multi-head fashion, allowing our model to explore rich context information from various subspaces. In addition, we build a powerful Contextual Bridger Module between the encoder and decoder, which can capture semantics from larger receptive fields and combine them with various decoder layers. Experiments demonstrate the proposed method with fewer parameters ($\sim 16\text{M}$) achieves SOTA on 4 public benchmarks. Our code will be released athttps://github.com/shijinghuihub/Polyp-Mixer
Jing-Hui Shi, Qing Zhang 0004, Yuhao Tang, Zhong-Qun Zhang
IEEE Trans. Circuits Syst. Video Technol.3
2023 Describe Fashion Products via Local Sparse Self-Attention Mechanism and Attribute-Based Re-Sampling Strategy
abstract
In this paper, we convert conventional image captioning task into a new paradigm in the fashion domain, fashion item captioning. This task requires the ability of associating a group of item angles and generating a longer and more fine-grained description. To link several images in a more efficient way, we propose a local sparse self-attention mechanism(LSAM), which only enables one region to interact with those in the adjacent area and allows a lighter architecture with fewer layers. To capture more subtle details of the item, an attribute-based re-sampling strategy(ARS) is introduced to enhance the learning on those low-frequency but content-related attribute words. Besides, existing fashion datasets are limited in the quality of annotations and single fashion style. To bridge this gap, we propose a novel dataset for fashion item captioning, termed Fashion Item Captioning Dataset(FICD). FICD provides a meaningful complement to the existing fashion datasets, which comprises 294K images, 62K real-world product descriptions with diverse linguistic styles, rich attributes and categories. It is worth mentioning that we also annotate an attribute-level sentence for each item, which enables the FICD to be used not only for fashion item captioning but also for other fashion-related tasks. Moreover, the complex structure is even more restrictive for application to real-world scenarios. Without bells and whistles, our framework is simply designed with an end-to-end manner. Extensive experiments demonstrate the effectiveness of our LSAM-ARS. More remarkably, LSAM-ARS achieves state-of-the-art performance on FACAD and FICD datasets, with the CIDEr-D score being increased from 65.4% to 81.8%, 69.8% to 77.4%, respectively.
Yuhao Tang, Liyan Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.1