VLDB 2026 Research / reviewers in the wild / expert
Tianyu Guo 0001
dblp:218/7273-1
· DBLP profile ↗
34ranked-venue papers
9as first author
27since 2021 · last 2025
0000-0001-5703-7064ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 9 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 4 first-author · 13 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | L-Man: A Large Multi-modal Model Unifying Human-centric TasksabstractLarge language models (LLMs) have recently shown notable progress in unifying various visual tasks with an open-ended form. However, when transferred to human-centric tasks, despite their remarkable multi-modal understanding ability in general domains, they lack further human-related domain knowledge and show unsatisfactory performance. Meanwhile, current human-centric unified models are mostly restricted to a pre-defined form and lack open-ended task capability. Therefore, it is necessary to propose a large multi-modal model which utilizes LLMs to unify various human-centric tasks. We forge ahead along this path from the aspects of dataset and model. Specifically, we first construct a large-scale language-image instruction-following dataset named HumanIns based on existing 20 open datasets from 6 diverse downstream tasks, which provides sufficient and diverse data to implement multi-modal training. Then, a model named L-Man including a query adapter is designed to extract the multi-grained semantics of image and align the cross-modal information between image and text. In practice, we introduce a two-stage training strategy, where the first stage extracts generic text-relevant visual information, and the second stage maps the visual features to the embedding space of the LLM. By tuning on HumanIns, our model shows significant superiority on human-centric tasks compared with existing large multi-modal models, and also achieves even better results on downstream datasets compared with respective task-specific models. Jialong Zuo, Tianyu Guo 0001, Huaxin Zhang, Jiahao Hong, Nong Sang, Changxin Gao, Kai Han 0002 |
AAAI | 3 |
| 2025 | MoRAgent: Parameter Efficient Agent Tuning with Mixture-of-RolesabstractDespite recent advancements of fine-tuning large language models (LLMs) to facilitate agent tasks, parameter-efficient fine-tuning (PEFT) methodologies for agent remain largely unexplored. In this paper, we introduce three key strategies for PEFT in agent tasks: 1) Inspired by the increasingly dominant \textit{Reason+Action} paradigm, we first decompose the capabilities necessary for the agent tasks into three distinct roles: reasoner, executor, and summarizer. The reasoner is responsible for comprehending the user's query and determining the next role based on the execution trajectory. The executor is tasked with identifying the appropriate functions and parameters to invoke. The summarizer conveys the distilled information from conversations back to the user. 2) We then propose the Mixture-of-Roles (MoR) framework, which comprises three specialized Low-Rank Adaptation (LoRA) groups, each designated to fulfill a distinct role. By focusing on their respective specialized capabilities and engaging in collaborative interactions, these LoRAs collectively accomplish the agent task. 3) To effectively fine-tune the framework, we develop a multi-role data generation pipeline based on publicly available datasets, incorporating role-specific content completion and reliability verification.
We conduct extensive experiments and thorough ablation studies on various LLMs and agent benchmarks, demonstrating the effectiveness of the proposed method. This project is publicly available at https://mor-agent.github.io Binwei Yan, Tianyu Guo 0001, Zheyuan Bai, Mengyu Zheng, Hanting Chen |
ICML | 3 |
| 2025 | DenseSSM: State Space Models with Dense Hidden Connection for Efficient Large Language ModelsabstractWei He, Kai Han, Yehui Tang, Chengcheng Wang, Yujie Yang, Tianyu Guo, Yunhe Wang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Wei He 0001, Kai Han 0002, Yehui Tang 0001, Tianyu Guo 0001, Yunhe Wang 0001 |
NAACL (Long Papers) | 6 |
| 2025 | CFinBench: A Comprehensive Chinese Financial Benchmark for Large Language ModelsabstractYing Nie, Binwei Yan, Tianyu Guo, Hao Liu, Haoyu Wang, Wei He, Binfan Zheng, Weihao Wang, Qiang Li, Weijian Sun, Yunhe Wang, Dacheng Tao. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Binwei Yan, Tianyu Guo 0001, Wei He 0001, Binfan Zheng, Qiang Li 0024, Weijian Sun, Yunhe Wang 0001, Dacheng Tao |
NAACL (Long Papers) | 3 |
| 2025 | BRACE: A Benchmark for Robust Audio Caption Quality EvaluationabstractAutomatic audio captioning is essential for audio understanding, enabling applications such as accessibility and content indexing. However, evaluating the quality of audio captions remains a major challenge, especially in reference-free settings where high-quality ground-truth captions are unavailable. While CLAPScore is currently the most widely used reference-free Audio Caption Evaluation Metric(ACEM), its robustness under diverse conditions has not been systematically validated. To address this gap, we introduce BRACE, a new benchmark designed to evaluate audio caption alignment quality in a reference-free setting. BRACE is primarily designed for assessing ACEMs, and can also be extended to measure the modality alignment abilities of Large Audio Language Model(LALM). BRACE consists of two sub-benchmarks: BRACE-Main for fine-grained caption comparison and BRACE-Hallucination for detecting subtle hallucinated content. We construct these datasets through high-quality filtering, LLM-based corruption, and human annotation. Given the widespread adoption of CLAPScore as a reference-free ACEM and the increasing application of LALMs in audio-language tasks, we evaluate both approaches using the BRACE benchmark, testing CLAPScore across various CLAP model variants and assessing multiple LALMs. Notably, even the best-performing CLAP-based ACEM achieves only a 70.01 F1-score on the BRACE-Main benchmark, while the best LALM reaches just 63.19. By revealing the limitations of CLAP models and LALMs, our BRACE benchmark offers valuable insights into the direction of future research. Our evaluation code and benchmark dataset are released in https://github.com/HychTus/BRACEEvaluation and https://huggingface.co/datasets/gtysssp/audiobenchmarks. Tianyu Guo 0001, Hao Liang 0017, Meiyi Qiang, Bohan Zeng, Linzhuang Sun, Bin Cui 0001, Wentao Zhang 0001 |
NeurIPS | 1 |
| 2025 | On Positive-Unlabeled Classification From Corrupted Data in GANsabstractThis paper defines a positive and unlabeled classification problem for standard GANs, which then leads to a novel technique to stabilize the training of the discriminator in GANs and deal with corrupted data. Traditionally, real data are taken as positive while generated data are negative. This positive-negative classification criterion was kept fixed all through the learning process of the discriminator without considering the gradually improved quality of generated data, even if they could be more realistic than real data at times. In contrast, it is more reasonable to treat the generated data as unlabeled, which could be positive or negative according to their quality. The discriminator is thus a classifier for this positive and unlabeled classification problem, and we derive a new Positive-Unlabeled GAN (PUGAN). We theoretically discuss the global optimality the proposed model will achieve and the equivalent optimization goal. Empirically, we find that PUGAN can achieve comparable or even better performance than those sophisticated discriminator stabilization methods. Considering the potential corrupted data problem in real-world scenarios, we further extend our approach to PUGAN-C, which treats real data as unlabeled that accounts for both clean and corrupted instances, and generated data as positive. The samples from generator could be closer to those corrupted data within unlabeled data at first, but within the framework of adversarial training, the generator will be optimized to cheat the discriminator and produce samples that are similar to those clean data. Experimental results on image generation from several corrupted datasets demonstrate the effectiveness and generalization of PUGAN-C. Yunke Wang, Chang Xu 0002, Tianyu Guo 0001, Bo Du 0001, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | GraphMLP: A graph MLP-like architecture for 3D human pose estimation
Wenhao Li 0002, Mengyuan Liu 0001, Hong Liu 0008, Tianyu Guo 0001, Ti Wang, Hao Tang 0005, Nicu Sebe |
Pattern Recognit. | 4 |
| 2024 | UFineBench: Towards Text-based Person Retrieval with Ultra-fine GranularityabstractExisting text-based person retrieval datasets often have relatively coarse-grained text annotations. This hinders the model to comprehend the fine-grained semantics of query texts in real scenarios. To address this problem, we con-tribute a new benchmark named UFineBench for text-based person retrieval with ultra-fine granularity. Firstly, we construct a new dataset named UFine6926. We collect a large number of person images and manually annotate each image with two detailed textual descriptions, averaging 80.8 words each. The average word count is three to four times that of the previous datasets. In addition of standard in-domain evaluation, we also propose a spe-cial evaluation paradigm more representative of real sce-narios. It contains a new evaluation set with cross domains, cross textual granularity and cross textual styles, named UFine3C, and a new evaluation metric for accurately mea-suring retrieval ability, named mean Similarity Distribution (mSD). Moreover, we propose CFAM, a more efficient al-gorithm especially designed for text-based person retrieval with ultra fine-grained texts. It achieves fine granularity mining by adopting a shared cross-modal granularity de-coder and hard negative match mechanism. With standard in-domain evaluation, CFAM establishes competitive performance across various datasets, espe-cially on our ultra fine-grained UFine6926. Furthermore, by evaluating on UFine3C, we demonstrate that training on our UFine6926 significantly improves generalization to real scenarios compared with other coarse-grained datasets. The dataset and code will be made publicly available at https://github.com/Zplusdragon/UFineBench. Jialong Zuo, Hanyu Zhou, Feng Zhang 0039, Tianyu Guo 0001, Nong Sang, Yunhe Wang 0001, Changxin Gao |
CVPR | 5 |
| 2024 | A Robust Audio Deepfake Detection System via Multi-View FeatureabstractWith the advancement of generative modeling techniques, synthetic human speech becomes increasingly indistinguishable from real, and tricky challenges are elicited for the audio deepfake detection (ADD) system. In this paper, we exploit audio features to improve the generalizability of ADD systems. Investigation of the ADD task performance is conducted over a broad range of audio features, including various handcrafted features and learning-based features. Experiments show that learning-based audio features pretrained on a large amount of data generalize better than hand-crafted features on out-of-domain scenarios. Subsequently, we further improve the generalizability of the ADD system using proposed multi-feature approaches to incorporate complimentary information from features of different views. The model trained on ASV2019 data achieves an equal error rate of 24.27% on the In-the-Wild dataset. The code will be released as soon1. Haochen Qin, Tianyu Guo 0001, Kai Han 0002, Yunhe Wang 0001 |
ICASSP | 5 |
| 2024 | Enhancing Large Language Models through Adaptive TokenizersabstractTokenizers serve as crucial interfaces between models and linguistic data, substantially influencing the efficacy and precision of large language models (LLMs). Traditional tokenization methods often rely on static frequency-based statistics and are not inherently synchronized with LLM architectures, which may limit model performance. In this study, we propose a simple but effective method to learn tokenizers specifically engineered for seamless integration with LLMs. Initiating with a broad initial vocabulary, we refine our tokenizer by monitoring changes in the model’s perplexity during training, allowing for the selection of a tokenizer that is closely aligned with the model’s evolving dynamics. Through iterative refinement, we develop an optimized tokenizer. Our empirical evaluations demonstrate that this adaptive approach significantly enhances accuracy compared to conventional methods, maintaining comparable vocabulary sizes and affirming its potential to improve LLM functionality. Mengyu Zheng, Hanting Chen, Tianyu Guo 0001, Chong Zhu, Binfan Zheng, Chang Xu 0002, Yunhe Wang 0001 |
NeurIPS | 3 |
| 2024 | Cross-video Identity Correlating for Person Re-identification Pre-trainingabstractRecent researches have proven that pre-training on large-scale person images extracted from internet videos is an effective way in learning better representations for person re-identification. However, these researches are mostly confined to pre-training at the instance-level or single-video tracklet-level. They ignore the identity-invariance in images of the same person across different videos, which is a key focus in person re-identification. To address this issue, we propose a Cross-video Identity-cOrrelating pre-traiNing (CION) framework. Defining a noise concept that comprehensively considers both intra-identity consistency and inter-identity discrimination, CION seeks the identity correlation from cross-video images by modeling it as a progressive multi-level denoising problem. Furthermore, an identity-guided self-distillation loss is proposed to implement better large-scale pre-training by mining the identity-invariance within person images. We conduct extensive experiments to verify the superiority of our CION in terms of efficiency and performance. CION achieves significantly leading performance with even fewer training samples. For example, compared with the previous state-of-the-art ISR, CION with the same ResNet50-IBN achieves higher mAP of 93.3% and 74.3% on Market1501 and MSMT17, while only utilizing 8% training samples. Finally, with CION demonstrating superior model-agnostic ability, we contribute a model zoo named ReIDZoo to meet diverse research and application needs in this field. It contains a series of CION pre-trained models with spanning structures and parameters, totaling 32 models with 10 different structures, including GhostNet, ConvNext, RepViT, FastViT and so on. The code and models will be open-sourced. Jialong Zuo, Hanyu Zhou, Huaxin Zhang, Haoyu Wang 0003, Tianyu Guo 0001, Nong Sang, Changxin Gao |
NeurIPS | 6 |
| 2024 | Improving self-supervised action recognition from extremely augmented skeleton sequences
Tianyu Guo 0001, Mengyuan Liu 0001, Hong Liu 0008, Wenhao Li 0002 |
Pattern Recognit. | 1 |
| 2024 | Cross-Model Cross-Stream Learning for Self-Supervised Human Action RecognitionabstractConsidering the instance-level discriminative ability, contrastive learning methods, including MoCo and SimCLR, have been adapted from the original image representation learning task to solve the self-supervised skeleton-based action recognition task. These methods usually use multiple data streams (i.e., joint, motion, and bone) for ensemble learning, meanwhile, how to construct a discriminative feature space within a single stream and effectively aggregate the information from multiple streams remains an open problem. To this end, this article first applies a new contrastive learning method called bootstrap your own latent (BYOL) to learn from skeleton data, and then formulate SkeletonBYOL as a simple yet effective baseline for self-supervised skeleton-based action recognition. Inspired by SkeletonBYOL, this article further presents a cross-model and cross-stream (CMCS) framework. This framework combines cross-model adversarial learning (CMAL) and cross-stream collaborative learning (CSCL). Specifically, CMAL learns single-stream representation by cross-model adversarial loss to obtain more discriminative features. To aggregate and interact with multistream information, CSCL is designed by generating similarity pseudolabel of ensemble learning as supervision and guiding feature generation for individual streams. Extensive experiments on three datasets verify the complementary properties between CMAL and CSCL and also verify that the proposed method can achieve better results than state-of-the-art methods using various evaluation protocols. Mengyuan Liu 0001, Hong Liu 0008, Tianyu Guo 0001 |
IEEE Trans. Hum. Mach. Syst. | 3 |
| 2024 | Feature Completion Transformer for Occluded Person Re-IdentificationabstractOccluded person re-identification is a challenging problem due to the destruction of occluders in different camera views. Most existing paradigms focus on visible human body parts through some external models to reduce noise interference. However, the feature misalignment problem caused by discarded occlusions negatively affects the performance of the network. Different from most previous works that discard the occluded regions, we present Feature Completion Transformer (FCFormer) that reduces noise interference and complements missing features in occluded parts. Specifically, Occlusion Instance Augmentation is proposed to simulate real and diverse occlusion situations on the holistic image, which enlarges the occlusion samples in the training set and forms aligned occluded-holistic pairs. To reduce the interference of noise, a two-stream architecture is proposed to learn pairwise discriminative features from aligned image pairs, while obtaining self-aligned occluded-holistic feature level sample-label pairs without additional auxiliary models. To complement the features of occluded regions, a Feature Completion Decoder is designed to aggregate possible information from self-generated occluded features in a self-supervised manner. Further, in order to correlate the completion features with identity information, Feature Completion Consistency loss is introduced to enforce the distribution of the generated completion features to be consistent with the real holistic feature distribution. In addition, we propose the Cross Hard Triplet loss to further bridge the gap between completion features and extracting features under the same ID. Extensive experiments over five challenging datasets demonstrate that the proposed FCFormer achieves superior performance and outperforms the state-of-theart methods by significant margins on Occluded-Duke dataset. Mengyuan Liu 0001, Hong Liu 0008, Wenhao Li 0002, Miaoju Ban, Tianyu Guo 0001, Yidi Li 0001 |
IEEE Trans. Multim. | 6 |
| 2023 | FSAR: Federated Skeleton-based Action Recognition with Adaptive Topology Structure and Knowledge DistillationabstractExisting skeleton-based action recognition methods typically follow a centralized learning paradigm, which can pose privacy concerns when exposing human-related videos. Federated Learning (FL) has attracted much attention due to its outstanding advantages in privacy-preserving. However, directly applying FL approaches to skeleton videos suffers from unstable training. In this paper, we investigate and discover that the heterogeneous human topology graph structure is the crucial factor hindering training stability. To address this limitation, we pioneer a novel Federated Skeleton-based Action Recognition (FSAR) paradigm, which enables the construction of a globally generalized model without accessing local sensitive data. Specifically, we introduce an Adaptive Topology Structure (ATS), separating generalization and personalization by learning a domain-invariant topology shared across clients and a domain-specific topology decoupled from global model aggregation. Furthermore, we explore Multi-grain Knowledge Distillation (MKD) to mitigate the discrepancy between clients and server caused by distinct updating patterns through aligning shallow block-wise motion features. Extensive experiments on multiple datasets demonstrate that FSAR outperforms state-of-the-art FL-based methods while inherently protecting privacy. Jingwen Guo, Hong Liu 0008, Shitong Sun, Tianyu Guo 0001, Min Zhang 0005, Chenyang Si |
ICCV | 4 |
| 2023 | Self-Supervised 3D Skeleton Representation Learning with Active Sampling and Adaptive Relabeling for Action RecognitionabstractSelf-supervised 3D skeleton representation learning has recently shown great potential for action recognition via contrastive learning. However, existing methods suffer from limited learning efficiency and the unreliability of representations, which is not conducive to action recognition. To this end, we propose an Active Sampling and Adaptive Relabeling (ASAR) contrastive learning method to achieve efficient and reliable learning of 3D skeleton representations. Specifically, the active sampling strategy is used to build a dictionary with informative samples for efficient representation learning. Additionally, the adaptive relabeling strategy is proposed to automatically modify the confidence scores of the extra positive samples and alleviate the unreliability of representations. Extensive experiments on NTU-60, NTU-120, and PKU-MMD datasets demonstrate the superiority of our approach. Hong Liu 0008, Tianyu Guo 0001, Jingwen Guo, Ti Wang, Yidi Li 0001 |
ICIP | 3 |
| 2023 | When Visual Prompt Tuning Meets Source-Free Domain Adaptive Semantic SegmentationabstractSource-free domain adaptive semantic segmentation aims to adapt a pre-trained source model to the unlabeled target domain
without accessing the private source data. Previous methods usually fine-tune the entire network, which suffers from expensive parameter tuning. To avoid this problem, we propose to utilize visual prompt tuning for parameter-efficient adaptation. However, the existing visual prompt tuning methods are unsuitable for source-free domain adaptive semantic segmentation due to the following two reasons: (1) Commonly used visual prompts like input tokens or pixel-level perturbations cannot reliably learn informative knowledge beneficial for semantic segmentation. (2) Visual prompts require sufficient labeled data to fill the gap between the pre-trained model and downstream tasks. To alleviate these problems, we propose a universal unsupervised visual prompt tuning (Uni-UVPT) framework, which is applicable to various transformer-based backbones. Specifically, we first divide the source pre-trained backbone with frozen parameters into multiple stages, and propose a lightweight prompt adapter for progressively encoding informative knowledge into prompts and enhancing the generalization of target features between adjacent backbone stages. Cooperatively, a novel adaptive pseudo-label correction strategy with a multiscale consistency loss is designed to alleviate the negative effect of target samples with noisy pseudo labels and raise the capacity of visual prompts to spatial perturbations. Extensive experiments demonstrate that Uni-UVPT achieves state-of-the-art performance on GTA5 $\to$ Cityscapes and SYNTHIA $\to$ Cityscapes tasks and can serve as a universal and parameter-efficient framework for large-model unsupervised knowledge transfer. Code will be available at https://gitee.com/mindspore/models/tree/master/research/cv/uni-uvpt and https://github.com/huawei-noah/noah-research/tree/master/uni-uvpt. Xinhong Ma, Tianyu Guo 0001, Yunhe Wang 0001 |
NeurIPS | 4 |
| 2023 | Towards Higher Ranks via Adversarial Weight PruningabstractConvolutional Neural Networks (CNNs) are hard to deploy on edge devices due to its high computation and storage complexities. As a common practice for model compression, network pruning consists of two major categories: unstructured and structured pruning, where unstructured pruning constantly performs better. However, unstructured pruning presents a structured pattern at high pruning rates, which limits its performance. To this end, we propose a Rank-based PruninG (RPG) method to maintain the ranks of sparse weights in an adversarial manner. In each step, we minimize the low-rank approximation error for the weight matrices using singular value decomposition, and maximize their distance by pushing the weight matrices away from its low rank approximation. This rank-based optimization objective guides sparse weights towards a high-rank topology. The proposed method is conducted in a gradual pruning fashion to stabilize the change of rank during training. Experimental results on various datasets and different tasks demonstrate the effectiveness of our algorithm in high sparsity. The proposed RPG outperforms the state-of-the-art performance by 1.13\% top-1 accuracy on ImageNet in ResNet-50 with 98\% sparsity. The codes are available at https://github.com/huawei-noah/Efficient-Computing/tree/master/Pruning/RPG and https://gitee.com/mindspore/models/tree/master/research/cv/RPG. Yuchuan Tian, Hanting Chen, Tianyu Guo 0001, Chao Xu 0006, Yunhe Wang 0001 |
NeurIPS | 3 |
| 2023 | PUe: Biased Positive-Unlabeled Learning Enhancement by Causal InferenceabstractPositive-Unlabeled (PU) learning aims to achieve high-accuracy binary classification with
limited labeled positive examples and numerous unlabeled ones. Existing cost-sensitive-based
methods often rely on strong assumptions that examples with an observed positive label were
selected entirely at random. In fact, the uneven distribution of labels is prevalent in
real-world PU problems, indicating that most actual positive and unlabeled data are subject
to selection bias. In this paper, we propose a PU learning enhancement (PUe) algorithm
based on causal inference theory, which employs normalized propensity scores and normalized
inverse probability weighting (NIPW) techniques to reconstruct the loss function, thus
obtaining a consistent, unbiased estimate of the classifier and enhancing the model's
performance. Moreover, we investigate and propose a method for estimating propensity scores
in deep learning using regularization techniques when the labeling mechanism is unknown.
Our experiments on three benchmark datasets demonstrate the proposed PUe algorithm significantly
improves the accuracy of classifiers on non-uniform label distribution datasets compared to
advanced cost-sensitive PU methods. Codes are available at https://github.com/huawei-noah/Noah-research/tree/master/PUe and https://gitee.com/mindspore/models/tree/master/research/cv/PUe. Xutao Wang, Hanting Chen, Tianyu Guo 0001, Yunhe Wang 0001 |
NeurIPS | 3 |
| 2022 | Contrastive Learning from Extremely Augmented Skeleton Sequences for Self-Supervised Action RecognitionabstractIn recent years, self-supervised representation learning for skeleton-based action recognition has been developed with the advance of contrastive learning methods. The existing contrastive learning methods use normal augmentations to construct similar positive samples, which limits the ability to explore novel movement patterns. In this paper, to make better use of the movement patterns introduced by extreme augmentations, a Contrastive Learning framework utilizing Abundant Information Mining for self-supervised action Representation (AimCLR) is proposed. First, the extreme augmentations and the Energy-based Attention-guided Drop Module (EADM) are proposed to obtain diverse positive samples, which bring novel movement patterns to improve the universality of the learned representations. Second, since directly using extreme augmentations may not be able to boost the performance due to the drastic changes in original identity, the Dual Distributional Divergence Minimization Loss (D3M Loss) is proposed to minimize the distribution divergence in a more gentle way. Third, the Nearest Neighbors Mining (NNM) is proposed to further expand positive samples to make the abundant information mining process more reasonable. Exhaustive experiments on NTU RGB+D 60, PKU-MMD, NTU RGB+D 120 datasets have verified that our AimCLR can significantly perform favorably against state-of-the-art methods under a variety of evaluation protocols with observed higher quality action representations. Our code is available at https://github.com/Levigty/AimCLR. Tianyu Guo 0001, Hong Liu 0008, Mengyuan Liu 0001, Runwei Ding |
AAAI | 1 |
| 2022 | Pose-Guided Feature Disentangling for Occluded Person Re-identification Based on TransformerabstractOccluded person re-identification is a challenging task as human body parts could be occluded by some obstacles (e.g. trees, cars, and pedestrians) in certain scenes. Some existing pose-guided methods solve this problem by aligning body parts according to graph matching, but these graph-based methods are not intuitive and complicated. Therefore, we propose a transformer-based Pose-guided Feature Disentangling (PFD) method by utilizing pose information to clearly disentangle semantic components (e.g. human body or joint parts) and selectively match non-occluded parts correspondingly. First, Vision Transformer (ViT) is used to extract the patch features with its strong capability. Second, to preliminarily disentangle the pose information from patch information, the matching and distributing mechanism is leveraged in Pose-guided Feature Aggregation (PFA) module. Third, a set of learnable semantic views are introduced in transformer decoder to implicitly enhance the disentangled body part features. However, those semantic views are not guaranteed to be related to the body without additional supervision. Therefore, Pose-View Matching (PVM) module is proposed to explicitly match visible body parts and automatically separate occlusion features. Fourth, to better prevent the interference of occlusions, we design a Pose-guided Push Loss to emphasize the features of visible body parts. Extensive experiments over five challenging datasets for two tasks (occluded and holistic Re-ID) demonstrate that our proposed PFD is superior promising, which performs favorably against state-of-the-art methods. Code is available at https://github.com/WangTaoAs/PFD_Net Hong Liu 0008, Pinhao Song, Tianyu Guo 0001, Wei Shi 0009 |
AAAI | 4 |
| 2022 | PDD-Net: A Precise Defect Detection Network Based on Point Set RepresentationabstractDefect detection has been widely studied in computer vision and used in industrial production. However, most existing methods for defect detection mainly suffer three drawbacks: i) Low-contrast problem between defects and background. ii) Large scale changes in defects size. iii) Extreme imbalance problem between defects and background classes during training. To address these issues, we propose a novel anchor-free defect detection network named PDD-Net. Specifically, a global-context FPN (GC-FPN) is designed to capture long-range dependency between defects and background. Simultaneously, to enhance feature extraction of defects at different scales, a receptive field pyramid block (RFPB) is proposed to provide various receptive field sizes. Furthermore, an equipped adaptive positive and negative samples allocation (APNSA) mechanism is built with statistical characteristics of defects, thus can select training samples automatically. We conduct experiments on MPSD dataset, DAGM2007 dataset, and NEU-DET dataset. Extensive experimental results on the three challenging datasets show that our PDD-Net achieves superior detection accuracy over the state-of-the-art methods. Miaoju Ban, Runwei Ding, Jian Zhang 0117, Tianyu Guo 0001 |
ICASSP | 4 |
| 2022 | FDSNeT: An Accurate Real-Time Surface Defect Segmentation NetworkabstractSurface defect detection is a common task for industrial quality control, which increasingly requires accuracy and real-time ability. However, the current segmentation networks are not effective in dealing with defect boundary details, local similarity of different defects and low contrast between defect and background. To this end, we propose a real-time surface defect segmentation network (FDSNet) based on two-branch architecture, in which two corresponding auxiliary tasks are introduced to encode more boundary details and semantic context. To handle the local similarity problem of different surface defects, we propose a Global Context Upsampling (GCU) module by capturing long-range context from multi-scales. Moreover, we present a representative Mobile phone screen Surface Defect (MSD) segmentation dataset to alleviate the lack of dataset in this field. Experiments on NEU-Seg, Magnetic-tile-defect-datasets and MSD dataset show that the proposed FDSNet achieves promising trade-off between accuracy and inference speed. The dataset and code are available at https://github.com/jianzhang96/fdsnet. Jian Zhang 0117, Runwei Ding, Miaoju Ban, Tianyu Guo 0001 |
ICASSP | 4 |
| 2022 | Optimizing Latent Distributions for Non-Adversarial Generative NetworksabstractThe generator in generative adversarial networks (GANs) is driven by a discriminator to produce high-quality images through an adversarial game. At the same time, the difficulty of reaching a stable generator has been increased. This paper focuses on non-adversarial generative networks that are trained in a plain manner without adversarial loss. The given limited number of real images could be insufficient to fully represent the real data distribution. We therefore investigate a set of distributions in a Wasserstein ball centred on the distribution induced by the training data and propose to optimize the generator over this Wasserstein ball. We theoretically discuss the solvability of the newly defined objective function and develop a tractable reformulation to learn the generator. The connections and differences between the proposed non-adversarial generative networks and GANs are analyzed. Experimental results on real-world datasets demonstrate that the proposed algorithm can effectively learn image generators in a non-adversarial approach, and the generated images are of comparable quality with those from GANs. Tianyu Guo 0001, Chang Xu 0002, Boxin Shi, Chao Xu 0006, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Adversarial Robustness through Disentangled RepresentationsabstractDespite the remarkable empirical performance of deep learning models, their vulnerability to adversarial examples has been revealed in many studies. They are prone to make a susceptible prediction to the input with imperceptible adversarial perturbation. Although recent works have remarkably improved the model's robustness under the adversarial training strategy, an evident gap between the natural accuracy and adversarial robustness inevitably exists. In order to mitigate this problem, in this paper, we assume that the robust and non-robust representations are two basic ingredients entangled in the integral representation. For achieving adversarial robustness, the robust representations of natural and adversarial examples should be disentangled from the non-robust part and the alignment of the robust representations can bridge the gap between accuracy and robustness. Inspired by this motivation, we propose a novel defense method called Deep Robust Representation Disentanglement Network (DRRDN). Specifically, DRRDN employs a disentangler to extract and align the robust representations from both adversarial and natural examples. Theoretical analysis guarantees the mitigation of the trade-off between robustness and accuracy with good disentanglement and alignment performance. Experimental results on benchmark datasets finally demonstrate the empirical superiority of our method. Shuo Yang 0006, Tianyu Guo 0001, Yunhe Wang 0001, Chang Xu 0002 |
AAAI | 2 |
| 2021 | Pre-Trained Image Processing TransformerabstractAs the computing power of modern hardware is increasing strongly, pre-trained deep learning models (e.g., BERT, GPT-3) learned on large-scale datasets have shown their effectiveness over conventional methods. The big progress is mainly contributed to the representation ability of transformer and its variant architectures. In this paper, we study the low-level computer vision task (e.g., denoising, super-resolution and deraining) and develop a new pre-trained model, namely, image processing transformer (IPT). To maximally excavate the capability of transformer, we present to utilize the well-known ImageNet benchmark for generating a large amount of corrupted image pairs. The IPT model is trained on these images with multi-heads and multi-tails. In addition, the contrastive learning is introduced for well adapting to different image processing tasks. The pre-trained model can therefore efficiently employed on desired task after fine-tuning. With only one pre-trained model, IPT outperforms the current state-of-the-art methods on various low-level benchmarks. Code is available at https://github.com/huawei-noah/Pretrained-IPT and https://gitee.com/mindspore/mindspore/tree/master/model_zoo/research/cv/IPT Hanting Chen, Yunhe Wang 0001, Tianyu Guo 0001, Chang Xu 0002, Yiping Deng, Zhenhua Liu 0003, Siwei Ma 0001, Chunjing Xu, Chao Xu 0006, Wen Gao 0001 |
CVPR | 3 |
| 2021 | Learning Student Networks in the WildabstractData-free learning for student networks is a new paradigm for solving users’ anxiety caused by the privacy problem of using original training data. Since the architectures of modern convolutional neural networks (CNNs) are compact and sophisticated, the alternative images or meta-data generated from the teacher network are often broken. Thus, the student network cannot achieve the comparable performance to that of the pre-trained teacher network especially on the large-scale image dataset. Different to previous works, we present to maximally utilize the massive available unlabeled data in the wild. Specifically, we first thoroughly analyze the output differences between teacher and student network on the original data and develop a data collection method. Then, a noisy knowledge distillation algorithm is proposed for achieving the performance of the student network. In practice, an adaptation matrix is learned with the student network for correcting the label noise produced by the teacher network on the collected unlabeled images. The effectiveness of our DFND (Data-Free Noisy Distillation) method is then verified on several benchmarks to demonstrate its superiority over state-of-the-art data-free distillation methods. Experiments on various datasets demonstrate that the student networks learned by the proposed method can achieve comparable performance with those using the original dataset. Code is available at https://github.com/huawei-noah/Data-Efficient-Model-Compression Hanting Chen, Tianyu Guo 0001, Chang Xu 0002, Chunjing Xu, Chao Xu 0006, Yunhe Wang 0001 |
CVPR | 2 |
| 2020 | Learning Student Networks with Few DataabstractRecently, the teacher-student learning paradigm has drawn much attention in compressing neural networks on low-end edge devices, such as mobile phones and wearable watches. Current algorithms mainly assume the complete dataset for the teacher network is also available for the training of the student network. However, for real-world scenarios, users may only have access to part of training examples due to commercial profits or data privacy, and severe over-fitting issues would happen as a result. In this paper, we tackle the challenge of learning student networks with few data by investigating the ground-truth data-generating distribution underlying these few data. Taking Wasserstein distance as the measurement, we assume this ideal data distribution lies in a neighborhood of the discrete empirical distribution induced by the training examples. Thus we propose to safely optimize the worst-case cost within this neighborhood to boost the generalization. Furthermore, with theoretical analysis, we derive a novel and easy-to-implement loss for training the student network in an end-to-end fashion. Experimental results on benchmark datasets validate the effectiveness of our proposed method. Shumin Kong, Tianyu Guo 0001, Shan You, Chang Xu 0002 |
AAAI | 2 |
| 2020 | On Positive-Unlabeled Classification in GANabstractThis paper defines a positive and unlabeled classification problem for standard GANs, which then leads to a novel technique to stabilize the training of the discriminator in GANs. Traditionally, real data are taken as positive while generated data are negative. This positive-negative classification criterion was kept fixed all through the learning process of the discriminator without considering the gradually improved quality of generated data, even if they could be more realistic than real data at times. In contrast, it is more reasonable to treat the generated data as unlabeled, which could be positive or negative according to their quality. The discriminator is thus a classifier for this positive and unlabeled classification problem, and we derive a new Positive-Unlabeled GAN (PUGAN). We theoretically discuss the global optimality the proposed model will achieve and the equivalent optimization goal. Empirically, we find that PUGAN can achieve comparable or even better performance than those sophisticated discriminator stabilization methods. Tianyu Guo 0001, Chang Xu 0002, Yunhe Wang 0001, Boxin Shi, Chao Xu 0006, Dacheng Tao |
CVPR | 1 |
| 2020 | EDD-Net: An Efficient Defect Detection NetworkabstractAs the most commonly used communication tool, the mobile phone has become an indispensable part of our daily life. The surface of the mobile phone as the main window of human-phone interaction directly affects the user experience. It is necessary to detect surface defects on the production line in order to ensure the high quality of the mobile phone. However, the existing mobile phone surface defect detection is mainly done manually. Currently, there are few automatic defect detection methods to replace human eyes. How to quickly and accurately detect the surface defects of the mobile phone is an urgent problem to be solved. Hence, an efficient defect detection network (EDD-Net) is proposed. Firstly, EfficientNet is used as the backbone network. Then, according to the small-scale of mobile phone surface defects, a feature pyramid module named GCSA-BiFPN is proposed to obtain more discriminative features. Finally, the box/class prediction network is used to achieve effective defect detection. We also build a mobile phone surface oil stain defect (MPSOSD) dataset to alleviate the lack of dataset in this field. The performance on the relevant datasets shows that the proposed network is effective and has practical significance for industrial production. Tianyu Guo 0001, Runwei Ding |
ICPR | 1 |
| 2020 | Robust Student Network LearningabstractDeep neural networks bring in impressive accuracy in various applications, but the success often relies on heavy network architectures. Taking well-trained heavy networks as teachers, classical teacher-student learning paradigm aims to learn a student network that is lightweight yet accurate. In this way, a portable student network with significantly fewer parameters can achieve considerable accuracy, which is comparable to that of a teacher network. However, beyond accuracy, the robustness of the learned student network against perturbation is also essential for practical uses. Existing teacher-student learning frameworks mainly focus on accuracy and compression ratios, but ignore the robustness. In this paper, we make the student network produce more confident predictions with the help of the teacher network, and analyze the lower bound of the perturbation that will destroy the confidence of the student network. Two important objectives regarding prediction scores and gradients of examples are developed to maximize this lower bound, to enhance the robustness of the student network without sacrificing the performance. Experiments on benchmark data sets demonstrate the efficiency of the proposed approach to learning robust student networks that have satisfying accuracy and compact sizes. Tianyu Guo 0001, Chang Xu 0002, Shiyi He, Boxin Shi, Chao Xu 0006, Dacheng Tao |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2019 | Smooth Deep Image Generator from NoisesabstractGenerative Adversarial Networks (GANs) have demonstrated a strong ability to fit complex distributions since they were presented, especially in the field of generating natural images. Linear interpolation in the noise space produces a continuously changing in the image space, which is an impressive property of GANs. However, there is no special consideration on this property in the objective function of GANs or its derived models. This paper analyzes the perturbation on the input of the generator and its influence on the generated images. A smooth generator is then developed by investigating the tolerable input perturbation. We further integrate this smooth generator with a gradient penalized discriminator, and design smooth GAN that generates stable and high-quality images. Experiments on real-world image datasets demonstrate the necessity of studying smooth generator and the effectiveness of the proposed algorithm. Tianyu Guo 0001, Chang Xu 0002, Boxin Shi, Chao Xu 0006, Dacheng Tao |
AAAI | 1 |
| 2019 | Learning from Bad Data via GenerationabstractBad training data would challenge the learning model from understanding the underlying data-generating scheme, which then increases the difficulty in achieving satisfactory performance on unseen test data. We suppose the real data distribution lies in a distribution set supported by the empirical distribution of bad data. A worst-case formulation can be developed over this distribution set, and then be interpreted as a generation task in an adversarial manner. The connections and differences between GANs and our framework have been thoroughly discussed. We further theoretically show the influence of this generation task on learning from bad data and reveal its connection with a data-dependent regularization. Given different distance measures (\eg, Wasserstein distance or JS divergence) of distributions, we can derive different objective functions for the problem. Experimental results on different kinds of bad training data demonstrate the necessity and effectiveness of the proposed method. Tianyu Guo 0001, Chang Xu 0002, Boxin Shi, Chao Xu 0006, Dacheng Tao |
NeurIPS | 1 |
| 2018 | Reinforced Multi-Label Image Classification by Exploring CurriculumabstractHumans and animals learn much better when the examples are not randomly presented but organized in a meaningful order which illustrates gradually more concepts, and gradually more complex ones. Inspired by this curriculum learning mechanism, we propose a reinforced multi-label image classification approach imitating human behavior to label image from easy to complex. This approach allows a reinforcement learning agent to sequentially predict labels by fully exploiting image feature and previously predicted labels. The agent discovers the optimal policies through maximizing the long-term reward which reflects prediction accuracies. Experimental results on PASCAL VOC2007 and 2012 demonstrate the necessity of reinforcement multi-label learning and the algorithm’s effectiveness in real-world multi-label image classification tasks. Shiyi He, Chang Xu 0002, Tianyu Guo 0001, Chao Xu 0006, Dacheng Tao |
AAAI | 3 |