VLDB 2026 Research / reviewers in the wild / expert
Yuntao Du 0001
dblp:231/8856-1
· DBLP profile ↗
36ranked-venue papers
10as first author
32since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 7 first-author · 24 since 2021Databases, data management, data science and information retrieval · 12 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Benchmarking Multimodal Knowledge Conflict for Large Multimodal ModelsabstractLarge Multimodal Models (LMMs) face notable challenges when encountering multimodal knowledge conflicts, particularly under retrieval-augmented generation (RAG) frameworks, where the contextual information from external sources may contradict the model’s internal parametric knowledge, leading to unreliable outputs. However, existing benchmarks fail to reflect such realistic conflict scenarios. Most focus solely on intra-memory conflicts, while context-memory and inter-context conflicts remain largely unaddressed. Furthermore, commonly used factual knowledge-based evaluations are often overlooked, and existing datasets lack a thorough investigation into conflict detection capabilities.To bridge this gap, we propose MMKC-Bench, a benchmark designed to evaluate factual knowledge conflicts in both context-memory and inter-context scenarios. MMKC-Bench encompasses four types of multimodal knowledge conflicts and includes 1,881 knowledge instances and 3,997 images across 32 broad types, collected through automated pipelines with human verification. We evaluate four representative series of LMMs on both model behavior analysis and conflict detection tasks. Our findings show that while current LMMs are capable of recognizing knowledge conflicts, they tend to favor internal parametric knowledge over external evidence. We hope MMKC-Bench will foster further research in multimodal knowledge conflict and enhance the development of multimodal RAG systems. Yuntao Du 0001, Kailin Jiang, Yuyang Liang, Qihan Ren, Yi Xin 0003, Fenze Feng, Mingcai Chen, Hengyang Lu, Haozhe Wang 0002, Xiaoye Qu, Qian Li 0043, Dongrui Liu |
AAAI | 2 |
| 2026 | A Novel Fine-Tuned CLIP-OOD Detection Method with Double Loss Constraint Through Optimal Transport Semantic AlignmentabstractDetecting Out-Of-Distribution (OOD) samples in image classification is crucial for model reliability. With the rise of Vision-Language Models (VLMs), CLIP-OOD has become a research hotspot. However, we observe the Low Focus Attention phenomenon from the image encoders of CLIP, which means the attention of image encoders often spreads to non-in-distribution regions. This phenomenon comes from the semantic mismalignment and inter-class feature confusion. To address these issues, we propose a novel fine-tuned OOD detection method with the Double loss constraint based on Optimal Transport (DOT-OOD). DOT-OOD integrates the Double Loss Constraint (DLC) module and Optimal Transport (OT) module. The DLC module comprises the Aligned Image-Text Concept Matching Loss and the Negative Sample Repulsion Loss, which respectively (1) focus on the core semantics of ID images and achieve cross-modal semantic alignment, (2) expand inter-class distances and enhance discriminative. While the OT module is introduced to obtain enhanced image feature representations. Extensive experimental results show that in the 16-shot scenario of the ImageNet-1k benchmark, DOT-OOD reduces the FPR95 by over 10% and improves the AUROC from 94.48% to 96.57% compared with SOTAs. Hengyang Lu, Yuntao Du 0001, Chenyou Fan |
AAAI | 5 |
| 2026 | GitTaskBench: A Benchmark for Code Agents Solving Real-World Tasks Through Code Repository LeveragingabstractBeyond scratch coding, exploiting large-scale code repositories (e.g., GitHub) for practical tasks is vital in real-world software development, yet current benchmarks rarely evaluate code agents in such authentic, workflow-driven scenarios. To bridge this gap, we introduce GitTaskBench, a benchmark designed to systematically assess this capability via 54 realistic tasks across 7 modalities and 7 domains. Each task pairs a relevant repository with an automated, human-curated evaluation harness specifying practical success criteria. Beyond measuring execution and task success, we also propose the alpha-value metric to quantify the economic benefit of agent performance, which integrates task success rates, token cost, and average developer salaries. Experiments across three state-of-the-art agent frameworks with multiple advanced LLMs show that leveraging code repositories for complex task solving remains challenging: even the best-performing system, OpenHands+Claude 3.7, solves only 48.15% of tasks. Error analysis attributes over half of failures to seemingly mundane yet critical steps like environment setup and dependency resolution, highlighting the need for more robust workflow management and increased timeout preparedness. By releasing GitTaskBench, we aim to drive progress and attention toward repository-aware code reasoning, execution, and deployment---moving agents closer to solving complex, end-to-end real-world tasks. Ziyi Ni, Huacan Wang, Shuo Lu, Wang You, Zhenheng Tang, Sen Hu 0005, Bo Li 0117, Binxing Jiao, Daxin Jiang, Yuntao Du 0001 |
AAAI | 13 |
| 2026 | Rough-to-precise ranking with coverage optimization for open link prediction
Qian Li 0043, Ning Liu 0014, Yuntao Du 0001, Daling Wang, Li-Zhen Cui 0001 |
Expert Syst. Appl. | 3 |
| 2026 | Parameter-Efficient Fine-Tuning for Pre-Trained Vision Models: A Survey and Benchmark
Yi Xin 0003, Jianjiang Yang, Yuntao Du 0001, Haoxing Chen, Kangrui Cen, Yangfan He, Yuewen Cao, Junjun He, Xiaokang Yang 0001, Guangtao Zhai, Ming-Hsuan Yang 0001, Xiaohong Liu 0001 |
Int. J. Comput. Vis. | 4 |
| 2026 | CE-CLIP: Cloud-edge collaborative fine-tuning for multimodal adaptation
Kejun Ren, Yuntao Du 0001, Lianming Xu, Yunxiang Yao, Lei Jin 0003, Li Wang 0039 |
Pattern Recognit. | 2 |
| 2025 | Robust Logit Adjustment for Learning with Long-Tailed Noisy DataabstractLearning with noisy labels (LNL) methods have enabled the deployment of machine learning systems with imperfectly labeled data. However, these methods often struggle to identify noise in the presence of long-tailed (LT) class distributions, where the memorization effect becomes class-dependent. Conversely, LT methods are suboptimal under label noise, as it hinders access to accurate label frequency statistics. This study aims to address the long-tailed noisy data by bridging the methodological gap between LNL and LT approaches. We propose a direct solution, termed Robust Logit Adjustment, which estimates ground-truth labels through label refurbishment, thereby mitigating the impact of label noise. Simultaneously, our method incorporates the distribution of training-time corrected target labels into the LT method logit adjustment, providing class-rebalanced supervision. Extensive experiments on both synthetic and real-world long-tailed noisy datasets demonstrate the superior performance of our method. Mingcai Chen, Yuntao Du 0001, Baoming Zhang, Yi Xin 0003, Chong-Jun Wang |
AAAI | 2 |
| 2025 | MMKE-Bench: A Multimodal Editing Benchmark for Diverse Visual KnowledgeabstractKnowledge editing techniques have emerged as essential tools for updating the factual knowledge of large language models (LLMs) and multimodal models (LMMs), allowing them to correct outdated or inaccurate information without retraining from scratch. However, existing benchmarks for multimodal knowledge editing primarily focus on entity-level knowledge represented as simple triplets, which fail to capture the complexity of real-world multimodal information. To address this issue, we introduce MMKE-Bench, a comprehensive **M**ulti**M**odal **K**nowledge **E**diting Benchmark, designed to evaluate the ability of LMMs to edit diverse visual knowledge in real-world scenarios. MMKE-Bench addresses these limitations by incorporating three types of editing tasks: visual entity editing, visual semantic editing, and user-specific editing. Besides, MMKE-Bench uses free-form natural language to represent and edit knowledge, offering a more flexible and effective format. The benchmark consists of 2,940 pieces of knowledge and 8,363 images across 33 broad categories, with evaluation questions automatically generated and human-verified. We assess five state-of-the-art knowledge editing methods on three prominent LMMs, revealing that no method excels across all criteria, and that visual and user-specific edits are particularly challenging. MMKE-Bench sets a new standard for evaluating the robustness of multimodal knowledge editing techniques, driving progress in this rapidly evolving field. Yuntao Du 0001, Kailin Jiang, Zhi Gao 0002, Chenrui Shi, Zilong Zheng, Siyuan Qi, Qing Li 0003 |
ICLR | 1 |
| 2025 | Test-Time Selective Adaptation for Uni-Modal Distribution Shift in Multi-Modal DataabstractModern machine learning applications are characterized by the increasing size of deep models and the growing diversity of data modalities. This trend underscores the importance of efficiently adapting pre-trained multi-modal models to the test distribution in real time, i.e., multi-modal test-time adaptation. In practice, the magnitudes of multi-modal shifts vary because multiple data sources interact with the impact factor in diverse manners. In this research, we investigate the the under-explored practical scenario uni-modal distribution shift, where the distribution shift influences only one modality, leaving the others unchanged. Through theoretical and empirical analyses, we demonstrate that the presence of such shift impedes multi-modal fusion and leads to the negative transfer phenomenon in existing test-time adaptation techniques. To flexibly combat this unique shift, we propose a selective adaptation schema that incorporates multiple modality-specific adapters to accommodate potential shifts and a “router” module that determines which modality requires adaptation. Finally, we validate the effectiveness of our proposed method through extensive experimental evaluations. Code available at https://github.com/chenmc1996/Uni-Modal-Distribution-Shift. Mingcai Chen, Baoming Zhang, Zongbo Han, Yanmeng Wang, Yuntao Du 0001, Bing-Kun Bao |
ICML | 7 |
| 2025 | RepoMaster: Autonomous Exploration and Understanding of GitHub Repositories for Complex Task SolvingabstractThe ultimate goal of code agents is to solve complex tasks autonomously.
Although large language models (LLMs) have made substantial progress in code generation, real-world tasks typically demand full-fledged code repositories rather than simple scripts. Building such repositories from scratch remains a major challenge. Fortunately, GitHub hosts a vast, evolving collection of open-source repositories, which developers frequently reuse as modular components for complex tasks. Yet, existing frameworks like OpenHands and SWE-Agent still struggle to effectively leverage these valuable resources.
Relying solely on README files provides insufficient guidance, and deeper exploration reveals two core obstacles: overwhelming information and tangled dependencies of repositories, both constrained by the limited context windows of current LLMs.
To tackle these issues, we propose RepoMaster, an autonomous agent framework designed to explore and reuse GitHub repositories for solving complex tasks.
For efficient understanding, RepoMaster constructs function-call graphs, module-dependency graphs, and hierarchical code trees to identify essential components, providing only identified core elements to the LLMs rather than the entire repository.
During autonomous execution, it progressively explores related components using our exploration tools and prunes information to optimize context usage.
Evaluated on the adjusted MLE-bench, RepoMaster achieves a 110\% relative boost in valid submissions over the strongest baseline OpenHands.
On our newly released GitTaskBench, RepoMaster lifts the task-pass rate from 40.7% to 62.9% while reducing token usage by 95%.
Our code and demonstration materials are publicly available at https://github.com/QuantaAlpha/RepoMaster. Huacan Wang, Ziyi Ni, Shuo Lu, Sen Hu 0005, Jiaye Lin, Yifu Guo, Yuntao Du 0001 |
NeurIPS | 10 |
| 2025 | LaplaceConfidence: A graph-based approach for learning with noisy labelsabstractReal-world machine learning applications seldom provide perfect labeled data, posing a challenge in developing models robust to noisy labels. Recent methods prioritize noise filtering based on the discrepancies between model predictions and the provided noisy labels, assuming samples with minimal classification losses to be clean. In this work, we capitalize on the consistency between the learned model and the complete noisy dataset, employing the data’s rich representational and topological information. We introduce LaplaceConfidence, a method that to obtain label confidence (i.e., clean probabilities) utilizing the Laplacian energy. Specifically, it first constructs graphs based on the feature representations of all noisy samples and minimizes the Laplacian energy to produce a low-energy graph. Clean labels should fit well into the low-energy graph while noisy ones should not, allowing our method to determine data’s clean probabilities. Furthermore, LaplaceConfidence is embedded into a holistic method for robust training, where co-training technique generates unbiased label confidence and label refurbishment technique better utilizes it. We also explore the dimensionality reduction technique to accommodate our method on large-scale noisy datasets. Our experiments demonstrate that LaplaceConfidence outperforms state-of-the-art methods on benchmark datasets under both synthetic and real-world noise. Code available at https://github.com/chenmc1996/LaplaceConfidence . Mingcai Chen, Yuntao Du 0001, Baoming Zhang, Chong-Jun Wang |
Intell. Data Anal. | 2 |
| 2025 | Enhancing few-shot out-of-distribution intent detection by reducing attention misallocation
Hengyang Lu, Jiaming Zhang 0006, Yuntao Du 0001, Chong-Jun Wang, Wei Fang 0001, Xiaojun Wu 0001 |
Neurocomputing | 3 |
| 2025 | Multi-source fully test-time adaptation
Yuntao Du 0001, Yi Xin 0003, Mingcai Chen, Mujie Zhang, Chong-Jun Wang |
Neural Networks | 1 |
| 2025 | MuSIA: Exploiting multi-source information fusion with abnormal activations for out-of-distribution detection
Hengyang Lu, Chenyou Fan, Yuntao Du 0001, Zhenhao Shao, Wei Fang 0001, Xiaojun Wu 0001 |
Neural Networks | 5 |
| 2024 | CLOVA: A Closed-LOop Visual Assistant with Tool Usage and UpdateabstractUtilizing large language models (LLMs) to compose off-the-shelf visual tools represents a promising avenue of research for developing robust visual assistants capable of addressing diverse visual tasks. However, these methods often overlook the potential for continual learning, typically by freezing the utilized tools, thus limiting their adaptation to environments requiring new knowledge. To tackle this challenge, we propose CLOVA, a Closed-LOop Visual Assistant, which operates within a framework encompassing inference, reflection, and learning phases. During the inference phase, LLMs generate programs and execute corresponding tools to complete assigned tasks. In the reflection phase, a multimodal global-local reflection scheme analyzes human feedback to determine which tools require updating. Lastly, the learning phase employs three flexible approaches to automatically gather training data and introduces a novel prompt tuning scheme to update the tools, allowing CLOVA to efficiently acquire new knowledge. Experimental findings demonstrate that CLOVA surpasses existing tool-usage methods by 5% in visual question answering and multiple-image reasoning, by 10% in knowledge tagging, and by 20% in image editing. These results under-score the significance of the continual learning capability in general visual assistants. Zhi Gao 0002, Yuntao Du 0001, Xiaojian Ma 0001, Wenjuan Han, Song-Chun Zhu, Qing Li 0003 |
CVPR | 2 |
| 2024 | 🤖 VideoAgent: A Memory-Augmented Multimodal Agent for Video Understanding
Xiaojian Ma 0001, Rujie Wu, Yuntao Du 0001, Jiaqi Li 0021, Zhi Gao 0002, Qing Li 0003 |
ECCV (22) | 4 |
| 2024 | V-PETL Bench: A Unified Visual Parameter-Efficient Transfer Learning BenchmarkabstractParameter-efficient transfer learning (PETL) methods show promise in adapting a pre-trained model to various downstream tasks while training only a few parameters. In the computer vision (CV) domain, numerous PETL algorithms have been proposed, but their direct employment or comparison remains inconvenient. To address this challenge, we construct a Unified Visual PETL Benchmark (V-PETL Bench) for the CV domain by selecting 30 diverse, challenging, and comprehensive datasets from image recognition, video action recognition, and dense prediction tasks. On these datasets, we systematically evaluate 25 dominant PETL algorithms and open-source a modular and extensible codebase for fair evaluation of these algorithms. V-PETL Bench runs on NVIDIA A800 GPUs and requires approximately 310 GPU days. We release all the benchmark, making it more efficient and friendly to researchers. Additionally, V-PETL Bench will be continuously updated for new PETL algorithms and CV tasks. Yi Xin 0003, Xuyang Liu 0002, Yuntao Du 0001, Haodi Zhou, Christina E. Lee, Junlong Du, Haozhe Wang 0002, Mingcai Chen, Ting Liu 0018, Guimin Hu, Zhongwei Wan, Rongchao Zhang, Aoxue Li, Mingyang Yi, Xiaohong Liu 0001 |
NeurIPS | 4 |
| 2024 | Generation, augmentation, and alignment: a pseudo-source domain based method for source-free domain adaptation
Yuntao Du 0001, Haiyang Yang, Mingcai Chen, Hongtao Luo, Juan Jiang, Yi Xin 0003, Chong-Jun Wang |
Mach. Learn. | 1 |
| 2023 | Two Wrongs Don't Make a Right: Combating Confirmation Bias in Learning with Label NoiseabstractNoisy labels damage the performance of deep networks. For robust learning, a prominent two-stage pipeline alternates between eliminating possible incorrect labels and semi-supervised training. However, discarding part of noisy labels could result in a loss of information, especially when the corruption has a dependency on data, e.g., class-dependent or instance-dependent. Moreover, from the training dynamics of a representative two-stage method DivideMix, we identify the domination of confirmation bias: pseudo-labels fail to correct a considerable amount of noisy labels, and consequently, the errors accumulate. To sufficiently exploit information from noisy labels and mitigate wrong corrections, we propose Robust Label Refurbishment (Robust LR)—a new hybrid method that integrates pseudo-labeling and confidence estimation techniques to refurbish noisy labels. We show that our method successfully alleviates the damage of both label noise and confirmation bias. As a result, it achieves state-of-the-art performance across datasets and noise types, namely CIFAR under different levels of synthetic noise and mini-WebVision and ANIMAL-10N with real-world noise. Mingcai Chen, Hao Cheng 0014, Yuntao Du 0001, Ming Xu 0014, Chong-Jun Wang |
AAAI | 3 |
| 2023 | Self-Training with Label-Feature-Consistency for Domain Adaptation
Yi Xin 0003, Pengsheng Jin, Yuntao Du 0001, Chong-Jun Wang |
DASFAA (4) | 4 |
| 2023 | CESED: Exploiting Hyperspherical Predefined Evenly-Distributed Class Centroids for OOD DetectionabstractOut-of-distribution (OOD) detection is critical for ensuring the safe deployment of machine learning models in the open world. Due to the simplicity and intuitiveness of distance- based methods, i.e., samples are detected as OOD if they are relatively far away from the centroids or prototypes of in-distribution (ID) classes, they have attracted widespread attention from researchers in the field of OOD detection. However, prior OOD detection methods directly take off-the- shelf loss functions, like widely used softmax cross-entropy (CE) loss, that suffices for classifying ID samples, but is not optimally designed for OOD detection. In this work, we propose CESED, an improved CE loss applied to the scalable Squared Euclidean Distance vector, which exploits hyper- spherical evenly-distributed class centroids for OOD detection. CESED can promote strong ID-OOD separability because it explicitly encourages maximization of inter-class distances and minimization of intra-class distances. Extensive experiments demonstrate that CESED achieves superior detection performance on a comprehensive suite of benchmark datasets. For the more challenging case where CIFAR-100 is used as ID, our method achieves a 31.98% reduction in average FPR95 and 6.20% reduction in ID test error compared to the baseline method using a softmax confidence score. Mingcai Chen, Yuntao Du 0001, Hao Cheng 0014, Yuxin Ge, Chong-Jun Wang |
SDM | 4 |
| 2022 | Semi-supervised Learning with Multi-Head Co-TrainingabstractCo-training, extended from self-training, is one of the frameworks for semi-supervised learning. Without natural split of features, single-view co-training works at the cost of training extra classifiers, where the algorithm should be delicately designed to prevent individual classifiers from collapsing into each other. To remove these obstacles which deter the adoption of single-view co-training, we present a simple and efficient algorithm Multi-Head Co-Training. By integrating base learners into a multi-head structure, the model is in a minimal amount of extra parameters. Every classification head in the unified model interacts with its peers through a “Weak and Strong Augmentation” strategy, in which the diversity is naturally brought by the strong data augmentation. Therefore, the proposed method facilitates single-view co-training by 1). promoting diversity implicitly and 2). only requiring a small extra computational overhead. The effectiveness of Multi-Head Co-Training is demonstrated in an empirical study on standard semi-supervised learning benchmarks. Mingcai Chen, Yuntao Du 0001, Yi Zhang 0073, Shuwei Qian, Chong-Jun Wang |
AAAI | 2 |
| 2022 | Joint Feature and Labeling Function Adaptation for Unsupervised Domain Adaptation
Fengli Cui, Yuntao Du 0001, Yikang Cao, Chong-Jun Wang |
PAKDD (1) | 3 |
| 2022 | InCo: Intermediate Prototype Contrast for Unsupervised Domain Adaptation
Yuntao Du 0001, Hongtao Luo, Haiyang Yang, Juan Jiang, Chong-Jun Wang |
ECML/PKDD (1) | 1 |
| 2022 | Learning transferable and discriminative features for unsupervised domain adaptationabstractAlthough achieving remarkable progress, it is very difficult to induce a supervised classifier without any labeled data. Unsupervised domain adaptation is able to overcome this challenge by transferring knowledge from a labeled source domain to an unlabeled target domain. Transferability and discriminability are two key criteria for characterizing the superiority of feature representations to enable successful domain adaptation. In this paper, a novel method called learning TransFerable and Discriminative Features for unsupervised domain adaptation (TFDF) is proposed to optimize these two objectives simultaneously. On the one hand, distribution alignment is performed to reduce domain discrepancy and learn more transferable representations. Instead of adopting Maximum Mean Discrepancy (MMD) which only captures the first-order statistical information to measure distribution discrepancy, we adopt a recently proposed statistic called Maximum Mean and Covariance Discrepancy (MMCD), which can not only capture the first-order statistical information but also capture the second-order statistical information in the reproducing kernel Hilbert space (RKHS). On the other hand, we propose to explore both local discriminative information via manifold regularization and global discriminative information via minimizing the proposed class confusion objective to learn more discriminative features, respectively. We integrate these two objectives into the Structural Risk Minimization (RSM) framework and learn a domain-invariant classifier. Comprehensive experiments are conducted on five real-world datasets and the results verify the effectiveness of the proposed method. Yuntao Du 0001, Ruiting Zhang, Yirong Yao, Hengyang Lu, Chong-Jun Wang |
Intell. Data Anal. | 1 |
| 2021 | AdaRNN: Adaptive Learning and Forecasting of Time SeriesabstractTime series has wide applications in the real world and is known to be difficult to forecast. Since its statistical properties change over time, its distribution also changes temporally, which will cause severe distribution shift problem to existing methods. However, it remains unexplored to model the time series in the distribution perspective. In this paper, we term this as Temporal Covariate Shift (TCS). This paper proposes Adaptive RNNs (AdaRNN) to tackle the TCS problem by building an adaptive model that generalizes well on the unseen test data. AdaRNN is sequentially composed of two novel algorithms. First, we propose Temporal Distribution Characterization to better characterize the distribution information in the TS. Second, we propose Temporal Distribution Matching to reduce the distribution mismatch in TS to learn the adaptive TS model. AdaRNN is a general framework with flexible distribution distances integrated. Experiments on human activity recognition, air quality prediction, and financial analysis show that AdaRNN outperforms the latest methods by a classification accuracy of 2.6% and significantly reduces the RMSE by 9.0%. We also show that the temporal distribution matching algorithm can be extended in Transformer structure to boost its performance. Yuntao Du 0001, Jindong Wang 0001, Wenjie Feng 0001, Sinno Jialin Pan, Tao Qin 0001, Renjun Xu, Chong-Jun Wang |
CIKM | 1 |
| 2021 | Adversarial Separation Network for Cross-Network Node ClassificationabstractNode classification is an important yet challenging task in various network applications, and many effective methods have been developed for a single network. While for cross-network scenarios, neither single network embedding nor traditional domain adaptation can directly solve the task. Existing approaches have been proposed to combine network embedding and domain adaptation for cross-network node classification. However, they only focus on domain-invariant features, ignoring the individual features of each network, and they only utilize 1-hop neighborhood information (local consistency), ignoring the global consistency information. To tackle the above problems, in this paper, we propose a novel model, Adversarial Separation Network(ASN), to learn effective node representations between source and target networks. We explicitly separate domain-private and domain-shared information. Two domain-private encoders are employed to extract the domain-specific features in each network and a shared encoder is employed to extract the domain-invariant shared features across networks. Moreover, in each encoder, we combine local and global consistency to capture network topology information more comprehensively. ASN integrates deep network embedding with adversarial domain adaptation to reduce the distribution discrepancy across domains. Extensive experiments on real-world datasets show that our proposed model achieves state-of-the-art performance in cross-network node classification tasks compared with existing algorithms. Yuntao Du 0001, Rongbiao Xie, Chong-Jun Wang |
CIKM | 2 |
| 2021 | Cross-Domain Error Minimization for Unsupervised Domain Adaptation
Yuntao Du 0001, Fengli Cui, Chong-Jun Wang |
DASFAA (2) | 1 |
| 2021 | Self Separation and Misseparation Impact Minimization for Open-Set Domain Adaptation
Yuntao Du 0001, Yikang Cao, Yumeng Zhou, Ruiting Zhang, Chong-Jun Wang |
DASFAA (2) | 1 |
| 2021 | Unsupervised Domain Adaptation with Unified Joint Distribution Alignment
Yuntao Du 0001, Zhiwen Tan, Yirong Yao, Hualei Yu, Chong-Jun Wang |
DASFAA (2) | 1 |
| 2021 | DMSPool: Dual Multi-Scale Pooling for Graph Representation Learning
Hualei Yu, Yuntao Du 0001, Hao Cheng 0014, Meng Cao 0004, Chong-Jun Wang |
DASFAA (1) | 3 |
| 2021 | Nested Dense Attention Network for Single Image Super-ResolutionabstractRecently, deep convolutional neural networks (CNNs) are widely used in single image super-resolution (SISR) and have recorded impressive performance. However, most of the existing CNNs architectures can not fully utilize the correlation of feature maps in the middle layers, and abundant features of different levels are lost. Furthermore, convolution operation is limited by processing one local neighborhood at a time, which lacks global information. To address these issues, we propose the nested dense attention network (NDAN) for generating more refined and structured high-resolution images. Specifically, we propose nested dense structure (NDS) to better integrate features of different levels extracted from different layers. Besides that, in order to capture inter-channel dependencies more efficiently, we propose the adaptive channel attention module (ACAM) to adaptively rescale channel-wise features by automatically adjusting the weights of different receptive fields. Furthermore, to better explore the global-level context information, we design hybrid non-local module (HNLM) and hybrid non-local up-sampler (HNLU) to upscale the images by capturing spatial-wise long-distance dependencies and channel-wise long-distance correlation. Numerous experiments demonstrate the effectiveness of our model by achieving higher PSNR and SSIM scores and generating images with better structures against the state-of-the-art methods. Yirong Yao, Yuntao Du 0001 |
ICMR | 3 |
| 2020 | Homogeneous Online Transfer Learning with Online Distribution Discrepancy MinimizationabstractTransfer learning has been demonstrated to be successful and essential in diverse applications, which transfers knowledge from related but different source domains to the target domain. Online transfer learning(OTL) is a more challenging problem where the target data arrive in an online manner. Most OTL methods combine source classifier and target classifier directly by assigning a weight to each classifier, and adjust the weights constantly. However, these methods pay little attention to reducing the distribution discrepancy between domains. In this paper, we propose a novel online transfer learning method which seeks to find a new feature representation, so that the marginal distribution and conditional distribution discrepancy can be online reduced simultaneously. We focus on online transfer learning with multiple source domains and use the Hedge strategy to leverage knowledge from source domains. We analyze the theoretical properties of the proposed algorithm and provide an upper mistake bound. Comprehensive experiments on two real-world datasets show that our method outperforms state-of-the-art methods by a large margin. Yuntao Du 0001, Zhiwen Tan, Yi Zhang 0073, Chong-Jun Wang |
ECAI | 1 |
| 2020 | Mining Knowledge within Categories in Global and Local Fashion for Multi-Label Text ClassificationabstractMulti-label text classification (MLTC) is an important task in natural language processing, which assigns multiple labels to each text in the dataset. Typical method like Binary Relevance (BR) is arguably the most intuitive solution for the task. It works by decomposing the multi-label learning task into a number of independent binary learning tasks while ignoring the correlation between labels. Recently, neural network models attract much attention. Researchers view the MLTC task as a sequence generation problem. Although some new methods based on generative model (e.g. sequence-to-sequence), such as novel decoder structure and various attention mechanisms, can improve the performance. These methods still have some short-comings, such as unreasonable loss function, unclear ordering of target labels. To address these limitations, we propose a simple and effective novel model, which combines the merits of neural network and BRs methods. Our model also takes into account the categories and levels of labels. We decompose the MLTC problem to binary classification, together with global and local extractor to avoid the impact of label ordering and cumulative error. Experimental results show that our model achieves an improvement of 3.0% micro-F1 and a reduction of 6.0% hamming loss on AAPD dataset compared with the state-of-the-art work. And obtained good performance on RCV1-V2 dataset. Yuntao Du 0001, Bin Jin, Lingshuang Yu |
IJCNN | 3 |
| 2020 | Unsupervised Domain Adaptation with Joint Domain-Adversarial Reconstruction Networks
Yuntao Du 0001, Zhiwen Tan, Yi Zhang 0073, Chong-Jun Wang |
ECML/PKDD (2) | 2 |
| 2018 | HetEOTL: An Algorithm for Heterogeneous Online Transfer LearningabstractTransfer learning is an important topic in machine learning and has been broadly studied for many years. However, most existing transfer learning methods assume the training sets are prepared in advance, which is often not the case in practice. Fortunately, online transfer learning (OTL), which addresses the transfer learning tasks in an online fashion, has been proposed to solve the problem. This paper mainly focuses on the heterogeneous OTL, which is in general very challenging because the feature space of target domain is different from that of the source domain. In order to enhance the learning performance, we designed the algorithm called Heterogeneous Ensembled Online Transfer Learning (HetEOTL) using ensemble learning strategy. Finally, we evaluate our algorithm on some benchmark datasets, and the experimental results show that HetEOTL has better performance than some other existing online learning and transfer learning algorithms, which proves the effectiveness of HetEOTL. Yuntao Du 0001, Ming Xu 0014, Chong-Jun Wang |
ICTAI | 2 |