EDBT 2026 Demo / reviewers in the wild / expert
Yong Luo 0002
dblp:57/5272-2
· DBLP profile ↗
146ranked-venue papers
22as first author
106since 2021 · last 2026
0000-0002-2296-6370ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 79 · 11 first-author · 64 since 2021Graphics, computer vision, multimedia, augmented reality and games · 74 · 13 first-author · 44 since 2021Databases, data management, data science and information retrieval · 9 · 4 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 8 since 2021Computer networks · 7 · 2 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SRD: Reinforcement-Learned Semantic Perturbation for Backdoor Defense in VLMsabstractVisual language models (VLMs) have made significant progress in image captioning tasks, yet recent studies have found they are vulnerable to backdoor attacks. Attackers can inject undetectable perturbations into the data during inference, triggering abnormal behavior and generating malicious captions. These attacks are particularly challenging to detect and defend against due to the stealthiness and cross-modal propagation of the trigger signals. In this paper, we identify two key vulnerabilities by analyzing existing attack patterns: (1) the model exhibits abnormal attention concentration on certain regions of the input image, and (2) backdoor attacks often induce semantic drift and sentence incoherence. Based on these insights, we propose Semantic Reward Defense (SRD), a reinforcement learning framework that mitigates backdoor behavior without requiring any prior knowledge of trigger patterns. SRD learns to apply discrete perturbations to sensitive contextual regions of image inputs via a deep Q-network policy, aiming to confuse attention and disrupt the activation of malicious paths. To guide policy optimization, we design a reward signal named semantic fidelity score, which jointly assesses the semantic consistency and linguistic fluency of the generated captions, encouraging the agent to achieve a robust yet faithful output. SRD offers a trigger-agnostic, policy-interpretable defense paradigm that effectively mitigates local (TrojVLM) and global (Shadowcast) backdoor attacks, reducing ASR to 3.6% and 5.6% respectively, with less than 15% average CIDEr drop on the clean inputs. Shuhan Xu, Siyuan Liang 0004, Hongling Zheng, Aishan Liu, Xinbiao Wang, Yong Luo 0002, Leszek Rutkowski, Dacheng Tao |
AAAI | 6 |
| 2026 | Rethinking Visual Token Reduction in LVLMs Under Cross-Modal MisalignmentabstractLarge Vision-Language Models (LVLMs) encode visual inputs as dense sequences of patch-level tokens to capture fine-grained semantics. These visual tokens often outnumber their textual counterparts by a large margin, leading to substantial computational overhead and limiting the scalability of LVLMs in practice. Previous efforts have explored visual token reduction either prior to or within the large language models (LLMs). However, most in-LLM reduction approaches rely on text-conditioned interactions, implicitly assuming that textual tokens can reliably capture the importance of visual tokens. In this work, we revisit this assumption and reveal causal, semantic, and spatial forms of cross-modal misalignment. These misalignments undermine the effectiveness of text-guided visual token reduction. To address this, we introduce VisionDrop, a training-free, visual-only pruning framework that selects informative visual tokens based on intra-modal (visual-to-visual) attention, without relying on textual signals. To further suppress redundancy throughout the model hierarchy, we treat the visual encoder and the LLM as a unified system and design a progressive pruning pipeline. Our method performs dominant token selection and lightweight contextual merging at multiple stages, enabling fine-grained visual information to be retained even under aggressive token budgets. Extensive experiments across diverse benchmarks show that VisionDrop achieves consistent improvements over existing approaches, despite requiring no additional training or complex modifications. Notably, when integrated with LLaVA-NeXT-7B, VisionDrop achieves a 2.7x reduction in inference latency and 6x in FLOPs, while retaining 95.71% of the original performance. Rui Xu 0031, Yunke Wang, Yong Luo 0002, Bo Du 0001 |
AAAI | 3 |
| 2026 | Physics-Informed Multi-Task Learning for Battery State of Health Prediction with Uncertainty QuantificationabstractExisting battery State of Health (SOH) prediction approaches often struggle to provide both accurate predictions and reliable uncertainty estimates. This paper presents a novel Multi-Task Learning (MTL) framework that jointly tackles SOH prediction and provides a proxy metric for uncertainty through a unified architecture. The framework combines a Physics-Informed Neural Network (PINN) for SOH prediction with a deep autoencoding Gaussian mixture model for uncertainty modeling. Particularly, the energy score from the Gaussian mixture model serves as a proxy metric for uncertainty, where a higher score indicates potential prediction unreliability. Moreover, to enhance task-specific learning, we employ a multi-head attention mechanism that adaptively captures distinct feature relationships. Our experiments show improvements in prediction performance compared to the state-of-the-art baseline. A comprehensive evaluation on six XJTU battery benchmark datasets demonstrates that our framework achieves a prediction accuracy of 99.50% (MAPE: 0.0050) while providing reliable uncertainty quantification through the proxy metric. Tianwen Zhu, Ruihang Wang, Jimin Jia, Yong Luo 0002, Yonggang Wen 0001 |
AAAI | 6 |
| 2026 | OACI: Object-aware contextual integration for image captioning
Shuhan Xu, Mengya Han, Wei Yu 0004, Zheng He 0001, Xin Zhou 0003, Yong Luo 0002 |
Knowl. Based Syst. | 6 |
| 2026 | Adaptive Batch Size Time Evolving Stochastic Gradient Descent for Federated LearningabstractVariance reduction has been shown to improve the performance of Stochastic Gradient Descent (SGD) in centralized machine learning. However, when it is extended to federated learning systems, many issues may arise, including (i) mega-batch size settings; (ii) additional noise introduced by the gradient difference between the current iteration and the snapshot point; and (iii) gradient (statistical) heterogeneity. In this paper, we propose a lightweight algorithm termed federated adaptive batch size time evolving variance reduction (FedATEVR) to tackle these issues, consisting of an adaptive batch size setting scheme and a time-evolving variance reduction gradient estimator. In particular, we use the historical gradient information to set an appropriate mega-batch size for each client, which can steadily accelerate the local SGD process and reduce the computation cost. The historical information involves both global and local gradient, which mitigates unstable varying in mega-batch size introduced by gradient heterogeneity among the clients. For each client, the gradient difference between the current iteration and the snapshot point is used to tune the time-evolving weight of the variance reduction term in the gradient estimator. This can avoid meaningless variance reduction caused by the out-of-date snapshot point gradient. We theoretically prove that our algorithm can achieve a linear speedup of of $\mathcal {O}(\frac{1}{\sqrt{SKT}})$O(1SKT) for non-convex objective functions under partial client participation. Extensive experiments demonstrate that our proposed method can achieve higher test accuracy than the baselines and decrease communication rounds greatly. Xuming An 0001, Li Shen 0008, Yong Luo 0002, Han Hu 0003, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | Efficient and Effective Weight-Ensembling Mixture of Experts for Multi-Task Model MergingabstractMulti-task learning (MTL) leverages a shared model to accomplish multiple tasks and facilitate knowledge transfer. Recent research on task arithmetic-based MTL demonstrates that merging the parameters of independently fine-tuned models can effectively achieve MTL. However, existing merging methods primarily seek a static optimal solution within the original model parameter space, which often results in performance degradation due to the inherent diversity among tasks and potential interferences. To address this challenge, in this paper, we propose a Weight-Ensembling Mixture of Experts (WEMoE) method for multi-task model merging. Specifically, we first identify critical (or sensitive) modules by analyzing parameter variations in core modules of Transformer-based models before and after fine-tuning. Then, our WEMoE statically merges non-critical modules while transforming critical modules into a mixture-of-experts (MoE) structure. During inference, expert modules in the MoE are dynamically merged based on input samples, enabling a more flexible and adaptive merging approach. Building on WEMoE, we further introduce an efficient-and-effective WEMoE (E-WEMoE) method, whose core mechanism involves eliminating non-essential elements in the critical modules of WEMoE and implementing shared routing across multiple MoE modules, thereby significantly reducing both the trainable parameters, the overall parameter count, and computational overhead of the merged model by WEMoE. Experimental results across various architectures and tasks demonstrate that both WEMoE and E-WEMoE outperform state-of-the-art (SOTA) model merging methods in terms of MTL performance, generalization, and robustness. Li Shen 0008, Anke Tang, Enneng Yang, Guibing Guo, Yong Luo 0002, Lefei Zhang, Xiaochun Cao, Bo Du 0001, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | Zero-Shot Sparse Mixture of Low-Rank Experts Construction From Pre-Trained Foundation ModelsabstractDeep model training on extensive datasets is increasingly becoming cost-prohibitive, prompting the widespread adoption of deep model fusion techniques to leverage knowledge from pre-existing models. From simple weight averaging to more sophisticated methods like AdaMerging, model fusion effectively improves model performance and accelerates the development of new models. However, potential interference between parameters of individual models and the lack of interpretability in the fusion progress remain significant challenges. Existing methods often try to resolve the parameter interference issue by evaluating attributes of parameters, such as their magnitude or sign, or by parameter pruning. In this study, we begin by examining the fine-tuning of linear layers through the lens of subspace analysis and explicitly define parameter interference as an optimization problem to shed light on this subject. Subsequently, we introduce an innovative approach to model fusion called zero-shot Sparse MIxture of Low-rank Experts (SMILE) construction, which allows for the upscaling of source models into an MoE model without extra data or further training. Our approach relies on the observation that fine-tuning mostly keeps the important parts from the pre-training, but it uses less significant or unused areas to adapt to new tasks. Additionally, the issue of parameter interference, which is intrinsically challenging in the original parameter space, can be managed by expanding the dimensions. We conduct extensive experiments across diverse scenarios, such as image classification and text generation tasks, using full fine-tuning and LoRA fine-tuning, and we apply our method to large language models (CLIP models, Flan-T5 models, and Mistral-7B models), highlighting the adaptability and scalability of SMILE. For full fine-tuned models, about 50% additional parameters can achieve around 98-99% of the performance of eight individual fine-tuned ViT models, while for LoRA fine-tuned Flan-T5 models, maintaining 99% performance with only 2% extra parameters. Code is available athttps://github.com/tanganke/fusion_bench. Anke Tang, Li Shen 0008, Yong Luo 0002, Shuai Xie, Han Hu 0003, Lefei Zhang, Bo Du 0001, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | Aligning Few-Step Diffusion Models With Dense Reward Difference LearningabstractFew-step diffusion models enable efficient high-resolution image synthesis but struggle to align with specific downstream objectives due to limitations of existing reinforcement learning (RL) methods in low-step regimes with limited state spaces and suboptimal sample quality. To address this, we propose Stepwise Diffusion Policy Optimization (SDPO), a novel RL framework tailored for few-step diffusion models. SDPO introduces a dual-state trajectory sampling mechanism, tracking both noisy and predicted clean states at each step to provide dense reward feedback and enable low-variance, mixed-step optimization. For further efficiency, we develop a latent similarity-based dense reward prediction strategy to minimize costly dense reward queries. Leveraging these dense rewards, SDPO optimizes a dense reward difference learning objective that enables more frequent and granular policy updates. Additional refinements, including stepwise advantage estimates, temporal importance weighting, and step-shuffled gradient updates, further enhance long-term dependency, low-step priority, and gradient stability. Our experiments demonstrate that SDPO consistently delivers superior reward-aligned results across diverse few-step settings and tasks. Ziyi Zhang 0001, Li Shen 0008, Sen Zhang 0006, Deheng Ye, Yong Luo 0002, Miaojing Shi, Dongjing Shan, Bo Du 0001, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | Towards to real world vehicle privacy protection: A new dataset and benchmark
Jiayi Lin 0010, Chengming Zou, Long Lan, Yong Luo 0002, Yue Yu 0001, Yaowei Wang 0001, Wei Zeng 0006, Yonghong Tian 0001 |
Pattern Recognit. | 4 |
| 2026 | MCF-UNet: Multi-level context fusion unet for immunohistochemical positive cell detection
Dixiao Tao, Yong Luo 0002, Zongjie Hao, Dehua Cao, Yamei Luo, Dongjing Shan |
Pattern Recognit. | 2 |
| 2026 | EEformer: Early Exiting for Transformer With Global-Local Exits and Progressive Fine-TuningabstractRecently, the efficient deployment and acceleration of transformer-based pre-trained models (TPMs) on resource-constrained edge devices for multimedia services have gained significant interest. Although early exiting is a feasible solution, it may lead to extra computational cost and substantial performance degradation compared to the original models. To tackle these issues, we propose a framework termed EEformer, which incorporates global-local heads (GLHs) into intermediate layers to construct the early exiting dynamic neural network (EDNN). The GLH can efficiently extract global and local information from hidden states produced by the backbone layer, thereby achieving a better performance-efficiency trade-off for the EDNN. Moreover, we propose a novel progressive fine-tuning strategy to steadily improve the efficiency of the EDNN while maintaining its performance comparable to the original mode through three fine-tuning stages. We conduct extensive experiments on image classification and natural language processing tasks, demonstrating the superiority of the proposed framework. In particular, the proposed framework achieves 1.87× speed-up while maintaining 99.0% performance on the CIFAR-100 dataset, and 3.05× speed-up while maintaining 98.5% performance on the SST-2 dataset. Guanyu Xu, Yong Luo 0002, Li Shen 0008, Han Hu 0003, Dan Zeng 0001 |
IEEE Trans. Multim. | 3 |
| 2026 | SigGen: Signal Generation for Wireless Sensing Based on Disentangled RepresentationabstractWith the thriving artificial intelligence-generated content (AIGC), it is becoming increasingly appealing to exploit generative AI to generate wireless signals for facilitating wireless sensing. However, this is a challenging task, as wireless signals are highly random in general and contain rich physical information. To tackle these challenges, we propose a novel signal disentanglement and generation framework termed SigGen, which is inspired by the Fourier Transform (FT) that converts signals to the frequency domain and accordingly separates objectives by distinct frequency bands. In our proposed framework, we first disentangle the features of objects embedded in the signal and subsequently modify these features to generate the desired signals. Specifically, we devise a neural network based on the vision transformer (ViT) to extract effective features for signal generation. In this neural network, we incorporate both local and global frequency attention modules to adaptively leverage frequency features, and introduce a hybrid patch embedding module to enhance information interaction for the ViT architecture. Furthermore, we propose a novel sequential training method to improve the disentanglement and generation capability of the neural network. Finally, extensive experiments on two benchmark public wireless sensing datasets demonstrate that our framework can effectively decouple wireless signals and generate diverse signals closely resembling real ones, surpassing state-of-the-art methods by 30.83%. A practical case study further demonstrates that our framework can be used as a data augmentation method to improve gesture recognition accuracy by 12.74%. Hanxiang He, Xintao Huan, Yong Luo 0002, Rongfei Fan, Jie Xu 0002, Han Hu 0003 |
IEEE Trans. Wirel. Commun. | 3 |
| 2025 | Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language ModelsabstractMultimodal large language models (MLLMs) have experienced significant advancements recently, but still struggle to recognize and interpret intricate details in high-resolution (HR) images effectively. While state-of-the-art (SOTA) MLLMs claim to process images at 4K resolution, existing MLLM benchmarks only support up to 2K, leaving the capabilities of SOTA models on true HR images largely untested. Furthermore, existing methods for enhancing HR image perception in MLLMs rely on computationally expensive visual instruction tuning. To address these limitations, we introduce HR-Bench, the first deliberately designed benchmark to rigorously evaluate MLLM performance on 4K & 8K images. Through extensive experiments, we demonstrate that while downsampling HR images leads to vision information loss, leveraging complementary modalities, e.g., text, can effectively compensate for this loss. Building upon this insight, we propose Divide, Conquer and Combine, a novel training-free framework for enhancing MLLM perception of HR images. Our method follows a three-staged approach: 1) Divide: recursively partitioning the HR image into patches and merging similar patches to minimize computational overhead, 2) Conquer: leveraging the MLLM to generate accurate textual descriptions for each image patch, and 3) Combine: utilizing the generated text descriptions to enhance the MLLM's understanding of the overall HR image. Extensive experiments show that: 1) the SOTA MLLM achieves 63% accuracy, which is markedly lower than the 87% accuracy achieved by humans on HR-Bench; 2) our method brings consistent and significant improvements (a relative increase of +6% on HR-Bench and +8% on general multimodal benchmarks). Liang Ding 0006, Minyan Zeng, Xiabin Zhou, Li Shen 0008, Yong Luo 0002, Wei Yu 0004, Dacheng Tao |
AAAI | 6 |
| 2025 | MTGA: Multi-View Temporal Granularity Aligned Aggregation for Event-Based Lip-ReadingabstractLip-reading is to utilize the visual information of the speaker’s lip movements to recognize words and sentences. Existing event-based lip-reading solutions integrate different frame rate branches to learn spatio-temporal features of varying granularities. However, aggregating events into event frames inevitably leads to the loss of fine-grained temporal information within frames. To remedy this drawback, we propose a novel framework termed Multi-view Temporal Granularity aligned Aggregation (MTGA). Specifically, we first present a novel event representation method, namely time-segmented voxel graph list, where the most significant local voxels are temporally connected into a graph list. Then we design a spatio-temporal fusion module based on temporal granularity alignment, where the global spatial features extracted from event frames, together with the local relative spatial and temporal features contained in voxel graph list are effectively aligned and integrated. Finally, we design a temporal aggregation module that incorporates positional encoding, which enables the capture of local absolute spatial and global temporal information. Experiments demonstrate that our method outperforms both the event-based and video-based lip-reading counterparts. Yong Luo 0002, Wei Yu 0004, Zheng He 0001, Jialie Shen 0001 |
AAAI | 3 |
| 2025 | ELBA-Bench: An Efficient Learning Backdoor Attacks Benchmark for Large Language ModelsabstractXuxu Liu, Siyuan Liang, Mengya Han, Yong Luo, Aishan Liu, Xiantao Cai, Zheng He, Dacheng Tao. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Xuxu Liu, Siyuan Liang 0004, Mengya Han, Yong Luo 0002, Aishan Liu, Xiantao Cai, Zheng He 0001, Dacheng Tao |
ACL (1) | 4 |
| 2025 | Masked Diffusion Models for Unsupervised Anomaly Detection in Brain ImagesabstractUnsupervised anomaly detection has gained significant attention in the field of medical imaging due to its capability of reducing the need for costly pixel-level annotation. To achieve this, existing approaches usually utilize generative models to produce healthy references of the diseased images and then identify the abnormalities by comparing healthy references and original diseased images. However, intrinsic characteristics of brain images, e.g., the low contrast and the intricate anatomical structure, make reconstruction challenging. To address those challenges, we propose a Masked Diffusion Model (MDiff), which incorporates a hierarchical patch partition strategy into the diffusion model for precise reconstruction of detailed content. Aligned with this strategy, we treat the perturbed upper-level patch as masked and introduce a masked modeling mechanism into MDiff's diffusion U-Net. This mechanism operates on sub-level patches, enhancing the model's ability to process contextual information surrounding the perturbed upper-level patch. To further improve the quality of healthy references, we integrate a memory module within the mechanism's encoder to retrieve the most relevant memory items as contextual information, while employing the learnable query embedding in its decoder to prevent the network from learning identical shortcuts. Experiments on tumor and multiple sclerosis lesion data demonstrate MDiff's effectiveness. Rui Xu 0031, Yunke Wang, Yong Luo 0002, Shu Yang 0004, Yihui Wang 0002, Bo Du 0001, Hao Chen 0011 |
BIBM | 4 |
| 2025 | EgoNet: An Unified Egocentric Active Speaker Detection Framework for both Camera Wearer and Visible CandidatesabstractActive Speaker Detection (ASD) aims to determine whether each candidate in a video frame is speaking. The egocentric dataset Ego4D introduces unique challenges for this task, such as dynamic shooting angles that cause candidates to frequently leave the sight, leading to temporal discontinuities. Additionally, Ego4D poses a novel task: detecting the speaking activities of the camera wearer, who never appears in the field of view. Existing methods treat these two tasks separately, and treat candidates out of sight as noise. In contrast, we propose EgoNet, a framework that uniformly models all candidates, including those not visible. By capturing interactions among all candidates and modeling broader temporal context, EgoNet reduces uncertainty and improves performance in egocentric active speaker detection. Yongqian Li, Xin Zhou 0003, Zheng He 0001, Wei Yu 0004, Yong Luo 0002 |
ICASSP | 5 |
| 2025 | CLEST-IQA: Contrastive Learning-Enhanced Swin Transformer for Image Quality Assessment
Yongqian Li, Dixiao Tao, Yong Luo 0002, Xin Zhou 0003, Dehua Cao |
ICIG (3) | 4 |
| 2025 | FIRING-Net: A filtered feature recycling network for speech enhancementabstractCurrent deep neural networks for speech enhancement (SE) aim to minimize the distance between the output signal and the clean target by filtering out noise features from input features. However, when noise and speech components are highly similar, SE models struggle to learn effective discrimination patterns. To address this challenge, we propose a Filter-Recycle-Interguide framework termed Filter-Recycle-INterGuide NETwork (FIRING-Net) for SE, which filters the input features to extract target features and recycles the filtered-out features as non-target features. These two feature sets then guide each other to refine the features, leading to the aggregation of speech information within the target features and noise information within the non-target features. The proposed FIRING-Net mainly consists of a Local Module (LM) and a Global Module (GM). The LM uses outputs of the speech extraction network as target features and the residual between input and output as non-target features. The GM leverages the energy distribution of self-attention map to extract target and non-target features guided by highest and lowest energy regions. Both LM and GM include interaction modules to leverage the two feature sets in an inter-guided manner for collecting speech from non-target features and filtering out noise from target features. Experiments confirm the effectiveness of the Filter-Recycle-Interguide framework, with FIRING-Net achieving a strong balance between SE performance and computational efficiency, surpassing comparable models across various SNR levels and noise environments. Xinmeng Xu, Jizhen Li, Yuhong Yang 0001, Yong Luo 0002, Weiping Tu |
ICLR | 5 |
| 2025 | Enhancing Multimodal Chain-of-Thought Reasoning with Tree-Searched Self-TrainingabstractDespite impressive performance on general visual benchmarks, Multimodal Large Language Models (MLLMs) still struggle with generating consistent and accurate reasoning processes for complex visual reasoning tasks. The limited availability of multimodal reasoning datasets further complicates fine-tuning efforts to improve reasoning capabilities. To address these challenges, we introduce Tree-Searched Self-Training (TSST), a novel framework that enhances multimodal reasoning through self-improvement without relying on extensive manual annotations. TSST introduces a hierarchical tree-search mechanism that combines stepwise rationales generation with value-guided path selection, enabling the model to explore and identify high-quality reasoning trajectories autonomously. Training on these self-generated data significantly improves the model’s reasoning ability, boosting consistency and accuracy. Extensive experiments on ScienceQA and VCR benchmarks show that TSST outperforms existing fine-tuning baselines by 6.0% in accuracy, demonstrating its effectiveness in enhancing multimodal reasoning capabilities. The code is publicly available at: https://github.com/laaambs/tsst. Yiwen Luo, Yong Luo 0002, Zengmao Wang |
ICME | 3 |
| 2025 | Targeted Low-rank Refinement: Enhancing Sparse Language Models with PrecisionabstractPruning is a widely used technique for compressing large neural networks that eliminates weights that have minimal impact on the model's performance. Current pruning methods, exemplified by magnitude pruning, assign an importance score to each weight based on its magnitude and remove weights with scores below a certain threshold. Nonetheless, these methods often create a gap between the original dense and the pruned sparse model, potentially impairing performance. Especially when the sparsity ratio is high, the gap becomes more pronounced. To mitigate this issue, we introduce a method to bridge the gap left by pruning by utilizing a low-rank approximation of the difference between the dense and sparse matrices. Our method entails the iterative refinement of the sparse weight matrix augmented by a low-rank adjustment. This technique captures and retains the essential information often lost during pruning, thereby improving the performance of the pruned model. Furthermore, we offer a comprehensive theoretical analysis of our approach, emphasizing its convergence properties and establishing a solid basis for its efficacy. Experimental results on LLaMa models validate its effectiveness on large language models across various pruning techniques and sparsity levels. Our method shows significant improvements: at 50\% sparsity, it reduces perplexity by 53.9\% compared to conventional magnitude pruning on LLaMa-7B. Furthermore, to achieve a specific performance target, our approach enables an 8.6\% reduction in model parameters while maintaining a sparsity ratio of about 50\%. Li Shen 0008, Anke Tang, Yong Luo 0002, Tao Sun 0005, Han Hu 0003, Xiaochun Cao |
ICML | 3 |
| 2025 | Retrieval-Augmented Perception: High-resolution Image Perception Meets Visual RAGabstractHigh-resolution (HR) image perception remains a key challenge in multimodal large language models (MLLMs). To drive progress beyond the limits of heuristic methods, this paper advances HR perception capabilities of MLLMs by harnessing cutting-edge long-context techniques such as retrieval-augmented generation (RAG). Towards this end, this paper presents the first study exploring the use of RAG to address HR perception challenges. Specifically, we propose Retrieval-Augmented Perception (RAP), a training-free framework that retrieves and fuses relevant image crops while preserving spatial context using the proposed Spatial-Awareness Layout. To accommodate different tasks, the proposed Retrieved-Exploration Search (RE-Search) dynamically selects the optimal number of crops based on model confidence and retrieval scores. Experimental results on HR benchmarks demonstrate the significant effectiveness of RAP, with LLaVA-v1.5-13B achieving a 43% improvement on $V^*$ Bench and 19% on HR-Bench. Code is available at https://github.com/DreamMr/RAP. Yongcheng Jing, Liang Ding 0006, Li Shen 0008, Yong Luo 0002, Bo Du 0001, Dacheng Tao |
ICML | 6 |
| 2025 | Decision Mixer: Integrating Long-term and Local Dependencies via Dynamic Token Selection for Decision-MakingabstractThe Conditional Sequence Modeling (CSM) paradigm, benefiting from the transformer’s powerful distribution modeling capabilities, has demonstrated considerable promise in offline Reinforcement Learning (RL) tasks. Depending on the task’s nature, it is crucial to carefully balance the interplay between inherent local features and long-term dependencies in Markov decision trajectories to mitigate potential performance degradation and unnecessary computational overhead. In this paper, we propose Decision Mixer (DM), which addresses the conflict between features of different scales in the modeling process from the perspective of dynamic integration. Drawing inspiration from conditional computation, we design a plug-and-play dynamic token selection mechanism to ensure the model can effectively allocate attention to different features based on task characteristics. Additionally, we employ an auxiliary predictor to alleviate the short-sightedness issue in the autoregressive sampling process. DM achieves state-of-the-art performance on various standard RL benchmarks while requiring significantly fewer computational resources, offering a viable solution for building efficient and scalable RL foundation models. Code is available at here. Hongling Zheng, Li Shen 0008, Yong Luo 0002, Deheng Ye, Bo Du 0001, Jialie Shen 0001, Dacheng Tao |
ICML | 3 |
| 2025 | Successive Mouse Movement: Dual-Phase Adaptation with Pre-balancing and Test-Time Learning
Jingyi Feng, Yong Luo 0002, Anke Tang, Xuxu Liu |
ICONIP (1) | 2 |
| 2025 | A Dual Stream Visual Tokenizer for LLM Image GenerationabstractWe proposes a novel visual tokenizer by combining high-level semantic tokens and low-level pixel tokens to represent images, aiming to address the challenges of image-to-sequence conversion for Large Language Models (LLMs). Existing visual tokenizers, such as VQ-VAE and diffusion-based models, either struggle with token explosion as image resolution increases or fail to capture detailed structural information. Our method introduces a dual-token system: high-level semantic tokens capture the main content of the image, while low-level pixel tokens preserve structural details. By integrating these tokens in a hybrid architecture, we leverage a VQ-VAE branch to generate low-resolution guidance and a diffusion process to reconstruct high-resolution images with both semantic coherence and structural accuracy. This approach significantly reduces the number of required tokens and enhances image reconstruction quality, offering an efficient solution for tasks like image generation and understanding based on LLMs. Yongqian Li, Yong Luo 0002, Xiantao Cai, Zheng He 0001, Zhennan Meng, Nidong Wang, Yunlin Chen |
IJCAI | 2 |
| 2025 | Open-Vocabulary Fine-Grained Hand Action DetectionabstractIn this work, we address the new challenge of open-vocabulary fine-grained hand action detection, which aims to recognize hand actions from both known and novel categories using textual descriptions. Traditional hand action detection methods are limited to closed-set detection, making it difficult for them to generalize to new, unseen hand action categories. While current open-vocabulary detection (OVD) methods are effective at detecting novel objects, they face challenges with fine-grained action recognition, particularly when data is limited and heterogeneous. This often leads to poor generalization and performance bias between base and novel categories. To address these issues, we propose a novel approach, Open-FGHA (Open-vocabulary Fine-Grained Hand Action), which learns to distinguish fine-grained features across multiple modalities from limited heterogeneous data. It then identifies optimal matching relationships among these features, enabling accurate open-vocabulary fine-grained hand action detection. Specifically, we introduce three key components: Hierarchical Heterogeneous Low-Rank Adaptation, Bidirectional Selection and Fusion Mechanism, and Cross-Modality Query Generator. These components work in unison to enhance the alignment and fusion of multimodal fine-grained features. Extensive experiments demonstrate that Open-FGHA outperforms existing OVD methods, showing its strong potential for open-vocabulary hand action detection. The source code is available at OV-FGHAD. Ting Zhe, Mengya Han, Xiaoshuai Hao, Yong Luo 0002, Zheng He 0001, Xiantao Cai, Jing Zhang 0037 |
IJCAI | 4 |
| 2025 | Anatomy-Conserving Unpaired CBCT-to-CT Translation via Schrödinger Bridge
Song Ouyang, Yong Luo 0002, Kehua Su, Zhiwen Liang, Bo Du 0001 |
MICCAI (4) | 4 |
| 2025 | Robust Active Speaker Detection in Challenging Environments Using GNN-Fused Multi-modal Cues and Body Language
Yongqian Li, Yong Luo 0002, Xin Zhou 0003 |
MMM (3) | 2 |
| 2025 | MixPrompt: Efficient Mixed Prompting for Multimodal Semantic SegmentationabstractRecent advances in multimodal semantic segmentation show that incorporating auxiliary inputs—such as depth or thermal images—can significantly improve performance over single-modality (RGB-only) approaches. However, most existing solutions rely on parallel backbone networks and complex fusion modules, greatly increasing model size and computational demands. Inspired by prompt tuning in large language models, we introduce \textbf{MixPrompt}: a prompting-based framework that integrates auxiliary modalities into a pretrained RGB segmentation model without modifying its architecture. MixPrompt uses a lightweight prompting module to extract and fuse information from auxiliary inputs into the main RGB backbone. This module is initialized using the early layers of a pretrained RGB feature extractor, ensuring a strong starting point. At each backbone layer, MixPrompt aligns RGB and auxiliary features in multiple low-rank subspaces, maximizing information use with minimal parameter overhead. An information mixing scheme enables cross-subspace interaction for further performance gains. During training, only the prompting module and segmentation head are updated, keeping the RGB backbone frozen for parameter efficiency. Experiments across NYU Depth V2, SUN-RGBD, MFNet, and DELIVER datasets show that MixPrompt achieves improvements of 4.3, 1.1, 0.4, and 1.1 mIoU, respectively, over two-branch baselines, while using nearly half the parameters. MixPrompt also outperforms recent prompting-based methods under similar compute budgets. Zhiwei Hao 0001, Zhongyu Xiao, Jianyuan Guo, Li Shen 0008, Yong Luo 0002, Han Hu 0003, Dan Zeng 0001 |
NeurIPS | 5 |
| 2025 | Merging on the Fly Without Retraining: A Sequential Approach to Scalable Continual Model MergingabstractDeep model merging represents an emerging research direction that combines multiple fine-tuned models to harness their specialized capabilities across different tasks and domains. Current model merging techniques focus on merging all available models simultaneously, with weight interpolation-based methods being the predominant approach. However, these conventional approaches are not well-suited for scenarios where models become available sequentially, and they often suffer from high memory requirements and potential interference between tasks. In this study, we propose a training-free projection-based continual merging method that processes models sequentially through orthogonal projections of weight matrices and adaptive scaling mechanisms. Our method operates by projecting new parameter updates onto subspaces orthogonal to existing merged parameter updates while using an adaptive scaling mechanism to maintain stable parameter distances, enabling efficient sequential integration of task-specific knowledge. Our approach maintains constant memory complexity to the number of models, minimizes interference between tasks through orthogonal projections, and retains the performance of previously merged models through adaptive task vector scaling. Extensive experiments on CLIP-ViT models demonstrate that our method achieves a 5-8% average accuracy improvement while maintaining robust performance in different task orderings. Code is publicly available at https://github.com/tanganke/opcm . Anke Tang, Enneng Yang, Li Shen 0008, Yong Luo 0002, Han Hu 0003, Lefei Zhang, Bo Du 0001, Dacheng Tao |
NeurIPS | 4 |
| 2025 | AiDE-Q: Synthetic Labeled Datasets Can Enhance Learning Models for Quantum Property EstimationabstractQuantum many-body problems are central to various scientific disciplines, yet their ground-state properties are intrinsically challenging to estimate. Recent advances in deep learning (DL) offer potential solutions in this field, complementing prior purely classical and quantum approaches. However, existing DL-based models typically assume access to a large-scale and noiseless labeled dataset collected by infinite sampling. This idealization raises fundamental concerns about their practical utility, especially given the limited availability of quantum hardware in the near term. To unleash the power of these DL-based models, we propose AiDE-Q (\underline{a}utomat\underline{i}c \underline{d}ata \underline{e}ngine for \underline{q}uantum property estimation), an effective framework that addresses this challenge by iteratively generating high-quality synthetic labeled datasets. Specifically, AiDE-Q utilizes a confidence-check method to assess the quality of synthetic labels and continuously improves the employed DL models with the identified high-quality synthetic dataset. To verify the effectiveness of AiDE-Q, we conduct extensive numerical simulations on a diverse set of quantum many-body and molecular systems, with up to 50 qubits. The results show that AiDE-Q enhances prediction performance for various reference learning models, with improvements of up to $14.2\\%$. Moreover, we exhibit that a basic supervised learning model integrated with AiDE-Q outperforms advanced reference models, highlighting the importance of a synthetic dataset. Our work paves the way for more efficient and practical applications of DL for quantum property estimation. Xinbiao Wang, Zihan Lou, Kaining Zhang, Yong Luo 0002, Bo Du 0001, Dacheng Tao |
NeurIPS | 6 |
| 2025 | Self-Evolving Pseudo-Rehearsal for Catastrophic Forgetting with Task Similarity in LLMsabstractContinual learning for large language models (LLMs) demands a precise balance between $\textbf{plasticity}$ - the ability to absorb new tasks - and $\textbf{stability}$ - the preservation of previously learned knowledge. Conventional rehearsal methods, which replay stored examples, are limited by long-term data inaccessibility; earlier pseudo-rehearsal methods require additional generation modules, while self-synthesis approaches often generate samples that poorly align with real tasks, suffer from unstable outputs, and ignore task relationships. We present $\textbf{\textit{Self-Evolving Pseudo-Rehearsal for Catastrophic Forgetting with Task Similarity}}(\textbf{SERS})$, a lightweight framework that 1) decouples pseudo-input synthesis from label creation, using semantic masking and template guidance to produce diverse, task-relevant prompts without extra modules; 2) applies label self-evolution, blending base-model priors with fine-tuned outputs to prevent over-specialization; and 3) introduces a dynamic regularizer driven by the Wasserstein distance between task distributions, automatically relaxing or strengthening constraints in proportion to task similarity. Experiments across diverse tasks on different LLMs show that our SERS reduces forgetting by over 2\% points against strong pseudo-rehearsal baselines, by ensuring efficient data utilization and wisely transferring knowledge.
The code will be released at https://github.com/JerryWangJun/LLM_CL_SERS/. Liang Ding 0006, Shuai Wang 0011, Hongyu Li 0004, Yong Luo 0002, Huangxuan Zhao, Han Hu 0003, Bo Du 0001 |
NeurIPS | 5 |
| 2025 | Value-Guided Decision Transformer: A Unified Reinforcement Learning Framework for Online and Offline SettingsabstractThe Conditional Sequence Modeling (CSM) paradigm, benefiting from the transformer's powerful distribution modeling capabilities, has demonstrated considerable promise in Reinforcement Learning (RL) tasks. However, much of the work has focused on applying CSM to single online or offline settings, with the general architecture rarely explored. Additionally, existing methods primarily focus on deterministic trajectory modeling, overlooking the randomness of state transitions and the diversity of future trajectory distributions. Fortunately, value-based methods offer a viable solution for CSM, further bridging the potential gap between offline and online RL. In this paper, we propose Value-Guided Decision Transformer (VDT), which leverages value functions to perform advantage-weighting and behavior regularization on the Decision Transformer (DT), guiding the policy toward upper-bound optimal decisions during the offline training phase. In the online tuning phase, VDT further integrates value-based policy improvement with behavior cloning under the CSM architecture through limited interaction and data collection, achieving performance improvement within minimal timesteps. The predictive capability of value functions for future returns is also incorporated into the sampling process. Our method achieves competitive performance on various standard RL benchmarks, providing a feasible solution for developing CSM architectures in general scenarios. Code is available at here. Hongling Zheng, Li Shen 0008, Yong Luo 0002, Deheng Ye, Shuhan Xu, Bo Du 0001, Jialie Shen 0001, Dacheng Tao |
NeurIPS | 3 |
| 2025 | M3Site: multiclass multimodal learning for protein active site identification and classificationabstractAccurately identifying and classifying protein active sites is crucial for understanding protein mechanisms, drug design, and synthetic biology. Current methods often rely on binary classification and single-modal data, limiting their scope. To address these limitations, we propose M$^{3}$Site, a multimodal framework that integrates protein sequence embeddings, structural graph representations, and functional text annotations for residue-level, multiclass active site prediction. Built upon a curated dataset of 25 883 proteins sourced from UniProt and AlphaFold2, M$^{3}$Site leverages pretrained protein language models, equivariant graph neural networks, and biomedical language models for feature extraction. The function informed cross-attention module enables cross-modal feature fusion, while the adaptive weighted fusion mechanism balances modality contributions. A compound loss function tackles class imbalance, ensuring robust performance. Experimental results show M$^{3}$Site significantly outperforms existing models, and an interactive application has been developed to enhance its practical utility for predictions and visualizations. The dataset, source code for experiments, and interactive application are publicly available at https://github.com/Gift-OYS/M3Site. Song Ouyang, Yong Luo 0002, Huiyu Cai, Kehua Su, Na Zhan, Huangxuan Zhao, Tailang Yin, Dongjing Shan |
Briefings Bioinform. | 2 |
| 2025 | DM-PCL: Text-Driven Dual-Modal Prototype Consistency Learning for Weakly-Supervised Few-Shot Part Segmentation
Mengya Han, Yong Luo 0002, Han Hu 0003, Zengmao Wang, Lefei Zhang, Bo Du 0001, Ling-Yu Duan, Dacheng Tao |
Int. J. Comput. Vis. | 2 |
| 2025 | ADEM-VL: Adaptive and Embedded Fusion for Efficient Vision-Language Tuning
Zhiwei Hao 0001, Jianyuan Guo, Li Shen 0008, Yong Luo 0002, Han Hu 0003, Yonggang Wen 0001 |
Int. J. Comput. Vis. | 4 |
| 2025 | Data-Adaptive Weight-Ensembling for Multi-task Model Fusion
Anke Tang, Li Shen 0008, Yong Luo 0002, Shiwei Liu 0003, Han Hu 0003, Bo Du 0001, Dacheng Tao |
Int. J. Comput. Vis. | 3 |
| 2025 | DoseNet: Dose-adaptive prediction of the parotid glands deformation for radiotherapy planning
Bohan Yang 0008, Yong Luo 0002, Bo Du 0001, Dongjing Shan, Chuan Cheng, Jingnan Liu |
Image Vis. Comput. | 2 |
| 2025 | FusionBench: A Unified Library and Comprehensive Benchmark for Deep Model FusionabstractDeep model fusion is an emerging technique that unifies the predictions or parameters of several deep neural networks into a single better-performing model in a cost-effective and data-efficient manner. Although a variety of deep model fusion techniques have been introduced, their evaluations tend to be inconsistent and often inadequate to validate their effectiveness and robustness. We present FusionBench, the first benchmark and a unified library designed specifically for deep model fusion. Our benchmark consists of multiple tasks, each with different settings of models and datasets. This variety allows us to compare fusion methods across different scenarios and model scales. Additionally, FusionBench serves as a unified library for easy implementation and testing of new fusion techniques. FusionBench is open source and actively maintained, with community contributions encouraged. Anke Tang, Li Shen 0008, Yong Luo 0002, Enneng Yang, Han Hu 0003, Lefei Zhang, Bo Du 0001, Dacheng Tao |
J. Mach. Learn. Res. | 3 |
| 2025 | ViF-SD2E: a robust weakly-supervised framework for neural decoding
Jingyi Feng, Yong Luo 0002, Han Hu 0003 |
Neural Comput. Appl. | 2 |
| 2025 | Aligning Text-to-Image Diffusion Models With Constrained Reinforcement LearningabstractReward finetuning has emerged as a powerful technique for aligning diffusion models with specific downstream objectives or user preferences. However, current approaches suffer from a persistent challenge of reward overoptimization, where models exploit imperfect reward feedback at the expense of overall performance. In this work, we identify three key contributors to overoptimization: (1) a granularity mismatch between the multi-step diffusion process and sparse rewards; (2) a loss of plasticity that limits the model's ability to adapt and generalize; and (3) an overly narrow focus on a single reward objective that neglects complementary performance criteria. Accordingly, we introduce Constrained Diffusion Policy Optimization (CDPO), a novel reinforcement learning framework that addresses reward overoptimization from multiple angles. Firstly, CDPO tackles the granularity mismatch through a temporal policy optimization strategy that delivers step-specific rewards throughout the entire diffusion trajectory, thereby reducing the risk of overfitting to sparse final-step rewards. Then we incorporate a neuron reset strategy that selectively resets overactive neurons in the model, preventing overoptimization induced by plasticity loss. Finally, to avoid overfitting to a narrow reward objective, we integrate constrained reinforcement learning with auxiliary reward objectives serving as explicit constraints, ensuring a balanced optimization across diverse performance metrics. Ziyi Zhang 0001, Sen Zhang 0006, Li Shen 0008, Yibing Zhan, Yong Luo 0002, Han Hu 0003, Bo Du 0001, Yonggang Wen 0001, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | PartSeg: Few-shot part segmentation via part-aware prompt learning
Mengya Han, Heliang Zheng, Yong Luo 0002, Han Hu 0003, Jing Zhang 0037, Bo Du 0001 |
Pattern Recognit. | 4 |
| 2025 | CoFormer: Collaborating With Heterogeneous Edge Devices for Scalable Transformer InferenceabstractThe impressive performance of transformer models has sparked the deployment of intelligent applications on resource-constrained edge devices. However, ensuring high-quality service for real-time edge systems is a significant challenge due to the considerable computational demands and resource requirements of these models. Existing strategies typically either offload transformer computations to other devices or directly deploy compressed models on individual edge devices. These strategies, however, result in either considerable communication overhead or suboptimal trade-offs between accuracy and efficiency. To tackle these challenges, we propose a collaborative inference system for general transformer models, termed CoFormer. The central idea behind CoFormer is to exploit the divisibility and integrability of transformer. An off-the-shelf large transformer can be decomposed into multiple smaller models for distributed inference, and their intermediate results are aggregated to generate the final output. We formulate an optimization problem to minimize both inference latency and accuracy degradation under heterogeneous hardware constraints. DeBo algorithm is proposed to first solve the optimization problem to derive the decomposition policy, and then progressively calibrate decomposed models to restore performance. We demonstrate the capability to support a wide range of transformer models on heterogeneous edge devices, achieving up to 3.1× inference speedup with large transformer models. Notably, CoFormer enables the efficient inference of GPT2-XL with 1.6 billion parameters on edge devices, reducing memory requirements by 76.3%. CoFormer can also reduce energy consumption by approximately 40% while maintaining satisfactory inference performance. Guanyu Xu, Zhiwei Hao 0001, Li Shen 0008, Yong Luo 0002, Fuhui Sun, Han Hu 0003, Yonggang Wen 0001 |
IEEE Trans. Computers | 4 |
| 2025 | Fuzzy-Assisted Contrastive Decoding Improving Code Generation of Large Language ModelsabstractLarge Language Models (LLMs) play a crucial role in intelligent code generation tasks. Most existing work focuses on pre-training or fine-tuning specialized code LLMs, e.g., CodeLlama. However, pre-training or fine-tuning a code LLM requires a vast corpus of data, significant computational resources, and considerable human effort. Compared to pre-training or fine-tuning LLMs, a simple and flexible method of contrastive decoding has garnered widespread attention to improve the text generation quality of LLMs. While contrastive decoding can indeed improve the text generation quality of LLMs, our research has found that directly using contrastive decoding: 1) introduces erroneous information into the logit distribution generated from normal prompts (i.e., user's input), particularly in the code generation of LLMs; 2) significantly impedes the inference and decoding time of LLMs. In this work, the limitations of using contrastive decoding directly are systematically highlighted, and a novel real-time fuzzy-assisted contrastive decoding (FCD) mechanism is proposed to improve the code generation quality of LLMs. The proposed FCD mechanism initially categorises prompts into high-quality and low-quality groups based on the results of the evaluator (i.e., unit test) before integrating the LLM. Next, feature values (e.g., standard deviation, peak value, etc.) related to the logit distribution of predicted tokens during the LLM's inference process for both high-quality and low-quality prompts are extracted. Finally, the extracted feature values are used to train the fuzzy neural network (i.e, fuzzy min-max neural network) offline, allowing for the prejudgement of the reliability of the logit distribution for normal prompt outputs. This prevents the direct use of erroneous information from contrastive decoding and improves the code generation quality of LLMs. Through extensive experiments, it has been demonstrated that the proposed FCD mechanism can significantly improve the code generation quality of LLMs through fuzzy-assisted contrastive decoding. Moreover, the FCD mechanism can also reduce the time required for inference and contrastive decoding. The code and data are publicly available on GitHub11https://github.com/LLMcodegen/Fuzzy_contrastive_decoding.and HuggingFace22https://huggingface.co/wangle123/Fuzzy_contrastive_decoding.. Shuai Wang 0011, Liang Ding 0006, Yibing Zhan, Yong Luo 0002, Shuai Liu 0002, Weiping Ding 0001 |
IEEE Trans. Fuzzy Syst. | 4 |
| 2025 | MHR: A Multi-Modal Hyperbolic Representation Framework for Fake News DetectionabstractThe rapid growth of the internet has led to an alarming increase in the dissemination of fake news, which has had many negative effects on society. Various methods have been proposed for detecting fake news. However, these approaches suffer from several limitations. First, most existing works only consider news as separate entities and do not consider the correlations between fake news and real news. Moreover, these works are usually conducted in the Euclidean space, which is unable to capture complex relationships between news, in particular the hierarchical relationships. To tackle these issues, we introduce a novelMulti-modalHyperbolicRepresentation framework (MHR) for fake news detection. Specifically, we capture the correlations between news for graph construction to arrange and analyze different news. To fully utilize the multi-modal characteristics, we first extract the textual and visual information, and then design a Lorentzian multi-modal fusion module to fuse them as the node information in the graph. By utilizing the fully hyperbolic graph neural networks, we learn the graph’s representation in hyperbolic space, followed by a detector for detecting fake news. The experimental results on three real-world datasets demonstrate that our proposed MHR model achieves state-of-the-art performance, indicating the benefits of hyperbolic representation. Shanshan Feng 0001, Guoxin Yu, Han Hu 0003, Yong Luo 0002, Yew-Soon Ong |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2025 | P3ID: A Privacy-Preserving Person Identification Framework Towards Multi-Environments Based on Transfer LearningabstractConcerns surrounding privacy leakages caused by prevalent vision-based person identifications are countless. A promising privacy-preserving solution is to identify the wireless signals reflecting persons, which, however, faces a major challenge of losing efficacy in multi-environments. In this paper, we work on person identification based on wireless signals using transfer learning, toward tackling the performance deterioration across environments. We investigate the feature variations induced by environmental shifts based on data measurements. Lay our foundation on the feature alignment concept, we propose a novel wireless-based person identification framework using transfer learning. In the framework, we integrate a series of signal processing methods including signal selection, pre-processing, and augmentation, where the first includes a reference environment to assist the feature extraction while the latter two respectively reduce the data noise and improve the data diversity. We also propose a model generalization method where a neural network is employed to align features from different environments, which facilitates the extraction of environment-independent features while incorporating both person and environment information. On a real wireless testbed consisting of an Impulse Radio Ultra-WideBand (IR-UWB) radar, we build and publicly release a dataset with 22,264 samples of ten individuals from three environments, varying in testing distance and obstruction condition. Extensive experimental evaluations demonstrate that the proposed framework can improve the identification accuracy across environments, and surpasses state-of-the-art methods by up to 18.06%. Hanxiang He, Xintao Huan, Jing Wang 0055, Yong Luo 0002, Han Hu 0003, Jianping An |
IEEE Trans. Mob. Comput. | 4 |
| 2025 | Sequential Federated Learning in Hierarchical Architecture on Non-IID DatasetsabstractIn a real federated learning (FL) system, communication overhead for passing model parameters between the clients and the parameter server (PS) is often a bottleneck. Hierarchical federated learning (HFL) that poses multiple edge servers (ESs) between clients and the PS can partially alleviate communication pressure but still needs the aggregation of model parameters from multiple ESs at the PS. To further reduce communication overhead, we remove the central PS, so that each iteration only completes model training by transmitting the global model between two adjacent ES. We call this serial learning method Sequential FL (SFL). For the first time, we introduced SFL into HFL and proposed a novel algorithm adapted to this combined framework, called Fed-CHS. Convergence results are derived for strongly convex and non-convex loss functions under various data heterogeneity setups, which show comparable convergence performance with the algorithms for HFL or SFL solely. Experimental results provide evidence of the superiority of our proposed Fed-CHS on both communication overhead saving and test accuracy over baseline methods. Xingrun Yan, Shiyuan Zuo, Rongfei Fan, Han Hu 0003, Li Shen 0008, Puning Zhao, Yong Luo 0002 |
IEEE Trans. Mob. Comput. | 7 |
| 2025 | Federated Learning With Only Positive Labels by Exploring Label CorrelationsabstractFederated learning (FL) aims to collaboratively learn a model by using the data from multiple users under privacy constraints. In this article, we study the multilabel classification (MLC) problem under the FL setting, where trivial solution and extremely poor performance may be obtained, especially when only positive data with respect to a single class label is provided for each client. This issue can be addressed by adding a specially designed regularizer on the server side. Although effective sometimes, the label correlations are simply ignored and thus suboptimal performance may be obtained. Besides, it is expensive and unsafe to exchange user's private embeddings between server and clients frequently, especially when training model in the contrastive way. To remedy these drawbacks, we propose a novel and generic method termed federated averaging (FedAvg) by exploring label correlations (FedALCs). Specifically, FedALC estimates the label correlations in the class embedding learning for different label pairs and utilizes it to improve the model training. To further improve the safety and also reduce the communication overhead, we propose a variant to learn fixed class embedding for each client, so that the server and clients only need to exchange class embeddings once. Extensive experiments on multiple popular datasets demonstrate that our FedALC can significantly outperform the existing counterparts. Xuming An 0001, Dui Wang, Li Shen 0008, Yong Luo 0002, Han Hu 0003, Bo Du 0001, Yonggang Wen 0001, Dacheng Tao |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Decomposing Semantic Shifts for Composed Image RetrievalabstractComposed image retrieval is a type of image retrieval task where the user provides a reference image as a starting point and specifies a text on how to shift from the starting point to the desired target image. However, most existing methods focus on the composition learning of text and reference images and oversimplify the text as a description, neglecting the inherent structure and the user's shifting intention of the texts. As a result, these methods typically take shortcuts that disregard the visual cue of the reference images. To address this issue, we reconsider the text as instructions and propose a Semantic Shift Network (SSN) that explicitly decomposes the semantic shifts into two steps: from the reference image to the visual prototype and from the visual prototype to the target image. Specifically, SSN explicitly decomposes the instructions into two components: degradation and upgradation, where the degradation is used to picture the visual prototype from the reference image, while the upgradation is used to enrich the visual prototype into the final representations to retrieve the desired target image. The experimental results show that the proposed SSN demonstrates a significant improvement of 5.42% and 1.37% on the CIRR and FashionIQ datasets, respectively, and establishes a new state-of-the-art performance. The code is available at https://github.com/starxing-yuu/SSN. Daqing Liu, Yong Luo 0002, Jing Zhang 0037 |
AAAI | 4 |
| 2024 | Cycle Self-Refinement for Multi-Source Domain AdaptationabstractMulti-source domain adaptation (MSDA) aims to transfer knowledge from multiple source domains to the unlabeled target domain. In this paper, we propose a cycle self-refinement domain adaptation method, which progressively attempts to learn the dominant transferable knowledge in each source domain in a cycle manner. Specifically, several source-specific networks and a domain-ensemble network are adopted in the proposed method. The source-specific networks are adopted to provide the dominant transferable knowledge in each source domain for instance-level ensemble on predictions of the samples in target domain. Then these samples with high-confidence ensemble predictions are adopted to refine the domain-ensemble network. Meanwhile, to guide each source-specific network to learn more dominant transferable knowledge, we force the features of the target domain from the domain-ensemble network and the features of each source domain from the corresponding source-specific network to be aligned with their predictions from the corresponding networks. Thus the adaptation ability of source-specific networks and the domain-ensemble network can be improved progressively. Extensive experiments on Office-31, Office-Home and DomainNet show that the proposed method outperforms the state-of-the-art methods for most tasks. Chaoyang Zhou, Zengmao Wang, Bo Du 0001, Yong Luo 0002 |
AAAI | 4 |
| 2024 | Improving Generalized Zero-Shot Learning by Exploring the Diverse Semantics from External Class NamesabstractGeneralized Zero-Shot Learning (GZSL) methods often assume that the unseen classes are similar to seen classes, and thus perform poor when unseen classes are dissimilar to seen classes. Although some existing GZSL approaches can alleviate this issue by leveraging additional semantic information from test unseen classes, their generalization ability to dissimilar unseen classes is still unsatisfactory. This motivates us to study GZSL in the more practical setting, where unseen classes can be either similar or dissimilar to seen classes. In this paper, we propose a simple yet effective GZSL framework by exploring diverse semantics from external class names (DSECN), which is simultaneously robust on the similar and dissimilar unseen classes. This is achieved by introducing diverse semantics from external class names and aligning the introduced semantics to visual space using the classification head of pretrained network. Furthermore, we show that the design idea of DSECN can easily be integrate into other advanced GZSL approaches, such as the generative-based ones, and enhance their robustness for dissimilar unseen classes. Extensive experiments in the practical setting including both similar and dissimilar unseen classes show that our method significantly outperforms the state-of-the-art approaches on all datasets and can be trained very efficiently. Yong Luo 0002, Zengmao Wang, Bo Du 0001 |
CVPR | 2 |
| 2024 | EDPS-SST: Enhanced Dynamic Path Stitching with Structural Similarity Thresholding for Large-Scale Medical Image Stitching Under Sparse Pixel Overlap
Zhuan Han, Dixiao Tao, Bohan Yang 0015, Yong Luo 0002, Dehua Cao, Xin Zhou 0003 |
ICANN (8) | 4 |
| 2024 | SCST: Spatial Consistent Swin Transformer for Multi-focus Biomedical Microscopic Image Fusion
Dengpan Liu, Bohan Yang 0015, Yong Luo 0002, Dehua Cao, Xin Zhou 0003 |
ICANN (8) | 4 |
| 2024 | Depression Diagnosis and Analysis via Multimodal Multi-order Factor Fusion
Chengbo Yuan, Xuxu Liu, Yongqian Li, Yong Luo 0002, Xin Zhou 0003 |
ICANN (8) | 5 |
| 2024 | Parameter-Efficient Multi-Task Model Fusion with Partial LinearizationabstractLarge pre-trained models have enabled significant advances in machine learning and served as foundation components.
Model fusion methods, such as task arithmetic, have been proven to be powerful and scalable to incorporate fine-tuned weights from different tasks into a multi-task model.
However, efficiently fine-tuning large pre-trained models on multiple downstream tasks remains challenging, leading to inefficient multi-task model fusion.
In this work, we propose a novel method to improve multi-task fusion for parameter-efficient fine-tuning techniques like LoRA fine-tuning.
Specifically, our approach partially linearizes only the adapter modules and applies task arithmetic over the linearized adapters.
This allows us to leverage the the advantages of model fusion over linearized fine-tuning, while still performing fine-tuning and inference efficiently.
We demonstrate that our partial linearization technique enables a more effective fusion of multiple tasks into a single model, outperforming standard adapter tuning and task arithmetic alone.
Experimental results demonstrate the capabilities of our proposed partial linearization technique to effectively construct unified multi-task models via the fusion of fine-tuned task vectors.
We evaluate performance over an increasing number of tasks and find that our approach outperforms standard parameter-efficient fine-tuning techniques. The results highlight the benefits of partial linearization for scalable and efficient multi-task model fusion. Anke Tang, Li Shen 0008, Yong Luo 0002, Yibing Zhan, Han Hu 0003, Bo Du 0001, Yixin Chen 0001, Dacheng Tao |
ICLR | 3 |
| 2024 | Training-Free Robust Neural Network Search Via PruningabstractAdversarial examples widely exist in visual tasks, which look almost the same as normal images but can induce neural networks to make completely wrong predictions. This exposes significant risks to deploying deep learning to visual tasks. The robustness of neural networks to adversarial images is shaped by multiple factors entangled together, including adversarial training strategies, model architectures, hyper-parameters, and optimization algorithms. This makes searching for robust architectures a significantly complex task. In this paper, we propose a Robust Training-free Pruning neural architecture search (RTP-NAS) framework that endeavors to disentangle the influence of architectures from other factors. We adopt Universal Adversarial Perturbations (UAPs) to construct a transferable adversarial data space across different network architectures. RTP-NAS further employs two adversarial training-free indicators, the condition number of the neural tangent kernel (NTK) and the number of linear regions calculated on the adversarial data space, to measure the intrinsic adversarial robustness of candidate architectures. Comprehensive experiments show that the proposed method achieves superior performance to other existing approaches, with a significantly lower computational cost. The code will be released publicly. Qiancheng Yang, Yong Luo 0002, Bo Du 0001 |
ICME | 2 |
| 2024 | Merging Multi-Task Models via Weight-Ensembling Mixture of ExpertsabstractMerging various task-specific Transformer-based vision models trained on different tasks into a single unified model can execute all the tasks concurrently. Previous methods, exemplified by task arithmetic, have been proven to be both effective and scalable. Existing methods have primarily focused on seeking a static optimal solution within the original model parameter space. A notable challenge is mitigating the interference between parameters of different models, which can substantially deteriorate performance. In this paper, we propose to merge most of the parameters while upscaling the MLP of the Transformer layers to a weight-ensembling mixture of experts (MoE) module, which can dynamically integrate shared and task-specific knowledge based on the input, thereby providing a more flexible solution that can adapt to the specific needs of each instance. Our key insight is that by identifying and separating shared knowledge and task-specific knowledge, and then dynamically integrating them, we can mitigate the parameter interference problem to a great extent. We conduct the conventional multi-task model merging experiments and evaluate the generalization and robustness of our method. The results demonstrate the effectiveness of our method and provide a comprehensive understanding of our method. The code is available at https://github.com/tanganke/weight-ensembling_MoE Anke Tang, Li Shen 0008, Yong Luo 0002, Lefei Zhang, Dacheng Tao |
ICML | 3 |
| 2024 | Confronting Reward Overoptimization for Diffusion Models: A Perspective of Inductive and Primacy BiasesabstractBridging the gap between diffusion models and human preferences is crucial for their integration into practical generative workflows. While optimizing downstream reward models has emerged as a promising alignment strategy, concerns arise regarding the risk of excessive optimization with learned reward models, which potentially compromises ground-truth performance. In this work, we confront the reward overoptimization problem in diffusion model alignment through the lenses of both inductive and primacy biases. We first identify a mismatch between current methods and the temporal inductive bias inherent in the multi-step denoising process of diffusion models, as a potential source of reward overoptimization. Then, we surprisingly discover that dormant neurons in our critic model act as a regularization against reward overoptimization while active neurons reflect primacy bias. Motivated by these observations, we propose Temporal Diffusion Policy Optimization with critic active neuron Reset (TDPO-R), a policy gradient algorithm that exploits the temporal inductive bias of diffusion models and mitigates the primacy bias stemming from active neurons. Empirical results demonstrate the superior efficacy of our methods in mitigating reward overoptimization. Code is avaliable at https://github.com/ZiyiZhang27/tdpo. Ziyi Zhang 0001, Sen Zhang 0006, Yibing Zhan, Yong Luo 0002, Yonggang Wen 0001, Dacheng Tao |
ICML | 4 |
| 2024 | Eliminating the Cross-Domain Misalignment in Text-guided Image Inpainting
Muqi Huang, Yong Luo 0002, Lefei Zhang |
IJCAI | 3 |
| 2024 | Zero-shot Learning for Preclinical Drug Screening
Kun Li 0009, Weiwei Liu 0003, Yong Luo 0002, Xiantao Cai, Jia Wu 0001, Wenbin Hu 0001 |
IJCAI | 3 |
| 2024 | Joint Input and Output Coordination for Class-Incremental Learning
Shuai Wang 0011, Yibing Zhan, Yong Luo 0002, Han Hu 0003, Wei Yu 0004, Yonggang Wen 0001, Dacheng Tao |
IJCAI | 3 |
| 2024 | PrimKD: Primary Modality Guided Multimodal Fusion for RGB-D Semantic SegmentationabstractThe recent advancements in cross-modal transformers have demonstrated their superior performance in RGB-D segmentation tasks by effectively integrating information from both RGB and depth modalities. However, existing methods often overlook the varying levels of informative content present in each modality, treating them equally and using models of the same architecture. This oversight can potentially hinder segmentation performance, especially considering that RGB images typically contain significantly more information than depth images. To address this issue, we propose PrimKD, a knowledge distillation based approach that focuses on guided multimodal fusion, with an emphasis on leveraging the primary RGB modality. In our approach, we utilize a model trained exclusively on the RGB modality as the teacher, guiding the learning process of a student model that fuses both RGB and depth modalities. To prioritize information from the primary RGB modality while leveraging the depth modality, we incorporate primary focused feature reconstruction and a selective alignment scheme. This integration enhances the overall freature fusion, resulting in improved segmentation results. We evaluate our proposed method on the NYU Depth V2 and SUN-RGBD datasets, and the experimental results demonstrate the effectiveness of PrimKD. Specifically, our approach achieves mIoU scores of 57.8 and 52.5 on these two datasets, respectively, surpassing existing counterparts by 1.5 and 0.4 mIoU. The code is available at https://github.com/xiaoshideta/PrimKD. Zhiwei Hao 0001, Zhongyu Xiao, Yong Luo 0002, Jianyuan Guo, Jing Wang 0055, Li Shen 0008, Han Hu 0003 |
ACM Multimedia | 3 |
| 2024 | WisdoM: Improving Multimodal Sentiment Analysis by Fusing Contextual World Knowledge
Liang Ding 0006, Li Shen 0008, Yong Luo 0002, Han Hu 0003, Dacheng Tao |
ACM Multimedia | 4 |
| 2024 | Multi-Granularity Hand Action DetectionabstractDetecting hand actions in videos is crucial for understanding video content and has diverse real-world applications. Existing approaches often focus on whole-body actions or coarse-grained action categories, lacking fine-grained hand-action localization information. To fill this gap, we introduce the FHA-Kitchens (Fine-Grained Hand Actions in Kitchen Scenes) dataset, providing both coarse- and fine-grained hand action categories along with localization annotations. This dataset comprises 2,377 video clips and 30,047 frames, annotated with approximately 200k bounding boxes and 880 action categories. Evaluation of existing action detection methods on FHA-Kitchens reveals varying generalization capabilities across different granularities. To handle multi-granularity in hand actions, we propose MG-HAD, an End-to-End Multi-Granularity Hand Action Detection method. It incorporates two new designs: Multi-dimensional Action Queries and Coarse-Fine Contrastive Denoising. Extensive experiments demonstrate MG-HAD's effectiveness for multi-granularity hand action detection, highlighting the significance of FHA-Kitchens for future research and real-world applications. The dataset and source code are available at MG-HAD. Ting Zhe, Jing Zhang 0037, Yongqian Li, Yong Luo 0002, Han Hu 0003, Dacheng Tao |
ACM Multimedia | 4 |
| 2024 | MMSite: A Multi-modal Framework for the Identification of Active Sites in ProteinsabstractThe accurate identification of active sites in proteins is essential for the advancement of life sciences and pharmaceutical development, as these sites are of critical importance for enzyme activity and drug design. Recent advancements in protein language models (PLMs), trained on extensive datasets of amino acid sequences, have significantly improved our understanding of proteins. However, compared to the abundant protein sequence data, functional annotations, especially precise per-residue annotations, are scarce, which limits the performance of PLMs. On the other hand, textual descriptions of proteins, which could be annotated by human experts or a pretrained protein sequence-to-text model, provide meaningful context that could assist in the functional annotations, such as the localization of active sites. This motivates us to construct a $\textbf{ProT}$ein-$\textbf{A}$ttribute text $\textbf{D}$ataset ($\textbf{ProTAD}$), comprising over 570,000 pairs of protein sequences and multi-attribute textual descriptions. Based on this dataset, we propose $\textbf{MMSite}$, a multi-modal framework that improves the performance of PLMs to identify active sites by leveraging biomedical language models (BLMs). In particular, we incorporate manual prompting and design a MACross module to deal with the multi-attribute characteristics of textual descriptions. MMSite is a two-stage ("First Align, Then Fuse") framework: first aligns the textual modality with the sequential modality through soft-label alignment, and then identifies active sites via multi-modal fusion. Experimental results demonstrate that MMSite achieves state-of-the-art performance compared to existing protein representation learning methods. The dataset and code implementation are available at https://github.com/Gift-OYS/MMSite. Song Ouyang, Huiyu Cai, Yong Luo 0002, Kehua Su, Lefei Zhang, Bo Du 0001 |
NeurIPS | 3 |
| 2024 | MG-Net: Learn to Customize QAOA with Circuit Depth AwarenessabstractQuantum Approximate Optimization Algorithm (QAOA) and its variants exhibit immense potential in tackling combinatorial optimization challenges. However, their practical realization confronts a dilemma: the requisite circuit depth for satisfactory performance is problem-specific and often exceeds the maximum capability of current quantum devices. To address this dilemma, here we first analyze the convergence behavior of QAOA, uncovering the origins of this dilemma and elucidating the intricate relationship between the employed mixer Hamiltonian, the specific problem at hand, and the permissible maximum circuit depth. Harnessing this understanding, we introduce the Mixer Generator Network (MG-Net), a unified deep learning framework adept at dynamically formulating optimal mixer Hamiltonians tailored to distinct tasks and circuit depths. Systematic simulations, encompassing Ising models and weighted Max-Cut instances with up to 64 qubits, substantiate our theoretical findings, highlighting MG-Net's superior performance in terms of both approximation ratio and efficiency. Xinbiao Wang, Yong Luo 0002, Dacheng Tao |
NeurIPS | 4 |
| 2024 | Decomposed Prompt Decision Transformer for Efficient Unseen Task GeneralizationabstractMulti-task offline reinforcement learning aims to develop a unified policy for diverse tasks without requiring real-time interaction with the environment. Recent work explores sequence modeling, leveraging the scalability of the transformer architecture as a foundation for multi-task learning. Given the variations in task content and complexity, formulating policies becomes a challenging endeavor, requiring careful parameter sharing and adept management of conflicting gradients to extract rich cross-task knowledge from multiple tasks and transfer it to unseen tasks. In this paper, we propose the Decomposed Prompt Decision Transformer (DPDT) that adopts a two-stage paradigm to efficiently learn prompts for unseen tasks in a parameter-efficient manner. We incorporate parameters from pre-trained language models (PLMs) to initialize DPDT, thereby providing rich prior knowledge encoded in language models. During the decomposed prompt tuning phase, we learn both cross-task and task-specific prompts on training tasks to achieve prompt decomposition. In the test time adaptation phase, the cross-task prompt, serving as a good initialization, were further optimized on unseen tasks through test time adaptation, enhancing the model's performance on these tasks. Empirical evaluation on a series of Meta-RL benchmarks demonstrates the superiority of our approach. The project is available at https://github.com/ruthless-man/DPDT. Hongling Zheng, Li Shen 0008, Yong Luo 0002, Tongliang Liu, Jialie Shen 0001, Dacheng Tao |
NeurIPS | 3 |
| 2024 | HPFL: hyper-network guided personalized federated learning for multi-center tuberculosis chest x-ray diagnosis
Chang Liu 0046, Yong Luo 0002, Yongchao Xu, Bo Du 0001 |
Multim. Tools Appl. | 2 |
| 2024 | SGDA: Towards 3-D Universal Pulmonary Nodule Detection via Slice Grouped Domain AttentionabstractLung cancer is the leading cause of cancer death worldwide. The best solution for lung cancer is to diagnose the pulmonary nodules in the early stage, which is usually accomplished with the aid of thoracic computed tomography (CT). As deep learning thrives, convolutional neural networks (CNNs) have been introduced into pulmonary nodule detection to help doctors in this labor-intensive task and demonstrated to be very effective. However, the current pulmonary nodule detection methods are usually domain-specific, and cannot satisfy the requirement of working in diverse real-world scenarios. To address this issue, we propose a slice grouped domain attention (SGDA) module to enhance the generalization capability of the pulmonary nodule detection networks. This attention module works in the axial, coronal, and sagittal directions. In each direction, we divide the input feature into groups, and for each group, we utilize a universal adapter bank to capture the feature subspaces of the domains spanned by all pulmonary nodule datasets. Then the bank outputs are combined from the perspective of domain to modulate the input group. Extensive experiments demonstrate that SGDA enables substantially better multi-domain pulmonary nodule detection performance compared with the state-of-the-art multi-domain learning methods. Rui Xu 0031, Zhi Liu 0002, Yong Luo 0002, Han Hu 0003, Li Shen 0008, Bo Du 0001, Kaiming Kuang, Jiancheng Yang |
IEEE Trans. Comput. Biol. Bioinform. | 3 |
| 2024 | MambaHSI: Spatial-Spectral Mamba for Hyperspectral Image ClassificationabstractTransformer has been extensively explored for hyperspectral image (HSI) classification. However, transformer poses challenges in terms of speed and memory usage because of its quadratic computational complexity. Recently, the Mamba model has emerged as a promising approach, which has strong long-distance modeling capabilities while maintaining a linear computational complexity. However, representing the HSI is challenging for the Mamba due to the requirement for an integrated spatial and spectral understanding. To remedy these drawbacks, we propose a novel HSI classification model based on a Mamba model, named MambaHSI, which can simultaneously model long-range interaction of the whole image and integrate spatial and spectral information in an adaptive manner. Specifically, we design a spatial Mamba block (SpaMB) to model the long-range interaction of the whole image at the pixel-level. Then, we propose a spectral Mamba block (SpeMB) to split the spectral vector into multiple groups, mine the relations across different spectral groups, and extract spectral features. Finally, we propose a spatial-spectral fusion module (SSFM) to adaptively integrate spatial and spectral features of a HSI. To our best knowledge, this is the first image-level HSI classification model based on the Mamba. We conduct extensive experiments on four diverse HSI datasets. The results demonstrate the effectiveness and superiority of the proposed model for HSI classification. This reveals the great potential of Mamba to be the next-generation backbone for HSI models. Codes are available athttps://github.com/li-yapeng/MambaHSI. Yong Luo 0002, Lefei Zhang, Zengmao Wang, Bo Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | UVaT: Uncertainty Incorporated View-Aware Transformer for Robust Multi-View ClassificationabstractExisting multi-view classification algorithms usually assume that all examples have observations on all views, and the data in different views are clean. However, in real-world applications, we are often provided with data that have missing representations or contain noise on some views (i.e., missing or noise views). This may lead to significant performance degeneration, and thus many algorithms are proposed to address the incomplete view or noisy view issues. However, most of existing algorithms deal with the two issues separately, and hence may fail when both missing and noisy views exist. They are also usually not flexible in that the view or feature significance cannot be adaptively identified. Besides, the view missing patterns may vary in the training and test phases, and such difference is often ignored. To remedy these drawbacks, we propose a novel multi-view classification framework that is simultaneously robust to both incomplete and noisy views. This is achieved by integrating early fusion and late fusion in a single framework. Specifically, in our early fusion module, we propose a view-aware transformer to mask the missing views and adaptively explore the relationships between views and target tasks to deal with missing views. Considering that view missing patterns may change from the training to the test phase, we also design single-view classification and category-consistency constraints to reduce the dependence of our model on view-missing patterns. In our late fusion module, we quantify the view uncertainty in an ensemble way to estimate the noise level of that view. Then the uncertainty and prediction logits of different views are integrated to make our model robust to noisy views. The framework is trained in an end-to-end manner. Experimental results on diverse datasets demonstrate the robustness and effectiveness of our model for both incomplete and noisy views. Codes are available at https://github.com/li-yapeng/UVaT. Yong Luo 0002, Bo Du 0001 |
IEEE Trans. Image Process. | 2 |
| 2024 | DeViT: Decomposing Vision Transformers for Collaborative Inference in Edge DevicesabstractRecent years have witnessed the great success of vision transformer (ViT), which has achieved state-of-the-art performance on multiple computer vision benchmarks. However, ViT models suffer from vast amounts of parameters and high computation cost, leading to difficult deployment on resource-constrained edge devices. Existing solutions mostly compress ViT models to a compact model but still cannot achieve real-time inference. To tackle this issue, we propose to explore the divisibility of transformer structure, and decompose the large ViT into multiple small models for collaborative inference at edge devices. Our objective is to achieve fast and energy-efficient collaborative inference while maintaining comparable accuracy compared with large ViTs. To this end, we first propose a collaborative inference framework termedDeViTto facilitate edge deployment by decomposing large ViTs. Subsequently, we design a decomposition-and-ensemble algorithm based on knowledge distillation, termed DEKD, to fuse multiple small decomposed models while dramatically reducing communication overheads, and handle heterogeneous models by developing a feature matching module to promote the imitations of decomposed models from the large ViT. Extensive experiments for three representative ViT backbones on four widely-used datasets demonstrate our method achieves efficient collaborative inference for ViTs and outperforms existing lightweight ViTs, striking a good trade-off between efficiency and accuracy. For example, our DeViTs improves end-to-end latency by 2.89× with only 1.65% accuracy sacrifice using CIFAR-100 compared to the large ViT, ViT-L/16, on the GPU server. DeDeiTs surpasses the recent efficient ViT, MobileViT-S, by 3.54% in accuracy on ImageNet-1 K, while running 1.72× faster and requiring 55.28% lower energy consumption on the edge device. Guanyu Xu, Zhiwei Hao 0001, Yong Luo 0002, Han Hu 0003, Jianping An, Shiwen Mao |
IEEE Trans. Mob. Comput. | 3 |
| 2024 | Textual Enhanced Adaptive Meta-Fusion for Few-Shot Visual RecognitionabstractFew-shot learning (FSL) is a challenging task that aims to train a classifier to recognize novel categories, where only a few annotated examples are available in each category. Recently, many FSL approaches have been proposed based on the meta-learning paradigm, which attempts to learn transferable knowledge from similar tasks by designing a meta-learner. However, most of these approaches only exploit the information from visual modality and do not utilize ones from additional modalities (e.g., textual description). Since the labeled examples in FSL are limited, increasing the information on the examples is a probable solution to improve the classification performance. This motivates us to propose a novel meta-learning method, termed textual enhanced adaptive meta-fusion FSL (TAMF-FSL), which leverages both the visual information from the visual image and semantic information from language supervision. Specifically, TAMF-FSL exploits the semantic information of textual description to improve the visual-based models. We first employ a text encoder to learn the semantic features of each visual category, and then design a modality alignment module and meta-fusion module to align and fuse the visual and semantic features for final prediction. Extensive experiments show that the proposed method outperforms many recent or competitive FSL counterparts on two popular datasets. Mengya Han, Yibing Zhan, Yong Luo 0002, Han Hu 0003, Kehua Su, Bo Du 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | Not All Instances Contribute Equally: Instance-Adaptive Class Representation Learning for Few-Shot Visual RecognitionabstractFew-shot visual recognition refers to recognize novel visual concepts from a few labeled instances. Many few-shot visual recognition methods adopt the metric-based meta-learning paradigm by comparing the query representation with class representations to predict the category of query instance. However, the current metric-based methods generally treat all instances equally and consequently often obtain biased class representation, considering not all instances are equally significant when summarizing the instance-level representations for the class-level representation. For example, some instances may contain unrepresentative information, such as too much background and information of unrelated concepts, which skew the results. To address the above issues, we propose a novel metric-based meta-learning framework termed instance-adaptive class representation learning network (ICRL-Net) for few-shot visual recognition. Specifically, we develop an adaptive instance revaluing network (AIRN) with the capability to address the biased representation issue when generating the class representation, by learning and assigning adaptive weights for different instances according to their relative significance in the support set of corresponding class. In addition, we design an improved bilinear instance representation and incorporate two novel structural losses, i.e., intraclass instance clustering loss and interclass representation distinguishing loss, to further regulate the instance revaluation process and refine the class representation. We conduct extensive experiments on four commonly adopted few-shot benchmarks: miniImageNet, tieredImageNet, CIFAR-FS, and FC100 datasets. The experimental results compared with the state-of-the-art approaches demonstrate the superiority of our ICRL-Net. Mengya Han, Yibing Zhan, Yong Luo 0002, Bo Du 0001, Han Hu 0003, Yonggang Wen 0001, Dacheng Tao |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Foundation models matter: federated learning for multi-center tuberculosis diagnosis via adaptive regularization and model-contrastive learning
Chang Liu 0046, Yong Luo 0002, Yongchao Xu, Bo Du 0001 |
World Wide Web (WWW) | 2 |
| 2023 | FedABC: Targeting Fair Competition in Personalized Federated LearningabstractFederated learning aims to collaboratively train models without accessing their client's local private data. The data may be Non-IID for different clients and thus resulting in poor performance. Recently, personalized federated learning (PFL) has achieved great success in handling Non-IID data by enforcing regularization in local optimization or improving the model aggregation scheme on the server. However, most of the PFL approaches do not take into account the unfair competition issue caused by the imbalanced data distribution and lack of positive samples for some classes in each client. To address this issue, we propose a novel and generic PFL framework termed Federated Averaging via Binary Classification, dubbed FedABC. In particular, we adopt the ``one-vs-all'' training strategy in each client to alleviate the unfair competition between classes by constructing a personalized binary classification problem for each class. This may aggravate the class imbalance challenge and thus a novel personalized binary classification loss that incorporates both the under-sampling and hard sample mining strategies is designed. Extensive experiments are conducted on two popular datasets under different settings, and the results demonstrate that our FedABC can significantly outperform the existing counterparts. Dui Wang, Li Shen 0008, Yong Luo 0002, Han Hu 0003, Kehua Su, Yonggang Wen 0001, Dacheng Tao |
AAAI | 3 |
| 2023 | FedARC: Federated Learning for Multi-Center Tuberculosis Chest X-ray Diagnosis with Adaptive Regularizing Contrastive RepresentationabstractTuberculosis (TB) poses a significant global health threat and leads to millions of deaths annually. While early diagnosis and treatment can substantially enhance survival prospects, it continues to present a major challenge, particularly in developing countries. In recent years, machine learning has emerged as a valuable tool for tuberculosis diagnosis. However, the training of a dependable diagnostic model necessitates a large volume of data, typically distributed across multiple medical centers. To safeguard data privacy across various centers, we have incorporated federated learning (FL) into TB diagnosis. However, conventional FL methods suffer from substantial performance degradation due to the considerable variation in TB data distribution across different centers. Consequently, we introduce a novel personalized FL approach, FedARC, to address this issue. To mitigate data heterogeneity across centers, we guide the objective function for each center with adaptive regularization to align it with the stationary point of the global loss, thereby enabling the model to converge towards the global optimum. Simultaneously, model-contrastive learning enables the exploration of the specific attributes of each client, enabling the local model to learn more generalizable features. Extensive experimental results on five publicly available chest X-ray image datasets demonstrate the significant outperformance of our proposed method over state-of-the-art methods in diverse settings. Chang Liu 0046, Yong Luo 0002, Yongchao Xu, Bo Du 0001 |
BIBM | 2 |
| 2023 | Spatially Invariant and Frequency-Aware CycleGAN for Unsupervised MR-to-CT Synthesis
Wenbin Hu 0001, Yong Luo 0002, Xin Zhou 0003 |
ICANN (9) | 4 |
| 2023 | MBMS-GAN: Multi-Band Multi-Scale Adversarial Learning for Enhancement of Coded Speech at Very Low Rate
Weiping Tu, Yong Luo 0002, Xin Zhou 0003, Li Xiao 0007, Youqiang Zheng |
ICANN (7) | 3 |
| 2023 | MFT: Multi-scale Fusion Transformer for Infrared and Visible Image Fusion
Chen-Ming Zhang, Chengbo Yuan, Yong Luo 0002, Xin Zhou 0003 |
ICANN (6) | 3 |
| 2023 | Symmetric Pruning in Quantum Neural Networks
Xinbiao Wang, Junyu Liu, Tongliang Liu, Yong Luo 0002, Dacheng Tao |
ICLR | 4 |
| 2023 | Audio-Visual Generalized Zero-Shot Learning Based on Variational Information BottleneckabstractAudio-visual generalized zero-shot learning (GZSL) aims to train a model on seen classes for classifying data samples from both seen classes and unseen classes. Due to the absence of unseen training samples, the model tends to misclassify unseen class samples into seen classes. To mitigate this problem, in this paper, we propose a method based on variational information bottleneck for audio-visual GZSL. Specifically, we model the joint representations as a product-of-experts over marginal representations to integrate the information of audio and visual. Besides, we introduce variational information bottleneck to the learning of audio-visual joint representations and marginal representations of audio, visual, and text label modalities. This helps our model reduce the negative impact of information that cannot be generalized to unseen classes. Experimental results conducted on the UCF-GZSL, VGGSound-GZSL, and ActivityNet-GZSL benchmarks demonstrate the effectiveness and superiority of the proposed model for audio-visual GZSL. Yong Luo 0002, Bo Du 0001 |
ICME | 2 |
| 2023 | Improving Heterogeneous Model Reuse by Density EstimationabstractThis paper studies multiparty learning, aiming to learn a model using the private data of different participants. Model reuse is a promising solution for multiparty learning, assuming that a local model has been trained for each party. Considering the potential sample selection bias among different parties, some heterogeneous model reuse approaches have been developed. However, although pre-trained local classifiers are utilized in these approaches, the characteristics of the local data are not well exploited. This motivates us to estimate the density of local data and design an auxiliary model together with the local classifiers for reuse. To address the scenarios where some local models are not well pre-trained, we further design a multiparty cross-entropy loss for calibration. Upon existing works, we address a challenging problem of heterogeneous model reuse from a decision theory perspective and take advantage of recent advances in density estimation. Experimental results on both synthetic and benchmark data demonstrate the superiority of the proposed method. Anke Tang, Yong Luo 0002, Han Hu 0003, Fengxiang He, Kehua Su, Bo Du 0001, Yixin Chen 0001, Dacheng Tao |
IJCAI | 2 |
| 2023 | Rethinking the Localization in Weakly Supervised Object LocalizationabstractWeakly supervised object localization (WSOL) is one of the most popular and challenging tasks in computer vision. This task is to localize the objects in the images given only the image-level supervision. Recently, dividing WSOL into two parts (class-agnostic object localization and object classification) has become the state-of-the-art pipeline for this task. However, existing solutions under this pipeline usually suffer from the following drawbacks: 1) they are not flexible since they can only localize one object for each image due to the adopted single-class regression (SCR) for localization; 2) the generated pseudo bounding boxes may be noisy, but the negative impact of such noise is not well addressed. To remedy these drawbacks, we first propose to replace SCR with a binary-class detector (BCD) for localizing multiple objects, where the detector is trained by discriminating the foreground and background. Then we design a weighted entropy (WE) loss using the unlabeled data to reduce the negative impact of noisy bounding boxes. Extensive experiments on the popular CUB-200-2011 and ImageNet-1K datasets demonstrate the effectiveness of our method. Rui Xu 0031, Yong Luo 0002, Han Hu 0003, Bo Du 0001, Jialie Shen 0001, Yonggang Wen 0001 |
ACM Multimedia | 2 |
| 2023 | LGViT: Dynamic Early Exiting for Accelerating Vision TransformerabstractRecently, the efficient deployment and acceleration of powerful vision transformers (ViTs) on resource-limited edge devices for providing multimedia services have become attractive tasks. Although early exiting is a feasible solution for accelerating inference, most works focus on convolutional neural networks (CNNs) and transformer models in natural language processing (NLP). Moreover, the direct application of early exiting methods to ViTs may result in substantial performance degradation. To tackle this challenge, we systematically investigate the efficacy of early exiting in ViTs and point out that the insufficient feature representations in shallow internal classifiers and the limited ability to capture target semantic information in deep internal classifiers restrict the performance of these methods. We then propose an early exiting framework for general ViTs termed LGViT, which incorporates heterogeneous exiting heads, namely, local perception head and global aggregation head, to achieve an efficiency-accuracy trade-off. In particular, we develop a novel two-stage training scheme, including end-to-end training and self-distillation with the backbone frozen to generate early exiting ViTs, which facilitates the fusion of global and local information extracted by the two types of heads. We conduct extensive experiments using three popular ViT backbones on three vision datasets. Results demonstrate that our LGViT can achieve competitive performance with approximately 1.8 × speed-up. Guanyu Xu, Li Shen 0008, Han Hu 0003, Yong Luo 0002, Jialie Shen 0001 |
ACM Multimedia | 5 |
| 2023 | Federated Learning with Manifold Regularization and Normalized Update ReaggregationabstractFederated Learning (FL) is an emerging collaborative machine learning framework where multiple clients train the global model without sharing their own datasets.
In FL, the model inconsistency caused by the local data heterogeneity across clients results in the near-orthogonality of client updates, which leads to the global update norm reduction and slows down the convergence. Most previous works focus on eliminating the difference of parameters (or gradients) between the local and global models, which may fail to reflect the model inconsistency due to the complex structure of the machine learning model and the Euclidean space's limitation in meaningful geometric representations.
In this paper, we propose FedMRUR by adopting the manifold model fusion scheme and a new global optimizer to alleviate the negative impacts.
Concretely, FedMRUR adopts a hyperbolic graph manifold regularizer enforcing the representations of the data in the local and global models are close to each other in a low-dimensional subspace.
Because the machine learning model has the graph structure, the distance in hyperbolic space can reflect the model bias better than the Euclidean distance.
In this way, FedMRUR exploits the manifold structures of the representations to significantly reduce the model inconsistency.
FedMRUR also aggregates the client updates norms as the global update norm, which can appropriately enlarge each client's contribution to the global update, thereby mitigating the norm reduction introduced by the near-orthogonality of client updates.
Furthermore, we theoretically prove that our algorithm can achieve a linear speedup property $\mathcal{O}(\frac{1}{\sqrt{SKT}})$ for non-convex setting under partial client participation, where $S$ is the participated clients number, $K$ is the local interval and $T$ is the total number of communication rounds.
Experiments demonstrate that FedMRUR can achieve a new state-of-the-art (SOTA) accuracy with less communication. Xuming An 0001, Li Shen 0008, Han Hu 0003, Yong Luo 0002 |
NeurIPS | 4 |
| 2023 | CLNode: Curriculum Learning for Node ClassificationabstractNode classification is a fundamental graph-based task that aims to predict the classes of unlabeled nodes, for which Graph Neural Networks (GNNs) are the state-of-the-art methods. Current GNNs assume that nodes in the training set contribute equally during training. However, the quality of training nodes varies greatly, and the performance of GNNs could be harmed by two types of low-quality training nodes: (1) inter-class nodes situated near class boundaries that lack the typical characteristics of their corresponding classes. Because GNNs are data-driven approaches, training on these nodes could degrade the accuracy. (2) mislabeled nodes. In real-world graphs, nodes are often mislabeled, which can significantly degrade the robustness of GNNs. To mitigate the detrimental effect of the low-quality training nodes, we present CLNode, which employs a selective training strategy to train GNN based on the quality of nodes. Specifically, we first design a multi-perspective difficulty measurer to accurately measure the quality of training nodes. Then, based on the measured qualities, we employ a training scheduler that selects appropriate training nodes to train GNN in each epoch. To evaluate the effectiveness of CLNode, we conduct extensive experiments by incorporating it in six representative backbone GNNs. Experimental results on real-world networks demonstrate that CLNode is a general framework that can be combined with various GNNs to improve their accuracy and robustness. Xiaowen Wei, Xiuwen Gong, Yibing Zhan, Bo Du 0001, Yong Luo 0002, Wenbin Hu 0001 |
WSDM | 5 |
| 2023 | Graph Regularized Structured Output SVM for Early Expression Detection With Online ExtensionabstractIn this study, a graph regularized algorithm for early expression detection (EED), called GraphEED, is proposed. EED is aimed at detecting the specified expression in the early stage of a video. Existing EED detectors fail to explicitly exploit the local geometrical structure of the data distribution, which may affect the prediction performance significantly. According to manifold learning, the data in real-world applications are likely to reside on a low-dimensional submanifold embedded in the high-dimensional ambient space. The proposed graph Laplacian consists of two parts: 1) a k -nearest neighbor graph is first constructed to encode the geometrical information under the manifold assumption and 2) the entire expressions are regarded as the must-link constraints since they all contain the complete duration information and it is shown that this can also be formulated as a graph regularization. GraphEED is to have a detection function representing these graph structures. Even with the inclusion of the graph Laplacian, the proposed GraphEED has the same computational complexity as that of the max-margin EED, which is a well-known learning-based EED, but the detection performance has been largely improved. To further make the model appropriate in large-scale applications, with the technique of online learning, the proposed GraphEED is extended to the so-called online GraphEED (OGraphEED). In OGraphEED, the buffering technique is employed to make the optimization practical by reducing the computation and storage cost. Extensive experiments on three video-based datasets have demonstrated the superiority of the proposed methods in terms of both effectiveness and efficiency. Yong Luo 0002, Shun-Feng Su, Haikun Wei |
IEEE Trans. Cybern. | 2 |
| 2023 | A Knowledge-Enriched Ensemble Method for Word Embedding and Multi-Sense EmbeddingabstractRepresenting words as embeddings has been proven to be successful in improving the performance in many natural language processing tasks. Different from the traditional methods that learn the embeddings from large text corpora, ensemble methods have been proposed to leverage the merits of pre-trained word embeddings as well as external semantic sources. In this paper, we propose a knowledge-enriched ensemble method to combine information from both knowledge graphs and pre-trained word embeddings. Specifically, we propose an attention network to retrofit the semantic information in the lexical knowledge graph into the pre-trained word embeddings. In addition, we further extend our method to contextual word embeddings and multi-sense embeddings. Extensive experiments demonstrate that the proposed word embeddings outperform the state-of-the-art models in word analogy, word similarity and several downstream tasks. The proposed word sense embeddings outperform the state-of-the-art models in word similarity and word sense induction tasks. Lanting Fang, Yong Luo 0002, Kaiyu Feng, Kaiqi Zhao 0001, Aiqun Hu |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2023 | Multi-Agent Collaborative Inference via DNN Decoupling: Intermediate Feature Compression and Edge LearningabstractRecently, deploying deep neural network (DNN) models via collaborative inference, which splits a pre-trained model into two parts and executes them on user equipment (UE) and edge server respectively, becomes attractive. However, the large intermediate feature of DNN impedes flexible decoupling, and existing approaches either focus on the single UE scenario or simply define tasks considering the required CPU cycles, but ignore the indivisibility of a single DNN layer. In this article, we study the multi-agent collaborative inference scenario, where a single edge server coordinates the inference of multiple UEs. Our goal is to achieve fast and energy-efficient inference for all UEs. To achieve this goal, we design a lightweight autoencoder-based method to compress the large intermediate feature at first. Then we define tasks according to the inference overhead of DNNs and formulate the problem as a Markov decision process (MDP). Finally, we propose a multi-agent hybrid proximal policy optimization (MAHPPO) algorithm to solve the optimization problem with a hybrid action space. We conduct extensive experiments with different types of networks, and the results show that our method can reduce up to 56% of inference latency and save up to 72% of energy consumption. Zhiwei Hao 0001, Guanyu Xu, Yong Luo 0002, Han Hu 0003, Jianping An, Shiwen Mao |
IEEE Trans. Mob. Comput. | 3 |
| 2023 | Two-Stream Prototype Learning Network for Few-Shot Face Recognition Under OcclusionsabstractFew-shot face recognition under occlusion (FSFRO) aims to recognize novel subjects given only a few, probably occluded face images, and it is challenging and common in real-world scenarios. Unknown occlusions may deteriorate the class prototypes, while an occluded image in the support set may be critical for recognition if the query image is occluded. This motivates us to propose a novel Two-stream Prototype Learning Network (TSPLN) for FSFR under occlusions by simultaneously considering the quality of support images and their relevance to the query i mage. Specifically, we design a two-stream architecture, which mainly consists of a support-centered stream and query-centered stream, to learn the optimal class prototypes. The former stream is to reduce the negative impact of occluded images on the prototype. This is achieved by exploring the similarities between different images in the support set. In the query-centered stream, we exploit the relevance between the query and support set based on feature alignment (FA). We conduct extensive experiments on two popular datasets: CASIA-WebFace and RMFRD. The experimental results show that our proposed method achieves the state-of-the-art performance for occluded face recognition in the few-shot setting. Mengya Han, Yong Luo 0002, Han Hu 0003, Yonggang Wen 0001 |
IEEE Trans. Multim. | 3 |
| 2023 | Unpaired Image Captioning by Image-Level Weakly-Supervised Visual Concept RecognitionabstractThe goal of unpaired image captioning (UIC) is to describe images without using image-caption pairs in the training phase. Although challenging, we expect the task can be accomplished by leveraging images aligned with visual concepts. Most existing studies use off-the-shelf algorithms to obtain the visual concepts because the Bounding Box (BBox) labels or relationship-triplet labels used for training are expensive to acquire. To avoid exhaustive annotations, we propose a novel approach to achieve cost-effective UIC. Specifically, we adopt image-level labels to optimize the UIC model in a weakly-supervised manner. For each image, we assume that only the image-level labels are available without specific locations and numbers. The image-level labels are utilized to train a weakly-supervised object recognition model to extract object information (e.g., instance), and the extracted instances are adopted to infer the relationships among different objects using an enhanced graph neural network (GNN). The proposed approach achieves comparable or even better performance compared with previous methods without expensive annotations. Furthermore, we design an unrecognized object (UnO) loss to improve the alignment of the inferred object and relationship information with the images. It can effectively alleviate the issue encountered by existing UIC models when generating sentences with nonexistent objects. To the best of our knowledge, this is the first attempt to address the problem of Weakly-Supervised visual concept recognition for UIC (WS-UIC) based only on image-level labels. Extensive experiments demonstrate that the proposed method achieves inspiring results on the COCO dataset while significantly reducing the labeling cost. Peipei Zhu, Xiao Wang 0014, Yong Luo 0002, Zhenglong Sun 0001, Wei-Shi Zheng 0001, Yaowei Wang 0001, Chang Wen Chen |
IEEE Trans. Multim. | 3 |
| 2023 | DRRNets: Dynamic Recurrent Routing via Low-Rank Regularization in Recurrent Neural NetworksabstractRecurrent neural networks (RNNs) continue to show outstanding performance in sequence learning tasks such as language modeling, but it remains difficult to train RNNs for long sequences. The main challenges lie in the complex dependencies, gradient vanishing or exploding, and low resource requirement in model deployment. In order to address these challenges, we propose dynamic recurrent routing neural networks (DRRNets), which can: 1) shorten the recurrent lengths by allocating recurrent routes dynamically for different dependencies and 2) reduce the number of parameters significantly by imposing low-rank constraints on the fully connected layers. A novel optimization algorithm via low-rank constraint and sparsity projection is developed to train the network. We verify the effectiveness of the proposed method by comparing it with multiple competitive approaches in several popular sequential learning tasks, such as language modeling and speaker recognition. The results in terms of different criteria demonstrate the superiority of our proposed method. Dongjing Shan, Yong Luo 0002, Xiongwei Zhang, Chao Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | Resistance Training Using Prior Bias: Toward Unbiased Scene Graph GenerationabstractScene Graph Generation (SGG) aims to build a structured representation of a scene using objects and pairwise relationships, which benefits downstream tasks. However, current SGG methods usually suffer from sub-optimal scene graph generation because of the long-tailed distribution of training data. To address this problem, we propose Resistance Training using Prior Bias (RTPB) for the scene graph generation. Specifically, RTPB uses a distributed-based prior bias to improve models' detecting ability on less frequent relationships during training, thus improving the model generalizability on tail categories. In addition, to further explore the contextual information of objects and relationships, we design a contextual encoding backbone network, termed as Dual Transformer (DTrans). We perform extensive experiments on a very popular benchmark, VG150, to demonstrate the effectiveness of our method for the unbiased scene graph generation. In specific, our RTPB achieves an improvement of over 10% under the mean recall when applied to current SGG methods. Furthermore, DTrans with RTPB outperforms nearly all state-of-the-art methods with a large margin. Code is available at https://github.com/ChCh1999/RTPB Yibing Zhan, Baosheng Yu, Liu Liu 0014, Yong Luo 0002, Bo Du 0001 |
AAAI | 5 |
| 2022 | LSSANet: A Long Short Slice-Aware Network for Pulmonary Nodule Detection
Rui Xu 0031, Yong Luo 0002, Bo Du 0001, Kaiming Kuang, Jiancheng Yang |
MICCAI (1) | 2 |
| 2022 | Leveraging GAN Priors for Few-Shot Part SegmentationabstractFew-shot part segmentation aims to separate different parts of an object given only a few annotated samples. Due to the challenge of limited data, existing works mainly focus on learning classifiers over pre-trained features, failing to learn task-specific features for part segmentation. In this paper, we propose to learn task-specific features in a "pre-training"-"fine-tuning" paradigm. We conduct prompt designing to reduce the gap between the pre-train task (i.e., image generation) and the downstream task (i.e., part segmentation), so that the GAN priors for generation can be leveraged for segmentation. This is achieved by projecting part segmentation maps into the RGB space and conducting interpolation between RGB segmentation maps and original images. Specifically, we design a fine-tuning strategy to progressively tune an image generator into a segmentation generator, where the supervision of the generator varying from images to segmentation maps by interpolation. Moreover, we propose a two-stream architecture, i.e., a segmentation stream to generate task-specific features, and an image stream to provide spatial constraints. The image stream can be regarded as a self-supervised auto-encoder, and this enables our model to benefit from large-scale support images. Overall, this work is an attempt to explore the internal relevance between generation tasks and perception tasks by prompt designing. Extensive experiments show that our model can achieve state-of-the-art performance on several part segmentation datasets. Mengya Han, Heliang Zheng, Yong Luo 0002, Han Hu 0003, Bo Du 0001 |
ACM Multimedia | 4 |
| 2022 | Knowledge Graph enhanced Multimodal Learning for Few-shot Visual RecognitionabstractFew-shot learning (FSL) aims to learn a classifier for novel classes with only a few labeled samples per category available. The mainstream FSL approaches fall in the meta-learning paradigm, where a meta-learner is used to learn transferable knowledge and generalize to new tasks. However, these approaches usually only leverage information from a single modality (e.g., visual image) and fail to explore the information from other modalities (e.g., the knowledge graph). Since the labeled samples are scarce in FSL, increasing the information for each example is a possible solution to improve the performance. This motivates us to develop a new meta-learning framework for few-shot visual recognition termed Knowledge Graph enhanced FSL (KGFSL), which combines the information from multiple modalities: 1) the visual information in images and 2) the rich semantics and structural information in a knowledge graph (KG). Specifically, KGFSL exploits the word embedding of the category and its relationship to other categories to improve the visual-based models. A graph convolutional network (GCN) is first introduced to learn the semantic embeddings for each node (a visual category) in KG. The visual and semantic embeddings are then aligned and combined for final prediction. Finally, the whole framework is trained in an end-to-end manner. We conduct extensive experiments on two widely-used FSL benchmarks: miniImageNet and tieredImageNet. Experimental results demonstrate the effectiveness of the multimodal information for few-shot learning, and our proposed method can significantly outperform the state-of-the-art approaches. Mengya Han, Yibing Zhan, Baosheng Yu, Yong Luo 0002, Bo Du 0001, Dacheng Tao |
MMSP | 4 |
| 2022 | Nonlinear Multi-Model ReuseabstractThe goal of model reuse is to build a model in a new target domain by reusing some pre-trained source models. It can significantly reduce the training costs and the data required for training, and hence has various potential applications. Most of the existing model reuse approaches only reuse the output features or labels of the source model, and more information contained in the model are ignored. Besides, only a single model can be utilized in these approaches. A recently proposed multi-model reuse method is able to remedy these drawbacks by utilizing the hidden layer representations of multiple source models to help improve the representations in the target model, but it assumes that there are linear connections between the source and target models. This assumption is too restrictive and may be not valid in real-world applications. In this paper, we relax this assumption by introducing the manifold regularization scheme to exploit arbitrary nonlinear relationships between the source and target models. Effectiveness of our method is demonstrated empirically by the extensive experiments in the popular person re-identification task for smart city application. Yong Luo 0002, Ling-Yu Duan, Tongliang Liu, Yihang Lou, Yonggang Wen 0001 |
MMSP | 1 |
| 2022 | Robust Metric Boosts TransferabstractTransfer metric learning (TML) aims to improve the metric learning in target domains by transferring knowledge from related tasks, where the distance metrics are strong and reliable. Existing TML approaches only focus on how to transfer the source metric knowledge, which is often prone to be over-fitting to the source domain. In this paper, we study how to train a source metric that is appropriate for transfer and then design a general deep TML method for effective metric transfer. In particular, we propose to learn the source metric parameterized by a deep neural network in an adversarial way and then transfer the metric to the target domain by embedding imitation, which allows the inputs of source and target domains to be heterogeneous. Besides, we restrict the size of the target metric network to be small so that the inference is efficient in the target domain. Results in the popular face verification application demonstrate the effectiveness of our method. Qiancheng Yang, Yong Luo 0002, Han Hu 0003, Xin Zhou 0003, Bo Du 0001, Dacheng Tao |
MMSP | 2 |
| 2022 | Covered Style Mining via Generative Adversarial Networks for Face Anti-spoofing
Yiqiang Wu, Dapeng Tao, Yong Luo 0002, Jun Cheng 0002, Xuelong Li 0001 |
Pattern Recognit. | 3 |
| 2022 | Intrinsic Performance Influence-based Participant Contribution Estimation for Horizontal Federated LearningabstractThe rapid development of modern artificial intelligence technique is mainly attributed to sufficient and high-quality data. However, in the data collection, personal privacy is at risk of being leaked. This issue can be addressed by federated learning, which is proposed to achieve efficient model training among multiple data providers without direct data access and aggregation. To encourage more parties owning high-quality data to participate in the federated learning, it is important to evaluate and reward the participant contribution in a reasonable, robust, and efficient manner. To achieve this goal, we propose a novel contribution estimation method: Intrinsic Performance Influence-based Contribution Estimation (IPICE). In particular, the class-level intrinsic performance influence is adopted as the contribution estimation criteria in IPICE, and a neural network is employed to exploit the non-linear relationship between the performance change and estimated contribution. Extensive experiments are conducted on various datasets, and the results demonstrate that IPICE is more accurate and stable than the counterpart in various data distribution settings. The computational complexity is significantly reduced in our IPICE, especially when a new party joins the federation. IPICE assigns small contributions to bad/garbage data and thus prevent them from participating and deteriorating the learning ecosystem. Lin Zhang 0014, Lixin Fan, Yong Luo 0002, Ling-Yu Duan |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2022 | CDFKD-MFS: Collaborative Data-Free Knowledge Distillation via Multi-Level Feature SharingabstractRecently, the compression and deployment of powerful deep neural networks (DNNs) on resource-limited edge devices to provide intelligent services have become attractive tasks. Although knowledge distillation (KD) is a feasible solution for compression, its requirement on the original dataset raises privacy concerns. In addition, it is common to integrate multiple pretrained models to achieve satisfactory performance. How to compress multiple models into a tiny model is challenging, especially when the original data are unavailable. To tackle this challenge, we propose a framework termed collaborative data-free knowledge distillation via multi-level feature sharing (CDFKD-MFS), which consists of a multi-header student module, an asymmetric adversarial data-free KD module, and an attention-based aggregation module. In this framework, the student model equipped with a multi-level feature-sharing structure learns from multiple teacher models and is trained together with a generator in an asymmetric adversarial manner. When some real samples are available, the attention module adaptively aggregates predictions of the student headers, which can further improve performance. We conduct extensive experiments on three popular computer visual datasets. In particular, compared with the most competitive alternative, the accuracy of the proposed framework is 1.18% higher on the CIFAR-100 dataset, 1.67% higher on the Caltech-101 dataset, and 2.99% higher on the mini-ImageNet dataset. Zhiwei Hao 0001, Yong Luo 0002, Zhi Wang 0001, Han Hu 0003, Jianping An |
IEEE Trans. Multim. | 2 |
| 2021 | Federated Learning for Non-IID Data via Unified Feature Learning and Optimization Objective AlignmentabstractFederated Learning (FL) aims to establish a shared model across decentralized clients under the privacy-preserving constraint. Despite certain success, it is still challenging for FL to deal with non-IID (non-independent and identical distribution) client data, which is a general scenario in real-world FL tasks. It has been demonstrated that the performance of FL will be reduced greatly under the non-IID scenario, since the discrepant data distributions will induce optimization inconsistency and feature divergence issues. Besides, naively minimizing an aggregate loss function in this scenario may have negative impacts on some clients and thus deteriorate their personal model performance. To address these issues, we propose a Unified Feature learning and Optimization objectives alignment method (FedUFO) for non-IID FL. In particular, an adversary module is proposed to reduce the divergence on feature representation among different clients, and two consensus losses are proposed to reduce the inconsistency on optimization objectives from two perspectives. Extensive experiments demonstrate that our FedUFO can outperform the state-of-the-art approaches, including the competitive one data-sharing method. Besides, FedUFO can enable more reasonable and balanced model performance among different clients. Lin Zhang 0014, Yong Luo 0002, Bo Du 0001, Ling-Yu Duan |
ICCV | 2 |
| 2021 | Model Compression via Collaborative Data-Free Knowledge Distillation for Edge IntelligenceabstractModel compression without the original data for fine-tuning is challenging for deploying large-size models on resource constrained edge devices. To this end, we propose a novel data-free model compression framework based on knowledge distillation (KD), where multiple teachers are utilized in a collaborative manner to enable reliable distillation. It mainly consists of three components: adversarial data generation, multi-teacher KD, and adaptive outputs aggregation. In particular, some synthesized data are generated in an adversarial manner to mimic the original data for model compression. Then a multi-header module is developed to simultaneously leverage diverse knowledge from multiple teachers. The distillation outputs are adaptively aggregated for final prediction. The experimental results demonstrate that our framework outperforms the data-free counterpart significantly (4.48% on MNIST and 2.96% on CIFAR-10). Effectiveness of different components of our method is also verified via carefully designed ablation study. Zhiwei Hao 0001, Yong Luo 0002, Zhi Wang 0001, Han Hu 0003, Jianping An |
ICME | 2 |
| 2021 | Data-Free Ensemble Knowledge Distillation for Privacy-conscious Multimedia Model CompressionabstractRecent advances in deep learning bring impressive performance for multimedia applications. Hence, compressing and deploying these applications on resource-limited edge devices via model compression becomes attractive. Knowledge distillation (KD) is one of the most popular model compression techniques. However, most well-behaved KD approaches require the original dataset, which is usually unavailable due to privacy issues, while existing data-free KD methods perform much worse than data-required counterparts. In this paper, we analyze previous data-free KD methods from the data perspective and point out that using a single pre-trained model limits the performance of these approaches. We then propose a Data-Free Ensemble knowledge Distillation (DFED) framework, which contains a student network, a generator network, and multiple pre-trained teacher networks. During training, the student mimics behaviors of the ensemble of teachers using samples synthesized by a generator, which aims to enlarge the prediction discrepancy between the student and teachers. A moment matching loss term assists the generator training by minimizing the distance between activations of synthesized samples and real samples. We evaluate DFED on three popular image classification datasets. Results demonstrate that our method achieves significant performance improvements compared with previous works. We also design an ablation study to verify the effectiveness of each component of the proposed framework. Zhiwei Hao 0001, Yong Luo 0002, Han Hu 0003, Jianping An, Yonggang Wen 0001 |
ACM Multimedia | 2 |
| 2021 | Near-Online Multi-Pedestrian Tracking via Combining Multiple Consistent Appearance CuesabstractAn important cue for multi-pedestrian tracking in video is the consistent appearance of an individual for quite a while. In this paper, we address multi-pedestrian tracking by learning a robust appearance model from the paradigm of tracking by detection. To separate detections of different pedestrians while assembling detections of the same pedestrian, we take advantage of the cue of consistent appearance and exploit three types of evidence from the recent, past and near-future. Existing online approaches only exploit the detection-to-detection and sequence-to-detection metrics, which focus on the recent and past appearance patterns respectively, while the future pedestrian appearance is simply ignored. This drawback is remedied in this paper by further considering the sequence-to-sequence metric, which resorts to near-future appearance presentation. Adaptive combination weights are learned to fuse these three different metrics. Moreover, we propose a novel Focal Triplet Loss to make the model focus more on hard examples than the easy ones. We demonstrate that this can significantly enhance the discriminating power of the model compared with treating every sample equally. Effectiveness and efficiency of the proposed method is verified by conducting comprehensive ablation studies and comparing with many competitive (offline/online/near-online) counterparts on the MOT16 and MOT17 Challenges. Weijiang Feng, Long Lan, Yong Luo 0002, Yue Yu 0001, Xiang Zhang 0008, Zhigang Luo |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Deep Heterogeneous Multi-Task Metric Learning for Visual Recognition and RetrievalabstractHow to estimate the distance between data instances is a fundamental problem in many artificial intelligence algorithms, and critical in diverse multimedia applications. A major challenge in the estimation is how to find an appropriate distance function when labeled data are insufficient for a certain task. Multi-task metric learning (MTML) is able to alleviate such data deficiency issue by learning distance metrics for multiple tasks together and sharing information between the different tasks. Recently, heterogeneous MTML (HMTML) has attracted much attention since it can handle multiple tasks with varied data representations. A major drawback of the current HMTML approaches is that only linear transformations are learned to connect different domains. This is suboptimal since the correlations between different domains may be very complex and highly nonlinear. To overcome this drawback, we propose a deep heterogeneous MTML (DHMTML) method, in which a nonlinear mapping is learned for each task by using a deep neural network. The correlations of different domains are exploited by sharing some parameters at the top layers of different networks. More importantly, the auto-encoder scheme and the adversarial learning mechanism are integrated and incorporated to help exploit the feature correlations in and between different tasks and the specific properties are preserved by learning additional task-specific layers together with the common layers. Experiments demonstrated that the proposed method outperforms single-task deep metric learning algorithms and other HMTML approaches consistently on several benchmark datasets. Shikang Gan, Yong Luo 0002, Yonggang Wen 0001, Tongliang Liu, Han Hu 0003 |
ACM Multimedia | 2 |
| 2020 | Semi-supervised Online Multi-Task Metric Learning for Visual Recognition and RetrievalabstractDistance metric learning (DML) is critial in many multimedia application tasks. However, it is hard to learn a satisfactory distance metric given only a few labeled samples for each task. In this paper, we proposed a novel semi-supervised online multi-task DML method termed SOMTML, which enables the models describing different tasks to help each other during the metric learning procedure and thus improving their respective performance. Besides, unlabeled data are leveraged to further help alleviate the data deficiency issue in different tasks by designing a novel regularization term, which also allows some prior information to be incorporated. More importantly, a quite efficient algorithm is developed to update the metrics of all tasks adaptively. The proposed SOMTML is experimentally validated in two popular visual analytic-based applications: handwriting digits recognition and face retrieval. We compared the proposed method with competitive single-task and multi-task metric learning approaches. Extensive experimental results demonstrate the effectiveness and efficiency of the proposed SOMTML. Yangxi Li, Han Hu 0003, Jin Li 0014, Yong Luo 0002, Yonggang Wen 0001 |
ACM Multimedia | 4 |
| 2020 | Look, Read and Feel: Benchmarking Ads Understanding with Multimodal Multitask LearningabstractGiven the massive market of advertising and the sharply increasing online multimedia content (such as videos), it is now fashionable to promote advertisements (ads) together with the multimedia content. However, manually finding relevant ads to match the provided content is labor-intensive, and hence some automatic advertising techniques are developed. Since ads are usually hard to understand only according to its visual appearance due to the contained visual metaphor, some other modalities, such as the contained texts, should be exploited for understanding. To further improve user experience, it is necessary to understand both the ads' topic and sentiment. This motivates us to develop a novel deep multimodal multitask framework that integrates multiple modalities to achieve effective topic and sentiment prediction simultaneously for ads understanding. In particular, in our framework termed Deep$M^2$Ad, we first extract multimodal information from ads and learn high-level and comparable representations. The visual metaphor of the ad is decoded in an unsupervised manner. The obtained representations are then fed into the proposed hierarchical multimodal attention modules to learn task-specific representations for final prediction. A multitask loss function is also designed to jointly train both the topic and sentiment prediction models in an end-to-end manner, where bottom-layer parameters are shared to alleviate over-fitting. We conduct extensive experiments on a large-scale advertisement dataset and achieve state-of-the-art performance for both prediction tasks. The obtained results could be utilized as a benchmark for ads understanding. Huaizheng Zhang, Yong Luo 0002, Qiming Ai, Yonggang Wen 0001, Han Hu 0003 |
ACM Multimedia | 2 |
| 2020 | Hysia: Serving DNN-Based Video-to-Retail Applications in CloudabstractCombining video streaming and online retailing (V2R) has been a growing trend recently. In this paper, we provide practitioners and researchers in multimedia with a cloud-based platform named Hysia for easy development and deployment of V2R applications. The system consists of: 1) a back-end infrastructure providing optimized V2R related services including data engine, model repository, model serving and content matching; and 2) an application layer which enables rapid V2R application prototyping. Hysia addresses industry and academic needs in large-scale multimedia by: 1) seamlessly integrating state-of-the-art libraries including NVIDIA video SDK, Facebook faiss, and gRPC; 2) efficiently utilizing GPU computation; and 3) allowing developers to bind new models easily to meet the rapidly changing deep learning (DL) techniques. On top of that, we implement an orchestrator for further optimizing DL model serving performance. Hysia has been released as an open source project on GitHub, and attracted considerable attention. We have published Hysia to DockerHub as an official image for seamless integration and deployment in current cloud environments. Huaizheng Zhang, Yuanming Li, Qiming Ai, Yong Luo 0002, Yonggang Wen 0001, Yichao Jin 0002, Ta Nguyen Binh Duong |
ACM Multimedia | 4 |
| 2020 | Transforming Device Fingerprinting for Wireless Security via Online Multitask Metric LearningabstractDevice fingerprinting is a crucial part in the Internet of Things applications. Existing device-fingerprinting solutions either ignore the influence of the type of network traffic or separately learn a fingerprinting model for each traffic type. This often leads to suboptimal solutions, especially when training data are limited. Considering that the data distributions of different traffic types may be different but related, we propose a novel multitask learning method to learn the fingerprinting models for several traffic types simultaneously. Specifically, we first design a system for device fingerprinting using the popular k-nearest neighbor (KNN) approach. Then, a novel distance metric learning (DML) algorithm termed online multitask metric learning (OMTML) is developed to improve the distance estimation in our system. OMTML enables the models describing different traffic types to help each other during the metric learning procedure, and thus improving their respective accuracies. OMTML can also be updated adaptively, and the updating process is efficient. The experimental results show that the proposed KNN-based system outperforms the artificial neural network (ANN)-based counterpart significantly. Besides, the comparisons of our OMTML and other representative online and multitask DML approaches demonstrate both effectiveness and efficiency of the proposed metric learning method. Yong Luo 0002, Han Hu 0003, Yonggang Wen 0001, Dacheng Tao |
IEEE Internet Things J. | 1 |
| 2020 | Towards Efficient Front-End Visual Sensing for Digital Retina: A Model-Centric ParadigmabstractThe digital retina excels at providing enhanced visual sensing and analysis capability for city brain in smart cities, and can feasibly convert the visual data from visual sensors into semantic features. With the deployment of deep learning or handcrafted models, these features are extracted on front-end devices, then delivered to back-end servers for advanced analysis. In this scenario, we propose a model generation, utilization and communication paradigm, aiming at strong front-end sensing capabilities for establishing better artificial visual systems in smart cities. In particular, we propose an integrated multiple deep learning models reuse and prediction strategy, which dramatically increases the feasibility of the digital retina in large-scale visual data analysis in smart cities. The proposed multi-model reuse scheme aims to reuse the knowledge from models cached and transmitted in digital retina to obtain more discriminative capability. To efficiently deliver these newly generated models, a model prediction scheme is further proposed by encoding and reconstructing model differences. Extensive experiments have been conducted to demonstrate the effectiveness of proposed model-centric paradigm. Yihang Lou, Ling-Yu Duan, Yong Luo 0002, Ziqian Chen, Tongliang Liu, Shiqi Wang 0001, Wen Gao 0001 |
IEEE Trans. Multim. | 3 |
| 2019 | ResumeGAN: An Optimized Deep Representation Learning Framework for Talent-Job Fit via Adversarial LearningabstractNowadays, it is popular to utilize online recruitment services for talent recruitment and job recommendation. Given the vast amounts of online talent profiles and job-posts, it is labor-intensive and exhausted for recruiters to manually select only a few potential candidates for further consideration, and also nontrivial for talents to find the most matched job positions. Recently, some deep learning-based approaches are developed to automatically matching the talent resumes and job requirements, and have achieved encouraging performance. In this paper, we propose a novel framework that targets the same task, but integrate different types of information in a more sophisticated way and introduce adversarial learning to learn more expressive representation. In addition, we build a dataset for model evaluation and the effectiveness of our framework is demonstrated by extensive experiments. Yong Luo 0002, Huaizheng Zhang, Yonggang Wen 0001, Xinwen Zhang |
CIKM | 1 |
| 2019 | Incorporating Category Taxonomy in Deep Reinforcement Learning Based Image HashingabstractImage hashing is critical for large-scale image analytic-based applications, such as image retrieval. Although there have been dozens of hashing approaches, few of them take the hierarchical structure of the image categories into consideration. In this paper, we propose to incorporate the category taxonomy information in a deep reinforcement learning (DRL) model for image hashing. In particular, we learn an agent to predict the hashing codes sequentially under the DRL theme. Each coordinate of the hashing function can take the errors incurred by previous ones into consideration and hence more reliable hashing codes can be obtained than learning them independently. Besides, we design a novel level-specific reward function to gradually refine the hashing function according to the taxonomy information. Extensive experiments on two popular datasets demonstrate effectiveness of the proposed method. Qiang Fu 0006, Linsen Dong, Yong Luo 0002, Yonggang Wen 0001, Ying Li 0012, Ling-Yu Duan |
ICME | 4 |
| 2019 | Towards Digital Retina in Smart Cities: A Model Generation, Utilization and Communication ParadigmabstractThe digital retina in smart cities is to select what the City Eye tells the City Brain, and convert the acquired visual data from front-end visual sensors to features in an intelligent sensing manner. By deploying deep learning and/or handcrafted models in front-end devices, the compact features can be extracted and subsequently delivered to back-end cloud for search and advanced analytics. In this context, we propose a model generation, utilization, and communication paradigm, aiming to address a set of unique challenges for better artificial intelligence services in smart cities. In particular, we present an integrated multiple deep learning models reuse and prediction strategy, which greatly increases the feasibility of the digital retina in processing and analyzing the large-scale visual data in smart cities. The promise of the proposed paradigm is demonstrated through a set of experiments. Yihang Lou, Ling-Yu Duan, Yong Luo 0002, Ziqian Chen, Tongliang Liu, Shiqi Wang 0001, Wen Gao 0001 |
ICME | 3 |
| 2019 | Toward Intelligent Visual Sensing and Low-cost Analysis: A Collaborative Computing ApproachabstractIn the big data era, there has been an increasing consensus that the label information, computational resources and communication bandwidth are particularly precious. State-of-the-art research is revolutionizing the vision systems of the smart city, which converts the visual signals from sensory input into feature representations and conveys the compact feature for analysis by using the computational resources of both front and back ends. To deploy a robust model, large amounts of labeled data are usually required, and thereby heavy computational and communication resources are incurred in model training as well as inference. However, the computational resources in front-end devices are usually constrained, and heavy transmission burden is imposed when leveraging multiple models amongst different ends. In this work, we propose a novel collaborative computing approach for intelligent sensing and low-cost analysis, which reduces the requirement of labeled data and communication cost, and balances the computational load in model training and inference. By incorporating the adversarial learning mechanism into collaborative model training, knowledge of different domains can be better exploited. Moreover, the learned models are deployed for inference in a collaborative manner, in which part of model is placed in front-ends for extracting intermediate feature maps, and part of the model remains in back ends for inference with received feature maps. The effectiveness of the proposed approach has been validated in the context of an emerging digital retina system for smart city intelligent applications. Ling-Yu Duan, Yong Luo 0002, Shiqi Wang 0001, Yonggang Wen 0001, Wen Gao 0001 |
VCIP | 3 |
| 2019 | Knowledge-Enhanced Ensemble Learning for Word EmbeddingsabstractRepresenting words as embeddings in a continuous vector space has been proven to be successful in improving the performance in many natural language processing (NLP) tasks. Beyond the traditional methods that learn the embeddings from large text corpora, ensemble methods have been proposed to leverage the merits from pre-trained word embeddings as well as external semantic sources. In this paper, we propose a knowledge-enhanced ensemble method to combine both knowledge graphs and pre-trained word embedding models. Specifically, we interpret relations in knowledge graphs as linear translation from one word to another. We also propose a novel weighting scheme to further distinguish edges in the knowledge graph with same type of relation. Extensive experiments demonstrate that our proposed method is up to 20% times better than state-of-the-art in word analogy task and up to 16% times better than state-of-the-art in word similarity task. Lanting Fang, Yong Luo 0002, Kaiyu Feng, Kaiqi Zhao 0001, Aiqun Hu |
WWW | 2 |
| 2019 | Transferring Knowledge Fragments for Learning Distance Metric from a Heterogeneous DomainabstractThe goal of transfer learning is to improve the performance of target learning task by leveraging information (or transferring knowledge) from other related tasks. In this paper, we examine the problem of transfer distance metric learning (DML), which usually aims to mitigate the label information deficiency issue in the target DML. Most of the current Transfer DML (TDML) methods are not applicable to the scenario where data are drawn from heterogeneous domains. Some existing heterogeneous transfer learning (HTL) approaches can learn target distance metric by usually transforming the samples of source and target domain into a common subspace. However, these approaches lack flexibility in real-world applications, and the learned transformations are often restricted to be linear. This motivates us to develop a general flexible heterogeneous TDML (HTDML) framework. In particular, any (linear/nonlinear) DML algorithms can be employed to learn the source metric beforehand. Then the pre-learned source metric is represented as a set of knowledge fragments to help target metric learning. We show how generalization error in the target domain could be reduced using the proposed transfer strategy, and develop novel algorithm to learn either linear or nonlinear target metric. Extensive experiments on various applications demonstrate the effectiveness of the proposed method. Yong Luo 0002, Yonggang Wen 0001, Tongliang Liu, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | ResumeNet: A Learning-Based Framework for Automatic Resume Quality AssessmentabstractRecruitment of appropriate people for certain positions is critical for any companies or organizations. Manually screening to select appropriate candidates from large amounts of resumes can be exhausted and time-consuming. However, there is no public tool that can be directly used for automatic resume quality assessment (RQA). This motivates us to develop a method for automatic RQA. Since there is also no public dataset for model training and evaluation, we build a dataset for RQA by collecting around 10K resumes, which are provided by a private resume management company. By investigating the dataset, we identify some factors or features that could be useful to discriminate good resumes from bad ones, e.g., the consistency between different parts of a resume. Then a neural-network model is designed to predict the quality of each resume, where some text processing techniques are incorporated. To deal with the label deficiency issue in the dataset, we propose several variants of the model by either utilizing the pair/triplet-based loss, or introducing some semi-supervised learning technique to make use of the abundant unlabeled data. Both the presented baseline model and its variants are general and easy to implement. Various popular criteria including the receiver operating characteristic (ROC) curve, F-measure and ranking-based average precision (AP) are adopted for model evaluation. We compare the different variants with our baseline model. Since there is no public algorithm for RQA, we further compare our results with those obtained from a website that can score a resume. Experimental results in terms of different criteria demonstrate effectiveness of the proposed method. We foresee that our approach would transform the way of future human resources management. Yong Luo 0002, Huaizheng Zhang, Yonggang Wen 0001, Xinwen Zhang |
ICDM | 1 |
| 2018 | Deep Discrete Prototype Multilabel LearningabstractkNN embedding methods, such as the state-of-the-art LM-kNN, have shown impressive results in multi-label learning. Unfortunately, these approaches suffer expensive computation and memory costs in large-scale settings. To fill this gap, this paper proposes a novel deep prototype compression, i.e., DBPC for fast multi-label prediction. DBPC compresses the database into a small set of short discrete prototypes, and uses the prototypes for prediction. The benefit of DBPC comes from two aspects: 1) The number of distance comparisons are reduced in the prototype; 2) The distance computation cost is significantly decreased in the reduced space. We propose to jointly learn the deep latent subspace and discrete prototypes within one framework. The encoding and decoding neural networks are employed to make deep discrete prototypes well represent the instances and labels. Extensive experiments on several large-scale datasets demonstrate that DBPC achieves several orders of magnitude lower storage and prediction complexity than state-of-the-art multi-label methods, while achieving competitive accuracy. Xiaobo Shen 0001, Weiwei Liu 0003, Yong Luo 0002, Yew-Soon Ong, Ivor W. Tsang |
IJCAI | 3 |
| 2018 | Online Heterogeneous Transfer Metric LearningabstractDistance metric learning (DML) has been demonstrated to be successful and essential in diverse applications. Transfer metric learning (TML) can help DML in the target domain with limited label information by utilizing information from some related source domains. The heterogeneous TML (HTML), where the feature representations vary from the source to the target domain, is general and challenging. However, current HTML approaches are usually conducted in a batch manner and cannot handle sequential data. This motivates the proposed online HTML (OHTML) method. In particular, the distance metric in the source domain is pre-trained using some existing DML algorithms. To enable knowledge transfer, we assume there are large amounts of unlabeled corresponding data that have representations in both the source and target domains. By enforcing the distances (between these unlabeled samples) in the target domain to agree with those in the source domain under the manifold regularization theme, we learn an improved target metric. We formulate the problem in the online setting so that the optimization is efficient and the model can be adapted to new coming data. Experiments in diverse applications demonstrate both effectiveness and efficiency of the proposed method. Yong Luo 0002, Tongliang Liu, Yonggang Wen 0001, Dacheng Tao |
IJCAI | 1 |
| 2018 | Cost-Sensitive Feature Selection by Optimizing F-MeasuresabstractFeature selection is beneficial for improving the performance of general machine learning tasks by extracting an informative subset from the high-dimensional features. Conventional feature selection methods usually ignore the class imbalance problem, thus the selected features will be biased towards the majority class. Considering that F-measure is a more reasonable performance measure than accuracy for imbalanced data, this paper presents an effective feature selection algorithm that explores the class imbalance issue by optimizing F-measures. Since F-measure optimization can be decomposed into a series of cost-sensitive classification problems, we investigate the cost-sensitive feature selection by generating and assigning different costs to each class with rigorous theory guidance. After solving a series of cost-sensitive feature selection problems, features corresponding to the best F-measure will be selected. In this way, the selected features will fully represent the properties of all classes. Experimental results on popular benchmarks and challenging real-world data sets demonstrate the significance of cost-sensitive feature selection for the imbalanced data setting and validate the effectiveness of the proposed method. Meng Liu 0003, Chang Xu 0002, Yong Luo 0002, Chao Xu 0006, Yonggang Wen 0001, Dacheng Tao |
IEEE Trans. Image Process. | 3 |
| 2018 | Toward Intelligent Product Retrieval for TV-to-Online (T2O) Application: A Transfer Metric Learning ApproachabstractIt is desired (especially for young people) to shop for the same or similar products shown in the multimedia contents (such as online TV programs). This indicates an urgent demand for improving the experience of TV-to-Online (T2O). In this paper, a transfer learning approach as well as a prototype system for effortless T2O experience is developed. In the system, a key component is high-precision product search, which is to fulfill exact matching between a query item and the database ones. The matching performance primarily relies on distance estimation, but the data characteristics cannot be well modeled and exploited by a simple Euclidean distance. This motivates us to introduce distance metric learning (DML) for improving the distance estimation. However, in traditional DML methods, the side information (such as the similar/dissimilar constraints or relevance/irrelevance judgements) in the target domain is leveraged. These methods may fail due to limited side information. Fortunately, this issue can be alleviated by utilizing transfer metric learning (TML) to exploit information from other related domains. In this paper, a novel manifold regularized heterogeneous multitask metric learning framework is proposed, in which each domain is treated equally. The proposed approach allows us to simultaneously exploit the information from other domains and the unlabeled information. Furthermore, the ranking-based loss is adopted to make our model more appropriate for search. Experiments on two challenging real-world datasets demonstrate the effectiveness of the proposed method. This TML approach is expected to impact the transformation of the emerging T2O trend in both TV and online video domains. Qiang Fu 0006, Yong Luo 0002, Yonggang Wen 0001, Dacheng Tao, Ying Li 0012, Ling-Yu Duan |
IEEE Trans. Multim. | 2 |
| 2018 | Data-Driven Lightweight Interest Point Selection for Large-Scale Visual SearchabstractWith the explosive increase of images and videos, visual analysis has become an essential technique in dealing with the big visual data, which utilizes the visual feature descriptors to search or recognize the images or frames with target objects or events. Subject to the constraints of resources (e.g., memory, bandwidth, storage, etc.), interest point selection is crucial to generate robust compact descriptors for high-efficiency visual analysis by selecting and aggregating the most discriminative local feature descriptors, which has been demonstrated in the state-of-the-art low bit rate visual search works. In this paper, we propose a data-driven lightweight interest point selection approach to significantly improve the performance of visual search, while ameliorating the efficiency of extracting feature descriptors. Comprehensive experimental results over benchmarks have shown that the proposed interest point selection algorithm has significantly improved image matching and retrieval performance in the completed MPEG Compact Descriptors for Visual Search (CDVS) standard as well as the emerging MPEG Compact Descriptors for Video Analytics (CDVA) standard, say 20% mAP gain by data-driven selection against random selection of interest points. In particular, the presented data-driven interest point selection has been adopted by MPEG-CDVS and MPEG-CDVA as a normative technique to improve the aggregation of handcrafted features, which has contributed to the combination of handcrafted features and deep learning (CNN) features as well. Feng Gao 0014, Xinfeng Zhang 0001, Yong Luo 0002, Xiaoming Li 0001, Ling-Yu Duan |
IEEE Trans. Multim. | 4 |
| 2018 | Heterogeneous Multitask Metric Learning Across Multiple DomainsabstractDistance metric learning plays a crucial role in diverse machine learning algorithms and applications. When the labeled information in a target domain is limited, transfer metric learning (TML) helps to learn the metric by leveraging the sufficient information from other related domains. Multitask metric learning (MTML), which can be regarded as a special case of TML, performs transfer across all related domains. Current TML tools usually assume that the same feature representation is exploited for different domains. However, in real-world applications, data may be drawn from heterogeneous domains. Heterogeneous transfer learning approaches can be adopted to remedy this drawback by deriving a metric from the learned transformation across different domains. However, they are often limited in that only two domains can be handled. To appropriately handle multiple domains, we develop a novel heterogeneous MTML (HMTML) framework. In HMTML, the metrics of all different domains are learned together. The transformations derived from the metrics are utilized to induce a common subspace, and the high-order covariance among the predictive structures of these domains is maximized in this subspace. There do exist a few heterogeneous transfer learning approaches that deal with multiple domains, but the high-order statistics (correlation information), which can only be exploited by simultaneously examining all domains, is ignored in these approaches. Compared with them, the proposed HMTML can effectively explore such high-order information, thus obtaining more reliable feature transformations and metrics. Effectiveness of our method is validated by the extensive and intensive experiments on text categorization, scene classification, and social image annotation. Yong Luo 0002, Yonggang Wen 0001, Dacheng Tao |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2017 | Cost-Sensitive Feature Selection via F-Measure Optimization ReductionabstractFeature selection aims to select a small subset from the high-dimensional features which can lead to better learning performance, lower computational complexity, and better model readability. The class imbalance problem has been neglected by traditional feature selection methods, therefore the selected features will be biased towards the majority classes. Because of the superiority of F-measure to accuracy for imbalanced data, we propose to use F-measure as the performance measure for feature selection algorithms. As a pseudo-linear function, the optimization of F-measure can be achieved by minimizing the total costs. In this paper, we present a novel cost-sensitive feature selection (CSFS) method which optimizes F-measure instead of accuracy to take class imbalance issue into account. The features will be selected according to optimal F-measure classifier after solving a series of cost-sensitive feature selection sub-problems. The features selected by our method will fully represent the characteristics of not only majority classes, but also minority classes. Extensive experimental results conducted on synthetic, multi-class and multi-label datasets validate the efficiency and significance of our feature selection method. Meng Liu 0003, Chang Xu 0002, Yong Luo 0002, Chao Xu 0006, Yonggang Wen 0001, Dacheng Tao |
AAAI | 3 |
| 2017 | Exploiting High-Order Information in Heterogeneous Multi-Task Feature LearningabstractMulti-task feature learning (MTFL) aims to improve the generalization performance of multiple related learning tasks by sharing features between them. It has been successfully applied to many pattern recognition and biometric prediction problems. Most of current MTFL methods assume that different tasks exploit the same feature representation, and thus are not applicable to the scenarios where data are drawn from heterogeneous domains. Existing heterogeneous transfer learning (including multi-task learning) approaches handle multiple heterogeneous domains by usually learning feature transformations across different domains, but they ignore the high-order statistics (correlation information) which can only be discovered by simultaneously exploring all domains. We therefore develop a tensor based heterogeneous MTFL (THMTFL) framework to exploit such high-order information. Specifically, feature transformations of all domains are learned together, and finally used to derive new representations. A connection between all domains is built by using the transformations to project the pre-learned predictive structures of different domains into a common subspace, and minimizing their divergence in the subspace. By exploring the high-order information, the proposed THMTFL can obtain more reliable feature transformations compared with existing heterogeneous transfer learning approaches. Extensive experiments on both text categorization and social image annotation demonstrate superiority of the proposed method. Yong Luo 0002, Dacheng Tao, Yonggang Wen 0001 |
IJCAI | 1 |
| 2017 | General Heterogeneous Transfer Distance Metric Learning via Knowledge Fragments TransferabstractTransfer learning aims to improve the performance of target learning task by leveraging information (or transferring knowledge) from other related tasks. Recently, transfer distance metric learning (TDML) has attracted lots of interests, but most of these methods assume that feature representations for the source and target learning tasks are the same. Hence, they are not suitable for the applications, in which the data are from heterogeneous domains (feature spaces, modalities and even semantics). Although some existing heterogeneous transfer learning (HTL) approaches is able to handle such domains, they lack flexibility in real-world applications, and the learned transformations are often restricted to be linear. We therefore develop a general and flexible heterogeneous TDML (HTDML) framework based on the knowledge fragment transfer strategy. In the proposed HTDML, any (linear or nonlinear) distance metric learning algorithms can be employed to learn the source metric beforehand. Then a set of knowledge fragments are extracted from the pre-learned source metric to help target metric learning. In addition, either linear or nonlinear distance metric can be learned for the target domain. Extensive experiments on both scene classification and object recognition demonstrate superiority of the proposed method. Yong Luo 0002, Yonggang Wen 0001, Tongliang Liu, Dacheng Tao |
IJCAI | 1 |
| 2016 | Toward Effortless TV-to-Online (T2O) Experience: A Novel Metric Learning ApproachabstractShopping of the same or similar types of products as shown in the online TV programs has been highly desired by many people, especially the youth. To meet this eminent market need, we develop a prototype system to enable effortless TV-to-Online (T2O) experience. A key component of this system is the product search that maps specific items embedded in the video into a list of online merchants. The search performance mainly depends on the estimation of the distance or similarity between the queried item and all the curated items in the database. The simple Euclidean (EU) distance cannot capture the data characteristics, we therefore introduce distance metric learning (DML) to improve the distance estimation. Traditional DML methods only utilize the side information (e.g., similar/dissimilar constraints or relevance/irrelevance judgements) in the target domain, and may fail when the side information is scarce. Transfer metric learning (TML) can be adopted to leverage the side information from related domains. In this paper, we treat each domain equally and propose a novel Ranking-based Heterogeneous Multi-Task Metric Learning (RHMTML) framework, which adopts ranking-based loss, so the learned metric is particularly suitable for search. Extensive experiments demonstrate the effectiveness of our proposed method. We foresee that our approach would transform the emerging T2O trend in both TV and online video market. Yong Luo 0002, Yonggang Wen 0001, Dacheng Tao, Qiang Fu 0006 |
GLOBECOM | 1 |
| 2016 | Tensor canonical correlation analysis for multi-view dimension reductionabstractCanonical correlation analysis (CCA) has proven an effective tool for two-view dimension reduction due to its profound theoretical foundation and success in practical applications. In respect of multi-view learning, however, it is limited by its capability of only handling data represented by two-view features, while in many real-world applications, the number of views is frequently many more. Although the ad hoc way of simultaneously exploring all possible pairs of features can numerically deal with multi-view data, it ignores the high order statistics (correlation information) which can only be discovered by simultaneously exploring all features. Therefore, in this work, we develop tensor CCA (TCCA) which straightforwardly yet naturally generalizes CCA to handle the data of an arbitrary number of views by analyzing the covariance tensor of the different views. TCCA aims to directly maximize the canonical correlation of multiple (more than two) views. Crucially, we prove that the main problem of multiview canonical correlation maximization is equivalent to finding the best rank-1 approximation of the data covariance tensor, which can be solved efficiently using the well-known alternating least squares (ALS) algorithm. As a consequence, the high order correlation information contained in the different views is explored and thus a more reliable common subspace shared by all features can be obtained. Yong Luo 0002, Dacheng Tao, Kotagiri Ramamohanarao, Chao Xu 0006, Yonggang Wen 0001 |
ICDE | 1 |
| 2016 | On Combining Side Information and Unlabeled Data for Heterogeneous Multi-Task Metric Learning
Yong Luo 0002, Yonggang Wen 0001, Dacheng Tao |
IJCAI | 1 |
| 2016 | Large Margin Multi-Modal Multi-Task Feature Extraction for Image ClassificationabstractThe features used in many image analysis-based applications are frequently of very high dimension. Feature extraction offers several advantages in high-dimensional cases, and many recent studies have used multi-task feature extraction approaches, which often outperform single-task feature extraction approaches. However, most of these methods are limited in that they only consider data represented by a single type of feature, even though features usually represent images from multiple modalities. We, therefore, propose a novel large margin multi-modal multi-task feature extraction (LM3FE) framework for handling multi-modal features for image classification. In particular, LM3FE simultaneously learns the feature extraction matrix for each modality and the modality combination coefficients. In this way, LM3FE not only handles correlated and noisy features, but also utilizes the complementarity of different modalities to further help reduce feature redundancy in each modality. The large margin principle employed also helps to extract strongly predictive features, so that they are more suitable for prediction (e.g., classification). An alternating algorithm is developed for problem optimization, and each subproblem can be efficiently solved. Experiments on two challenging real-world image data sets demonstrate the effectiveness and superiority of the proposed method. Yong Luo 0002, Yonggang Wen 0001, Dacheng Tao, Jie Gui, Chao Xu 0006 |
IEEE Trans. Image Process. | 1 |
| 2015 | Low-Rank Multi-View Learning in Matrix Completion for Multi-Label Image ClassificationabstractMulti-label image classification is of significant interest due to its major role in real-world web image analysis applications such as large-scale image retrieval and browsing. Recently, matrix completion (MC) has been developed to deal with multi-label classification tasks. MC has distinct advantages, such as robustness to missing entries in the feature and label spaces and a natural ability to handle multi-label problems. However, current MC-based multi-label image classification methods only consider data represented by a single-view feature, therefore, do not precisely characterize images that contain several semantic concepts. An intuitive way to utilize multiple features taken from different views is to concatenate the different features into a long vector; however, this concatenation is prone to over-fitting and leads to high time complexity in MC-based image classification. Therefore, we present a novel multi-view learning model for MC-based image classification, called low-rank multi-view matrix completion (lrMMC), which first seeks a low-dimensional common representation of all views by utilizing the proposed low-rank multi-view learning (lrMVL) algorithm. In lrMVL, the common subspace is constrained to be low rank so that it is suitable for MC. In addition, combination weights are learned to explore complementarity between different views. An efficient solver based on fixed-point continuation (FPC) is developed for optimization, and the learned low-rank representation is then incorporated into MC-based image classification. Extensive experimentation on the challenging PASCAL VOC' 07 dataset demonstrates the superiority of lrMMC compared to other multi-label image classification approaches. Meng Liu 0003, Yong Luo 0002, Dacheng Tao, Chao Xu 0006, Yonggang Wen 0001 |
AAAI | 2 |
| 2015 | Manifold Regularized Transfer Distance Metric LearningabstractThe performance of many computer vision and machine learning algorithms are heavily depend on the distance metric between samples. It is necessary to exploit abundant of side information like pairwise constraints to learn a robust and reliable distance metric[2, 3]. Let D = {(xl i ,xj,yi j)} l i, j=1 denotes the labeled training set for the target task, wherein xi, x j ∈ Rd and yi j = ±1 indicates xl i and xl i are similar/dissimilar to each other. Then, a metric is usually learned to minimize the distance between the data from the same class and maximize their distance otherwise. This leads to the following loss function for learning the metric A: Haibo Shi, Yong Luo 0002, Chao Xu 0006, Yonggang Wen 0001 |
BMVC | 2 |
| 2015 | Multiview Matrix Completion for Multilabel Image ClassificationabstractThere is growing interest in multilabel image classification due to its critical role in web-based image analytics-based applications, such as large-scale image retrieval and browsing. Matrix completion (MC) has recently been introduced as a method for transductive (semisupervised) multilabel classification, and has several distinct advantages, including robustness to missing data and background noise in both feature and label space. However, it is limited by only considering data represented by a single-view feature, which cannot precisely characterize images containing several semantic concepts. To utilize multiple features taken from different views, we have to concatenate the different features as a long vector. However, this concatenation is prone to over-fitting and often leads to very high time complexity in MC-based image classification. Therefore, we propose to weightedly combine the MC outputs of different views, and present the multiview MC (MVMC) framework for transductive multilabel image classification. To learn the view combination weights effectively, we apply a cross-validation strategy on the labeled set. In particular, MVMC splits the labeled set into two parts, and predicts the labels of one part using the known labels of the other part. The predicted labels are then used to learn the view combination coefficients. In the learning process, we adopt the average precision (AP) loss, which is particular suitable for multilabel image classification, since the ranking-based criteria are critical for evaluating a multilabel classification system. A least squares loss formulation is also presented for the sake of efficiency, and the robustness of the algorithm based on the AP loss compared with the other losses is investigated. Experimental evaluation on two real-world data sets (PASCAL VOC' 07 and MIR Flickr) demonstrate the effectiveness of MVMC for transductive (semisupervised) multilabel image classification, and show that MVMC can exploit complementary properties of different features and output-consistent labels for improved multilabel image classification. Yong Luo 0002, Tongliang Liu, Dacheng Tao, Chao Xu 0006 |
IEEE Trans. Image Process. | 1 |
| 2015 | Tensor Canonical Correlation Analysis for Multi-View Dimension ReductionabstractCanonical correlation analysis (CCA) has proven an effective tool for two-view dimension reduction due to its profound theoretical foundation and success in practical applications. In respect of multi-view learning, however, it is limited by its capability of only handling data represented by two-view features, while in many real-world applications, the number of views is frequently many more. Although the ad hoc way of simultaneously exploring all possible pairs of features can numerically deal with multi-view data, it ignores the high order statistics (correlation information) which can only be discovered by simultaneously exploring all features. Therefore, in this work, we develop tensor CCA (TCCA) which straightforwardly yet naturally generalizes CCA to handle the data of an arbitrary number of views by analyzing the covariance tensor of the different views. TCCA aims to directly maximize the canonical correlation of multiple (more than two) views. Crucially, we prove that the main problem of multi-view canonical correlation maximization is equivalent to finding the best rank-1 approximation of the data covariance tensor, which can be solved efficiently using the well-known alternating least squares (ALS) algorithm. As a consequence, the high order correlation information contained in the different views is explored and thus a more reliable common subspace shared by all features can be obtained. In addition, a non-linear extension of TCCA is presented. Experiments on various challenge tasks, including large scale biometric structure prediction, internet advertisement classification, and web image annotation, demonstrate the effectiveness of the proposed method. Yong Luo 0002, Dacheng Tao, Kotagiri Ramamohanarao, Chao Xu 0006, Yonggang Wen 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2014 | Pre-Trained Multi-View Word Embedding Using Two-Side Neural NetworkabstractWord embedding aims to learn a continuous representation for each word. It attracts increasing attention due to its effectiveness in various tasks such as named entity recognition and language modeling. Most existing word embedding results are generally trained on one individual data source such as news pages or Wikipedia articles. However, when we apply them to other tasks such as web search, the performance suffers. To obtain a robust word embedding for different applications, multiple data sources could be leveraged. In this paper, we proposed a two-side multimodal neural network to learn a robust word embedding from multiple data sources including free text, user search queries and search click-through data. This framework takes the word embeddings learned from different data sources as pre-train, and then uses a two-side neural network to unify these embeddings. The pre-trained embeddings are obtained by adapting the recently proposed CBOW algorithm. Since the proposed neural network does not need to re-train word embeddings for a new task, it is highly scalable in real world problem solving. Besides, the network allows weighting different sources differently when applied to different application tasks. Experiments on two real-world applications including web search ranking and word similarity measuring show that our neural network with multiple sources outperforms state-of-the-art word embedding algorithm with each individual source. It also outperforms other competitive baselines using multiple sources. Yong Luo 0002, Jian Tang 0005, Jun Yan 0001, Chao Xu 0006, Zheng Chen 0001 |
AAAI | 1 |
| 2014 | Multi-view Multi-task Feature Extraction for Web Image ClassificationabstractThe features used in many multimedia analysis-based applications are frequently of very high dimension. Feature extraction offers several advantages in highly dimensional cases, and many recent studies have used multi-task feature extraction approaches, which often outperform single-task feature extraction approaches. However, most of these methods are limited in that they only consider data represented by a single type of feature, even though features usually represent images from multiple views. We therefore propose a novel multi-view multi-task feature extraction (MVMTFE) framework for handling multi-view features for image classification. In particular, MVMTFE simultaneously learns the feature extraction matrix for each view and the view combination coefficients. In this way, MVMTFE not only handles correlated and noisy features, but also utilizes the complementarity of different views to further help reduce feature redundancy in each view. An alternating algorithm is developed for problem optimization and each sub-problem can be efficiently solved. Experiments on an real-world web image dataset demonstrate the effectiveness and superiority of the proposed method. Zhiqiang Zuo 0001, Yong Luo 0002, Dacheng Tao, Chao Xu 0006 |
ACM Multimedia | 2 |
| 2014 | Group Sparse Multiview Patch Alignment Framework With View Consistency for Image ClassificationabstractNo single feature can satisfactorily characterize the semantic concepts of an image. Multiview learning aims to unify different kinds of features to produce a consensual and efficient representation. This paper redefines part optimization in the patch alignment framework (PAF) and develops a group sparse multiview patch alignment framework (GSM-PAF). The new part optimization considers not only the complementary properties of different views, but also view consistency. In particular, view consistency models the correlations between all possible combinations of any two kinds of view. In contrast to conventional dimensionality reduction algorithms that perform feature extraction and feature selection independently, GSM-PAF enjoys joint feature extraction and feature selection by exploiting l(2,1)-norm on the projection matrix to achieve row sparsity, which leads to the simultaneous selection of relevant features and learning transformation, and thus makes the algorithm more discriminative. Experiments on two real-world image data sets demonstrate the effectiveness of GSM-PAF for image classification. Jie Gui, Dacheng Tao, Zhenan Sun, Yong Luo 0002, Xinge You, Yuan Yan Tang |
IEEE Trans. Image Process. | 4 |
| 2014 | Decomposition-Based Transfer Distance Metric Learning for Image ClassificationabstractDistance metric learning (DML) is a critical factor for image analysis and pattern recognition. To learn a robust distance metric for a target task, we need abundant side information (i.e., the similarity/dissimilarity pairwise constraints over the labeled data), which is usually unavailable in practice due to the high labeling cost. This paper considers the transfer learning setting by exploiting the large quantity of side information from certain related, but different source tasks to help with target metric learning (with only a little side information). The state-of-the-art metric learning algorithms usually fail in this setting because the data distributions of the source task and target task are often quite different. We address this problem by assuming that the target distance metric lies in the space spanned by the eigenvectors of the source metrics (or other randomly generated bases). The target metric is represented as a combination of the base metrics, which are computed using the decomposed components of the source metrics (or simply a set of random bases); we call the proposed method, decomposition-based transfer DML (DTDML). In particular, DTDML learns a sparse combination of the base metrics to construct the target metric by forcing the target metric to be close to an integration of the source metrics. The main advantage of the proposed method compared with existing transfer metric learning approaches is that we directly learn the base metric coefficients instead of the target metric. To this end, far fewer variables need to be learned. We therefore obtain more reliable solutions given the limited side information and the optimization tends to be faster. Experiments on the popular handwritten image (digit, letter) classification and challenge natural image annotation tasks demonstrate the effectiveness of the proposed method. Yong Luo 0002, Tongliang Liu, Dacheng Tao, Chao Xu 0006 |
IEEE Trans. Image Process. | 1 |
| 2013 | Vector-Valued Multi-View Semi-Supervsed Learning for Multi-Label Image ClassificationabstractImages are usually associated with multiple labels and comprised of multiple views, due to each image containing several objects (e.g. a pedestrian, bicycle and tree) and multiple visual features (e.g. color, texture and shape). Currently available tools tend to use either labels or features for classification, but both are necessary to describe the image properly. There have been recent successes in using vector-valued functions, which construct matrix-valued kernels, to explore the multi-label structure in the output space. This has motivated us to develop multi-view vector-valued manifold regularization (MV$^3$MR) in order to integrate multiple features. MV$^3$MR exploits the complementary properties of different features, and discovers the intrinsic local geometry of the compact support shared by different features, under the theme of manifold regularization. We validate the effectiveness of the proposed MV$^3$MR methodology for image classification by conducting extensive experiments on two challenge datasets, PASCAL VOC' 07 and MIR Flickr. Yong Luo 0002, Dacheng Tao, Chang Xu 0002, Chao Xu 0006 |
AAAI | 1 |
| 2013 | FIM: A Real-Time Content Based Sample Image Matching SystemabstractSample Image Matching is to decide if a queried image is belongs to the database or not. In this paper, we focus on real-time image matching, which is critical in many real world applications. Although traditional image retrieval methods can be directly utilized for image matching, they usually suffer the high computational cost problem and thus is not applicable here. To resolve this problem, we first introduce ORB, a recently proposed and well-established image feature, for image matching. Then we compare several variants of the descriptors, different size of the codebook, and two approaches to compute the matching scores, based on which we propose a strategy for final matching decision. According to the comparison results, we finally present a real-time image matching system, fast image matching (FIM), which can process about 33 images per second, with a satisfactory accuracy. Yong Luo 0002, Yangxi Li, Jinhui Tu, Chao Xu 0006 |
ICIG | 1 |
| 2013 | Manifold Regularized Multitask Learning for Semi-Supervised Multilabel Image ClassificationabstractIt is a significant challenge to classify images with multiple labels by using only a small number of labeled samples. One option is to learn a binary classifier for each label and use manifold regularization to improve the classification performance by exploring the underlying geometric structure of the data distribution. However, such an approach does not perform well in practice when images from multiple concepts are represented by high-dimensional visual features. Thus, manifold regularization is insufficient to control the model complexity. In this paper, we propose a manifold regularized multitask learning (MRMTL) algorithm. MRMTL learns a discriminative subspace shared by multiple classification tasks by exploiting the common structure of these tasks. It effectively controls the model complexity because different tasks limit one another's search volume, and the manifold regularization ensures that the functions in the shared hypothesis space are smooth along the data manifold. We conduct extensive experiments, on the PASCAL VOC'07 dataset with 20 classes and the MIR dataset with 38 classes, by comparing MRMTL with popular image classification algorithms. The results suggest that MRMTL is effective for image classification. Yong Luo 0002, Dacheng Tao, Bo Geng, Chao Xu 0006, Stephen J. Maybank |
IEEE Trans. Image Process. | 1 |
| 2013 | Multiview Vector-Valued Manifold Regularization for Multilabel Image ClassificationabstractIn computer vision, image datasets used for classification are naturally associated with multiple labels and comprised of multiple views, because each image may contain several objects (e.g., pedestrian, bicycle, and tree) and is properly characterized by multiple visual features (e.g., color, texture, and shape). Currently, available tools ignore either the label relationship or the view complementarily. Motivated by the success of the vector-valued function that constructs matrix-valued kernels to explore the multilabel structure in the output space, we introduce multiview vector-valued manifold regularization (MV(3)MR) to integrate multiple features. MV(3)MR exploits the complementary property of different features and discovers the intrinsic local geometry of the compact support shared by different features under the theme of manifold regularization. We conduct extensive experiments on two challenging, but popular, datasets, PASCAL VOC' 07 and MIR Flickr, and validate the effectiveness of the proposed MV(3)MR for image classification. Yong Luo 0002, Dacheng Tao, Chang Xu 0002, Chao Xu 0006, Hong Liu 0008, Yonggang Wen 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2011 | Shared feature extraction for semi-supervised image classificationabstractMulti-task learning (MTL) plays an important role in image analysis applications, e.g. image classification, face recognition and image annotation. That is because MTL can estimate the latent shared subspace to represent the common features given a set of images from different tasks. However, the geometry of the data probability distribution is always supported on an intrinsic image sub-manifold that is embedded in a high dimensional Euclidean space. Therefore, it is improper to directly apply MTL to multiclass image classification. In this paper, we propose a manifold regularized MTL (MRMTL) algorithm to discover the latent shared subspace by treating the high-dimensional image space as a sub-manifold embedded in an ambient space. We conduct experiments on the PASCAL VOC'07 dataset with 20 classes and the MIR dataset with 38 classes by comparing MRMTL with conventional MTL and several representative image classification algorithms. The results suggest that MRMTL can properly extract the common features for image representation and thus improve the generalization performance of the image classification models. Yong Luo 0002, Dacheng Tao, Bo Geng, Chao Xu 0006, Stephen J. Maybank |
ACM Multimedia | 1 |
| 2011 | Query Difficulty Guided Image Retrieval System
Yangxi Li, Yong Luo 0002, Dacheng Tao, Chao Xu 0006 |
MMM (2) | 2 |