EDBT 2026 Demo / reviewers in the wild / expert
Yang Li 0104
dblp:37/4190-104
· DBLP profile ↗
50ranked-venue papers
3as first author
37since 2021 · last 2026
0000-0002-2053-6393ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 3 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 16 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 3 first-author · 11 since 2021Databases, data management, data science and information retrieval · 9 · 3 first-author · 5 since 2021Computer networks · 5 · 3 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning Optimal Prompt Ensemble for Multi-source Visual Prompt TransferabstractPrompt tuning has emerged as a lightweight strategy for adapting foundation models to downstream tasks, particularly for resource-constrained systems. As pre-trained prompts become valuable assets, combining multiple source prompts offers a promising approach to enhance generalization for new tasks by leveraging complementary knowledge. However, naive aggregation often overlooks different source prompts have different contribution potential to the target task. To address this, we propose HGPrompt, a dynamic framework that learns optimal ensemble weights. These weights are optimized by jointly maximizing an information-theoretic metric for transferability and minimizing gradient conflicts via a novel regularization strategy. Specifically, we propose a differentiable prompt transferability metric to captures the discriminability of prompt-induced features on the target task. Meanwhile, HGPrompt match the gradient variances with respect to different source prompts based on Hessian and Fisher Information, ensuring stable and coherent knowledge transfer while suppressing gradient conflicts among them. Extensive experiments on the large-scale VTAB benchmark demonstrate the state-of-the-art performance of HGPrompt, validating its effectiveness in learning an optimal ensemble for effective multi-source prompt transfer. Enming Zhang, Liwen Cao, Yanru Wu, Yang Li 0104 |
AAAI | 5 |
| 2026 | ACE-ProtoNet: Adaptive covariance eigen-gate and uncertainty-aware prototype learning for coronary artery segmentation
Caixia Dong, Duwei Dai, Guowei Dai 0001, Linyun Zhou, Yang Li 0104 |
Medical Image Anal. | 7 |
| 2026 | High-quality coronary artery segmentation via fuzzy logic modeling coupled with dynamic graph convolutional network
Caixia Dong, Duwei Dai, Yang Li 0104, Songhua Xu |
Pattern Recognit. | 3 |
| 2026 | S2ML: Spatio-Spectral Mutual Learning for Depth CompletionabstractThe raw depth images captured by RGB-D cameras using Time-of-Flight (TOF) or structured light often suffer from incomplete depth values due to weak reflections, boundary shadows, and artifacts, which limit their applications in downstream vision tasks. Existing methods address this problem through depth completion in the image domain, but they overlook the physical characteristics of raw depth images. It has been observed that the presence of invalid depth areas alters the frequency distribution pattern. In this work, we propose a Spatio-Spectral Mutual Learning framework (S2ML) to harmonize the advantages of both spatial and frequency domains for depth completion. Specifically, we consider the distinct properties of amplitude and phase spectra and devise a dedicated spectral fusion module. Meanwhile, the local and global correlations between spatial-domain and frequency-domain features are calculated in a unified embedding space. The gradual mutual representation and refinement encourage the network to fully explore complementary physical characteristics and priors for more accurate depth completion. Extensive experiments demonstrate the effectiveness of our proposed S2ML method, outperforming the state-of-the-art method CFormer by 0.828 dB and 0.834 dB on the NYU-Depth V2 and SUN RGB-D datasets, respectively. Zihui Zhao, Zheng Wang 0007, Yang Li 0104, Kui Jiang, Zihan Geng, Chia-Wen Lin |
IEEE Trans. Multim. | 4 |
| 2025 | pFedGPA: Diffusion-based Generative Parameter Aggregation for Personalized Federated LearningabstractFederated Learning (FL) offers a decentralized approach to model training, where data remains local and only model parameters are shared between the clients and the central server. Traditional methods, such as Federated Averaging (FedAvg), linearly aggregate these parameters which are usually trained on heterogeneous data distributions, potentially overlooking the complex, high-dimensional nature of the parameter space. This can result in degraded performance of the aggregated model. While personalized FL approaches can mitigate the heterogeneous data issue to some extent, the limitation of linear aggregation remains unresolved. To alleviate this issue, we investigate the generative approach of diffusion model and propose a novel generative parameter aggregation framework for personalized FL, pFedGPA. In this framework, we deploy a diffusion model on the server to integrate the diverse parameter distributions and propose a parameter inversion method to efficiently generate a set of personalized parameters for each client. This inversion method transforms the uploaded parameters into a latent code, which is then aggregated through denoising sampling to produce the final personalized parameters. By encoding the dependence of a client's model parameters on the specific data distribution using the high-capacity diffusion model, pFedGPA can effectively decouple the complexity of the overall distribution of all clients' model parameters from the complexity of each individual client's parameter distribution. Our experimental results consistently demonstrate the superior performance of the proposed method across multiple datasets, surpassing baseline approaches. Jiahao Lai, Jiaqi Li 0028, Jian Xu 0016, Yanru Wu, Boshi Tang, Wenbo Ding 0001, Yang Li 0104 |
AAAI | 9 |
| 2025 | Transfer Risk Map: Mitigating Pixel-level Negative Transfer in Medical SegmentationabstractHow to mitigate negative transfer in transfer learning is a long-standing and challenging issue, especially in the application of medical image segmentation. Existing methods for reducing negative transfer focuses on classification or regression tasks, ignoring the non-uniform negative transfer risk in different image regions. In this work, we propose a simple yet effective weighted fine-tuning method that directs the model’s attention towards regions with significant transfer risk for medical semantic segmentation. Specifically, we compute a transferability-guided transfer risk map to quantify the transfer hardness for each pixel and the potential risks of negative transfer. During the fine-tuning phase, we introduce a map-weighted loss function, normalized with image foreground size to counter class imbalance. Extensive experiments on brain segmentation datasets show our method significantly improves the target task performance, with gains of 4.37% on FeTS2021 and 1.81% on iSeg-2019, avoiding negative transfer across modalities and tasks. Meanwhile, a 2.9% gain under a few-shot scenario validates the robustness of our approach. Shutong Duan, Yang Tan 0004, Yang Li 0104, Xiao-Ping Zhang 0002 |
ICASSP | 5 |
| 2025 | Reinforced Domain Selection for Continuous Domain AdaptationabstractContinuous Domain Adaptation (CDA) effectively bridges significant domain shifts by progressively adapting from the source domain through intermediate domains to the target domain. However, selecting intermediate domains without explicit metadata remains a substantial challenge that has not been extensively explored in existing studies. To tackle this issue, we propose a novel framework that combines reinforcement learning with feature disentanglement to conduct domain path selection in an unsupervised CDA setting. Our approach introduces an innovative unsupervised reward mechanism that leverages the distances between latent domain embeddings to facilitate the identification of optimal transfer paths. Furthermore, by disentangling features, our method facilitates the calculation of unsupervised rewards using domain-specific features and promotes domain adaptation by aligning domain-invariant features. This integrated strategy is designed to simultaneously optimize transfer paths and target task performance, enhancing the effectiveness of domain adaptation processes. Extensive empirical evaluations on datasets such as Rotated MNIST and ADNI demonstrate substantial improvements in prediction accuracy and domain selection efficiency, establishing our method’s superiority over traditional CDA approaches. Huaze Tang, Yanru Wu, Yang Li 0104, Xiao-Ping Zhang 0002 |
ICASSP | 4 |
| 2025 | Causal-aware Graph Neural Architecture Search under Distribution ShiftsabstractGraph neural architecture search (NAS) has emerged as a promising approach for autonomously designing graph neural network architectures by leveraging correlations between graphs and architectures. However, existing methods merely rely on correlations, which may be spurious and vary across distributions. This reliance, without considering causal graph-architecture relationships, limits their ability to generalize under distribution shifts that are ubiquitous in real-world graph scenarios. In this paper, we propose to handle the distribution shifts in NAS process by exploiting the causal graph-architecture relationship to search for optimal architectures that can generalize under distribution shifts. Key challenges remain unexplored: discovering causal graph-architecture relationships with stable cross-distribution predictive abilities, and leveraging them to handle distribution shifts. To address these challenges, we propose a novel approach, Causal-aware Graph Neural Architecture Search (CARNAS), which is capable of capturing causal graph-architecture relationship during NAS process and discovering optimal graph architecture under distribution shifts. We propose Disentangled Causal Subgraph Identification to extract causal subgraphs with stable predictive power across distributions, followed by Graph Embedding Intervention to intervene on these subgraphs in latent space by preserving essential features while filtering out non-causal elements, and Invariant Architecture Customization to enhance their causal invariance for optimizing graph architectures. Extensive experiments on synthetic and real-world datasets show that CARNAS enhances out-of-distribution generalization by uncovering causal graph-architecture relationships during NAS. Peiwen Li, Xin Wang 0019, Zeyang Zhang 0001, Ziwei Zhang 0001, Fang Shen, Jialong Wang 0001, Yang Li 0104, Wenwu Zhu 0001 |
KDD (2) | 7 |
| 2025 | Hierarchical Part-Based Generative Model for Realistic 3D Blood Vessel
Jiahao Lai, Bingzhi Shen, Sihong Zhang, Caixia Dong, Xuejin Chen, Yang Li 0104 |
MICCAI (3) | 8 |
| 2025 | Hierarchical Feature Learning for Medical Point Clouds via State Space Model
Yang Li 0104 |
MICCAI (10) | 3 |
| 2025 | Exploiting Task Relationships in Continual Learning via Transferability-Aware Task EmbeddingsabstractContinual learning (CL) has been a critical topic in contemporary deep neural network applications, where higher levels of both forward and backward transfer are desirable for an effective CL performance. Existing CL strategies primarily focus on task models — either by regularizing model updates or by separating task-specific and shared components — while often overlooking the potential of leveraging inter-task relationships to enhance transfer. To address this gap, we propose a transferability-aware task embedding, termed H-embedding, and construct a hypernet framework under its guidance to learn task-conditioned model weights for CL tasks. Specifically, H-embedding is derived from an information theoretic measure of transferability and is designed to be online and easy to compute. Our method is also characterized by notable practicality, requiring only the storage of a low-dimensional task embedding per task and supporting efficient end-to-end training. Extensive evaluations on benchmarks including CIFAR-100, ImageNet-R, and DomainNet show that our framework performs prominently compared to various baseline and SOTA approaches, demonstrating strong potential in capturing and utilizing intrinsic task relationships. Our code is publicly available at \url{https://github.com/viki760/Hembedding_Guided_Hypernet}. Yanru Wu, Jianning Wang, Aurora, Yang Li 0104 |
NeurIPS | 7 |
| 2025 | A High-Dimensional Statistical Method for Optimizing Transfer Quantities in Multi-Source Transfer LearningabstractMulti-source transfer learning provides an effective solution to data scarcity in real-world supervised learning scenarios by leveraging multiple source tasks. In this field, existing works typically use all available samples from sources in training, which constrains their training efficiency and may lead to suboptimal results. To address this, we propose a theoretical framework that answers the question: what is the optimal quantity of source samples needed from each source task to jointly train the target model? Specifically, we introduce a generalization error measure based on K-L divergence, and minimize it based on high-dimensional statistical analysis to determine the optimal transfer quantity for each source task. Additionally, we develop an architecture-agnostic and data-efficient algorithm OTQMS to implement our theoretical results for target model training in multi-source transfer learning. Experimental studies on diverse architectures and two real-world benchmark datasets show that our proposed algorithm significantly outperforms state-of-the-art approaches in both accuracy and data efficiency. The code is available at https://github.com/zqy0126/OTQMS. Qingyue Zhang 0003, Haohao Fu, Guanbo Huang, Yaoyuan Liang, Chang Chu, Tianren Peng, Yanru Wu, Qi Li 0002, Yang Li 0104, Shao-Lun Huang |
NeurIPS | 9 |
| 2025 | Cauchy Graph Convolutional NetworksabstractA common approach to learning Bayesian networks involves specifying an appropriately chosen family of parameterized probability density such as Gaussian. However, the distribution of most real-life data is leptokurtic and may not necessarily be best described by a Gaussian process. In this work we introduce Cauchy Graphical Models (CGM), a class of multivariate Cauchy densities that can be represented as directed acyclic graphs with arbitrary network topologies, the edges of which encode linear dependencies between random variables. We develop CGLearn, the resultant algorithm for learning the structure and Cauchy parameters based on Minimum Dispersion Criterion (MDC). Experiments using simulated datasets on benchmark network topologies demonstrate the efficacy of our approach when compared to Gaussian Graphical Models (GGM). Most Graph Convolutional Neural Networks (GCN) process input graphs as ground-truth representations of node relationships, yet these graphs are constructed based on modeling assumptions and noisy data and their use may lead to suboptimal performance on downstream prediction tasks. We propose Cauchy GCN which leverages CGM to infer graph topology that depicts latent relationships between nodes. We evaluate the effectiveness and quality of the structural graphs learned by CGM, and demonstrate that Cauchy-GCN achieves superior performance compared to widely used graph construction methods. Taurai Muvunza, Yang Li 0104, Ercan E. Kuruoglu |
Int. J. Approx. Reason. | 2 |
| 2025 | Transferability-Guided Cross-Domain Cross-Task Transfer LearningabstractWe propose two novel transferability metrics fast optimal transport-based conditional entropy (F-OTCE) and joint correspondence OTCE (JC-OTCE) to evaluate how much the source model (task) can benefit the learning of the target task and to learn more generalizable representations for cross-domain cross-task transfer learning. Unlike the original OTCE metric that requires evaluating the empirical transferability on auxiliary tasks, our metrics are auxiliary-free such that they can be computed much more efficiently. Specifically, F-OTCE estimates transferability by first solving an optimal transport (OT) problem between source and target distributions and then uses the optimal coupling to compute the negative conditional entropy (NCE) between the source and target labels. It can also serve as an objective function to enhance downstream transfer learning tasks including model finetuning and domain generalization (DG). Meanwhile, JC-OTCE improves the transferability accuracy of F-OTCE by including label distances in the OT problem, though it incurs additional computation costs. Extensive experiments demonstrate that F-OTCE and JC-OTCE outperform state-of-the-art auxiliary-free metrics by 21.1% and 25.8%, respectively, in correlation coefficient with the ground-truth transfer accuracy. By eliminating the training cost of auxiliary tasks, the two metrics reduce the total computation time of the previous method from 43 min to 9.32 and 10.78 s, respectively, for a pair of tasks. When applied in the model finetuning and DG tasks, F-OTCE shows significant improvements in the transfer accuracy in few-shot classification experiments, with up to 4.41% and 2.34% accuracy gains, respectively. Yang Tan 0004, Enming Zhang, Yang Li 0104, Shao-Lun Huang, Xiao-Ping Zhang 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Learning Persistent Community Structures in Dynamic Networks via Topological Data AnalysisabstractDynamic community detection methods often lack effective mechanisms to ensure temporal consistency, hindering the analysis of network evolution. In this paper, we propose a novel deep graph clustering framework with temporal consistency regularization on inter-community structures, inspired by the concept of minimal network topological changes within short intervals. Specifically, to address the representation collapse problem, we first introduce MFC, a matrix factorization-based deep graph clustering algorithm that preserves node embedding. Based on static clustering results, we construct probabilistic community networks and compute their persistence homology, a robust topological measure, to assess structural similarity between them. Moreover, a novel neural network regularization TopoReg is introduced to ensure the preservation of topological similarity between inter-community structures over time intervals. Our approach enhances temporal consistency and clustering accuracy on real-world datasets with both fixed and varying numbers of communities. It is also a pioneer application of TDA in temporally persistent community detection, offering an insightful contribution to field of network analysis. Code and data are available at the public git repository: https://github.com/kundtx/MFC-TopoReg. Dexu Kong, Anping Zhang, Yang Li 0104 |
AAAI | 3 |
| 2024 | H-ensemble: An Information Theoretic Approach to Reliable Few-Shot Multi-Source-Free TransferabstractMulti-source transfer learning is an effective solution to data scarcity by utilizing multiple source tasks for the learning of the target task. However, access to source data and model details is limited in the era of commercial models, giving rise to the setting of multi-source-free (MSF) transfer learning that aims to leverage source domain knowledge without such access. As a newly defined problem paradigm, MSF transfer learning remains largely underexplored and not clearly formulated. In this work, we adopt an information theoretic perspective on it and propose a framework named H-ensemble, which dynamically learns the optimal linear combination, or ensemble, of source models for the target task, using a generalization of maximal correlation regression. The ensemble weights are optimized by maximizing an information theoretic metric for transferability. Compared to previous works, H-ensemble is characterized by: 1) its adaptability to a novel and realistic MSF setting for few-shot target tasks, 2) theoretical reliability, 3) a lightweight structure easy to interpret and adapt. Our method is empirically validated by ablation studies, along with extensive comparative analysis with other task ensemble and transfer learning methods. We show that the H-ensemble can successfully learn the optimal task ensemble, as well as outperform prior arts. Yanru Wu, Jianning Wang, Weida Wang, Yang Li 0104 |
AAAI | 4 |
| 2024 | Graph-guided Source Selection with Sequential Transfer for Medical Image SegmentationabstractTransfer learning is a critical technique in training deep neural networks to tackle downstream tasks with better or faster solutions, by leveraging previously acquired knowledge. This is especially useful in the low-resource medical image analysis field. However, transferring knowledge from a less related source task can negatively impact the target task’s performance, making the selection of appropriate source tasks vital. To make up for its deficiency when applying transfer learning to medical image segmentation, we propose a novel source selection framework to identify the landmark source with an effective sequential transfer path for the given target task. Specifically, we first construct a comprehensive graph that reflects the relatedness among source tasks, using a medical-image-tailored task affinity metric. Guided by the graph, we assess the informativeness and representativeness of each node and identify the landmark source for the target medical segmentation task. To ensure a positive transfer, we pinpoint a sequential transfer path to the target by minimizing both transfer and search costs. By optimizing the use of source tasks, we consequently improve the transfer performance on the target task. Extensive experiments on three brain MRI medical datasets demonstrate the efficacy of the proposed framework in finding the best source sequence. The results show that our method outperforms other transfer learning approaches by a considerable margin, improving state-of-the-art performance by 6.61% for FeTS 2022, 0.66% for iSeg-2019, and 1.70% for WMH in terms of segmentation Dice score. Jingge Wang, Yang Li 0104 |
BIBM | 4 |
| 2024 | Flemme: A Flexible and Modular Learning Platform for Medical ImagesabstractWith the rapid development of computer vision and the emergence of powerful network backbones and architectures, the application of deep learning in medical imaging has become increasingly significant. Unlike natural images, medical images lack huge volumes of data but feature more modalities, making it difficult to train a general model that has satisfactory performance across various datasets. In practice, practitioners often suffer from manually creating and testing models combining independent backbones and architectures, which is a laborious and time-consuming process. We propose Flemme, a FLExible and Modular learning platform for MEdical images. Our platform separates encoders from the model architectures so that different models can be constructed via various combinations of supported encoders and architectures. We construct encoders using building blocks based on convolution, transformer, and state-space model (SSM) to process both 2D and 3D image patches. A base architecture is implemented following an encoder-decoder style, with several derived architectures for image segmentation, reconstruction, and generation tasks. In addition, we propose a general hierarchical architecture incorporating a pyramid loss to optimize and fuse vertical features. Experiments demonstrate that this simple design leads to an average improvement of 5.60% in Dice score and 7.81% in mean interaction of units (mIoU) for segmentation models, as well as an enhancement of 5.57% in peak signal-to-noise ratio (PSNR) and 8.22% in structural similarity (SSIM) for reconstruction models. We further utilize Flemme as an analytical tool to assess the effectiveness and efficiency of various encoders across different tasks. Code is available at https://github.com/wlsdzyzl/flemme. Yang Li 0104 |
BIBM | 3 |
| 2024 | RealTCD: Temporal Causal Discovery from Interventional Data with Large Language ModelabstractIn the field of Artificial Intelligence for Information Technology Operations, causal discovery is pivotal for operation and maintenance of systems, facilitating downstream industrial tasks such as root cause analysis. Temporal causal discovery, as an emerging method, aims to identify temporal causal relations between variables directly from observations by utilizing interventional data. However, existing methods mainly focus on synthetic datasets with heavy reliance on interventional targets and ignore the textual information hidden in real-world systems, failing to conduct causal discovery for real industrial scenarios. To tackle this problem, in this paper we investigate temporal causal discovery in industrial scenarios, which faces two critical challenges: how to discover causal relations without the interventional targets that are costly to obtain in practice, and how to discover causal relations via leveraging the textual information in systems which can be complex yet abundant in industrial contexts. To address these challenges, we propose the RealTCD framework, which is able to leverage domain knowledge to discover temporal causal relations without interventional targets. We first develop a score-based temporal causal discovery method capable of discovering causal relations without relying on interventional targets through strategic masking and regularization. Then, by employing Large Language Models (LLMs) to handle texts and integrate domain knowledge, we introduce LLM-guided meta-initialization to extract the meta-knowledge from textual information hidden in systems to boost the quality of discovery. We conduct extensive experiments on both simulation datasets and our real-world application scenario to show the superiority of our proposed RealTCD over existing baselines in temporal causal discovery. Peiwen Li, Xin Wang 0019, Zeyang Zhang 0001, Fang Shen, Yue Li 0053, Jialong Wang 0001, Yang Li 0104, Wenwu Zhu 0001 |
CIKM | 8 |
| 2024 | Enhancing Implicit Shape Generators Using Topological RegularizationsabstractA fundamental problem in learning 3D shapes generative models is that when the generative model is simply fitted to the training data, the resulting synthetic 3D models can present various artifacts. Many of these artifacts are topological in nature, e.g., broken legs, unrealistic thin structures, and small holes. In this paper, we introduce a principled approach that utilizes topological regularization losses on an implicit shape generator to rectify topological artifacts. The objectives are two-fold. The first is to align the persistent diagram (PD) distribution of the training shapes with that of synthetic shapes. The second ensures that the PDs are smooth among adjacent synthetic shapes. We show how to achieve these two objectives using two simple but effective formulations. Specifically, distribution alignment is achieved to learn a generative model of PDs and align this generator with PDs of synthetic shapes. We show how to handle discrete and continuous variabilities of PDs by using a shape-regularization term when performing PD alignment. Moreover, we enforce the smoothness of the PDs using a smoothness loss on the PD generator, which further improves the behavior of PD distribution alignment. Experimental results on ShapeNet show that our approach leads to much better generalization behavior than state-of-the-art implicit shape generators. Yang Li 0104, Lohit Anirudh Jagarapu, Hao Kang, Gang Hua 0001, Qixing Huang |
ICML | 3 |
| 2024 | A Geometric Algorithm for Blood Vessel Reconstruction from Skeletal Representation
Yang Li 0104 |
ISBRA (1) | 2 |
| 2024 | Safe Routes, Safer Rides: A Multi-Tiered Approach to Trajectory Anomaly DetectionabstractThe swift expansion of the ride-hailing industry has given rise to pressing safety issues. This study investigates enhancing safety in the ride-hailing industry by developing an anomaly detection system for detecting abnormal vehicle trajectories. This system defines what constitutes an abnormal trajectory and establishes evaluation metrics for detection methods. Traditional methods for anomaly detection, such as real-time monitoring of speed and direction, are inadequate in the face of data heterogeneity and the sheer volume of information. And the proposed anomaly detection system is designed to overcome these limitations: We introduce a multi-tiered strategy, dividing the anomaly detection task into critical points (including origin and destination) analysis and path analyses to effectively identify and categorize abnormal patterns. By concentrating different features or patterns in each tier, the detection capability of this system is effectively enhanced. Combining our evaluation metrics, the system can provide risk assessment for abnormal critical points and trajectories, thus ensuring the safety of both drivers and passengers. Chenye Wu, Zuxin Li, Yang Li 0104 |
MobiCom | 4 |
| 2024 | Predicting community case transfer path and processing time using decoder modelsabstractGovernment agencies and non-profit organizations often rely on case management systems to process the large influx of community request cases. To improve the efficiency of community case management, it's important to model how a community request case is transferred between different departments within the organization and how long it takes to resolve the case. In this paper, we propose two decoder models to predict the departmental transfer path of a given community case and estimate the total processing time based on the predicted path, trained on historical community case records. We compared our prediction results with those obtained using other common machine learning models on a dataset collected from multiple community platforms in Shenzhen, China. Experiments show that our proposed method significantly outperforms the baselines in transfer path and total processing time prediction. Yuanbo Tang, Qingmin Liao, Yang Li 0104 |
MobiCom | 7 |
| 2024 | Enhancing Continuous Domain Adaptation with Multi-path Transfer Curriculum
Jingge Wang, Xuan Zhang 0004, Yang Li 0104 |
PAKDD (2) | 5 |
| 2024 | Cross-Sign Language Transfer Learning Using Domain Adaptation with Multi-scale Temporal Alignment
Keren Artiaga, Yang Li 0104, Ercan E. Kuruoglu, Wai Kin Chan |
Multim. Tools Appl. | 2 |
| 2024 | A Transferability-Based Method for Evaluating the Protein Representation LearningabstractSelf-supervised pre-trained language models have recently risen as a powerful approach in learning protein representations, showing exceptional effectiveness in various biological tasks, such as drug discovery. Amidst the evolving trend in protein language model development, there is an observable shift towards employing large-scale multimodal and multitask models. However, the predominant reliance on empirical assessments using specific benchmark datasets for evaluating these models raises concerns about the comprehensiveness and efficiency of current evaluation methods. Addressing this gap, our study introduces a novel quantitative approach for estimating the performance of transferring multi-task pre-trained protein representations to downstream tasks. This transferability-based method is designed to quantify the similarities in latent space distributions between pre-trained features and those fine-tuned for downstream tasks. It encompasses a broad spectrum, covering multiple domains and a variety of heterogeneous tasks. To validate this method, we constructed a diverse set of protein-specific pre-training tasks. The resulting protein representations were then evaluated across several downstream biological tasks. Our experimental results demonstrate a robust correlation between the transferability scores obtained using our method and the actual transfer performance observed. This significant correlation highlights the potential of our method as a more comprehensive and efficient tool for evaluating protein representation learning. Weihong Zhang 0002, Huazhen Huang, Yang Li 0104 |
IEEE J. Biomed. Health Informatics | 5 |
| 2023 | Investigating Consistency Constraints in Heterogeneous Multi-task Learning for Medical Image ProcessingabstractMulti-task learning (MTL), a learning paradigm that leverages the relationships among different tasks to concurrently improve their prediction performance, has been widely applied in the medical image processing field. To tackle the challenge of inconsistent task-specific predictions in early MTL methods, recent works primarily focus on incorporating certain consistency constraints into model training for specific MTL settings, which is difficult to generalize to heterogeneous medical image processing tasks. Meanwhile, the synergy of different consistency objectives has not been well explored. In this paper, we address these gaps by proposing a novel multi-task consistency regularization scheme. It consists of two complementary consistency constraints: Task-level Output Consistency (TOC) which constrains different tasks to give consistent predictions and Feature-level Representation Consistency (FRC) which enforces feature representation consistency between target and auxiliary tasks. Furthermore, we present an optimal strategy to select appropriate constraints for any given set of tasks. Extensive experiments on three popular benchmark datasets were conducted in both the few-shot and semi-supervised settings. The results demonstrate our proposed multi-task consistency regularization scheme improves MTL performance by 13.7% in Dice score on average. Additionally, we show that FRC is effective under all MTL settings, achieving an average 6.3% gain in accuracy. Yicong Li 0003, Yang Tan 0004, Yang Li 0104 |
BIBM | 5 |
| 2023 | Topology-Preserving Hard Pixel Mining for Tubular Structure Segmentation
Caixia Dong, Yang Li 0104 |
BMVC | 3 |
| 2023 | Explainable Trajectory Representation through Dictionary LearningabstractTrajectory representation learning on a network enhances our understanding of vehicular traffic patterns and benefits numerous downstream applications. Existing approaches using classic machine learning or deep learning embed trajectories as dense vectors, which lack interpretability and are inefficient to store and analyze in downstream tasks. In this paper, an explainable trajectory representation learning framework through dictionary learning is proposed. Given a collection of trajectories on a network, it extracts a compact dictionary of commonly used subpaths called "pathlets", which optimally reconstruct each trajectory by simple concatenations. The resulting representation is naturally sparse and encodes strong spatial semantics. Theoretical analysis of our proposed algorithm is conducted to provide a probabilistic bound on the estimation error of the optimal dictionary. A hierarchical dictionary learning scheme is also proposed to ensure the algorithm's scalability on large networks, leading to a multi-scale trajectory representation. Our framework is evaluated on two large-scale real-world taxi datasets. Compared to previous work, the dictionary learned by our method is more compact and has better reconstruction rate for new trajectories. We also demonstrate the promising performance of this method in downstream tasks including trip time prediction task and data compression. Yuanbo Tang, Yang Li 0104 |
SIGSPATIAL/GIS | 3 |
| 2023 | SDG-L: A Semiparametric Deep Gaussian Process based Framework for Battery Capacity PredictionabstractLithium-ion batteries are becoming increasingly omnipresent in energy supply. However, the durability of energy storage using lithium-ion batteries is threatened by their dropping capacity with the growing number of charging/discharging cycles. An accurate capacity prediction is the key to ensure system efficiency and reliability, where the exploitation of battery state information in each cycle has been largely undervalued. In this paper, we propose a semiparametric deep Gaussian process regression framework named SDG-L to give predictions based on the modeling of time series battery state data. By introducing an LSTM feature extractor, the SDG-L is specially designed to better utilize the auxiliary profiling information during charging/discharging process. In experimental studies based on NASA dataset, our proposed method obtains an average test MSE error of 1.2‰. We also show that SDG-L achieves better performance compared to existing works and validate the framework using ablation studies. Yanru Wu, Yang Li 0104, Ercan E. Kuruoglu, Xuan Zhang 0004 |
ICASSP | 3 |
| 2023 | Efficient Prediction of Model Transferability in Semantic Segmentation TasksabstractHow to efficiently select highly transferable pretrained models remains a challenging problem in few-shot semantic segmentation tasks. Existing transferability metrics for classification tasks are difficult to compute on segmentation data due to the high-dimensional output of the segmentation model. In this work, we generalize existing transferability metrics to efficiently predict the transferability of semantic segmentation models, by calculating transferability scores over the sampled pixel-wise features. Then with the help of transferability, we propose a transferability-weighted finetuning method which puts more importance on those low-transferability regions to improve the overall transfer accuracy on the target task. Experiments on a challenging benchmark show that the transferability scores produced by our adaptation method are highly correlated with the groundtruth transfer accuracy, achieving 0.718 Spearman’s correlation coefficient on average and at least 67× gain on efficiency. In addition, our transferability-weighted finetuning method outperforms vanilla fine-tuning by 4% in transfer accuracy. Yang Tan 0004, Yicong Li 0003, Yang Li 0104, Xiao-Ping Zhang 0003 |
ICIP | 3 |
| 2023 | Session-based recommendation with temporal dynamics for large volunteer networks
Taurai Muvunza, Yang Li 0104 |
J. Intell. Inf. Syst. | 2 |
| 2022 | Finding the Most Transferable Tasks for Brain Image SegmentationabstractAlthough many studies have successfully applied transfer learning to medical image segmentation, very few of them have investigated the selection strategy when multiple source tasks are available for transfer. In this paper, we propose a prior knowledge guided and transferability based framework to select the best source tasks among a collection of brain image segmentation tasks, to improve the transfer learning performance on the given target task. The framework consists of modality analysis, RoI (region of interest) analysis, and transferability estimation, such that the source task selection can be refined step by step. Specifically, we adapt the state-of-the-art analytical transferability estimation metrics to medical image segmentation tasks and further show that their performance can be significantly boosted by filtering candidate source tasks based on modality and RoI characteristics. Our experiments on brain matter, brain tumor, and white matter hyperintensities segmentation datasets reveal that transferring from different tasks under the same modality is often more successful than transferring from the same task under different modalities. Furthermore, within the same modality, transferring from the source task that has stronger RoI shape similarity with the target task can significantly improve the final transfer performance. And such similarity can be captured using the Structural Similarity index in the label space. Yicong Li 0003, Yang Tan 0004, Yang Li 0104, Xiao-Ping Zhang 0003 |
BIBM | 4 |
| 2022 | Riemannian Geometric Instance Filtering for Transfer Learning in Brain-Computer InterfacesabstractDue to the inter-subject variability of Electroencephalogram(EEG) signals, a long calibration time is required to collect a large number of labeled trials to calibrate classifier parameters before using the Brain-computer Interface(BCI). This challenge greatly limits the practical roll-out of BCIs. To address this problem, we propose a novel instance-based transfer learning framework named Riemannian Geometric Instance Filtering (RGIF) to reduce calibration time without sacrificing accuracy. A new inter-subject similarity metric based on Riemannian geometry is proposed to measure the similarity between a few trials from the target subject and adequate trials from source subjects. The classification model for the target subject is then trained with the help of abundant trials from similar source subjects with high similarity to the target subject. We evaluate our method on two open-source EEG datasets. The results show that our approach improves significantly compared with other baselines. Furthermore, compared with using all source subjects data, our method reduces the training time by at least half and achieves slightly better accuracy. Qianxin Hui, Yang Li 0104, Susu Xu, Shuailei Zhang, Ying Sun 0012, Shuai Wang 0049, Xinlei Chen, Dezhi Zheng |
SenSys | 3 |
| 2021 | OTCE: A Transferability Metric for Cross-Domain Cross-Task RepresentationsabstractTransfer learning across heterogeneous data distributions (a.k.a. domains) and distinct tasks is a more general and challenging problem than conventional transfer learning, where either domains or tasks are assumed to be the same. While neural network based feature transfer is widely used in transfer learning applications, finding the optimal transfer strategy still requires time-consuming experiments and domain knowledge. We propose a transferability metric called Optimal Transport based Conditional Entropy (OTCE), to analytically predict the transfer performance for supervised classification tasks in such cross-domain and cross-task feature transfer settings. Our OTCE score characterizes transferability as a combination of domain difference and task difference, and explicitly evaluates them from data in a unified framework. Specifically, we use optimal transport to estimate domain difference and the optimal coupling between source and target distributions, which is then used to derive the conditional entropy of the target task (task difference). Experiments on the largest cross-domain dataset DomainNet and Office31 demonstrate that OTCE shows an average of 21% gain in the correlation with the ground truth transfer accuracy compared to state-of-the-art methods. We also investigate two applications of the OTCE score including source model selection and multi-source feature fusion. Yang Tan 0004, Yang Li 0104, Shao-Lun Huang |
CVPR | 2 |
| 2021 | Semi-Supervised Multimodal Image Translation for Missing Modality ImputationabstractMissing data is a common problem in multimodal and multi-view learning. It raises a critical challenge for most multimodal algorithms, which are unable to deal with incomplete datasets. Rather than discarding entries with missing modalities, this paper aims to reconstruct the complete image-based multimodal data by imputing missing modalities. We solve the imputation problem as an image translation task, which transforms images in one domain to other domains. Existing image translation techniques either can not fully utilize the information contained in partially complete entries or are limited to the bimodal situation. We propose a semi-supervised algorithm for multimodal learning with missing data, namely Cyclic Autoencoder (CycAE). Specifically, a novel cyclical structure, as well as the correlation among modalities, is integrated to leverage infoπnation from complete entries to incomplete ones. Experiments on two multimodal datasets show that our model outperforms state-of-the-art models. Downstream tasks can also benefit from the completed datasets. Wangbin Sun, Fei Ma 0006, Yang Li 0104, Shao-Lun Huang, Shiguang Ni, Lin Zhang 0001 |
ICASSP | 3 |
| 2021 | Joint PVL Detection and Manual Ability Classification Using Semi-supervised Multi-task Learning
Yicong Li 0003, Yang Li 0104 |
MICCAI (7) | 5 |
| 2020 | Semantically Supervised Maximal Correlation For Cross-Modal RetrievalabstractWith the rapid growth of multimedia data, the cross-modal retrieval problem has attracted a lot of interest in both research and industry in recent years. However, the inconsistency of data distribution from different modalities makes such task challenging. In this paper, we propose Semantically Supervised Maximal Correlation (S2MC) method for cross-modal retrieval by incorporating semantic label information into the traditional maximal correlation framework. Combining with maximal correlation based method for extracting unsupervised pairing information, our method effectively exploits supervised semantic information on both common feature space and label space. Extensive experiments show that our method outperforms other current state-of-the-art methods on cross-modal retrieval tasks on three widely used datasets. Yang Li 0104, Shao-Lun Huang, Lin Zhang 0001 |
ICIP | 2 |
| 2020 | Person Recognition with HGR Maximal Correlation on Multimodal DataabstractMultimodal person recognition is a common task in video analysis and public surveillance, where information from multiple modalities, such as images and audio extracted from videos, are used to jointly determine the identity of a person. Previous person recognition techniques either use only uni-modal data or only consider shared representations between different input modalities, while leaving the extraction of their relationship with identity information to downstream tasks. Furthermore, real-world data often contain noise, which makes recognition more challenging practical situations. In our work, we propose a novel correlation-based multimodal person recognition framework that is relatively simple but can efficaciously learn supervised information in multimodal data fusion and resist noise. Specifically, our framework learns a discriminative embeddings of persons by joint learning visual features and audio features while maximizing HGR maximal correlation among multimodal input and persons' identities. Experiments are done on a subset of Voxceleb2. Compared with state-of-the-art methods, the proposed method demonstrates an improvement of accuracy and robustness to noise. Yihua Liang, Fei Ma 0006, Yang Li 0104, Shao-Lun Huang |
ICPR | 3 |
| 2020 | Mining Regional Mobility Patterns for Urban Dynamic Analytics
Jing Lian 0003, Yang Li 0104, Weixi Gu, Shao-Lun Huang, Lin Zhang 0001 |
Mob. Networks Appl. | 2 |
| 2019 | HTTE: A Hybrid Technique For Travel Time Estimation In Sparse Data EnvironmentsabstractTravel time estimation is a critical task, useful to many urban applications at the individual citizen and the stakeholder level. This paper presents a novel hybrid algorithm for travel time estimation that leverages historical and sparse real-time trajectory data. Given a path and a departure time we estimate the travel time taking into account the historical information, the real-time trajectory data and the correlations among different road segments. We detect similar road segments using historical trajectories, and use a latent representation to model the similarities. Our experimental evaluation demonstrates the effectiveness of our approach. Nikolaos Zygouras, Nikolaos Panagiotou, Yang Li 0104, Dimitrios Gunopulos, Leonidas J. Guibas |
SIGSPATIAL/GIS | 3 |
| 2019 | An Information-Theoretic Approach to Transferability in Task Transfer LearningabstractTask transfer learning is a popular technique in image processing applications that uses pre-trained models to reduce the supervision cost of related tasks. An important question is to determine task transferability, i.e. given a common input domain, estimating to what extent representations learned from a source task can help in learning a target task. Typically, transferability is either measured experimentally or inferred through task relatedness, which is often defined without a clear operational meaning. In this paper, we present a novel metric, H-score, an easily-computable evaluation function that estimates the performance of transferred representations from one task to another in classification problems using statistical and information theoretic principles. Experiments on real image data show that our metric is not only consistent with the empirical transferability measurement, but also useful to practitioners in applications such as source model selection and task transfer curriculum learning. Yajie Bao, Yang Li 0104, Shao-Lun Huang, Lin Zhang 0001, Lizhong Zheng, Amir Zamir, Leonidas J. Guibas |
ICIP | 2 |
| 2019 | Maximal Correlation Embedding Network for Multilabel Learning with Missing LabelsabstractMultilabel learning, the problem of mapping each data instance to a subset of labels, appears frequently in many real-world applications. However, obtaining complete label annotation for every instance requires tremendous efforts, especially when the label set is large. As a result, multilabel learning with missing labels remains as a common challenge. Existing works either cannot handle missing labels or lack nonlinear expressiveness and scalability to large label set. In this paper, we present a novel end-to-end solution for multilabel learning with missing labels. Our algorithm, Maximal Correlation Embedding Network learns a low dimensional label embedding using an encoder-decoder architecture. It exploits label similarity through a maximal correlation regularization in the embedded label space to reduce the classification bias due to missing labels. A series of experiments on popular multilabel datasets demonstrate that our approach outperforms state of the art, both in complete data and partially observed data. Yang Li 0104, Xiangxiang Xu 0001, Shao-Lun Huang, Lin Zhang 0001 |
ICME | 2 |
| 2019 | An End-to-End Learning Approach for Multimodal Emotion Recognition: Extracting Common and Private InformationabstractMultimodal emotion recognition is important for facilitating efficient interaction between humans and machines. To better detect emotional states from multimodal data, we need to effectively extract both the common information that captures dependencies among different modalities, and the private information that characterizes variations in each modality. However, existing works are mostly designed to pursue either one of these objectives but not both. In our work, we propose an end-to-end learning approach to simultaneously extract the common and private information for multimodal emotion recognition. Specifically, we use a correlation loss based on Hirschfeld-Gebelein-Renyi (HGR) maximal correlation and a reconstruction loss based on autoencoders to preserve the common and private information, respectively. Experimental results on eNTERFACE'05 database and RML database demonstrate the effectiveness of our proposed approach. Fei Ma 0006, Wei Zhang 0185, Yang Li 0104, Shao-Lun Huang, Lin Zhang 0001 |
ICME | 3 |
| 2019 | Info-Detection: An Information-Theoretic Approach to Detect Outlier
Fei Ma 0006, Yang Li 0104, Shao-Lun Huang, Lin Zhang 0001 |
ICONIP (5) | 3 |
| 2019 | A maximal correlation embedding method for multilabel human context recognition: poster abstractabstractReal-time human context recognition is one of the most exciting emerging technologies in sensing nowadays. Compared with most recognition problems in machine learning, the challenge lies in the complexity and incompleteness of labels, in other words, each sample can have several label concepts simultaneously but some of them could be missing. This poster proposes an effective approach for multilabel human context recognition with signals from sensors embedded in the wearable devices. The proposed algorithm demonstrates to be very robust to incomplete labels. Yang Li 0104, Xiangxiang Xu 0001, Lin Zhang 0001 |
IPSN | 2 |
| 2018 | Joint Mobility Pattern Mining with Urban Region PartitionsabstractMobility pattern mining answers the fundamental question of where people are likely to go from a given location. It plays an important role in city planning, public transport management and location-based mobile applications. Among these applications, many concern the mobility pattern over contiguous spatial regions as a whole. Traditional ways of mobility pattern mining either result in trip clusters with overlapped origin and destination regions, or require an extra step to partition the city into discrete regions, which may not be optimal for mobility pattern extraction. In this paper, we present a region-aware mobility pattern mining framework to jointly extract trip clusters while maintaining non-overlapping partitions of trip origins and destinations. We developed kernelized ACE, a novel extension to a classic algorithm in statistics to compute the optimal mobility clusters under spatial constraints. Experimental results using Beijing taxi trip data show that our approach outperforms other methods with only ~ 0.3% spatial overlap and 86.43% origin-destination correlation. Our case studies on New York City's and Beijing's taxi datasets also yield insightful findings that reveal city-scale mobility patterns and propose potential improvement for public transportation. Jing Lian 0003, Yang Li 0104, Weixi Gu, Shao-Lun Huang, Lin Zhang 0001 |
MobiQuitous | 2 |
| 2017 | Urban Travel Time Prediction using a Small Number of GPS Floating CarsabstractPredicting the travel time of a path is an important task in route planning and navigation applications. As more GPS floating car data has been collected to monitor urban traffic, GPS trajectories of floating cars have been frequently used to predict path travel time. However, most trajectory-based methods rely on deploying GPS devices and collect real-time data on a large taxi fleet, which can be expensive and unreliable in smaller cities. This work deals with the problem of predicting path travel time when only a small number of GPS floating cars are available. We developed an algorithm that learns local congestion patterns of a compact set of frequently shared paths from historical data. Given a travel time prediction query, we identify the current congestion patterns around the query path from recent trajectories, then infer its travel time in the near future. Experimental results using 10-15 taxis tracked for 11 months in urban areas of Shenzhen, China show that our prediction has on average 5.4 minutes of error on trips of duration 10-75 minutes. This result improves the baseline approach of using purely historical trajectories by 2-30% on regions with various degree of path regularity. It also outperforms a state-of-the-art travel time prediction method that uses both historical trajectories and real-time trajectories. Yang Li 0104, Dimitrios Gunopulos, Cewu Lu, Leonidas J. Guibas |
SIGSPATIAL/GIS | 1 |
| 2016 | Knowledge-based trajectory completion from sparse GPS samplesabstractTraffic trajectories collected from GPS-enabled mobile devices or vehicles are widely used in urban planning, traffic management, and location based services. Their performance often relies on having dense trajectories. However, due to the power and bandwidth limitation on these devices, collecting dense trajectory is too costly on a large scale. We show that by exploiting structural regularity in large trajectory data, the complete geometry of trajectories can be inferred from sparse GPS samples without information about the underlying road network - a process called trajectory completion. In this paper, we present a knowledge-based approach for completing traffic trajectories. Our method extracts a network of road junctions and estimates traffic flows across junctions. GPS samples within each flow cluster are then used to achieve fine-level completion of individual trajectories. Finally, we demonstrate that our method is effective for trajectory completion on both synthesized and real traffic trajectories. On average 72.7% of real trajectories with sampling rate of 60 seconds/sample are completed without map information. Comparing to map matching, over 89% of points on completed trajectories are within 15 meters from the map matched path. Yang Li 0104, Yangyan Li, Dimitrios Gunopulos, Leonidas J. Guibas |
SIGSPATIAL/GIS | 1 |
| 2013 | Large-scale joint map matching of GPS tracesabstractWe present a robust method for solving the map matching problem exploiting massive GPS trace data. Map matching is the problem of determining the path of a user on a map from a sequence of GPS positions of that user --- what we call a trajectory. Commonly obtained from GPS devices, such trajectory data is often sparse and noisy. As a result, the accuracy of map matching is limited due to ambiguities in the possible routes consistent with trajectory samples. Our approach is based on the observation that many regularity patterns exist among common trajectories of human beings or vehicles as they normally move around. Among all possible connected k-segments on the road network (i.e., consecutive edges along the network whose total length is approximately k units), a typical trajectory collection only utilizes a small fraction. This motivates our data-driven map matching method, which optimizes the projected paths of the input trajectories so that the number of the k-segments being used is minimized. We present a formulation that admits efficient computation via alternating optimization. Furthermore, we have created a benchmark for evaluating the performance of our algorithm and others alike. Experimental results demonstrate that the proposed approach is superior to state-of-art single trajectory map matching techniques. Moreover, we also show that the extracted popular k-segments can be used to process trajectories that are not present in the original trajectory set. This leads to a map matching algorithm that is as efficient as existing single trajectory map matching algorithms, but with much improved map matching accuracy. Yang Li 0104, Qixing Huang, Michael Kerber, Lin Zhang 0001, Leonidas J. Guibas |
SIGSPATIAL/GIS | 1 |